ABCDEFGHIJKLMNOPQRSTUVWXYZ
1
2
To use this spreadsheet click on "File" and "Make a copy" for yourself. Then you can edit the respective values.
3
4
Data size (in TB)9.5Calculations
5
# Documents5225000# Paragraphs156,750,000
6
# Paragraphs per document30Embedding size1. B
7
Average paragraph size (text + metadata)500 B
8
Time/paragraph0.00007442483557Mins
Based on https://docs.google.com/spreadsheets/d/1X0dFmL5jyikzS-A0-Ba--Q9PJZpeHr6XkS8bNKsz0IU/edit?gid=0#gid=0
9
EmbeddingsTotal Time11666.09298Mins
10
# Dimensions256Total Time8.101453455Days
11
Data typebyte2x Buffer16.20290691Days
12
13
14
Requirements
15
Storage110.36 GiB
16
Memory39.12 GiB
17
18
For a better explanation of the calculations presented in this spreadsheet see: https://docs.squirro.com/en/latest/technical/search/features/semantic-search.html#hardware-requirements
19
20
Recommendations based on ECB dataset
21
22
- Either 500k documents or a TB of data for each agency.
- No big excels or zip files to be ingested.
- One GPU over the full lifetime, and 2-3 GPUs for the first month.
- State out all the assumptions about index document size, number of paragraphs, embedding model use, and time consumption on the GPU
- Staggered rollout recommendation per agency.
- Add a rhein-insight machine windows machine.
- Do keywords first, validate the data and then in low priority ingester pipelines with embedding
23
jFrog Artifactory for mirror
24
25
Total assumed Index Size
26
with a per document size contribution of
27
512 kb1.1TB
28
29
30
31
32
33
### Rhein
34
Hard-disk/1m documents100GB
35
36
Number of CPUs4
37
RAM32GB
38
Hard disk needed522.5GB
39
OSLinux
40
DBPostgres
41
DeploymentContainer
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100