Resource requirements
This article describes the infrastructure the Data and Search Platform depends on and what drives its resource footprint, so you can plan capacity for your deployment.
Components
The Data and Search Platform is composed of a retrieval service (handles the HTTP API: search stores, uploads, search) and an ingestion worker pool (runs document-processing workflows), backed by:
| Component | Concrete container | Purpose |
|---|---|---|
Retrieval service |
Sherlock |
HTTP API: search-store/file/document CRUD, hybrid vector search, file upload/download, ingestion workflow runs, MCP endpoint. |
Content delivery |
Sherlock CDN |
Content-serving deployment of the same Sherlock image; authorized file downloads streamed from object storage with ETag/Range support. |
Ingestion worker pool |
Holmes worker |
Temporal worker running ingestion workflows: collects files, parses/OCRs, chunks, embeds, writes pages/chunks to Postgres and vectors to Qdrant. |
Vector database |
Qdrant |
Stores chunk embeddings; queried for every search request. One collection per search store. |
Relational database |
PostgreSQL |
Stores metadata, extracted page text, and workflow/run state. |
Object storage |
Object storage (S3/MinIO) |
Stores uploaded files and export bundles (S3-compatible). |
Identity and authorization |
IAM sidecar (+ Envoy) and OpenFGA |
The IAM sidecar supplies user identity on every request; OpenFGA enforces per-search-store access control. |
Workflow orchestration |
Temporal |
Schedules and tracks ingestion workflow runs. |
Distributed compute (optional) |
Dask scheduler + workers |
Scales out ingestion processing for high document volumes. |
Document conversion (optional) |
Gotenberg |
Converts Office documents to PDF. Off by default. |
External dependencies
These are the services a deployment has to be able to reach — the identity pieces integrate it with the rest of the platform, the rest is infrastructure provisioned alongside it. Dashed edges are optional.
%%{init: {"theme": "base", "flowchart": {"curve": "basis", "padding": 18, "nodeSpacing": 45, "rankSpacing": 55}}}%%
flowchart LR
client([Your application])
subgraph identity["Platform identity"]
sidecar["Auth sidecar<br/>validates the token,<br/>stamps the caller"]
idp["Identity provider<br/>Pharia IAM on-prem"]
fga["OpenFGA<br/>store grants and roles"]
end
subgraph dsp["Data and Search Platform"]
direction TB
api["Retrieval service<br/>search stores, uploads, search"]
workers["Ingestion workers<br/>document workflows"]
end
subgraph state["State"]
pg[("PostgreSQL<br/>metadata, page text, run state")]
vectors[("Vector database<br/>chunk embeddings")]
files[("Object storage<br/>uploaded files")]
end
subgraph compute["Execution and models"]
temporal["Temporal<br/>workflow orchestration"]
inference["Inference endpoints<br/>embeddings, vision, rerank"]
dask["Dask cluster<br/>ingestion scale-out"]
end
subgraph obs["Observability"]
otlp["OTLP collector<br/>audit records and traces"]
metrics["Prometheus and Grafana"]
end
client --> sidecar
sidecar --> idp
sidecar -->|"identity stamped"| api
api --> fga
api --> pg
api --> vectors
api --> files
api --> inference
api --> temporal
temporal --> workers
workers --> pg
workers --> vectors
workers --> files
workers --> inference
workers -.-> dask
api --> otlp
workers --> otlp
metrics -.->|scrapes| api
classDef optional stroke-dasharray: 5 4
class dask,metrics optional
The same set, with the details that matter when you provision them:
| Dependency | Needed for | Operator | Location | Notes |
|---|---|---|---|---|
Identity provider |
Authenticating every request |
self |
eu01 |
The service never validates tokens itself: a sidecar (IAM sidecar + Envoy) checks the bearer token against the deployment’s identity provider — Pharia IAM on-prem, or Zitadel userinfo — and stamps the caller’s identity ( |
OpenFGA |
Per-search-store authorization |
self |
eu01 |
Holds the owner/writer/viewer grants and service-level admin roles. Only deployed where |
PostgreSQL |
Metadata, extracted text, workflow and run state |
STACKIT (managed) |
eu01 |
Pages and chunks store full extracted document text, not just metadata, plus workflow run inputs and error messages. Reached through the platform pgdog pooler. Open: managed-Postgres legal entity per environment still to be confirmed. |
Vector database (Qdrant) |
Chunk embeddings and search |
self (+ Qdrant Cloud where used) |
eu01 |
Runs in-cluster via the Qdrant operator. Payload carries |
S3-compatible object storage |
Uploaded files |
STACKIT (managed) |
eu01 |
Uploaded files carry the |
Temporal |
Ingestion workflow orchestration |
self |
eu01 |
Needs a namespace and task queue. Histories persist run inputs and activity failure payloads, which can embed document text from stack traces. Production runs self-hosted Temporal; staging runs on Temporal Cloud via api_key auth. Planned mitigations: self-hosting and encryption. |
Inference endpoints |
Embeddings, vision/OCR extraction, reranking, transcription |
self (Pharia) |
eu01 |
Which provider serves a given call is configured per search store and per workflow; a deployment needs credentials for each provider it offers. Search query text and document-derived content are forwarded here. Open: whether the third-party inference provider should be listed as a separate operator. |
OTLP collector |
Audit events and traces |
self |
eu01 |
Administrative actions are emitted as OpenTelemetry log records to an OTLP endpoint rather than written to a database, so audit retention is a property of your log pipeline, not of this service. |
Prometheus and Grafana (optional) |
Metrics and dashboards |
self |
eu01 |
|
Dask cluster (optional) |
Distributed ingestion compute |
self |
eu01 |
Only needed to scale ingestion beyond what the worker pool handles in-process. Enabled by default in the chart; holds the same document content as the Holmes worker in memory per task. |
Container inventory and data processing
| Container | What it does | Operator | Location | Personal data | Notes |
|---|---|---|---|---|---|
Sherlock |
Retrieval API: search-store/file/document CRUD, hybrid vector search, file upload/download, ingestion workflow runs and MCP endpoint |
self |
eu01 |
prompt, upload, user-id, credential, trace, ip |
Query text is forwarded to the inference provider (Pharia) for embedding/rerank. The end-user bearer token transits the pod but only the iam-sidecar reads it; the app trusts the stamped |
Sherlock CDN |
Content-serving deployment of the Sherlock image; authorized file downloads from object storage with ETag/Range support |
self |
eu01 |
upload, user-id, trace, ip |
Serves raw uploaded bytes after a per-request authorization check. Read-only filesystem, no upload path, no disk cache — nothing retained in the container. |
IAM sidecar (+ Envoy) |
Auth sidecar next to Sherlock/CDN pods: validates the bearer token against Pharia IAM or Zitadel userinfo and stamps |
self |
eu01 |
credential, user-id, ip |
Bearer token sent verbatim to the userinfo endpoint, cached in memory against the resolved identity for 60s. Userinfo responses may carry email/name claims — unconfirmed whether anything beyond id and roles is logged. |
Holmes worker |
Temporal worker running ingestion workflows: collects files, parses/OCRs, chunks, embeds, writes pages/chunks to Postgres and vectors to Qdrant (cpu/io/gpu worker groups) |
self |
eu01 |
prompt, completion, upload, user-id, trace |
Page images and extracted text go to the inference endpoint for VLM OCR and embeddings; output returns as page text. Credentials are service keys (inference API key / TVM-issued JWT), never the end user’s token. Activity failures can embed document snippets in Temporal histories. |
Dask scheduler + workers |
Distributed compute backend for Holmes modules (same Holmes image); document content passes through scheduler/worker memory during module execution |
self |
eu01 |
upload, user-id, trace |
Enabled by default in the chart. Same data as the Holmes worker, held in memory per task. Open: whether workers can spill to local disk under memory pressure. Spooling currently active only for the audio modality. |
Gotenberg |
Office document-to-PDF conversion API: receives raw uploaded file bytes, returns the converted PDF |
self |
eu01 |
upload |
Off by default, no workflow consumes it yet. Chainguard image pinned by digest, stateless, in-cluster only. |
Postgres |
Metadata and content store: search stores, files, folders, documents, pages (full extracted text), chunks, workflow runs |
STACKIT |
eu01 |
upload, user-id |
"upload" here is the extracted document text itself — pages/chunks store full content — plus workflow run inputs and error messages. Reached through the platform pgdog pooler. Open: managed-Postgres legal entity per environment. |
Qdrant |
Vector database: chunk embeddings with a payload of |
self (sheet’s first table: "self + Qdrant") |
eu01 |
user-id, upload |
Vectors computed from chunk text; payload can carry document metadata fields. In-cluster via the Qdrant operator. Qdrant Cloud / API-key mode supported by config but not deployed. |
Object storage (S3/MinIO) |
Binary file storage for uploaded files and export bundles |
STACKIT |
eu01 |
upload, user-id |
Uploaded files carry the |
Temporal |
Workflow orchestration for ingestion runs: task queues plus persisted workflow histories (run inputs, source URIs, failure payloads) |
self |
eu01 |
user-id, trace (sheet’s first table also lists "upload") |
Histories persist run inputs and failure payloads, which can embed document text from stack traces. Production self-hosted; staging on Temporal Cloud via api_key auth — if that stays, Temporal Technologies Inc. needs its own operator row. Planned mitigations: self-hosting and encryption. |
OpenFGA |
Authorization store: owner/viewer/writer relationship tuples between user ids and search stores |
self |
eu01 |
user-id |
Only deployed where |
What drives sizing
Exact CPU/memory figures depend heavily on your workload and are best confirmed with your platform team as part of deployment sizing; the drivers to plan around are:
Vector database
Memory scales with the number of indexed chunks, the embedding dimension, and the number of replicas configured for availability. More chunks (from more documents, or a smaller chunk_size producing more chunks per document) and higher-dimensional embeddings both increase memory usage per replica.
Relational database
Storage scales roughly with the volume of extracted text: page content, chunk metadata, and workflow/run history all accumulate here as you ingest more documents.
Object storage
Storage scales directly with the total size of uploaded files. Deduplication by content hash means re-uploading the same file to the same search store doesn’t add to this.
Ingestion throughput
Ingestion worker CPU/memory scales with concurrent workflow runs and document complexity — OCR and vision-based extraction (see Workflows and presets) are more resource-intensive per document than plain text extraction.
Best practices
-
Start conservative: begin with your platform team’s recommended minimums for a new deployment.
-
Monitor usage: watch vector database memory and ingestion queue depth once real document volume starts flowing through.
-
Scale gradually: increase resources based on observed usage rather than projected peak.
-
Plan for growth: as covered in What is the Data and Search Platform?, the same search store and workflow configuration should carry you from initial experimentation through production — only the infrastructure sizing needs to grow.