Resource requirements

This article describes the infrastructure the Data and Search Platform depends on and what drives its resource footprint, so you can plan capacity for your deployment.


Components

The Data and Search Platform is composed of a retrieval service (handles the HTTP API: search stores, uploads, search) and an ingestion worker pool (runs document-processing workflows), backed by:

Component Concrete container Purpose

Retrieval service

Sherlock

HTTP API: search-store/file/document CRUD, hybrid vector search, file upload/download, ingestion workflow runs, MCP endpoint.

Content delivery

Sherlock CDN

Content-serving deployment of the same Sherlock image; authorized file downloads streamed from object storage with ETag/Range support.

Ingestion worker pool

Holmes worker

Temporal worker running ingestion workflows: collects files, parses/OCRs, chunks, embeds, writes pages/chunks to Postgres and vectors to Qdrant.

Vector database

Qdrant

Stores chunk embeddings; queried for every search request. One collection per search store.

Relational database

PostgreSQL

Stores metadata, extracted page text, and workflow/run state.

Object storage

Object storage (S3/MinIO)

Stores uploaded files and export bundles (S3-compatible).

Identity and authorization

IAM sidecar (+ Envoy) and OpenFGA

The IAM sidecar supplies user identity on every request; OpenFGA enforces per-search-store access control.

Workflow orchestration

Temporal

Schedules and tracks ingestion workflow runs.

Distributed compute (optional)

Dask scheduler + workers

Scales out ingestion processing for high document volumes.

Document conversion (optional)

Gotenberg

Converts Office documents to PDF. Off by default.

External dependencies

These are the services a deployment has to be able to reach — the identity pieces integrate it with the rest of the platform, the rest is infrastructure provisioned alongside it. Dashed edges are optional.

%%{init: {"theme": "base", "flowchart": {"curve": "basis", "padding": 18, "nodeSpacing": 45, "rankSpacing": 55}}}%%
flowchart LR
  client([Your application])

  subgraph identity["Platform identity"]
    sidecar["Auth sidecar<br/>validates the token,<br/>stamps the caller"]
    idp["Identity provider<br/>Pharia IAM on-prem"]
    fga["OpenFGA<br/>store grants and roles"]
  end

  subgraph dsp["Data and Search Platform"]
    direction TB
    api["Retrieval service<br/>search stores, uploads, search"]
    workers["Ingestion workers<br/>document workflows"]
  end

  subgraph state["State"]
    pg[("PostgreSQL<br/>metadata, page text, run state")]
    vectors[("Vector database<br/>chunk embeddings")]
    files[("Object storage<br/>uploaded files")]
  end

  subgraph compute["Execution and models"]
    temporal["Temporal<br/>workflow orchestration"]
    inference["Inference endpoints<br/>embeddings, vision, rerank"]
    dask["Dask cluster<br/>ingestion scale-out"]
  end

  subgraph obs["Observability"]
    otlp["OTLP collector<br/>audit records and traces"]
    metrics["Prometheus and Grafana"]
  end

  client --> sidecar
  sidecar --> idp
  sidecar -->|"identity stamped"| api
  api --> fga
  api --> pg
  api --> vectors
  api --> files
  api --> inference
  api --> temporal
  temporal --> workers
  workers --> pg
  workers --> vectors
  workers --> files
  workers --> inference
  workers -.-> dask
  api --> otlp
  workers --> otlp
  metrics -.->|scrapes| api

  classDef optional stroke-dasharray: 5 4
  class dask,metrics optional

The same set, with the details that matter when you provision them:

Dependency Needed for Operator Location Notes

Identity provider

Authenticating every request

self

eu01

The service never validates tokens itself: a sidecar (IAM sidecar + Envoy) checks the bearer token against the deployment’s identity provider — Pharia IAM on-prem, or Zitadel userinfo — and stamps the caller’s identity (X-User-ID / X-User-Is-Admin) onto the request the service then trusts. Where traffic is fronted by Envoy, Envoy delegates its ext_authz check to that same sidecar. The bearer token is sent verbatim to the userinfo endpoint and cached in memory for 60s.

OpenFGA

Per-search-store authorization

self

eu01

Holds the owner/writer/viewer grants and service-level admin roles. Only deployed where AUTHORIZATION_MODE=fga; a deployment without it runs allow-all, appropriate only for local development. Tuples contain user and group ids only.

PostgreSQL

Metadata, extracted text, workflow and run state

STACKIT (managed)

eu01

Pages and chunks store full extracted document text, not just metadata, plus workflow run inputs and error messages. Reached through the platform pgdog pooler. Open: managed-Postgres legal entity per environment still to be confirmed.

Vector database (Qdrant)

Chunk embeddings and search

self (+ Qdrant Cloud where used)

eu01

Runs in-cluster via the Qdrant operator. Payload carries user_id, searchstore_id and filterable document metadata. Qdrant Cloud / API-key mode is supported by config but not what we deploy.

S3-compatible object storage

Uploaded files

STACKIT (managed)

eu01

Uploaded files carry the user_id in their object key (users/{user_id}/search-stores/…), so key listings alone are personal data. Tenant export zips are the exception: they sit outside that tree, at tenant-exports/{run_id}/{store_id}.zip. MinIO in dev, S3-compatible object storage in production.

Temporal

Ingestion workflow orchestration

self

eu01

Needs a namespace and task queue. Histories persist run inputs and activity failure payloads, which can embed document text from stack traces. Production runs self-hosted Temporal; staging runs on Temporal Cloud via api_key auth. Planned mitigations: self-hosting and encryption.

Inference endpoints

Embeddings, vision/OCR extraction, reranking, transcription

self (Pharia)

eu01

Which provider serves a given call is configured per search store and per workflow; a deployment needs credentials for each provider it offers. Search query text and document-derived content are forwarded here. Open: whether the third-party inference provider should be listed as a separate operator.

OTLP collector

Audit events and traces

self

eu01

Administrative actions are emitted as OpenTelemetry log records to an OTLP endpoint rather than written to a database, so audit retention is a property of your log pipeline, not of this service.

Prometheus and Grafana (optional)

Metrics and dashboards

self

eu01

Dask cluster (optional)

Distributed ingestion compute

self

eu01

Only needed to scale ingestion beyond what the worker pool handles in-process. Enabled by default in the chart; holds the same document content as the Holmes worker in memory per task.

Container inventory and data processing

Container What it does Operator Location Personal data Notes

Sherlock

Retrieval API: search-store/file/document CRUD, hybrid vector search, file upload/download, ingestion workflow runs and MCP endpoint

self

eu01

prompt, upload, user-id, credential, trace, ip

Query text is forwarded to the inference provider (Pharia) for embedding/rerank. The end-user bearer token transits the pod but only the iam-sidecar reads it; the app trusts the stamped X-User-ID. Uploaded bytes stream through to object storage.

Sherlock CDN

Content-serving deployment of the Sherlock image; authorized file downloads from object storage with ETag/Range support

self

eu01

upload, user-id, trace, ip

Serves raw uploaded bytes after a per-request authorization check. Read-only filesystem, no upload path, no disk cache — nothing retained in the container.

IAM sidecar (+ Envoy)

Auth sidecar next to Sherlock/CDN pods: validates the bearer token against Pharia IAM or Zitadel userinfo and stamps X-User-ID / X-User-Is-Admin; Envoy delegates ext_authz to it

self

eu01

credential, user-id, ip

Bearer token sent verbatim to the userinfo endpoint, cached in memory against the resolved identity for 60s. Userinfo responses may carry email/name claims — unconfirmed whether anything beyond id and roles is logged.

Holmes worker

Temporal worker running ingestion workflows: collects files, parses/OCRs, chunks, embeds, writes pages/chunks to Postgres and vectors to Qdrant (cpu/io/gpu worker groups)

self

eu01

prompt, completion, upload, user-id, trace

Page images and extracted text go to the inference endpoint for VLM OCR and embeddings; output returns as page text. Credentials are service keys (inference API key / TVM-issued JWT), never the end user’s token. Activity failures can embed document snippets in Temporal histories.

Dask scheduler + workers

Distributed compute backend for Holmes modules (same Holmes image); document content passes through scheduler/worker memory during module execution

self

eu01

upload, user-id, trace

Enabled by default in the chart. Same data as the Holmes worker, held in memory per task. Open: whether workers can spill to local disk under memory pressure. Spooling currently active only for the audio modality.

Gotenberg

Office document-to-PDF conversion API: receives raw uploaded file bytes, returns the converted PDF

self

eu01

upload

Off by default, no workflow consumes it yet. Chainguard image pinned by digest, stateless, in-cluster only.

Postgres

Metadata and content store: search stores, files, folders, documents, pages (full extracted text), chunks, workflow runs

STACKIT

eu01

upload, user-id

"upload" here is the extracted document text itself — pages/chunks store full content — plus workflow run inputs and error messages. Reached through the platform pgdog pooler. Open: managed-Postgres legal entity per environment.

Qdrant

Vector database: chunk embeddings with a payload of user_id, searchstore_id and filterable document metadata; one collection per search store

self (sheet’s first table: "self + Qdrant")

eu01

user-id, upload

Vectors computed from chunk text; payload can carry document metadata fields. In-cluster via the Qdrant operator. Qdrant Cloud / API-key mode supported by config but not deployed.

Object storage (S3/MinIO)

Binary file storage for uploaded files and export bundles

STACKIT

eu01

upload, user-id

Uploaded files carry the user_id in their object key (users/{user_id}/search-stores/…), so key listings alone are personal data. Tenant export zips are the exception: they sit outside that tree, at tenant-exports/{run_id}/{store_id}.zip. MinIO in dev, S3-compatible storage in production. Open: operating legal entity per environment.

Temporal

Workflow orchestration for ingestion runs: task queues plus persisted workflow histories (run inputs, source URIs, failure payloads)

self

eu01

user-id, trace (sheet’s first table also lists "upload")

Histories persist run inputs and failure payloads, which can embed document text from stack traces. Production self-hosted; staging on Temporal Cloud via api_key auth — if that stays, Temporal Technologies Inc. needs its own operator row. Planned mitigations: self-hosting and encryption.

OpenFGA

Authorization store: owner/viewer/writer relationship tuples between user ids and search stores

self

eu01

user-id

Only deployed where AUTHORIZATION_MODE=fga. Platform-shared instance; tuples contain user and group ids only.

What drives sizing

Exact CPU/memory figures depend heavily on your workload and are best confirmed with your platform team as part of deployment sizing; the drivers to plan around are:

Vector database

Memory scales with the number of indexed chunks, the embedding dimension, and the number of replicas configured for availability. More chunks (from more documents, or a smaller chunk_size producing more chunks per document) and higher-dimensional embeddings both increase memory usage per replica.

Relational database

Storage scales roughly with the volume of extracted text: page content, chunk metadata, and workflow/run history all accumulate here as you ingest more documents.

Object storage

Storage scales directly with the total size of uploaded files. Deduplication by content hash means re-uploading the same file to the same search store doesn’t add to this.

Ingestion throughput

Ingestion worker CPU/memory scales with concurrent workflow runs and document complexity — OCR and vision-based extraction (see Workflows and presets) are more resource-intensive per document than plain text extraction.

Best practices

  1. Start conservative: begin with your platform team’s recommended minimums for a new deployment.

  2. Monitor usage: watch vector database memory and ingestion queue depth once real document volume starts flowing through.

  3. Scale gradually: increase resources based on observed usage rather than projected peak.

  4. Plan for growth: as covered in What is the Data and Search Platform?, the same search store and workflow configuration should carry you from initial experimentation through production — only the infrastructure sizing needs to grow.