Ingesting files

This guide describes how to upload files into a search store and run an ingestion workflow to make them searchable.


Prerequisites

  • The Data and Search Platform is deployed: https://sherlock.{ingressDomain}

  • You can reach it with one of: curl, the Sherlock Python SDK (sherlock-client), or the platform CLI (pharia).

  • You have created a search store — see Creating search stores.

  • You have sample documents in a supported format: PDF, Office files (DOCX, PPTX, XLSX, and the legacy DOC, PPT, XLS), HTML, Markdown, plain text, and images. Each file type is routed to its own reader — see the reader groups in Configuring workflows; any other text file falls back to the plain-text reader.

Client setup

Every example below is shown three ways. They all drive the same REST API at https://sherlock.{ingressDomain}/api/v1 — the SDK and the CLI are wrappers over the requests in the curl tab. Every request needs a bearer token; the deployment’s IAM sidecar validates it and stamps the user identity that Sherlock authorizes against.

  • curl

  • Python SDK

  • Platform CLI

Log in once with the platform CLI, then export the deployment URL and a token:

pharia --env aleph-alpha iam login    # one-time OAuth login via your browser

export SHERLOCK_URL="https://sherlock.{ingressDomain}"
export AA_TOKEN="$(pharia --env aleph-alpha iam token)"

Every call then carries -H "Authorization: Bearer $AA_TOKEN". Any other source of a platform token works just as well — the API only sees the header.

Install the client from the Aleph Alpha package index:

pip install sherlock-client --extra-index-url https://alephalpha.jfrog.io/artifactory/api/pypi/holmes/simple

Then point it at the deployment and take a handle on the store you’re ingesting into:

import os

from sherlock import SyncSherlock

client = SyncSherlock(
    url="https://sherlock.{ingressDomain}",
    token=os.environ["AA_TOKEN"],
)

store = client.v1.search_stores("<store-id>")

url and token default to the SHERLOCK_BASE_URL and SHERLOCK_TOKEN environment variables, so SyncSherlock() with no arguments works once those are set. Sherlock is the async twin with an identical API. Use either as a context manager (with SyncSherlock(...) as client:) to close the connection pool when you’re done.

The platform CLI (pharia) handles authentication for you — log in once, and every command reuses the cached token:

pharia iam login    # one-time OAuth login via your browser

Use the global --env flag to select the environment (for example, pharia --env aleph-alpha iam login); tokens are cached per environment and refreshed automatically. To reach a deployment the CLI doesn’t know, pass --base-url https://sherlock.{ingressDomain} or set PHARIA_SHERLOCK_URL. For your own tooling, print the raw token with TOKEN=$(pharia iam token).

Overview of the pharia iam commands

1. Upload a file

  • curl

  • Python SDK

  • Platform CLI

curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/files" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -F "files=@sample.pdf"

The upload is multipart/form-data with one files part per document — repeat -F "files=@..." to send several at once.

files = store.files.upload("sample.pdf")

print(files[0].id)

upload is variadic and also accepts Path objects, open file handles, and glob patterns (store.files.upload("docs/**/*.pdf")). It returns one handle per uploaded file.

pharia --env aleph-alpha sherlock files upload <store-id> sample.pdf

The response includes the file’s id. You can pass several paths at once. Files are deduplicated by content hash within a search store — uploading the same file twice returns the existing file rather than creating a duplicate.

Accepted file types are the ones listed under Prerequisites. The type is sniffed from the file’s bytes, not its name; anything else is rejected with 415.

Attach metadata to uploads

The upload takes one optional metadata part: a JSON object (max 8 KiB) applied to every file in the request.

  • curl

  • Python SDK

  • Platform CLI

curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/files" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -F "files=@sample.pdf" \
  -F 'metadata={"site": "dresden", "vendor": "acme"}'
files = store.files.upload("sample.pdf", metadata={"site": "dresden", "vendor": "acme"})

The CLI has no flag for the metadata part yet — use the curl or SDK form.

Once the file is ingested, the object lands on every one of its chunks under metadata.user.<field> — filterable at search after you declare the fields on the store with the user. prefix, for example {"path": "user.site", "type": "string"}; see declaring filterable metadata.

To change a file’s metadata later, replace the whole object — no re-upload, no re-ingestion; the change cascades to the file’s documents, pages, and chunks:

  • curl

  • Python SDK

  • Platform CLI

curl -X PUT "$SHERLOCK_URL/api/v1/search-stores/<store-id>/files/<file-id>/metadata" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{"site": "leipzig", "vendor": "acme"}'
store.files("<file-id>").update_metadata({"site": "leipzig", "vendor": "acme"})

Not in the CLI yet — use the curl or SDK form.

Version availability

Upload metadata and the metadata PUT are available from release 0.9.0.

2. Find the workflow to run

Every search store has a default ingestion workflow ready to use:

  • curl

  • Python SDK

  • Platform CLI

curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows" \
  -H "Authorization: Bearer $AA_TOKEN"
for workflow in store.workflows.list():
    print(workflow.id, workflow.name, workflow.workflow_name)
pharia --env aleph-alpha sherlock workflows list <store-id>

Use the id of the workflow you want (the default preset, or one you’ve configured — see Workflows and presets).

3. Run the workflow

  • curl

  • Python SDK

  • Platform CLI

curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{"input": {"source": ["files://<file-id>"]}}'
run = store.workflows("<workflow-id>").runs.execute(
    input={"source": ["files://<file-id>"]},
)

print(run.id, run.status)
pharia --env aleph-alpha sherlock runs execute <store-id> <workflow-id> \
  --input '{"source": ["files://<file-id>"]}'

This extracts text, splits it into chunks, generates embeddings, and stores the results — all in one run. The response includes a run_id.

source accepts multiple file references at once, so you can process a batch of uploads in a single run.

4. Stream the run status

Rather than polling, watch the run’s event stream until a terminal event arrives:

  • curl

  • Python SDK

  • Platform CLI

curl -N "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs/<run-id>/event-stream" \
  -H "Accept: text/event-stream" \
  -H "Authorization: Bearer $AA_TOKEN"

-N disables curl’s output buffering so events print as they arrive.

for event in run.stream():
    print(event.event, event.data)

stream() yields WorkflowEventData and stops on its own at complete or error. To block until the run finishes without handling each event, use data = run.wait(raise_on_failure=True) and read data.status.

pharia --env aleph-alpha sherlock runs watch <store-id> <workflow-id> <run-id>

The stream replays every event recorded so far, then stays live for new ones:

  • status — the run’s current state.

  • complete — the run finished; your content is now searchable.

  • error — the run failed, with the error detail.

  • keepalive — a periodic comment that keeps the connection open through proxies; ignore it.

The stream ends when the run completes (runs watch exits non-zero on failure). The server drops idle connections after 10 minutes — watching again is safe, since the stream always replays the full history first.

For a one-off snapshot instead of a live stream, get the run directly:

  • curl

  • Python SDK

  • Platform CLI

curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs/<run-id>" \
  -H "Authorization: Bearer $AA_TOKEN"
data = store.workflows("<workflow-id>").runs.get("<run-id>")

print(data.status, data.error)
pharia --env aleph-alpha sherlock runs get <store-id> <workflow-id> <run-id>

Storage and page limits

A deployment can limit what you upload and what a search store holds. Four limits apply to uploads and runs:

Limit Applies to Default

max_file_size

One uploaded file, in plaintext bytes

100 MiB

max_upload_request_size

One upload request, with all the parts it carries

100 MiB

max_storage_bytes

The bytes one search store holds

No limit unless configured

max_pages_per_search_store

The pages one search store holds

No limit unless configured

The two request limits are always in force. A platform admin sets the two store limits, and can change them without a redeploy.

The platform checks a store limit before each operation. The operation that crosses the limit completes, and the platform refuses the next one. An upload past max_storage_bytes and a run submission past max_pages_per_search_store both return 409 Conflict with the type /errors/usage-limit-exceeded. The response gives limit_key, limit, current_usage, scope, and unit. The upload route puts them in details, and the run route puts them in errors[0].value. The platform does not use 429 for usage limits. 429 is reserved for rate limiting.

Read a store’s usage

Read a store’s usage before a large upload or run:

curl -s "$SHERLOCK_URL/api/v1/search-stores/$STORE_ID/usage" \
  -H "Authorization: Bearer $AA_TOKEN"

For each per-store limit, the response gives what the store holds, the enforced limit (null when none is set), and an over_limit flag. The flag tells you that the platform refuses the next operation against that limit. You need can_view on the store. The Python SDK and the platform CLI do not cover this endpoint yet.

What counts against a limit

The counters measure what the platform stores physically. This has three consequences:

  • A deleted file keeps its bytes until the retention purge erases it. If a store is full, deleting a file and uploading it again fails.

  • The pages of a failed ingestion count until the purge. Each retry counts again.

  • A second upload of the same file into the same store costs nothing.

Export zips do not count against any store.

Validation

Once a run completes, confirm your content is searchable — see Searching your data.

Troubleshooting

Workflow run stays pending

  • Error: the event stream keeps sending status events and never reaches complete or error

  • Solution: Large files or high concurrent load can extend processing time; keep the connection open (reconnecting is safe — the stream replays history). If the run doesn’t progress, check the file format is supported.

Input validation error

  • Error: 422 Unprocessable Entity when starting a run

  • Solution: Your run input doesn’t match the workflow’s expected schema. Read the workflow’s schema first — GET /api/v1/search-stores/<store-id>/workflows/<workflow-id>, store.workflows("<workflow-id>").get(), or pharia sherlock workflows get <store-id> <workflow-id> — before submitting a run.

Search store is full

  • Error: 409 Conflict with the type /errors/usage-limit-exceeded, on an upload or a run submission

  • Solution: The store reached max_storage_bytes or max_pages_per_search_store. Read limit_key, limit, and current_usage in the response to see which limit stopped you. Erase content you no longer need, or ask a platform admin to raise the limit — see Storage and page limits.

File or request is too large

  • Error: 413 on upload. A file past max_file_size returns a problem detail with the type /errors/file-too-large. A request past max_upload_request_size is refused by the HTTP server before it reaches the API, so that 413 carries no body.

  • Solution: Send fewer files per request, or ask a platform admin to raise the limit — see Storage and page limits.

Unsupported file type

  • Error: The run fails immediately after starting

  • Solution: Confirm the file is a supported format (see Prerequisites). An unsupported type is rejected at upload with 415; a run that fails on a specific file names it in the run’s failed list.