Ingesting files
This guide describes how to upload files into a search store and run an ingestion workflow to make them searchable.
Prerequisites
-
The Data and Search Platform is deployed:
https://sherlock.{ingressDomain} -
You can reach it with one of:
curl, the Sherlock Python SDK (sherlock-client), or the platform CLI (pharia). -
You have created a search store — see Creating search stores.
-
You have sample documents in a supported format: PDF, Office files (DOCX, PPTX, XLSX, and the legacy DOC, PPT, XLS), HTML, Markdown, plain text, and images. Each file type is routed to its own reader — see the reader groups in Configuring workflows; any other text file falls back to the plain-text reader.
Client setup
Every example below is shown three ways. They all drive the same REST API at https://sherlock.{ingressDomain}/api/v1 — the SDK and the CLI are wrappers over the requests in the curl tab. Every request needs a bearer token; the deployment’s IAM sidecar validates it and stamps the user identity that Sherlock authorizes against.
-
curl
-
Python SDK
-
Platform CLI
Log in once with the platform CLI, then export the deployment URL and a token:
pharia --env aleph-alpha iam login # one-time OAuth login via your browser
export SHERLOCK_URL="https://sherlock.{ingressDomain}"
export AA_TOKEN="$(pharia --env aleph-alpha iam token)"
Every call then carries -H "Authorization: Bearer $AA_TOKEN". Any other source of a platform token works just as well — the API only sees the header.
Install the client from the Aleph Alpha package index:
pip install sherlock-client --extra-index-url https://alephalpha.jfrog.io/artifactory/api/pypi/holmes/simple
Then point it at the deployment and take a handle on the store you’re ingesting into:
import os
from sherlock import SyncSherlock
client = SyncSherlock(
url="https://sherlock.{ingressDomain}",
token=os.environ["AA_TOKEN"],
)
store = client.v1.search_stores("<store-id>")
url and token default to the SHERLOCK_BASE_URL and SHERLOCK_TOKEN environment variables, so SyncSherlock() with no arguments works once those are set. Sherlock is the async twin with an identical API. Use either as a context manager (with SyncSherlock(...) as client:) to close the connection pool when you’re done.
The platform CLI (pharia) handles authentication for you — log in once, and every command reuses the cached token:
pharia iam login # one-time OAuth login via your browser
Use the global --env flag to select the environment (for example, pharia --env aleph-alpha iam login); tokens are cached per environment and refreshed automatically. To reach a deployment the CLI doesn’t know, pass --base-url https://sherlock.{ingressDomain} or set PHARIA_SHERLOCK_URL. For your own tooling, print the raw token with TOKEN=$(pharia iam token).
1. Upload a file
-
curl
-
Python SDK
-
Platform CLI
curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/files" \
-H "Authorization: Bearer $AA_TOKEN" \
-F "files=@sample.pdf"
The upload is multipart/form-data with one files part per document — repeat -F "files=@..." to send several at once.
files = store.files.upload("sample.pdf")
print(files[0].id)
upload is variadic and also accepts Path objects, open file handles, and glob patterns (store.files.upload("docs/**/*.pdf")). It returns one handle per uploaded file.
pharia --env aleph-alpha sherlock files upload <store-id> sample.pdf
The response includes the file’s id. You can pass several paths at once. Files are deduplicated by content hash within a search store — uploading the same file twice returns the existing file rather than creating a duplicate.
Accepted file types are the ones listed under Prerequisites. The type is sniffed from the file’s bytes, not its name; anything else is rejected with 415.
Attach metadata to uploads
The upload takes one optional metadata part: a JSON object (max 8 KiB) applied to every file in the request.
-
curl
-
Python SDK
-
Platform CLI
curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/files" \
-H "Authorization: Bearer $AA_TOKEN" \
-F "files=@sample.pdf" \
-F 'metadata={"site": "dresden", "vendor": "acme"}'
files = store.files.upload("sample.pdf", metadata={"site": "dresden", "vendor": "acme"})
The CLI has no flag for the metadata part yet — use the curl or SDK form.
Once the file is ingested, the object lands on every one of its chunks under metadata.user.<field> — filterable at search after you declare the fields on the store with the user. prefix, for example {"path": "user.site", "type": "string"}; see declaring filterable metadata.
To change a file’s metadata later, replace the whole object — no re-upload, no re-ingestion; the change cascades to the file’s documents, pages, and chunks:
-
curl
-
Python SDK
-
Platform CLI
curl -X PUT "$SHERLOCK_URL/api/v1/search-stores/<store-id>/files/<file-id>/metadata" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AA_TOKEN" \
-d '{"site": "leipzig", "vendor": "acme"}'
store.files("<file-id>").update_metadata({"site": "leipzig", "vendor": "acme"})
Not in the CLI yet — use the curl or SDK form.
|
Version availability
Upload metadata and the metadata |
2. Find the workflow to run
Every search store has a default ingestion workflow ready to use:
-
curl
-
Python SDK
-
Platform CLI
curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows" \
-H "Authorization: Bearer $AA_TOKEN"
for workflow in store.workflows.list():
print(workflow.id, workflow.name, workflow.workflow_name)
pharia --env aleph-alpha sherlock workflows list <store-id>
Use the id of the workflow you want (the default preset, or one you’ve configured — see Workflows and presets).
3. Run the workflow
-
curl
-
Python SDK
-
Platform CLI
curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AA_TOKEN" \
-d '{"input": {"source": ["files://<file-id>"]}}'
run = store.workflows("<workflow-id>").runs.execute(
input={"source": ["files://<file-id>"]},
)
print(run.id, run.status)
pharia --env aleph-alpha sherlock runs execute <store-id> <workflow-id> \
--input '{"source": ["files://<file-id>"]}'
This extracts text, splits it into chunks, generates embeddings, and stores the results — all in one run. The response includes a run_id.
source accepts multiple file references at once, so you can process a batch of uploads in a single run.
4. Stream the run status
Rather than polling, watch the run’s event stream until a terminal event arrives:
-
curl
-
Python SDK
-
Platform CLI
curl -N "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs/<run-id>/event-stream" \
-H "Accept: text/event-stream" \
-H "Authorization: Bearer $AA_TOKEN"
-N disables curl’s output buffering so events print as they arrive.
for event in run.stream():
print(event.event, event.data)
stream() yields WorkflowEventData and stops on its own at complete or error. To block until the run finishes without handling each event, use data = run.wait(raise_on_failure=True) and read data.status.
pharia --env aleph-alpha sherlock runs watch <store-id> <workflow-id> <run-id>
The stream replays every event recorded so far, then stays live for new ones:
-
status— the run’s current state. -
complete— the run finished; your content is now searchable. -
error— the run failed, with the error detail. -
keepalive— a periodic comment that keeps the connection open through proxies; ignore it.
The stream ends when the run completes (runs watch exits non-zero on failure). The server drops idle connections after 10 minutes — watching again is safe, since the stream always replays the full history first.
For a one-off snapshot instead of a live stream, get the run directly:
-
curl
-
Python SDK
-
Platform CLI
curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs/<run-id>" \
-H "Authorization: Bearer $AA_TOKEN"
data = store.workflows("<workflow-id>").runs.get("<run-id>")
print(data.status, data.error)
pharia --env aleph-alpha sherlock runs get <store-id> <workflow-id> <run-id>
Storage and page limits
A deployment can limit what you upload and what a search store holds. Four limits apply to uploads and runs:
| Limit | Applies to | Default |
|---|---|---|
|
One uploaded file, in plaintext bytes |
100 MiB |
|
One upload request, with all the parts it carries |
100 MiB |
|
The bytes one search store holds |
No limit unless configured |
|
The pages one search store holds |
No limit unless configured |
The two request limits are always in force. A platform admin sets the two store limits, and can change them without a redeploy.
The platform checks a store limit before each operation. The operation that crosses the limit completes, and the platform refuses the next one. An upload past max_storage_bytes and a run submission past max_pages_per_search_store both return 409 Conflict with the type /errors/usage-limit-exceeded. The response gives limit_key, limit, current_usage, scope, and unit. The upload route puts them in details, and the run route puts them in errors[0].value. The platform does not use 429 for usage limits. 429 is reserved for rate limiting.
Read a store’s usage
Read a store’s usage before a large upload or run:
curl -s "$SHERLOCK_URL/api/v1/search-stores/$STORE_ID/usage" \
-H "Authorization: Bearer $AA_TOKEN"
For each per-store limit, the response gives what the store holds, the enforced limit (null when none is set), and an over_limit flag. The flag tells you that the platform refuses the next operation against that limit. You need can_view on the store. The Python SDK and the platform CLI do not cover this endpoint yet.
What counts against a limit
The counters measure what the platform stores physically. This has three consequences:
-
A deleted file keeps its bytes until the retention purge erases it. If a store is full, deleting a file and uploading it again fails.
-
The pages of a failed ingestion count until the purge. Each retry counts again.
-
A second upload of the same file into the same store costs nothing.
Export zips do not count against any store.
Validation
Once a run completes, confirm your content is searchable — see Searching your data.
Troubleshooting
Workflow run stays pending
-
Error: the event stream keeps sending
statusevents and never reachescompleteorerror -
Solution: Large files or high concurrent load can extend processing time; keep the connection open (reconnecting is safe — the stream replays history). If the run doesn’t progress, check the file format is supported.
Input validation error
-
Error:
422 Unprocessable Entitywhen starting a run -
Solution: Your run input doesn’t match the workflow’s expected schema. Read the workflow’s
schemafirst —GET /api/v1/search-stores/<store-id>/workflows/<workflow-id>,store.workflows("<workflow-id>").get(), orpharia sherlock workflows get <store-id> <workflow-id>— before submitting a run.
Search store is full
-
Error:
409 Conflictwith the type/errors/usage-limit-exceeded, on an upload or a run submission -
Solution: The store reached
max_storage_bytesormax_pages_per_search_store. Readlimit_key,limit, andcurrent_usagein the response to see which limit stopped you. Erase content you no longer need, or ask a platform admin to raise the limit — see Storage and page limits.
File or request is too large
-
Error:
413on upload. A file pastmax_file_sizereturns a problem detail with the type/errors/file-too-large. A request pastmax_upload_request_sizeis refused by the HTTP server before it reaches the API, so that413carries no body. -
Solution: Send fewer files per request, or ask a platform admin to raise the limit — see Storage and page limits.
Unsupported file type
-
Error: The run fails immediately after starting
-
Solution: Confirm the file is a supported format (see Prerequisites). An unsupported type is rejected at upload with
415; a run that fails on a specific file names it in the run’sfailedlist.