Configuring workflows
A workflow attached to a search store is a named instance of a workflow class, usually the DocumentAI preset. Two JSON documents configure it:
- Client setup
- Inspecting what exists
- The parameter model
- Recipe: text-first ingestion
- Recipe: vision (OCR) ingestion
- Recipe: figure descriptions (
figure_extractor+image_describer) - Inference providers
- Sampling temperature
- Choosing a chunker
- Embeddings (
vectorize) - Enrichments (
page_enrichments,chunk_enrichments) - Cleaning page headers and footers
- The input schema
- Updating a workflow
- Running and watching
- Related pages
-
parameters— overrides for the workflow’s module configuration. The workflow class has working defaults for every module. Specify only the fields that you want to change. -
schema— the JSON schema of the run input. It is currently required at creation but is planned for deprecation. See The input schema.
This page describes the real parameter surface of DocumentAI and how to manage workflows. We validated all examples on this page against a live deployment.
Client setup
Every example below is shown three ways. They all drive the same REST API at https://sherlock.{ingressDomain}/api/v1 — the SDK and the CLI are wrappers over the requests in the curl tab. Every request needs a bearer token; the deployment’s IAM sidecar validates it and stamps the user identity that Sherlock authorizes against.
-
curl
-
Python SDK
-
Platform CLI
Log in once with the platform CLI, then export the deployment URL and a token:
pharia --env aleph-alpha iam login # one-time OAuth login via your browser
export SHERLOCK_URL="https://sherlock.{ingressDomain}"
export AA_TOKEN="$(pharia --env aleph-alpha iam token)"
Every call then carries -H "Authorization: Bearer $AA_TOKEN". Any other source of a platform token works just as well — the API only sees the header.
Install the client from the Aleph Alpha package index:
pip install sherlock-client --extra-index-url https://alephalpha.jfrog.io/artifactory/api/pypi/holmes/simple
Then point it at the deployment and take a handle on the store whose workflows you’re configuring:
import os
from sherlock import SyncSherlock
client = SyncSherlock(
url="https://sherlock.{ingressDomain}",
token=os.environ["AA_TOKEN"],
)
store = client.v1.search_stores("<store-id>")
url and token default to the SHERLOCK_BASE_URL and SHERLOCK_TOKEN environment variables, so SyncSherlock() with no arguments works once those are set. Sherlock is the async twin with an identical API. Use either as a context manager (with SyncSherlock(...) as client:) to close the connection pool when you’re done.
The pharia CLI has a sherlock subcommand that covers the full workflow lifecycle. Log in one time; each subsequent command uses the cached token:
pharia --env aleph-alpha iam login
JSON flags accept inline JSON, @file, or @- for stdin. Use the global --env flag to select the environment. To reach a different tenant, use --base-url or PHARIA_SHERLOCK_URL.
Inspecting what exists
List your stores, the workflows on one, and a single workflow’s current configuration:
-
curl
-
Python SDK
-
Platform CLI
curl "$SHERLOCK_URL/api/v1/search-stores" \
-H "Authorization: Bearer $AA_TOKEN"
curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows" \
-H "Authorization: Bearer $AA_TOKEN"
curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id-or-name>" \
-H "Authorization: Bearer $AA_TOKEN"
for s in client.v1.search_stores.list():
print(s.id, s.name)
for workflow in store.workflows.list():
print(workflow.id, workflow.name, workflow.workflow_name)
current = store.workflows("<workflow-id-or-name>").get()
print(current.parameters)
pharia --env aleph-alpha sherlock stores list
pharia --env aleph-alpha sherlock workflows list <store-id>
pharia --env aleph-alpha sherlock workflows get <store-id> <workflow-id-or-name>
The workflow read accepts either the workflow’s id or its name.
The parameter model
DocumentAI is a composition of modules. Each top-level key in parameters configures one module:
| Group | Purpose |
|---|---|
|
Reader for PDF files. |
|
Readers for Word ( |
|
One reader per file type; the workflow routes each document to the matching reader. |
|
Uses a vision model to annotate rasterized PDF pages. Required when |
|
Optional modules that extract figures from PDF files and describe them with a vision model. |
|
Splits pages into chunks. One of |
|
Optional lists of enrichment modules that add metadata to pages or chunks, for example keywords or generated questions. |
|
Embeds chunks and writes vectors to the vector store. |
Groups with more than one implementation (pdf_reader, docx_reader, xlsx_reader, chunker) select it with a "kind" field. Every module also accepts the execution fields batch_size, max_in_flight, and timeout together with its own parameters.
Recipe: text-first ingestion
This is the standard configuration for digital documents. It extracts the embedded text, chunks it recursively, and embeds the chunks through the platform’s inference gateway.
{
"pdf_reader": {"kind": "pdf_text"},
"chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600},
"vectorize": {"provider": "pharia"}
}
Create a workflow with those parameters, plus the standard run-input schema (see The input schema):
-
curl
-
Python SDK
-
Platform CLI
curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AA_TOKEN" \
-d '{
"name": "text-ingestion",
"workflow_name": "DocumentAI",
"parameters": {
"pdf_reader": {"kind": "pdf_text"},
"chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600},
"vectorize": {"provider": "pharia"}
},
"schema": {
"type": "object",
"additionalProperties": true,
"properties": {
"source": {
"anyOf": [{"type": "string"}, {"type": "array", "items": {"type": "string"}}],
"description": "Source path, connector URI, or list thereof"
}
}
}
}'
workflow = store.workflows.create(
name="text-ingestion",
workflow_name="DocumentAI",
parameters={
"pdf_reader": {"kind": "pdf_text"},
"chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600},
"vectorize": {"provider": "pharia"},
},
schema={
"type": "object",
"additionalProperties": True,
"properties": {
"source": {
"anyOf": [
{"type": "string"},
{"type": "array", "items": {"type": "string"}},
],
"description": "Source path, connector URI, or list thereof",
}
},
},
)
print(workflow.id)
pharia --env aleph-alpha sherlock workflows create <store-id> \
--name text-ingestion \
--workflow-name DocumentAI \
--parameters @parameters.json \
--schema @input-schema.json
Recipe: vision (OCR) ingestion
For scanned documents or complex layouts, rasterize each PDF page and annotate it with a vision model. This is a two-module pipeline. pdf_rasterizer renders each page into an image. vision_reader turns each image into a structured page — one page in, one page out — in a separate stage that fans out per page. pdf_rasterizer only renders, so a vision_reader with an explicit provider and model is required; there is no default endpoint. A run without one fails immediately and does not retry.
{
"pdf_reader": {"kind": "pdf_rasterizer", "pdf_backend": "pdfboss"},
"vision_reader": {"provider": "pharia_next", "model": "rednote-hilab/dots.mocr", "timeout": 120},
"chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600}
}
Useful pdf_rasterizer fields: pdf_backend (pdf_oxide default, or pdfboss), dpi (default 150), and image_format (default png).
Useful vision_reader fields: max_tokens, temperature (see Sampling temperature; use 0 for reproducible OCR), ignore_headers_footers (default true), prompt, and extra_body. prompt replaces the built-in layout-extraction prompt; dots-OCR models automatically get a dots-format prompt, and other models get a normalized-coordinate prompt. extra_body carries vendor-specific payload fields — for example, {"chat_template_kwargs": {"thinking": false}} switches off a reasoning model’s thinking mode. The annotated pages flow into the chunker exactly like text-extracted pages.
Recipe: figure descriptions (figure_extractor + image_describer)
Charts, diagrams, and photos carry meaning that neither text extraction nor OCR turns into searchable text. The figure pipeline closes that gap in two stages. A figure_extractor produces one image crop per figure. image_describer sends each crop to a vision model and stores the returned description.
|
Not yet active
Current releases accept these parameter groups but do not run the figure-extraction stage yet, so the pipeline produces no descriptions. The parameter surface below is stable and documented before activation. |
{
"pdf_reader": {"kind": "pdf_rasterizer", "dpi": 200},
"vision_reader": {"provider": "pharia_next", "model": "rednote-hilab/dots.mocr"},
"figure_extractor": {"kind": "bbox_figure_extractor", "dpi": 200},
"image_describer": {
"provider": "pharia_next",
"model": "rednote-hilab/dots.mocr",
"max_tokens": 4096,
"temperature": 0,
"extra_body": {"chat_template_kwargs": {"thinking": false}}
}
}
Two extractors fill the figure_extractor slot. Select one with kind:
| Kind | Strategy | Pair with |
|---|---|---|
|
Extracts the images that are embedded in the PDF file. |
|
|
Renders each page again and crops it at the figure bounding boxes that the vision reader detected. Its |
|
image_describer fields: provider (default pharia_next) and model (default moonshotai/kimi-k2.6) name the vision deployment. prompt replaces the built-in description prompt. max_tokens (default 2048), temperature, and extra_body shape the call. max_in_flight (default 4) limits how many describe calls run at the same time.
The platform stores descriptions append-only, keyed by page, image, model, and prompt. A re-run with the same model and prompt is a no-op. A different model or prompt adds new descriptions and does not change existing ones.
Inference providers
Modules that call models (vision_reader, vectorize, image_describer, …) select their client by name. The provider parameter is a lookup key; the endpoint, credentials, and retry policy behind it are server-side deployment configuration. Workflow parameters never carry secrets or URLs.
These providers are available only when the data platform runs on the Aleph Alpha platform. On-premise installations have their own, different providers.
| Provider | Route |
|---|---|
|
The platform inference gateway (default). Routes to the legacy backend. |
|
The same gateway, pinned to the next-generation inference backend. The intended long-term route. |
|
A direct connection to the next-generation backend that bypasses the gateway. A temporary route until service-account authentication is available. |
|
External OpenAI-compatible endpoints, when the deployment configures them. |
You can select a provider only if the deployment configures it. An unconfigured provider fails the run.
Sampling temperature
Every module that calls a chat or vision model accepts a temperature field (0.0--2.0):
-
vision_readerandimage_reader— page annotation and image-file reading -
image_describer— figure descriptions -
pdf_readerwithkind: "parsy_ocr"— OCR extraction -
page_enrichments/chunk_enrichmentsentries that call an LLM, such aspage_questions
Embeddings (vectorize) and non-LLM modules (chunkers, keyword enrichment) have no temperature.
If you do not set temperature (the default), the request omits the field and the model server’s own default applies. 0 is a real value, and the module sends it. Use 0 for reproducible extraction — the usual choice for OCR and structured layout parsing:
{
"vision_reader": {"provider": "pharia_next", "model": "rednote-hilab/dots.mocr", "temperature": 0}
}
|
Version availability
|
Choosing a chunker
| Kind | Strategy | Fields (defaults) |
|---|---|---|
|
Fixed-size word groups with no overlap. Fast, but it can break sentences. |
|
|
Character-based with overlap. It prefers sentence boundaries. |
|
|
Recursive separator hierarchy: paragraphs before line breaks before spaces and punctuation. |
|
|
Splits at top-level markdown blocks. It never splits fenced code, and it prefixes each chunk with its heading breadcrumb. |
|
Smaller chunks are more precise but lose surrounding context. Overlap keeps context at chunk boundaries.
Embeddings (vectorize)
The vectorize module calls an OpenAI-compatible embeddings endpoint through a named provider (default pharia; see Inference providers).
The search store is the source of truth for the vector schema. When you do not set model and dimension, the module adopts them from the store’s models entry, so {"provider": "pharia"} is usually sufficient. If you set them, they must match a slot that the store provisions. If they do not match, the run fails; it does not write vectors that search cannot find.
Enrichments (page_enrichments, chunk_enrichments)
Both groups take a list of enrichment modules, selected by kind. Enrichments add metadata rows next to pages or chunks. The page and chunk content does not change.
| Kind | What it does | Fields (defaults) |
|---|---|---|
|
Extracts the top keyword lemmas per item with spaCy. No model call. |
|
|
Asks an LLM for questions that each page can answer. Useful for HyDE-style retrieval and Q&A dataset generation. Page-level only. |
|
{
"page_enrichments": [
{"kind": "keyword", "language": "de"},
{"kind": "page_questions", "num_questions": 3, "temperature": 0.7}
]
}
Stripping repeated page headers and footers is not an enrichment — it happens at read time; see Cleaning page headers and footers.
Cleaning page headers and footers
Repeated headers and footers — running titles, page numbers — pollute chunks and search hits. The readers remove them:
-
vision_reader—ignore_headers_footers(defaulttrue) skips the header and footer regions the layout model classifies; see Recipe: vision (OCR) ingestion. -
pdf_readerwithkind: "docling"— the same switch, off by default:
{
"pdf_reader": {"kind": "docling", "ignore_headers_footers": true}
}
Affected pages record the number of dropped lines as metadata.headers_footers_removed.
|
Version availability
|
The input schema
To create a workflow, you currently must send schema: the JSON schema of the input that each run accepts.
|
Deprecation ahead
The |
For DocumentAI the standard schema declares the source property:
{
"type": "object",
"additionalProperties": true,
"properties": {
"source": {
"anyOf": [{"type": "string"}, {"type": "array", "items": {"type": "string"}}],
"description": "Source path, connector URI, or list thereof"
}
}
}
Updating a workflow
-
curl
-
Python SDK
-
Platform CLI
curl -X PUT "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AA_TOKEN" \
-d '{
"parameters": {
"pdf_reader": {"kind": "pdf_rasterizer", "pdf_backend": "pdfboss"},
"vision_reader": {"provider": "pharia_next", "model": "rednote-hilab/dots.mocr", "timeout": 120},
"chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600}
}
}'
This body moves the workflow created above onto the vision path: the whole parameters document, with pdf_reader switched and a vision_reader group added.
store.workflows("<workflow-id>").update(
parameters={
"pdf_reader": {"kind": "pdf_rasterizer", "pdf_backend": "pdfboss"},
"vision_reader": {
"provider": "pharia_next",
"model": "rednote-hilab/dots.mocr",
"timeout": 120,
},
"chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600},
},
)
pharia --env aleph-alpha sherlock workflows update <store-id> <workflow-id> \
--parameters @parameters.json
|
Updates replace, not merge
|
Only future runs use the change. The platform does not automatically reprocess content that is already indexed.
Running and watching
To start a run, name a source (file IDs from Ingesting files). Then stream its progress:
-
curl
-
Python SDK
-
Platform CLI
curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AA_TOKEN" \
-d '{"input": {"source": ["files://<file-id>"]}}'
curl -N "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs/<run-id>/event-stream" \
-H "Accept: text/event-stream" \
-H "Authorization: Bearer $AA_TOKEN"
curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs" \
-H "Authorization: Bearer $AA_TOKEN"
runs = store.workflows("<workflow-id>").runs
run = runs.execute(input={"source": ["files://<file-id>"]})
for event in run.stream():
print(event.event, event.data)
for past in runs.list():
print(past.id, past.status)
pharia --env aleph-alpha sherlock runs execute <store-id> <workflow-id> \
--input '{"source": ["files://<file-id>"]}'
pharia --env aleph-alpha sherlock runs watch <store-id> <workflow-id> <run-id>
pharia --env aleph-alpha sherlock runs list <store-id> <workflow-id>
The event stream carries server-sent events until the run completes or fails. Deleting a workflow deletes all its runs and their derived data with it.
Related pages
-
Workflows and presets — the built-in presets and how to select one.
-
For workflow authors — how to write new workflow classes and modules in code.