Configuring workflows

A workflow attached to a search store is a named instance of a workflow class, usually the DocumentAI preset. Two JSON documents configure it:


  • parameters — overrides for the workflow’s module configuration. The workflow class has working defaults for every module. Specify only the fields that you want to change.

  • schema — the JSON schema of the run input. It is currently required at creation but is planned for deprecation. See The input schema.

This page describes the real parameter surface of DocumentAI and how to manage workflows. We validated all examples on this page against a live deployment.

Client setup

Every example below is shown three ways. They all drive the same REST API at https://sherlock.{ingressDomain}/api/v1 — the SDK and the CLI are wrappers over the requests in the curl tab. Every request needs a bearer token; the deployment’s IAM sidecar validates it and stamps the user identity that Sherlock authorizes against.

  • curl

  • Python SDK

  • Platform CLI

Log in once with the platform CLI, then export the deployment URL and a token:

pharia --env aleph-alpha iam login    # one-time OAuth login via your browser

export SHERLOCK_URL="https://sherlock.{ingressDomain}"
export AA_TOKEN="$(pharia --env aleph-alpha iam token)"

Every call then carries -H "Authorization: Bearer $AA_TOKEN". Any other source of a platform token works just as well — the API only sees the header.

Install the client from the Aleph Alpha package index:

pip install sherlock-client --extra-index-url https://alephalpha.jfrog.io/artifactory/api/pypi/holmes/simple

Then point it at the deployment and take a handle on the store whose workflows you’re configuring:

import os

from sherlock import SyncSherlock

client = SyncSherlock(
    url="https://sherlock.{ingressDomain}",
    token=os.environ["AA_TOKEN"],
)

store = client.v1.search_stores("<store-id>")

url and token default to the SHERLOCK_BASE_URL and SHERLOCK_TOKEN environment variables, so SyncSherlock() with no arguments works once those are set. Sherlock is the async twin with an identical API. Use either as a context manager (with SyncSherlock(...) as client:) to close the connection pool when you’re done.

The pharia CLI has a sherlock subcommand that covers the full workflow lifecycle. Log in one time; each subsequent command uses the cached token:

pharia --env aleph-alpha iam login

JSON flags accept inline JSON, @file, or @- for stdin. Use the global --env flag to select the environment. To reach a different tenant, use --base-url or PHARIA_SHERLOCK_URL.

Inspecting what exists

List your stores, the workflows on one, and a single workflow’s current configuration:

  • curl

  • Python SDK

  • Platform CLI

curl "$SHERLOCK_URL/api/v1/search-stores" \
  -H "Authorization: Bearer $AA_TOKEN"

curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows" \
  -H "Authorization: Bearer $AA_TOKEN"

curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id-or-name>" \
  -H "Authorization: Bearer $AA_TOKEN"
for s in client.v1.search_stores.list():
    print(s.id, s.name)

for workflow in store.workflows.list():
    print(workflow.id, workflow.name, workflow.workflow_name)

current = store.workflows("<workflow-id-or-name>").get()
print(current.parameters)
pharia --env aleph-alpha sherlock stores list
pharia --env aleph-alpha sherlock workflows list <store-id>
pharia --env aleph-alpha sherlock workflows get <store-id> <workflow-id-or-name>

The workflow read accepts either the workflow’s id or its name.

The parameter model

DocumentAI is a composition of modules. Each top-level key in parameters configures one module:

Group Purpose

pdf_reader

Reader for PDF files. pdf_text (the default) extracts the embedded text. docling runs local layout analysis and can drop classified headers and footers. pdf_rasterizer renders pages to images for vision models; pair it with vision_reader.

docx_reader, xlsx_reader

Readers for Word (.docx) and Excel (.xlsx/.xlsm) files. Each has two implementations. The default doc and excel kinds use anydoc to parse the file into markdown content blocks. The previous docx (python-docx) and xlsx (openpyxl) kinds stay available.

doc_reader, excel_reader, pptx_reader, markdown_reader, html_reader, image_reader, legacy_office_reader, plaintext_reader

One reader per file type; the workflow routes each document to the matching reader. doc_reader reads legacy .doc files, and excel_reader reads legacy .xls/.xlsb workbooks. plaintext_reader is the fallback for all other text files.

vision_reader

Uses a vision model to annotate rasterized PDF pages. Required when pdf_reader is pdf_rasterizer; not used otherwise.

figure_extractor, image_describer

Optional modules that extract figures from PDF files and describe them with a vision model.

chunker

Splits pages into chunks. One of naive, fixed_size, recursive, or markdown.

page_enrichments, chunk_enrichments

Optional lists of enrichment modules that add metadata to pages or chunks, for example keywords or generated questions.

vectorize

Embeds chunks and writes vectors to the vector store.

Groups with more than one implementation (pdf_reader, docx_reader, xlsx_reader, chunker) select it with a "kind" field. Every module also accepts the execution fields batch_size, max_in_flight, and timeout together with its own parameters.

Recipe: text-first ingestion

This is the standard configuration for digital documents. It extracts the embedded text, chunks it recursively, and embeds the chunks through the platform’s inference gateway.

{
  "pdf_reader": {"kind": "pdf_text"},
  "chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600},
  "vectorize": {"provider": "pharia"}
}

Create a workflow with those parameters, plus the standard run-input schema (see The input schema):

  • curl

  • Python SDK

  • Platform CLI

curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{
    "name": "text-ingestion",
    "workflow_name": "DocumentAI",
    "parameters": {
      "pdf_reader": {"kind": "pdf_text"},
      "chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600},
      "vectorize": {"provider": "pharia"}
    },
    "schema": {
      "type": "object",
      "additionalProperties": true,
      "properties": {
        "source": {
          "anyOf": [{"type": "string"}, {"type": "array", "items": {"type": "string"}}],
          "description": "Source path, connector URI, or list thereof"
        }
      }
    }
  }'
workflow = store.workflows.create(
    name="text-ingestion",
    workflow_name="DocumentAI",
    parameters={
        "pdf_reader": {"kind": "pdf_text"},
        "chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600},
        "vectorize": {"provider": "pharia"},
    },
    schema={
        "type": "object",
        "additionalProperties": True,
        "properties": {
            "source": {
                "anyOf": [
                    {"type": "string"},
                    {"type": "array", "items": {"type": "string"}},
                ],
                "description": "Source path, connector URI, or list thereof",
            }
        },
    },
)

print(workflow.id)
pharia --env aleph-alpha sherlock workflows create <store-id> \
  --name text-ingestion \
  --workflow-name DocumentAI \
  --parameters @parameters.json \
  --schema @input-schema.json

Recipe: vision (OCR) ingestion

For scanned documents or complex layouts, rasterize each PDF page and annotate it with a vision model. This is a two-module pipeline. pdf_rasterizer renders each page into an image. vision_reader turns each image into a structured page — one page in, one page out — in a separate stage that fans out per page. pdf_rasterizer only renders, so a vision_reader with an explicit provider and model is required; there is no default endpoint. A run without one fails immediately and does not retry.

{
  "pdf_reader": {"kind": "pdf_rasterizer", "pdf_backend": "pdfboss"},
  "vision_reader": {"provider": "pharia_next", "model": "rednote-hilab/dots.mocr", "timeout": 120},
  "chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600}
}

Useful pdf_rasterizer fields: pdf_backend (pdf_oxide default, or pdfboss), dpi (default 150), and image_format (default png).

Useful vision_reader fields: max_tokens, temperature (see Sampling temperature; use 0 for reproducible OCR), ignore_headers_footers (default true), prompt, and extra_body. prompt replaces the built-in layout-extraction prompt; dots-OCR models automatically get a dots-format prompt, and other models get a normalized-coordinate prompt. extra_body carries vendor-specific payload fields — for example, {"chat_template_kwargs": {"thinking": false}} switches off a reasoning model’s thinking mode. The annotated pages flow into the chunker exactly like text-extracted pages.

Recipe: figure descriptions (figure_extractor + image_describer)

Charts, diagrams, and photos carry meaning that neither text extraction nor OCR turns into searchable text. The figure pipeline closes that gap in two stages. A figure_extractor produces one image crop per figure. image_describer sends each crop to a vision model and stores the returned description.

Not yet active

Current releases accept these parameter groups but do not run the figure-extraction stage yet, so the pipeline produces no descriptions. The parameter surface below is stable and documented before activation.

{
  "pdf_reader": {"kind": "pdf_rasterizer", "dpi": 200},
  "vision_reader": {"provider": "pharia_next", "model": "rednote-hilab/dots.mocr"},
  "figure_extractor": {"kind": "bbox_figure_extractor", "dpi": 200},
  "image_describer": {
    "provider": "pharia_next",
    "model": "rednote-hilab/dots.mocr",
    "max_tokens": 4096,
    "temperature": 0,
    "extra_body": {"chat_template_kwargs": {"thinking": false}}
  }
}

Two extractors fill the figure_extractor slot. Select one with kind:

Kind Strategy Pair with

figure_extractor

Extracts the images that are embedded in the PDF file.

pdf_text

bbox_figure_extractor

Renders each page again and crops it at the figure bounding boxes that the vision reader detected. Its dpi (default 200) must match the rasterizer’s dpi, or the coordinates do not align. The rasterizer default is 150, so set both values explicitly, as in the example above.

pdf_rasterizer + vision_reader

image_describer fields: provider (default pharia_next) and model (default moonshotai/kimi-k2.6) name the vision deployment. prompt replaces the built-in description prompt. max_tokens (default 2048), temperature, and extra_body shape the call. max_in_flight (default 4) limits how many describe calls run at the same time.

The platform stores descriptions append-only, keyed by page, image, model, and prompt. A re-run with the same model and prompt is a no-op. A different model or prompt adds new descriptions and does not change existing ones.

Inference providers

Modules that call models (vision_reader, vectorize, image_describer, …​) select their client by name. The provider parameter is a lookup key; the endpoint, credentials, and retry policy behind it are server-side deployment configuration. Workflow parameters never carry secrets or URLs.

These providers are available only when the data platform runs on the Aleph Alpha platform. On-premise installations have their own, different providers.

Provider Route

pharia

The platform inference gateway (default). Routes to the legacy backend.

pharia_next

The same gateway, pinned to the next-generation inference backend. The intended long-term route.

d_inference

A direct connection to the next-generation backend that bypasses the gateway. A temporary route until service-account authentication is available.

openrouter, openai

External OpenAI-compatible endpoints, when the deployment configures them.

You can select a provider only if the deployment configures it. An unconfigured provider fails the run.

Sampling temperature

Every module that calls a chat or vision model accepts a temperature field (0.0--2.0):

  • vision_reader and image_reader — page annotation and image-file reading

  • image_describer — figure descriptions

  • pdf_reader with kind: "parsy_ocr" — OCR extraction

  • page_enrichments / chunk_enrichments entries that call an LLM, such as page_questions

Embeddings (vectorize) and non-LLM modules (chunkers, keyword enrichment) have no temperature.

If you do not set temperature (the default), the request omits the field and the model server’s own default applies. 0 is a real value, and the module sends it. Use 0 for reproducible extraction — the usual choice for OCR and structured layout parsing:

{
  "vision_reader": {"provider": "pharia_next", "model": "rednote-hilab/dots.mocr", "temperature": 0}
}
Version availability

temperature is available from release 0.8.7. Older deployments accept the field but ignore it; the call uses the server default and does not fail.

Choosing a chunker

Kind Strategy Fields (defaults)

naive

Fixed-size word groups with no overlap. Fast, but it can break sentences.

chunk_size (1024 words)

fixed_size

Character-based with overlap. It prefers sentence boundaries.

chunk_size (1000 chars), chunk_overlap (200), respect_sentences (true)

recursive

Recursive separator hierarchy: paragraphs before line breaks before spaces and punctuation.

chunk_size (1024 chars), chunk_overlap (200)

markdown

Splits at top-level markdown blocks. It never splits fenced code, and it prefixes each chunk with its heading breadcrumb.

chunk_size (1024 chars), chunk_overlap (200)

Smaller chunks are more precise but lose surrounding context. Overlap keeps context at chunk boundaries.

Embeddings (vectorize)

The vectorize module calls an OpenAI-compatible embeddings endpoint through a named provider (default pharia; see Inference providers).

The search store is the source of truth for the vector schema. When you do not set model and dimension, the module adopts them from the store’s models entry, so {"provider": "pharia"} is usually sufficient. If you set them, they must match a slot that the store provisions. If they do not match, the run fails; it does not write vectors that search cannot find.

Enrichments (page_enrichments, chunk_enrichments)

Both groups take a list of enrichment modules, selected by kind. Enrichments add metadata rows next to pages or chunks. The page and chunk content does not change.

Kind What it does Fields (defaults)

keyword

Extracts the top keyword lemmas per item with spaCy. No model call.

language (en or de), max_keywords (10), min_word_length (3)

page_questions

Asks an LLM for questions that each page can answer. Useful for HyDE-style retrieval and Q&A dataset generation. Page-level only.

provider (pharia_next), model (moonshotai/kimi-k2.6), num_questions (5), max_chars (4000), max_tokens (500), prompt, temperature

{
  "page_enrichments": [
    {"kind": "keyword", "language": "de"},
    {"kind": "page_questions", "num_questions": 3, "temperature": 0.7}
  ]
}

Stripping repeated page headers and footers is not an enrichment — it happens at read time; see Cleaning page headers and footers.

Cleaning page headers and footers

Repeated headers and footers — running titles, page numbers — pollute chunks and search hits. The readers remove them:

  • vision_reader — ignore_headers_footers (default true) skips the header and footer regions the layout model classifies; see Recipe: vision (OCR) ingestion.

  • pdf_reader with kind: "docling" — the same switch, off by default:

{
  "pdf_reader": {"kind": "docling", "ignore_headers_footers": true}
}

Affected pages record the number of dropped lines as metadata.headers_footers_removed.

Version availability

ignore_headers_footers on vision_reader is already available. On the docling reader it follows in the next release.

The input schema

To create a workflow, you currently must send schema: the JSON schema of the input that each run accepts.

Deprecation ahead

The schema property is planned for deprecation. Workflow creation will then not require it. Until then, pass the standard schema below.

For DocumentAI the standard schema declares the source property:

{
  "type": "object",
  "additionalProperties": true,
  "properties": {
    "source": {
      "anyOf": [{"type": "string"}, {"type": "array", "items": {"type": "string"}}],
      "description": "Source path, connector URI, or list thereof"
    }
  }
}

Updating a workflow

  • curl

  • Python SDK

  • Platform CLI

curl -X PUT "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{
    "parameters": {
      "pdf_reader": {"kind": "pdf_rasterizer", "pdf_backend": "pdfboss"},
      "vision_reader": {"provider": "pharia_next", "model": "rednote-hilab/dots.mocr", "timeout": 120},
      "chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600}
    }
  }'

This body moves the workflow created above onto the vision path: the whole parameters document, with pdf_reader switched and a vision_reader group added.

store.workflows("<workflow-id>").update(
    parameters={
        "pdf_reader": {"kind": "pdf_rasterizer", "pdf_backend": "pdfboss"},
        "vision_reader": {
            "provider": "pharia_next",
            "model": "rednote-hilab/dots.mocr",
            "timeout": 120,
        },
        "chunker": {"kind": "recursive", "chunk_size": 4096, "chunk_overlap": 600},
    },
)
pharia --env aleph-alpha sherlock workflows update <store-id> <workflow-id> \
  --parameters @parameters.json
Updates replace, not merge

update replaces the entire parameters object. First read the current configuration back (see Inspecting what exists). Then edit it and send the full document back. If you send only the changed group, you lose every other override.

Only future runs use the change. The platform does not automatically reprocess content that is already indexed.

Running and watching

To start a run, name a source (file IDs from Ingesting files). Then stream its progress:

  • curl

  • Python SDK

  • Platform CLI

curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{"input": {"source": ["files://<file-id>"]}}'

curl -N "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs/<run-id>/event-stream" \
  -H "Accept: text/event-stream" \
  -H "Authorization: Bearer $AA_TOKEN"

curl "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>/runs" \
  -H "Authorization: Bearer $AA_TOKEN"
runs = store.workflows("<workflow-id>").runs

run = runs.execute(input={"source": ["files://<file-id>"]})

for event in run.stream():
    print(event.event, event.data)

for past in runs.list():
    print(past.id, past.status)
pharia --env aleph-alpha sherlock runs execute <store-id> <workflow-id> \
  --input '{"source": ["files://<file-id>"]}'

pharia --env aleph-alpha sherlock runs watch <store-id> <workflow-id> <run-id>
pharia --env aleph-alpha sherlock runs list <store-id> <workflow-id>

The event stream carries server-sent events until the run completes or fails. Deleting a workflow deletes all its runs and their derived data with it.