Workflows and presets

This guide describes the built-in ingestion presets of the Data and Search Platform and how to configure them. The platform calls this "bring your own workflow" (BYOW).


Prerequisites

  • The Data and Search Platform is deployed: https://sherlock.{ingressDomain}

  • You can reach it with one of: curl, the Sherlock Python SDK (sherlock-client), or the platform CLI (pharia).

  • You have created a search store — see Creating search stores.

Client setup

Every example below is shown three ways. They all drive the same REST API at https://sherlock.{ingressDomain}/api/v1 — the SDK and the CLI are wrappers over the requests in the curl tab. Every request needs a bearer token; the deployment’s IAM sidecar validates it and stamps the user identity that Sherlock authorizes against.

  • curl

  • Python SDK

  • Platform CLI

Log in once with the platform CLI, then export the deployment URL and a token:

pharia --env aleph-alpha iam login    # one-time OAuth login via your browser

export SHERLOCK_URL="https://sherlock.{ingressDomain}"
export AA_TOKEN="$(pharia --env aleph-alpha iam token)"

Every call then carries -H "Authorization: Bearer $AA_TOKEN". Any other source of a platform token works just as well — the API only sees the header.

Install the client from the Aleph Alpha package index:

pip install sherlock-client --extra-index-url https://alephalpha.jfrog.io/artifactory/api/pypi/holmes/simple

Then point it at the deployment and take a handle on the store whose workflows you’re configuring:

import os

from sherlock import SyncSherlock

client = SyncSherlock(
    url="https://sherlock.{ingressDomain}",
    token=os.environ["AA_TOKEN"],
)

store = client.v1.search_stores("<store-id>")

url and token default to the SHERLOCK_BASE_URL and SHERLOCK_TOKEN environment variables, so SyncSherlock() with no arguments works once those are set. Sherlock is the async twin with an identical API. Use either as a context manager (with SyncSherlock(...) as client:) to close the connection pool when you’re done.

The platform CLI (pharia) handles authentication for you. Log in one time; each subsequent command uses the cached token:

pharia iam login    # one-time OAuth login via your browser

Use the global --env flag to select the environment, for example pharia --env aleph-alpha iam login. The CLI caches tokens per environment and refreshes them automatically. To reach a deployment that the CLI does not know, pass --base-url https://sherlock.{ingressDomain} or set PHARIA_SHERLOCK_URL. For your own tooling, print the raw token with TOKEN=$(pharia iam token).

Overview of the pharia iam commands

Built-in presets

Every ingestion workflow is a preset with sensible defaults for extraction, chunking, and embedding. Select one when you create a workflow:

Preset Description

DocumentAI

Full pipeline: extraction, chunking, and embeddings. The default for most use cases.

DocumentAIFull

DocumentAI, plus page-level enrichment such as keyword extraction.

Configuring a workflow

Create a workflow against your search store. Name a preset and the parameters to use. Parameters are grouped per pipeline module: one group per reader, plus chunker and vectorize. The create call currently also requires the run-input schema:

  • curl

  • Python SDK

  • Platform CLI

parameters and schema are plain objects inside the request body. This one is the text-extraction document from Text extraction below:

curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{
    "name": "text-ingestion",
    "workflow_name": "DocumentAI",
    "parameters": {
      "chunker": {
        "batch_size": 1000,
        "chunk_overlap": 600,
        "chunk_size": 4096,
        "include_image_descriptions": false,
        "kind": "recursive"
      },
      "docx_reader": {
        "batch_size": 100,
        "kind": "doc",
        "split_level": 1
      },
      "markdown_reader": {
        "batch_size": 100,
        "kind": "markdown"
      },
      "pdf_reader": {
        "batch_size": 10,
        "kind": "pdf_text",
        "pdf_backend": "pdfboss"
      },
      "plaintext_reader": {
        "batch_size": 100,
        "kind": "plaintext",
        "max_content_size": 500000,
        "min_space_ratio": 0.02
      },
      "pptx_reader": {
        "batch_size": 100,
        "include_notes": true,
        "kind": "pptx"
      },
      "vectorize": {
        "batch_size": 10,
        "distance": "cosine",
        "kind": "openai_vectorize",
        "modality": "text",
        "provider": "pharia",
        "vector_type": "dense"
      },
      "xlsx_reader": {
        "batch_size": 100,
        "kind": "excel",
        "split_level": 2
      }
    },
    "schema": {
      "type": "object",
      "properties": {
        "source": {
          "description": "Source path, connector URI, or list thereof",
          "anyOf": [
            {"type": "string"},
            {"type": "array", "items": {"type": "string"}}
          ]
        }
      },
      "additionalProperties": true
    }
  }'
import json
from pathlib import Path

workflow = store.workflows.create(
    name="text-ingestion",
    workflow_name="DocumentAI",
    parameters=json.loads(Path("parameters.json").read_text()),
    schema=json.loads(Path("input-schema.json").read_text()),
)

print(workflow.id)
pharia --env aleph-alpha sherlock workflows create <store-id> \
  --name text-ingestion \
  --workflow-name DocumentAI \
  --parameters @parameters.json \
  --schema @input-schema.json

The schema property (the run-input schema) is currently required but planned for deprecation.

Below are two complete parameters documents that run without changes. Always send the whole document. Updates replace the entire object and do not merge it, so workflows update takes the same shape.

Text extraction

pdf_reader.kind: "pdf_text" extracts the text layer directly from the PDF. This is fast, cheap, and the correct default for digital-born documents. This path uses no vision model, so the document has no vision_reader group:

{
  "chunker": {
    "batch_size": 1000,
    "chunk_overlap": 600,
    "chunk_size": 4096,
    "include_image_descriptions": false,
    "kind": "recursive"
  },
  "docx_reader": {
    "batch_size": 100,
    "kind": "doc",
    "split_level": 1
  },
  "markdown_reader": {
    "batch_size": 100,
    "kind": "markdown"
  },
  "pdf_reader": {
    "batch_size": 10,
    "kind": "pdf_text",
    "pdf_backend": "pdfboss"
  },
  "plaintext_reader": {
    "batch_size": 100,
    "kind": "plaintext",
    "max_content_size": 500000,
    "min_space_ratio": 0.02
  },
  "pptx_reader": {
    "batch_size": 100,
    "include_notes": true,
    "kind": "pptx"
  },
  "vectorize": {
    "batch_size": 10,
    "distance": "cosine",
    "kind": "openai_vectorize",
    "modality": "text",
    "provider": "pharia",
    "vector_type": "dense"
  },
  "xlsx_reader": {
    "batch_size": 100,
    "kind": "excel",
    "split_level": 2
  }
}

docx_reader and xlsx_reader select their implementation with kind, like pdf_reader. The default kinds are doc and excel, backed by anydoc. They parse the file into markdown content blocks and cut pages at headings. split_level sets the heading level that starts a new page; max_page_chars (default 10000) is the backstop for documents without headings. Each Excel sheet surfaces as a level-2 heading, so the default split_level: 2 gives one page per sheet. The previous implementations stay available as "kind": "docx" (python-docx) and "kind": "xlsx" (openpyxl).

Version availability

The doc kind is available from release 0.8.10. The excel kind follows in the next release. Older deployments accept only the docx and xlsx kinds.

Scanned pages, and pages whose content is locked in diagrams, come out empty or garbled this way. The vision path corrects this.

Vision and OCR

Switch pdf_reader.kind to pdf_rasterizer to render every page to an image instead of reading its text layer. A vision_reader group is then required to turn the images back into text. The rest of the document does not change:

{
  "chunker": {
    "batch_size": 1000,
    "chunk_overlap": 600,
    "chunk_size": 4096,
    "include_image_descriptions": false,
    "kind": "recursive"
  },
  "docx_reader": {
    "batch_size": 100,
    "kind": "doc",
    "split_level": 1
  },
  "markdown_reader": {
    "batch_size": 100,
    "kind": "markdown"
  },
  "pdf_reader": {
    "batch_size": 10,
    "kind": "pdf_rasterizer",
    "pdf_backend": "pdfboss"
  },
  "plaintext_reader": {
    "batch_size": 100,
    "kind": "plaintext",
    "max_content_size": 500000,
    "min_space_ratio": 0.02
  },
  "pptx_reader": {
    "batch_size": 100,
    "include_notes": true,
    "kind": "pptx"
  },
  "vectorize": {
    "batch_size": 10,
    "distance": "cosine",
    "kind": "openai_vectorize",
    "modality": "text",
    "provider": "pharia",
    "vector_type": "dense"
  },
  "vision_reader": {
    "batch_size": 1,
    "extra_body": {
      "chat_template_kwargs": {
        "thinking": false
      }
    },
    "max_in_flight": 16,
    "model": "rednote-hilab/dots.mocr",
    "provider": "d_inference",
    "temperature": 0,
    "timeout": 120
  },
  "xlsx_reader": {
    "batch_size": 100,
    "kind": "excel",
    "split_level": 2
  }
}

vision_reader does not need a kind. Only one module fills that slot, and its kind defaults to vision. For pdf_reader, in contrast, the kind field selects pdf_text or pdf_rasterizer. vision_reader does need provider and model: both are mandatory, and there is no implicit default endpoint. They name an inference deployment, not an embedding deployment — here a vision-OCR model served through d_inference.

extra_body goes to that model verbatim on every call; this one turns off its thinking mode. temperature: 0 pins the sampling, so the same page always extracts to the same text. This is the correct choice for OCR. If you do not set it, the model server’s default applies (Holmes builds before 0.8.7 ignore the field). batch_size: 1 with max_in_flight: 16 means one page per call and sixteen calls in flight. This is the correct shape when each unit of work is a model request, not local compute.

This path is slower and more expensive per page than pdf_text. Size the timeout for the slowest page that you care about, not for the average.

See Configuring workflows for the full parameter surface and the CLI-based lifecycle.

Run the resulting workflow the same way as the default one. See Ingesting files.

Updating parameters

  • curl

  • Python SDK

  • Platform CLI

Send the whole parameters document, not just the groups you changed. This body switches the workflow created above onto the vision path — pdf_reader.kind is now pdf_rasterizer and a vision_reader group is present:

curl -X PUT "$SHERLOCK_URL/api/v1/search-stores/<store-id>/workflows/<workflow-id>" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{
    "parameters": {
      "chunker": {
        "batch_size": 1000,
        "chunk_overlap": 600,
        "chunk_size": 4096,
        "include_image_descriptions": false,
        "kind": "recursive"
      },
      "docx_reader": {
        "batch_size": 100,
        "kind": "doc",
        "split_level": 1
      },
      "markdown_reader": {
        "batch_size": 100,
        "kind": "markdown"
      },
      "pdf_reader": {
        "batch_size": 10,
        "kind": "pdf_rasterizer",
        "pdf_backend": "pdfboss"
      },
      "plaintext_reader": {
        "batch_size": 100,
        "kind": "plaintext",
        "max_content_size": 500000,
        "min_space_ratio": 0.02
      },
      "pptx_reader": {
        "batch_size": 100,
        "include_notes": true,
        "kind": "pptx"
      },
      "vectorize": {
        "batch_size": 10,
        "distance": "cosine",
        "kind": "openai_vectorize",
        "modality": "text",
        "provider": "pharia",
        "vector_type": "dense"
      },
      "vision_reader": {
        "batch_size": 1,
        "extra_body": {
          "chat_template_kwargs": {
            "thinking": false
          }
        },
        "max_in_flight": 16,
        "model": "rednote-hilab/dots.mocr",
        "provider": "d_inference",
        "temperature": 0,
        "timeout": 120
      },
      "xlsx_reader": {
        "batch_size": 100,
        "kind": "excel",
        "split_level": 2
      }
    }
  }'
import json
from pathlib import Path

store.workflows("<workflow-id>").update(
    parameters=json.loads(Path("parameters.json").read_text()),
)
pharia --env aleph-alpha sherlock workflows update <store-id> <workflow-id> \
  --parameters @parameters.json

Updates replace the entire parameters object and do not merge it. Send one of the full documents above with your edit applied. If you post only the changed group, every other group falls back to its default. To switch a text workflow to the vision path: change pdf_reader.kind to pdf_rasterizer, add the vision_reader group, and send the whole document.

Only future runs use the change. The platform does not reprocess content that is already indexed.

Advanced: custom workflows and modules

Beyond the built-in presets, the workflow engine supports custom workflows and processing modules that you write in code — for example, an enrichment module for a specific document type. This is a code-level extension point for platform engineering teams, not for application developers who consume the API. See For workflow authors.