What is the Data and Search Platform?

The Data and Search Platform is Aleph Alpha’s ingestion and retrieval platform: it turns raw documents into searchable, citation-ready content for your AI applications. This article explains what it does, how it differs from the previous generation of the platform, and when you’d reach for it.


What the Data and Search Platform does

Building document search from scratch is deceptively hard: extracting structure from PDFs and Office documents, chunking text without losing context, generating and storing embeddings at scale, and combining keyword and semantic search all need to work together — and keep working as your document volume grows.

The Data and Search Platform provides one integrated pipeline for this:

Documents -> Ingestion -> Embeddings -> Retrieval -> Search results

You upload documents, run an ingestion workflow against them, and query the result over HTTP — with hybrid search (dense vectors and keyword/BM25) and per-request metadata filtering built in.

When to reach for it

The Data and Search Platform is designed to support a project across its lifecycle, from initial experimentation, through scaling up usage, to running in production — the same search store and workflow configuration carries over; only the parameters and infrastructure sizing change as you grow.

How this differs from the previous generation

The presets the platform ships with today cover document ingestion and retrieval — turning files into searchable chunks. That’s a starting point, not a ceiling on the platform: workflows are a general, user-extensible pipeline (see For workflow authors), and the ambition is to cover any data workflow need this way, with users authoring the workflows their integration requires. If you need a workflow the built-in presets don’t cover, talk to your data team about building it.

Within its scope, the Data and Search Platform also simplifies the object model:

  • One HTTP API surface instead of two separate services (previously pharia-data-api for ingestion and document-index for search).

  • One end-user object, the Search Store, instead of separately managing namespaces, collections, and indexes — see Core concepts.

  • Workflow-based ingestion with built-in presets, instead of manually wiring stages, transformations, and repositories — see Workflows and presets.

Next steps