Searching your data

This guide describes how to query a search store for content relevant to a user’s question, including filtering and citations.


Prerequisites

  • You have a search store with at least one completed ingestion run — see Ingesting files.

  • You can reach the deployment with one of: curl, the Sherlock Python SDK (sherlock-client), or the platform CLI (pharia).

Client setup

Every example below is shown three ways. They all drive the same REST API at https://sherlock.{ingressDomain}/api/v1 — the SDK and the CLI are wrappers over the requests in the curl tab. Every request needs a bearer token; the deployment’s IAM sidecar validates it and stamps the user identity that Sherlock authorizes against.

  • curl

  • Python SDK

  • Platform CLI

Log in once with the platform CLI, then export the deployment URL and a token:

pharia --env aleph-alpha iam login    # one-time OAuth login via your browser

export SHERLOCK_URL="https://sherlock.{ingressDomain}"
export AA_TOKEN="$(pharia --env aleph-alpha iam token)"

Every call then carries -H "Authorization: Bearer $AA_TOKEN". Any other source of a platform token works just as well — the API only sees the header.

Install the client from the Aleph Alpha package index:

pip install sherlock-client --extra-index-url https://alephalpha.jfrog.io/artifactory/api/pypi/holmes/simple

Then point it at the deployment and take a handle on the store you’re searching:

import os

from sherlock import SyncSherlock

client = SyncSherlock(
    url="https://sherlock.{ingressDomain}",
    token=os.environ["AA_TOKEN"],
)

store = client.v1.search_stores("<store-id>")

url and token default to the SHERLOCK_BASE_URL and SHERLOCK_TOKEN environment variables, so SyncSherlock() with no arguments works once those are set. Sherlock is the async twin with an identical API. Use either as a context manager (with SyncSherlock(...) as client:) to close the connection pool when you’re done.

The platform CLI (pharia) handles authentication for you — log in once, and every command reuses the cached token:

pharia iam login    # one-time OAuth login via your browser

Use the global --env flag to select the environment (for example, pharia --env aleph-alpha iam login); tokens are cached per environment and refreshed automatically. To reach a deployment the CLI doesn’t know, pass --base-url https://sherlock.{ingressDomain} or set PHARIA_SHERLOCK_URL. For your own tooling, print the raw token with TOKEN=$(pharia iam token).

Overview of the pharia iam commands
  • curl

  • Python SDK

  • Platform CLI

curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/searches" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{
    "requests": [
      {"query": "What is our refund policy?", "limit": 10}
    ]
  }'
results = store.search(query="What is our refund policy?", limit=10)

for hit in results:
    print(f"{hit.score:.2f} {hit.chunk_id} {hit.text[:100]}")
pharia --env aleph-alpha sherlock search <store-id> \
  --query "What is our refund policy?" --limit 10

By default, each hit returns the matching chunk’s text and chunk_id, a relevance score, and any metadata attached during ingestion. requests is a list, so you can batch multiple independent queries into a single call, each getting its own result list in the response, in the same order. The Python search() above is a convenience over that batch body — it sends one request and hands back its result list; store.batch_search(requests=[...]) sends several and returns all the lists. On the CLI, --query builds the same one-request batch and takes --limit, --offset, and --filters; everything else goes through --requests, the full batch body.

Per-request options: citations, rerank ({"model": ..., "candidates": ...} — fetches a larger candidate pool, scores it with a cross-encoder, returns the top limit), fusion to override the store’s fusion strategy for this query (see Hybrid search), vectors to search with pre-computed embeddings instead of a query string, and provider/rerank_provider to override which inference provider serves the embedding and rerank calls. The response reports per request whether reranking actually ran.

Set citations to true on a request to additionally get the source document, file, page number, and layout position for every hit — useful for showing users where an answer came from:

  • curl

  • Python SDK

  • Platform CLI

curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/searches" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{
    "requests": [
      {
        "query": "What is our refund policy?",
        "limit": 10,
        "citations": true
      }
    ]
  }'
results = store.search(
    query="What is our refund policy?",
    limit=10,
    citations=True,
)

for hit in results:
    if hit.citation:
        print(hit.citation.file_name, hit.citation.page_number)
pharia --env aleph-alpha sherlock search <store-id> --requests '{
  "requests": [
    {
      "query": "What is our refund policy?",
      "limit": 10,
      "citations": true
    }
  ]
}'

Hybrid search (dense + BM25)

A store that pairs one dense entry with one bm25 entry runs every query over both legs — the query string is embedded and BM25-encoded in-process — and fuses the results; see Creating search stores for declaring the bm25 entry.

The store’s fusion_strategy decides how the legs combine. A per-request fusion overrides it for one query — useful for comparing strategies on the same store.

Strategy How the legs combine

rrf

Rank-based union

convex

Normalized score blend — the derived default for a dense + bm25 pair whose dense leg uses cosine or dot

anchored

The dense ranking stays authoritative; lexical evidence only reorders within a narrow score band

dbsf

Distribution-based score fusion

dense_only

Disables the lexical leg

When a hybrid store answers without its lexical leg — a vectors-only request, or a query whose lexical encoding is empty — the batch response sets that request’s entry in degraded (the field is omitted when nothing in the batch degraded; an empty string means not degraded). An explicitly requested dense_only is not marked.

Version availability

Hybrid bm25 stores are available from Data and Search Platform release 0.9.0, bundled in PhariaAI 1.260900.0.

Filtering results

Filters are Qdrant-native: must (all must match), should (at least one should match), and must_not (none must match). Each condition names a payload field with field, plus a match or range:

  • curl

  • Python SDK

  • Platform CLI

curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/searches" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AA_TOKEN" \
  -d '{
    "requests": [
      {
        "query": "quarterly revenue",
        "limit": 10,
        "filters": {
          "must": [
            {"field": "file_id", "match": {"value": "01JBRZ..."}},
            {"field": "metadata.language", "match": {"value": "en"}}
          ],
          "must_not": [
            {"field": "metadata.draft", "match": {"value": true}}
          ]
        }
      }
    ]
  }'
results = store.search(
    query="quarterly revenue",
    limit=10,
    filters={
        "must": [
            {"field": "file_id", "match": {"value": "01JBRZ..."}},
            {"field": "metadata.language", "match": {"value": "en"}},
        ],
        "must_not": [
            {"field": "metadata.draft", "match": {"value": True}},
        ],
    },
)
pharia --env aleph-alpha sherlock search <store-id> \
  --query "quarterly revenue" --limit 10 \
  --filters '{
    "must": [
      {"field": "file_id", "match": {"value": "01JBRZ..."}},
      {"field": "metadata.language", "match": {"value": "en"}}
    ],
    "must_not": [
      {"field": "metadata.draft", "match": {"value": true}}
    ]
  }'
Operator Example

match.value

{"field": "metadata.language", "match": {"value": "en"}} — exact match on a string, integer, or boolean

match.any

{"field": "metadata.tags", "match": {"any": ["a", "b"]}} — match any value in a list

match.prefix

{"field": "metadata.user.site", "match": {"prefix": "dres"}} — prefix match on a string field whose index was declared with "prefix": true; see declaring filterable metadata

range

{"field": "metadata.year", "range": {"gte": 2024}} — any of gt, gte, lt, lte. Bounds are numbers, or RFC 3339 timestamps for a datetime field: {"field": "metadata.published_at", "range": {"gte": "2026-01-01T00:00:00Z", "lt": "2026-02-01T00:00:00Z"}}. One kind per range — mixing them, or sending an unparsable timestamp, returns 400 naming the bound.

nested

{"nested": {"must": [...]}} — a sub-filter as one condition

Two shorthands are also accepted: the field name as the key ({"metadata.language": {"match": {"value": "en"}}}, or a bare bound like {"metadata.year": {"gte": 2024}}) and a bare value for exact match ({"metadata.language": "en"}).

Name the field with field

A condition written as {"key": "metadata.language", "match": {...}} does not filter on metadata.language. With no field present, the first key in the object is taken as the field name, so the condition becomes a filter on a payload field literally named key (matching nothing) or an invalid condition (400) — which one you get is not stable. Use field, or the field name as the key.

What you can filter on

The store’s own isolation filter is added for you, so a filter only ever narrows within your own store.

System fields — six ids carried by every chunk, indexed in every collection. You never declare these; they are filterable in any store. Each one narrows the search to the vectors derived from that entity. The listings that carry each value are relative to /api/v1/search-stores/<store-id>, and each has an SDK method and a CLI command of the same name:

Field Narrows the search to Get the value from

file_id

one uploaded file, across all its versions and pages

GET /files

document_id

one extracted document version of a file

GET /documents

page_id

a single page of a document

GET /documents/<document-id>/pages

chunk_id

one specific chunk

the chunk_id on any search hit

workflow_id

everything ingested by one workflow

GET /workflows

workflow_run_id

everything ingested by one run of that workflow

GET /workflows/<workflow-id>/runs

Searching within a single document is the common case — {"field": "file_id", "match": {"value": "<file-id>"}}. The workflow pair is what lets you compare two ingestion configurations in the same store: run both, then filter each search to one workflow_run_id.

Your own metadata — anything your ingestion attached, filtered under the metadata. prefix: metadata.language, metadata.uri, metadata.draft. The prefix is what keeps your keys from colliding with the system fields, so a metadata key you happened to name document_id is filtered as metadata.document_id and never interferes with the real one. Metadata attached at file upload lands under the user sub-namespace and is filtered as metadata.user.<field>.

A metadata field is only filterable if the store declared it at creation — see declaring filterable metadata. Referencing an undeclared one is rejected with 400 and a hint naming the field, since the vector database refuses to filter on an unindexed field. The check covers must, should, must_not, and nested filters, in both the field and shorthand forms.

A malformed filter returns 400 Bad Request, not a server error — see Troubleshooting.

Troubleshooting