Searching your data
This guide describes how to query a search store for content relevant to a user’s question, including filtering and citations.
Prerequisites
-
You have a search store with at least one completed ingestion run — see Ingesting files.
-
You can reach the deployment with one of:
curl, the Sherlock Python SDK (sherlock-client), or the platform CLI (pharia).
Client setup
Every example below is shown three ways. They all drive the same REST API at https://sherlock.{ingressDomain}/api/v1 — the SDK and the CLI are wrappers over the requests in the curl tab. Every request needs a bearer token; the deployment’s IAM sidecar validates it and stamps the user identity that Sherlock authorizes against.
-
curl
-
Python SDK
-
Platform CLI
Log in once with the platform CLI, then export the deployment URL and a token:
pharia --env aleph-alpha iam login # one-time OAuth login via your browser
export SHERLOCK_URL="https://sherlock.{ingressDomain}"
export AA_TOKEN="$(pharia --env aleph-alpha iam token)"
Every call then carries -H "Authorization: Bearer $AA_TOKEN". Any other source of a platform token works just as well — the API only sees the header.
Install the client from the Aleph Alpha package index:
pip install sherlock-client --extra-index-url https://alephalpha.jfrog.io/artifactory/api/pypi/holmes/simple
Then point it at the deployment and take a handle on the store you’re searching:
import os
from sherlock import SyncSherlock
client = SyncSherlock(
url="https://sherlock.{ingressDomain}",
token=os.environ["AA_TOKEN"],
)
store = client.v1.search_stores("<store-id>")
url and token default to the SHERLOCK_BASE_URL and SHERLOCK_TOKEN environment variables, so SyncSherlock() with no arguments works once those are set. Sherlock is the async twin with an identical API. Use either as a context manager (with SyncSherlock(...) as client:) to close the connection pool when you’re done.
The platform CLI (pharia) handles authentication for you — log in once, and every command reuses the cached token:
pharia iam login # one-time OAuth login via your browser
Use the global --env flag to select the environment (for example, pharia --env aleph-alpha iam login); tokens are cached per environment and refreshed automatically. To reach a deployment the CLI doesn’t know, pass --base-url https://sherlock.{ingressDomain} or set PHARIA_SHERLOCK_URL. For your own tooling, print the raw token with TOKEN=$(pharia iam token).
Basic search
-
curl
-
Python SDK
-
Platform CLI
curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/searches" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AA_TOKEN" \
-d '{
"requests": [
{"query": "What is our refund policy?", "limit": 10}
]
}'
results = store.search(query="What is our refund policy?", limit=10)
for hit in results:
print(f"{hit.score:.2f} {hit.chunk_id} {hit.text[:100]}")
pharia --env aleph-alpha sherlock search <store-id> \
--query "What is our refund policy?" --limit 10
By default, each hit returns the matching chunk’s text and chunk_id, a relevance score, and any metadata attached during ingestion. requests is a list, so you can batch multiple independent queries into a single call, each getting its own result list in the response, in the same order. The Python search() above is a convenience over that batch body — it sends one request and hands back its result list; store.batch_search(requests=[...]) sends several and returns all the lists. On the CLI, --query builds the same one-request batch and takes --limit, --offset, and --filters; everything else goes through --requests, the full batch body.
Per-request options: citations, rerank ({"model": ..., "candidates": ...} — fetches a larger candidate pool, scores it with a cross-encoder, returns the top limit), fusion to override the store’s fusion strategy for this query (see Hybrid search), vectors to search with pre-computed embeddings instead of a query string, and provider/rerank_provider to override which inference provider serves the embedding and rerank calls. The response reports per request whether reranking actually ran.
Set citations to true on a request to additionally get the source document, file, page number, and layout position for every hit — useful for showing users where an answer came from:
-
curl
-
Python SDK
-
Platform CLI
curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/searches" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AA_TOKEN" \
-d '{
"requests": [
{
"query": "What is our refund policy?",
"limit": 10,
"citations": true
}
]
}'
results = store.search(
query="What is our refund policy?",
limit=10,
citations=True,
)
for hit in results:
if hit.citation:
print(hit.citation.file_name, hit.citation.page_number)
pharia --env aleph-alpha sherlock search <store-id> --requests '{
"requests": [
{
"query": "What is our refund policy?",
"limit": 10,
"citations": true
}
]
}'
Hybrid search (dense + BM25)
A store that pairs one dense entry with one bm25 entry runs every query over both legs — the query string is embedded and BM25-encoded in-process — and fuses the results; see Creating search stores for declaring the bm25 entry.
The store’s fusion_strategy decides how the legs combine. A per-request fusion overrides it for one query — useful for comparing strategies on the same store.
| Strategy | How the legs combine |
|---|---|
|
Rank-based union |
|
Normalized score blend — the derived default for a dense + bm25 pair whose dense leg uses |
|
The dense ranking stays authoritative; lexical evidence only reorders within a narrow score band |
|
Distribution-based score fusion |
|
Disables the lexical leg |
When a hybrid store answers without its lexical leg — a vectors-only request, or a query whose lexical encoding is empty — the batch response sets that request’s entry in degraded (the field is omitted when nothing in the batch degraded; an empty string means not degraded). An explicitly requested dense_only is not marked.
|
Version availability
Hybrid |
Filtering results
Filters are Qdrant-native: must (all must match), should (at least one should match), and must_not (none must match). Each condition names a payload field with field, plus a match or range:
-
curl
-
Python SDK
-
Platform CLI
curl -X POST "$SHERLOCK_URL/api/v1/search-stores/<store-id>/searches" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AA_TOKEN" \
-d '{
"requests": [
{
"query": "quarterly revenue",
"limit": 10,
"filters": {
"must": [
{"field": "file_id", "match": {"value": "01JBRZ..."}},
{"field": "metadata.language", "match": {"value": "en"}}
],
"must_not": [
{"field": "metadata.draft", "match": {"value": true}}
]
}
}
]
}'
results = store.search(
query="quarterly revenue",
limit=10,
filters={
"must": [
{"field": "file_id", "match": {"value": "01JBRZ..."}},
{"field": "metadata.language", "match": {"value": "en"}},
],
"must_not": [
{"field": "metadata.draft", "match": {"value": True}},
],
},
)
pharia --env aleph-alpha sherlock search <store-id> \
--query "quarterly revenue" --limit 10 \
--filters '{
"must": [
{"field": "file_id", "match": {"value": "01JBRZ..."}},
{"field": "metadata.language", "match": {"value": "en"}}
],
"must_not": [
{"field": "metadata.draft", "match": {"value": true}}
]
}'
| Operator | Example |
|---|---|
|
|
|
|
|
|
|
|
|
|
Two shorthands are also accepted: the field name as the key ({"metadata.language": {"match": {"value": "en"}}}, or a bare bound like {"metadata.year": {"gte": 2024}}) and a bare value for exact match ({"metadata.language": "en"}).
|
Name the field with
fieldA condition written as |
What you can filter on
The store’s own isolation filter is added for you, so a filter only ever narrows within your own store.
System fields — six ids carried by every chunk, indexed in every collection. You never declare these; they are filterable in any store. Each one narrows the search to the vectors derived from that entity. The listings that carry each value are relative to /api/v1/search-stores/<store-id>, and each has an SDK method and a CLI command of the same name:
| Field | Narrows the search to | Get the value from |
|---|---|---|
|
one uploaded file, across all its versions and pages |
|
|
one extracted document version of a file |
|
|
a single page of a document |
|
|
one specific chunk |
the |
|
everything ingested by one workflow |
|
|
everything ingested by one run of that workflow |
|
Searching within a single document is the common case — {"field": "file_id", "match": {"value": "<file-id>"}}. The workflow pair is what lets you compare two ingestion configurations in the same store: run both, then filter each search to one workflow_run_id.
Your own metadata — anything your ingestion attached, filtered under the metadata. prefix: metadata.language, metadata.uri, metadata.draft. The prefix is what keeps your keys from colliding with the system fields, so a metadata key you happened to name document_id is filtered as metadata.document_id and never interferes with the real one. Metadata attached at file upload lands under the user sub-namespace and is filtered as metadata.user.<field>.
A metadata field is only filterable if the store declared it at creation — see declaring filterable metadata. Referencing an undeclared one is rejected with 400 and a hint naming the field, since the vector database refuses to filter on an unindexed field. The check covers must, should, must_not, and nested filters, in both the field and shorthand forms.
A malformed filter returns 400 Bad Request, not a server error — see Troubleshooting.
Troubleshooting
See Troubleshooting.