Data and Search Platform: Core concepts
This article introduces the concepts behind the Data and Search Platform: the entities it manages and how access is scoped per user.
Entity model
The Search Store is the platform’s central object: it is the unit of storage, transformation, search, and access control. You create a search store, upload files into it, run ingestion workflows over them, and query the result — every other entity is scoped inside it:
-
Search Store — the top-level container you create and own. It holds the embedding model configuration used for everything inside it. A search store is private by default, but can be shared with other users or made public.
-
File — the raw bytes you upload, under a MIME type, identified by hash and URI. Files remain unchanged throughout all transformations: workflows read them and derive new entities from them, but never modify the original.
-
Document — an extracted, versioned representation of a file. Each document version corresponds 1:1 with a file version and is scoped under its workflow run. A document has one or more pages (or other derived entities, depending on the file type).
-
Page — a larger semantic unit of text within a document (for example, one PDF page).
-
Chunk — a semantic subset of a page’s text. Chunks are what search retrieves and ranks against your query.
Multi-user isolation and sharing
Every search store belongs to the user who created it and is private by default — only the owner can see or use it. Access to a store is role-based:
-
Owner — the creator: full control, including deleting the store and changing its visibility and permissions.
-
Writer — can ingest and modify content in the store.
-
Viewer — read-only access: listing and searching.
Owners can grant viewer or writer access to individual users or groups, or make the store public, which gives every authenticated user read access. Beyond these per-store roles, a service-level admin role can access any store; administrative actions emit audit events.
This isolation applies to everything scoped inside a search store: its files, documents, pages, and chunks all inherit the store’s access rules. There is no separate per-file permission model. This design is intentional: it reduces strain on the authorization machinery and allows the system to scale out to a larger set of users.
Advanced: how search stores are stored
You don’t need this section to use the API — it explains what sits behind a search store, why several stores share one collection, and why store creation can fail with 422 Collection Not Provisioned.
Collections are sharded by geometry, not by store
Behind a search store sits a Collection: the physical vector-database primitive (Qdrant), holding one named vector per entry in the store’s models configuration. Its name is derived from that configuration — one code per vector, joined with _ in type order (dense, sparse, bm25, multi, quantized):
| Part | Values |
|---|---|
modality |
|
vector type |
|
distance |
|
dimension |
appended when fixed, omitted when variable |
A store that combines dense text embeddings at 1024 dimensions scored by cosine distance with a sparse text vector scored by dot product lands in tdc1024_tsd. A store with no vector models has no backing collection at all.
Because that name is derived rather than allocated per store, every store with the same geometry shares one collection:
store "docs" (t·d·c·1024 + t·s·d) ┐ store "contracts" (t·d·c·1024 + t·s·d) ├─▶ collection tdc1024_tsd store "wiki" (t·d·c·1024 + t·s·d) ┘ points tagged searchstore_id, filtered on every query store "screenshots" (i·d·c·512) ──────────▶ collection idc512
This is what makes a search store a virtual view: every point carries the searchstore_id and user_id it belongs to, and each query has the searchstore_id filter injected before it reaches the vector database — a search that would run without store isolation is rejected outright, so stores sharing a collection cannot see each other’s vectors. The collection indexes those tenant keys alongside the entity ids (file, document, page, chunk, workflow, run) and any custom metadata.* fields declared by the stores using it — custom indexes are the additive union across the sharing stores.
Why stores share collections
Qdrant is not built to hold a collection per tenant: its multitenancy guide recommends "a single collection per embedding model with payload-based partitioning", warns that creating "hundreds and thousands of collections per cluster" raises resource overhead unsustainably, and notes that "in Qdrant Cloud, we limit the amount of collections per cluster to 1000". Sharding by geometry keeps the collection count bounded by the number of embedding configurations in use rather than the number of search stores — which is what lets a store stay cheap enough to create one per project or per user.
Provisioning collections
Collections are global infrastructure, so provisioning them is a platform-admin action — but it happens through the API, and it can happen at any time, not only during environment setup:
| Endpoint | Who | Behaviour |
|---|---|---|
|
platform admin |
Provision the collection and its standard payload indexes for a set of vector models. Idempotent: provisioning an existing collection ensures its indexes and returns it. |
|
any authenticated user |
List provisioned collections. Use it to check which embedding configurations are available before creating a store. |
|
any authenticated user |
Get one collection by name. |
|
platform admin |
Idempotent. Affects every store backed by that collection. |
Admin mutations are audited. Deployments can also pre-provision collections at rollout time through the server’s migrate command.
Creating a search store checks up front that its derived collection exists — otherwise the failure would surface much later, on ingest or search. If it doesn’t, the API returns 422 Collection Not Provisioned together with the derived collection name; ask an admin to provision it, or pick a configuration that already exists. See Troubleshooting.
Beyond the vector database, the platform is backed by a relational database for metadata and extracted text, object storage for uploaded files, and a workflow orchestration layer that runs ingestion pipelines. See Resource requirements for what this means for sizing your deployment.