Connecting AI agents via MCP

This guide describes how to connect IDE assistants and AI agents (for example, Cursor) to your search stores on the Data and Search Platform using the Model Context Protocol (MCP), so they can search and reason over your indexed documents directly.


What MCP is for

MCP gives an agent read-mostly access to your indexed content: it can search, list stores/files/workflows, run an existing workflow, poll its status, and generate a temporary download link. It cannot create search stores or upload files — do those with the platform CLI first (see Creating search stores and Ingesting files).

Prerequisites

  • The Data and Search Platform is deployed: https://sherlock.{ingressDomain}

  • You have a search store with at least one completed ingestion run.

  • You have a valid authorisation token (see Get an authorisation token).

Get an authorisation token

To use the Aleph Alpha APIs, you need a valid authorisation token. Use the platform CLI (pharia) to log in and obtain one:

pharia iam login    # one-time OAuth login via your browser
pharia iam token    # print the cached access token

The token is cached per environment and refreshed automatically, so you can pipe it straight into your MCP client configuration:

TOKEN=$(pharia iam token)

Use the global --env flag to select the environment (for example, pharia --env aleph-alpha iam login).

Overview of the pharia iam commands

Connect Cursor

Add the Data and Search Platform as an MCP server in Cursor’s MCP settings:

{
  "mcpServers": {
    "data-search": {
      "url": "https://sherlock.{ingressDomain}/api/v1/mcp",
      "headers": {
        "Authorization": "Bearer {your-token}"
      }
    }
  }
}

Restart Cursor or reload its MCP servers after saving. Other MCP-compatible clients (Claude Desktop, custom agent clients) can connect the same way.

Available tools

Tool Description Mutates data

search

Semantic search over a search store, with optional citations

No

list_search_stores

List the caller’s search stores

No

list_files

List files in a search store

No

get_doc

Get a document’s metadata by ID

No

get_doc_pages

Get a document’s extracted pages

No

list_workflows

List workflows registered on a store

No

get_workflow

Get a workflow’s details and expected input schema

No

run_workflow

Execute a workflow against existing files

Yes

get_run

Poll a workflow run’s status

No

create_public_url

Generate a temporary, presigned download URL for a file

No

Since there’s no upload tool, a typical agent-driven ingestion flow looks like:

  1. list_files — find the ID of a file already uploaded via the CLI.

  2. get_workflow — read the workflow’s input_schema.

  3. run_workflow — pass {"source": ["files://<file-id>"]} as the tool’s input argument (the same value the CLI takes via runs execute --input).

  4. get_run — poll until status is completed or failed.

  5. search — query the newly indexed content.

Once content is indexed, prompts like "Search my documents for mentions of revenue" use the search tool automatically.

Troubleshooting

Tools not visible in the client

  • Error: The MCP server doesn’t appear, or has no tools

  • Solution: Reload the MCP server configuration; double-check the URL and that the JSON is valid.

Unauthorized

  • Error: 401 Unauthorized

  • Solution: Verify the bearer token in your MCP client configuration is current.

Search returns no results

  • Error: search returns an empty result list

  • Solution: The store has no indexed chunks yet. Run an ingestion workflow first — see Ingesting files.