View as Markdown

How vaultspec-rag runsLink to How vaultspec-rag runs

vaultspec-rag is two programs you set up separately. You install a search service and start it yourself. Your agent launches a small MCP server that reaches that service. Both run on your machine, and neither listens beyond loopback.

The four parts of a vaultspec-rag setup and who starts each Four boxes in a row. Your agent reads its MCP configuration and starts the MCP server. The MCP server, vaultspec-search-mcp, talks to the agent over stdio, loads no models, and forwards every call to the search service. You start the search service; it listens on 127.0.0.1:8766 and runs the models on the GPU. It keeps its index in a store that is either a managed server on 127.0.0.1:8765 or, with --local-only, a directory on disk. The agent starts the first two; you start the last two. vaultspec-core's own MCP server is a separate entry in the same configuration and never calls this service. your agent MCP server search service index store reads .mcp.json and starts the MCP server stdio, loads no models, forwards calls 127.0.0.1:8766, models on the GPU managed on :8765, or on disk: --local-only started by your agent started by you: server start vaultspec-core's MCP server is a separate entry in the same .mcp.json and never calls this service
  1. your agentreads its MCP configuration and starts the MCP server
  2. MCP servervaultspec-search-mcp over stdio; it loads no models and forwards every call
  3. search serviceyou start it; it listens on 127.0.0.1:8766 and runs the models on the GPU
  4. index storea managed server on 127.0.0.1:8765, or a directory on disk with --local-only

vaultspec-core's MCP server is a separate entry in the same .mcp.json and never calls this service.

Your agent starts its MCP server. You start the search service, once, and it serves every project on the machine.

What you installLink to What you install

One package holds the command, the service and the MCP server. The service needs the [gpu] extra and the MCP server needs the [mcp] extra, so a machine that runs all of it takes vaultspec-rag[gpu,mcp]. The release binaries carry both. An environment that only talks to a running service can leave [gpu] out; choose what each environment runs has the combinations.

vaultspec-rag install sets up the project and writes the MCP entry your agent reads. A host with a usable accelerator also provisions the models and the store’s server binary; a client skips every download. Host server start provisions any missing files before startup. Check the requirements first. The service runs only on a supported GPU. Its default models are public and need no Hugging Face token; provisioning explains the downloads.

You start the serviceLink to You start the service

Command

uv run vaultspec-rag server start

That starts the managed store on 127.0.0.1:8765, loads the models once, and serves on 127.0.0.1:8766 until you stop it. Later searches reuse the loaded models, which is why it stays running. Drop the uv run prefix if you installed a standalone tool or a binary.

The service runs in the Python environment that launched it, and that environment’s PyTorch build decides the GPU. Start it from an installation with the [gpu] extra; one without it cannot start the service. Which Python environment runs the service explains why. To keep the index in a directory instead of a managed server, add --local-only. That changes storage only; the models still run on the GPU. See storage backends.

Nothing starts the service for you at login. vaultspec-rag ships no system service, so wrap server start yourself if you want that; see running it automatically.

Your agent starts the MCP serverLink to Your agent starts the MCP server

The agent reads the entry install wrote and launches the MCP server as a child process over stdio. The entry runs it through uvx or through the project’s environment, depending on how vaultspec-rag is installed; what your agent launches shows each form. That process loads no models. It resolves the project, forwards each search to the running service, and tags the call with the project so one service can answer for many.

The service does not speak MCP, and no client connects to it directly. How the connection works has the detail, and Use vaultspec-rag with MCP clients covers client setup and the tools it exposes.

vaultspec-core’s MCP server is separateLink to vaultspec-core’s MCP server is separate

vaultspec-core has its own MCP server for records, plans and checks. It sits beside vaultspec-rag’s entry in the same .mcp.json and never calls the search service. Connect your agent shows the two entries together and what happens when they are installed in different modes.

One service for every projectLink to One service for every project

Only one service runs per machine, because it owns the GPU and the managed store. It keeps a slot for each project it has served and drops idle ones on its own. A new git worktree of a project it has already indexed reuses the vectors it can, rather than encoding everything again.

Day to dayLink to Day to day

To

Run

Start it

uv run vaultspec-rag server start

See whether it is up

uv run vaultspec-rag server status

Find out why it is not healthy

uv run vaultspec-rag server doctor

Read its log

uv run vaultspec-rag server logs

Stop it

uv run vaultspec-rag server stop

On Windows, server stop cannot signal the detached service, so it ends it with a bounded force-stop and cleans up after it. Stop and restart the service has the details.

When the service is not runningLink to When the service is not running

A search does not quietly load the models into the calling process. It fails, names the problem and suggests the fix, so a stopped service is never mistaken for an empty result. The same happens when the command-line tool and the running service are different releases. Troubleshooting covers ports in use, a service that will not stop, and a store that will not start.

Where to go nextLink to Where to go next