How vaultspec-rag runsLink to How vaultspec-rag runs
vaultspec-rag is two programs you set up separately. You install a search service and start it yourself. Your agent launches a small MCP server that reaches that service. Both run on your machine, and neither listens beyond loopback.
- your agentreads its MCP configuration and starts the MCP server
- MCP server
vaultspec-search-mcpover stdio; it loads no models and forwards every call - search serviceyou start it; it listens on 127.0.0.1:8766 and runs the models on the GPU
- index storea managed server on 127.0.0.1:8765, or a directory on disk with
--local-only
vaultspec-core's MCP server is a separate entry in the same .mcp.json and never calls this service.
What you installLink to What you install
One package holds the command, the service and the MCP server. The service needs
the [gpu] extra and the MCP server needs the [mcp] extra, so a machine that
runs all of it takes vaultspec-rag[gpu,mcp]. The release binaries carry both.
An environment that only talks to a running service can leave [gpu] out;
choose what each environment runs
has the combinations.
vaultspec-rag install sets up the project and writes the MCP entry your agent
reads. A host with a usable accelerator also provisions the models and the store’s
server binary; a client skips every download. Host server start provisions any
missing files before startup. Check the requirements first.
The service runs only on a supported GPU. Its default models are public and need
no Hugging Face token; provisioning explains the downloads.
You start the serviceLink to You start the service
Command
uv run vaultspec-rag server start
That starts the managed store on 127.0.0.1:8765, loads the models once, and serves
on 127.0.0.1:8766 until you stop it. Later searches reuse the loaded models, which
is why it stays running. Drop the uv run prefix if you installed a standalone tool
or a binary.
The service runs in the Python environment that launched it, and that environment’s
PyTorch build decides the GPU. Start it from an installation with the [gpu] extra;
one without it cannot start the service.
Which Python environment runs the service
explains why. To keep the index in a directory instead of a managed server, add
--local-only. That changes storage only; the models still run on the GPU. See
storage backends.
Nothing starts the service for you at login. vaultspec-rag ships no system service,
so wrap server start yourself if you want that; see
running it automatically.
Your agent starts the MCP serverLink to Your agent starts the MCP server
The agent reads the entry install wrote and launches the MCP server as a child
process over stdio. The entry runs it through uvx or through the project’s
environment, depending on how vaultspec-rag is installed;
what your agent launches shows
each form. That process loads no models. It resolves the project,
forwards each search to the running service, and tags the call with the project so
one service can answer for many.
The service does not speak MCP, and no client connects to it directly. How the connection works has the detail, and Use vaultspec-rag with MCP clients covers client setup and the tools it exposes.
vaultspec-core’s MCP server is separateLink to vaultspec-core’s MCP server is separate
vaultspec-core has its own MCP server for records, plans and checks. It sits beside
vaultspec-rag’s entry in the same .mcp.json and never calls the search service.
Connect your agent shows the two entries together and
what happens when they are installed in different modes.
One service for every projectLink to One service for every project
Only one service runs per machine, because it owns the GPU and the managed store. It keeps a slot for each project it has served and drops idle ones on its own. A new git worktree of a project it has already indexed reuses the vectors it can, rather than encoding everything again.
Day to dayLink to Day to day
To |
Run |
|---|---|
Start it |
|
See whether it is up |
|
Find out why it is not healthy |
|
Read its log |
|
Stop it |
|
On Windows, server stop cannot signal the detached service, so it ends it with a
bounded force-stop and cleans up after it.
Stop and restart the service
has the details.
When the service is not runningLink to When the service is not running
A search does not quietly load the models into the calling process. It fails, names the problem and suggests the fix, so a stopped service is never mistaken for an empty result. The same happens when the command-line tool and the running service are different releases. Troubleshooting covers ports in use, a service that will not stop, and a store that will not start.