---
generator: "vaultspec-marketing discovery"
title: "How vaultspec-rag runs - Vaultspec documentation"
description: "vaultspec-rag is two programs you set up separately. You install a search service and start it yourself."
url: "https://vaultspec.neve.md/docs/guides/rag-service.html"
component: "site"
date_modified: "2026-10-08T14:11:28+02:00"

---

[View as Markdown](<https://vaultspec.neve.md/docs/guides/rag-service.md>)

# How vaultspec-rag runs

vaultspec-rag is two programs you set up separately. You install a search service and start it yourself. Your agent launches a small MCP server that reaches that service. Both run on your machine, and neither listens beyond loopback.

1. **your agent** reads its MCP configuration and starts the MCP server
2. **MCP server** `vaultspec-search-mcp` over stdio; it loads no models and forwards every call
3. **search service** you start it; it listens on 127.0.0.1:8766 and runs the models on the GPU
4. **index store** a managed server on 127.0.0.1:8765, or a directory on disk with `--local-only`

vaultspec-core's MCP server is a separate entry in the same `.mcp.json` and never calls this service.

Your agent starts its MCP server. You start the search service, once, and it serves every project on the machine.

## What you install

One package holds the command, the service and the MCP server. The service needs the `[gpu]` extra and the MCP server needs the `[mcp]` extra, so a machine that runs all of it takes `vaultspec-rag[gpu,mcp]`. The release binaries carry both. An environment that only talks to a running service can leave `[gpu]` out; [choose what each environment runs](<https://vaultspec.neve.md/docs/rag/installation.html#choose-what-this-environment-runs>) has the combinations.

`vaultspec-rag install` sets up the project and writes the MCP entry your agent reads. A host with a usable accelerator also provisions the models and the store’s server binary; a client skips every download. Host `server start` provisions any missing files before startup. Check the [requirements](<https://vaultspec.neve.md/docs/requirements.html>) first. The service runs only on a supported GPU. Its default models are public and need no Hugging Face token; [provisioning](<https://vaultspec.neve.md/docs/rag/provisioning.html>) explains the downloads.

## You start the service

**Command**

```bash
uv run vaultspec-rag server start
```

That starts the managed store on `127.0.0.1:8765`, loads the models once, and serves on `127.0.0.1:8766` until you stop it. Later searches reuse the loaded models, which is why it stays running. Drop the `uv run` prefix if you installed a standalone tool or a binary.

The service runs in the Python environment that launched it, and that environment’s PyTorch build decides the GPU. Start it from an installation with the `[gpu]` extra; one without it cannot start the service. [Which Python environment runs the service](<https://vaultspec.neve.md/docs/rag/service-mode.html#which-python-environment-runs-the-service>) explains why. To keep the index in a directory instead of a managed server, add `--local-only`. That changes storage only; the models still run on the GPU. See [storage backends](<https://vaultspec.neve.md/docs/rag/backends.html>).

Nothing starts the service for you at login. vaultspec-rag ships no system service, so wrap `server start` yourself if you want that; see [running it automatically](<https://vaultspec.neve.md/docs/rag/service-mode.html#running-it-automatically>).

## Your agent starts the MCP server

The agent reads the entry `install` wrote and launches the MCP server as a child process over stdio. The entry runs it through `uvx` or through the project’s environment, depending on how vaultspec-rag is installed; [what your agent launches](<https://vaultspec.neve.md/docs/rag/installation.html#what-your-agent-launches>) shows each form. That process loads no models. It resolves the project, forwards each search to the running service, and tags the call with the project so one service can answer for many.

The service does not speak MCP, and no client connects to it directly. [How the connection works](<https://vaultspec.neve.md/docs/rag/mcp.html#how-the-connection-works>) has the detail, and [Use vaultspec-rag with MCP clients](<https://vaultspec.neve.md/docs/rag/mcp.html>) covers client setup and the tools it exposes.

## vaultspec-core’s MCP server is separate

vaultspec-core has its own MCP server for records, plans and checks. It sits beside vaultspec-rag’s entry in the same `.mcp.json` and never calls the search service. [Connect your agent](<https://vaultspec.neve.md/docs/guides/mcp.html#two-products-one-file>) shows the two entries together and what happens when they are installed in different modes.

## One service for every project

Only one service runs per machine, because it owns the GPU and the managed store. It keeps a slot for each project it has served and drops idle ones on its own. A new git worktree of a project it has already indexed reuses the vectors it can, rather than encoding everything again.

- [Manage projects](<https://vaultspec.neve.md/docs/rag/service-mode.html#manage-projects>)
- [Reusing vectors across worktrees](<https://vaultspec.neve.md/docs/rag/indexing.html#reusing-vectors-across-worktrees>)

## Day to day

| To | Run |
| --- | --- |
| Start it | `uv run vaultspec-rag server start` |
| See whether it is up | `uv run vaultspec-rag server status` |
| Find out why it is not healthy | `uv run vaultspec-rag server doctor` |
| Read its log | `uv run vaultspec-rag server logs` |
| Stop it | `uv run vaultspec-rag server stop` |

On Windows, `server stop` cannot signal the detached service, so it ends it with a bounded force-stop and cleans up after it. [Stop and restart the service](<https://vaultspec.neve.md/docs/rag/service-mode.html#stop-and-restart-the-service>) has the details.

## When the service is not running

A search does not quietly load the models into the calling process. It fails, names the problem and suggests the fix, so a stopped service is never mistaken for an empty result. The same happens when the command-line tool and the running service are different releases. [Troubleshooting](<https://vaultspec.neve.md/docs/rag/service-mode.html#troubleshooting>) covers ports in use, a service that will not stop, and a store that will not start.

## Where to go next

- [Install vaultspec-rag](<https://vaultspec.neve.md/docs/rag/installation.html>)
- [Your first search](<https://vaultspec.neve.md/docs/rag/getting-started.html>)
- [The background service in full](<https://vaultspec.neve.md/docs/rag/service-mode.html>)
