vaultspec-core
Vaultspec MCP serverLink to Vaultspec MCP server
The server exposes Vaultspec document and workflow tools to Model Context Protocol (MCP) clients using JSON-RPC over standard input and output.
For the workflow and document types, see workflow stages.
SetupLink to Setup
Vaultspec keeps provider-neutral MCP definitions in .vaultspec/mcps/*.json.
Installation and vaultspec-core spec mcps sync render those definitions into each
selected supported MCP provider’s native configuration:
Provider |
Project scope |
Broader scope |
|---|---|---|
Claude |
|
Local or user enrollment in |
Codex |
|
User enrollment in |
Antigravity |
|
Broader scopes are unsupported |
Project scope is the safe default. Select user or local scope explicitly. Native host
files contain only host-valid configuration; vaultspec records project and local
ownership in the workspace’s .vaultspec/mcp-ownership.json and user ownership in
~/.vaultspec/mcp-ownership.json, so unrelated host entries remain external. Use
vaultspec-core install --skip mcp if you manage enrollment yourself.
Configure your client to launch the server with the project root as its working directory, or set an explicit workspace.
Keep the generated python -m vaultspec_core.mcp_server.app invocation. On Windows,
running the vaultspec-core-mcp console executable can keep it locked while the client
is connected, blocking package updates that replace it.
In tool mode, the generated entry in Claude’s .mcp.json and Antigravity’s
.agents/mcp_config.json uses this configuration:
File contents
{
"mcpServers": {
"vaultspec-core": {
"command": "uvx",
"args": [
"--from",
"vaultspec-core",
"python",
"-m",
"vaultspec_core.mcp_server.app"
]
}
}
}
Codex takes the same invocation as TOML, inside a marked region the installer owns:
File contents
# <vaultspec type="mcps">
[mcp_servers."vaultspec-core"]
args = ["--from", "vaultspec-core", "python", "-m", "vaultspec_core.mcp_server.app"]
command = "uvx"
# </vaultspec>
The launch shape follows the install mode rather than being fixed: a project that
carries vaultspec-core as a dependency gets uv run --no-sync with no --from, because
the environment already has the package. Compare against these when you manage
enrollment yourself, and read the markers as ownership boundaries: edit inside them and
the next sync declines to touch the entry rather than repairing it.
Serving a read-only surfaceLink to Setup, Serving a read-only surface
Use --read-only to limit the tools the server advertises:
Command
vaultspec-core-mcp --read-only
This exposes only status, find, check, and discover to connected clients. In
this mode, check has no fix parameter.
Install modesLink to Setup, Install modes
For Core’s generated MCP entry, the install mode determines how the server starts:
Tool:
uvxuses an isolated tool environment and may download dependencies or create that environment at launch. See uv’s tool environments.Dev or dependency:
uv run --no-syncskips project environment synchronization. Prepare that environment using your project’suv syncprocedure before connecting.
See install mode selection for precedence and dependency declarations.
Convergence on upgradeLink to Setup, Convergence on upgrade
Launch entries that vaultspec wrote converge to the current standard automatically. A
managed entry whose bytes still match the fingerprint recorded when vaultspec last wrote
it is provably untouched, so vaultspec-core sync, vaultspec-core spec mcps sync, and
vaultspec-core install --upgrade all refresh it in place - no --force required - and
print exactly what changed: the entry name, the old launch command, the new launch
command, and why. A registered migration applies the same refresh, so a launch rendered
before the --no-sync guard existed (a bare uv run shape) converges whatever its
provisioned version - but it converges when a converging command runs, not on any
contact with the CLI. Those commands are vaultspec-core install --upgrade,
vaultspec-core migrations run, vaultspec-core vault repair,
vaultspec-core vault add, vaultspec-core vault feature index, and the MCP create
tool. Everything else, reads included, observes the workspace as it finds it and reports
any pending migrations as a warning.
Two kinds of entry never converge automatically. An entry you edited by hand (its bytes
no longer match the recorded fingerprint) is skipped with a warning and requires an
explicit vaultspec-core spec mcps sync --force, which overwrites your edit. An entry
recorded by a release that predates fingerprinting cannot be verified as untouched and
keeps the same --force-only behavior. External entries vaultspec never wrote are never
adopted or modified without --force. To opt out of enrollment management entirely,
provision with vaultspec-core install --skip mcp; a workspace without recorded
enrollment is never touched by the convergence migration.
vaultspec-core spec doctor additionally warns - without failing - about two states
vaultspec cannot converge itself: canonical hooks that cannot be refreshed because
prek.toml is present (transplant the entries manually), and a package-bundled seed
definition still in a static pre-mode shape (re-run that package’s installer with
--upgrade).
Select a mode with vaultspec-core install --mode tool,
vaultspec-core install --mode dependency, or vaultspec-core install --mode dev.
vaultspec consumes the package, module, and optional tool requirement while rendering;
that metadata never reaches native host configuration.
The chosen mode is recorded per package in the workspace’s committed workspace.json,
which keys each provisioned package to its own mode and version floor. A workspace that
provisions both vaultspec-core and a companion package can therefore declare each in a
different mode - one as a dependency, the other as a tool - without either overwriting
the other’s choice.
A workspace set up before install modes existed carries no recorded choice. Running
vaultspec-core install --upgrade infers the mode from how the workspace already runs -
a uv run-shaped setup that lists vaultspec-core as a dependency stays in dependency
mode, and anything else moves to tool mode - and records it, so the launch command stays
stable from then on.
Point the server at a different workspaceLink to Setup, Point the server at a different workspace
When the client’s working directory isn’t the workspace you want - for example in a
standalone setup outside an editor - set VAULTSPEC_TARGET_DIR in the canonical MCP
definition’s env map, then run vaultspec-core spec mcps sync. The provider adapter
renders normalized command, args, and env fields into each selected native target.
VAULTSPEC_TARGET_DIR is the environment equivalent of the CLI’s --target flag.
Your first callLink to Setup, Your first call
Claude reads project enrollment from .mcp.json. Codex reads .codex/config.toml for a
trusted project. Antigravity uses .agents/mcp_config.json. Reload or reconnect the
provider as required by that host. vaultspec writes enrollment only: it does not start
the server, grant trust, or bypass provider approval.
Call the status tool with no arguments. It returns a rollup report:
tool_schema_version, the workspace’s features with document counts and lifecycle
status, and any plans in flight with their completion and next open step. A populated
report confirms the server found the workspace.
Call find with no arguments. It returns the feature listing, for example
[{"name": "auth", "doc_count": 4, "weight": 7}, ...]. A non-empty listing confirms the
vault is readable. Follow up with find(feature=...) to fetch a feature’s documents,
each with a blob_hash ready for a later edit call.
EnvironmentLink to Environment
Variable |
Default |
Controls |
|---|---|---|
|
current working directory |
The workspace root containing |
|
enabled |
Set to |
Those are the two VAULTSPEC_ variables the MCP server reads directly. See
CLI reference for the full VAULTSPEC_ variable family.
VerificationLink to Verification
Command
vaultspec-core spec mcps status
vaultspec-core spec mcps status codex --scope project
vaultspec-core spec mcps status --json
Status reports aggregate and per-provider configuration health. It identifies managed,
missing, drifted, stale managed, and external entries, plus warnings. It does not start
or probe an MCP process. The command exits 0 only when aggregate status is "ok".
If status isn’t "ok", inspect the missing, drifted, stale_managed, and
warnings fields, run vaultspec-core spec mcps sync to reconcile enrollment, then
re-run the status check. Project scope is the default; user or local enrollment must be
requested explicitly, and unsupported provider/scope combinations fail.
Run vaultspec-core spec doctor --json for a broader workspace diagnosis.
ToolsLink to Tools
Most tools cover the everyday path: status orients you in the workspace, find
locates documents, create scaffolds new ones, edit makes body-prose changes, check
validates and repairs the vault, plan_progress marks steps complete, plan_edit
authors step content, and log appends a step’s rows to the plan’s execution ledger.
discover and invoke form a gateway that reaches every remaining CLI verb.
The table below is generated from the running server: each tool’s purpose is its own
handler’s summary and each annotation column its declared MCP hints. Run
vaultspec-core spec reference generate to refresh it; do not hand-edit between the
markers.
The server exposes 10 tools.
Tool |
Purpose |
Annotations |
|---|---|---|
|
Find vault documents or list features. |
read-only, idempotent |
|
Scaffold one or more vault documents from templates. |
non-destructive, not idempotent |
|
Apply one or more body-prose edits to vault documents. |
destructive, not idempotent |
|
Orient in a vaultspec project, project-wide or targeted. |
read-only, idempotent |
|
Run the vault health-check suite, optionally repairing. |
non-destructive, idempotent |
|
Mark plan steps closed or open by canonical identifier. |
non-destructive, idempotent |
|
Author plan steps: add, insert, edit, or remove. |
destructive, not idempotent |
|
Append one Step’s rows to its plan’s execution ledger. |
non-destructive, idempotent |
|
Search the long-tail verb catalog and return ranked schemas. |
read-only, idempotent |
|
Execute one cataloged verb. A verb that runs and fails is still a successful call; its exit code and stderr arrive in the error payload. |
destructive, not idempotent |
Surface provenanceLink to Tools, Surface provenance
Which of the tools above a host installing the published server actually sees. Generated from the recorded surface of that release, so it empties itself when the next one ships.
The latest published release is 0.2.4, and every tool above is in it.
Two bands of operations sit outside the tool surface.
The gateway reaches some operations, but a CLI command carries them by default:
synchronizing generated surfaces (vaultspec-core sync,
vaultspec-core spec <resource> sync) and restructuring a plan above the step level
(vaultspec-core vault plan tier promote, tier demote, and the wave, phase, and
epic intent verbs). Every invoke call carries the destructive-operation annotation,
so a connected host still asks you to confirm each one; the CLI skips that round-trip.
Other operations have no MCP path at all. Regenerating a feature index runs through
vaultspec-core vault feature index. For MCP configuration, list, add, and remove
inspect or mutate canonical definitions; status inspects native enrollment and
ownership health; sync reconciles definitions into selected targets; and uninstall
removes vaultspec-owned enrollment while preserving canonical definitions and external
host entries. Removing vaultspec from a workspace runs through
vaultspec-core uninstall.
TermsLink to Tools, Terms
Canonical identifier: the stable
W##,P##, orS##label a plan container keeps for its entire life. A retired identifier is never reused, so gaps in the sequence are expected. The display path you see, such asP01.S02, is computed for reading; the canonical identifier is the durable handle underneath it.Body prose versus CLI-owned structure: frontmatter, filenames, and plan-table structure belong to the CLI and its owning verbs. Body prose in a scaffolded document is the surface you edit by hand.
Feature index: an auto-generated document at
.vault/index/<feature>.index.mdthat lists every document belonging to a feature. It regenerates automatically when acreatebatch succeeds.Destructive-operation annotation: the MCP
destructiveHintmarked on a tool. A connected host reads this hint and asks you to confirm before it runs a call carrying it.
findLink to Tools, find
Find vault documents or list features. Read-only, idempotent.
With no arguments, find lists features: each row gives a feature’s name, document
count, and graph weight. Pass any filter (feature, type, date, or text) and
find switches to search mode, returning matching documents instead.
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
string or null |
|
Feature filter, without the |
|
list of strings or null |
|
Document-type filter (for example, |
|
string or null |
|
Exact ISO-8601 date filter. Passing it switches |
|
string or null |
|
Case-insensitive substring over the document stem and feature tag. Composes with the other filters, and matches identifiers rather than body prose. Passing it switches |
|
|
|
In search mode, how much document text to inline in each row. This is an enum, not a boolean: passing |
|
boolean |
|
In feature-listing mode, enrich each row with |
|
integer, 1 to 100 |
|
Maximum number of rows to return. |
Behavior notes:
Feature-listing rows report
name,doc_count, andweight(a graph-ranking score).weightis0when the dependency graph can’t be built, and the row notes that the ranking is unavailable.Document-search rows report
name(the file stem),type,feature,date,path(relative to the vault),blob_hash(the git blob object ID (OID) of the document’s current bytes, ornullif the file can’t be read),resource_uri(afile://locator), andbodywhen you setbodytoexcerptorfull.resource_uriis a path for your host to read directly. The server registers no MCP resources, so there is nothing toresources/readand the URI is not a protocolresource_link. Setbodywhen you want the text in the response instead.In search mode,
limitis a global cap applied across all matched types combined, not a per-type cap. If an early type fills the cap, later types can be crowded out entirely. Callfindonce per type when you need a fair spread across types.findnever raises an error for filters that match nothing; an empty result comes back as an empty list.
Example feature-listing response:
File contents
[
{"name": "auth", "doc_count": 4, "weight": 7},
{"name": "search-api", "doc_count": 2, "weight": 3}
]
Example document-search response row:
File contents
{
"name": "2026-07-11-search-api-research",
"type": "research",
"feature": "search-api",
"date": "2026-07-11",
"path": ".vault/research/2026-07-11-search-api-research.md",
"blob_hash": "<git blob OID>",
"resource_uri": "file:///.../2026-07-11-search-api-research.md"
}
The batch envelopeLink to Tools, The batch envelope
create, edit, plan_progress, and plan_edit take a list and return one envelope
describing the whole call:
Field |
Type |
Meaning |
|---|---|---|
|
|
|
|
list |
Item rows. See the sampling rule below. |
|
object |
How many items landed in each status. Exact regardless of sampling. |
|
integer |
How many items were in the request. |
|
integer |
Uneventful successes left out of |
items is a sample, not the full list. It carries every failure and every item that
produced a warning, plus some plain successes; the rest are counted in items_omitted.
A client that tallies items to learn what happened will undercount on a large batch.
Read counts for the outcome and items for the detail.
A partially failed batch is still a successful call. It returns status: "mixed" and
reports each item’s outcome rather than raising a protocol error. What fails the call
itself is a malformed request and, for create only, a schema convergence that could
not be applied before the first write; see create.
createLink to Tools, create
Scaffold one or more vault documents from templates. Mutates the vault; not idempotent.
create takes a single parameter, documents, a list of document specifications with
at least one entry.
Field |
Type |
Default |
Description |
|---|---|---|---|
|
string |
none |
Required. Feature tag, in kebab-case. A leading |
|
string or null |
|
Document type: |
|
string or null |
|
Rendered into the document’s heading. |
|
string or null |
today’s date (UTC) |
ISO-8601 date. |
|
string or null |
|
Seed prose, appended under a |
|
list of strings or null |
|
Documents to link in the |
|
list of strings or null |
|
Only the document’s required directory and feature tags. Duplicates are ignored; other tags fail the item before writing. Omit for ordinary creation. |
|
string or null |
|
Plan complexity tier, one of |
|
string or null |
|
Kebab-case filename infix that distinguishes a second document of the same type for one feature ( |
Behavior notes:
createreads the template at.vaultspec/templates/{type}.md, fills its placeholders, and writes the document to.vault/{type}/{date}-{feature}-{type}.md; you don’t set the filename directly.Items in a batch apply in order, and one item’s failure doesn’t abort the rest of the batch. Because each item validates against a vault that already includes the batch’s earlier writes, a single call can create a full dependency chain, such as a research document followed by the ADR and plan that depend on it.
On success,
createregenerates the feature index for every feature that received at least one successfully created document.An empty
documentslist is a whole-call error, distinct from a per-item failure.So is a failed schema convergence.
createdecides where it writes from the current schema, so it applies any pending schema migrations before the first item. If that fails, the whole call fails with a protocol error naming the cause and nothing is written - a client that submitted three specs gets one call-level failure rather than three item rows. Refusing is deliberate: writing part of the batch into a layout the schema no longer describes leaves a split brain that is harder to recover from than a clean refusal. Resolve it withvaultspec-core migrations statusandvaultspec-core migrations run, then retry the call.
create can fail an item for these reasons:
an invalid feature tag, extra tag, tier, or document type
a request for type
index(index documents are auto-generated; use the feature-index verb instead)an unresolvable
relatedreferencea lifecycle-dependency error (for example, requesting an
execdocument before its plan exists)a missing template
a file that already exists
Example item, from the items list:
File contents
{
"index": 0,
"target": "adr:search-api",
"status": "created",
"path": ".vault/adr/2026-07-12-search-api-adr.md",
"blob_hash": "<oid>",
"error": null,
"warnings": []
}
Example failure item, in a batch whose overall status is "mixed":
File contents
{
"index": 1,
"target": "exec:search-api",
"status": "failed",
"path": null,
"blob_hash": null,
"error": {"message": "... requires a plan ..."},
"warnings": []
}
editLink to Tools, edit
Apply one or more body-prose edits to existing vault documents. Mutates the vault; not
idempotent, and set_body and replace_section overwrite existing prose.
edit takes a single parameter, operations, a list of edit operations with at least
one entry.
Field |
Type |
Default |
Description |
|---|---|---|---|
|
string |
none |
Required. The document to edit: a stem, filename, path, or |
|
string |
none |
Required. One of |
|
string |
none |
Required. For |
|
string or null |
|
Exact heading-line text, for example |
|
string or null |
|
Optimistic-concurrency guard. If the document’s current blob hash doesn’t match, the item fails as a conflict and |
edit changes body prose only: it strips frontmatter before composing the new body,
never touches frontmatter, and never renames files. set_body replaces the whole body
while preserving frontmatter. append_section inserts content after the matched
section; replace_section keeps the heading and swaps the prose beneath it. Section
matching is exact against the heading-line text; a heading that doesn’t match fails the
item with "section_not_found": true.
Batches apply sequentially, and one item’s failure doesn’t abort the rest. Multiple
operations against the same document chain correctly: only the first operation needs
expected_blob_hash, and each later operation validates against what the prior
operation just wrote, so a blob_hash returned by one call can guard a later call.
edit can fail an item for these reasons:
an unknown operation
a missing
sectionfield on a section operationan unresolvable target (reported as
"Cannot resolve document: '...'")a blob-hash conflict (reported with
"conflict": true, and the file left untouched)
An empty operations list is a whole-call error, distinct from a per-item failure.
Example success item:
File contents
{
"index": 0,
"target": "2026-07-12-search-api-adr",
"status": "updated",
"path": ".vault/adr/2026-07-12-search-api-adr.md",
"blob_hash": "<post-write oid>",
"error": null,
"warnings": []
}
Example conflict item:
File contents
{
"index": 0,
"target": "...",
"status": "failed",
"blob_hash": null,
"error": {"message": "...", "conflict": true},
"warnings": []
}
statusLink to Tools, status
Orient in a vaultspec project, project-wide or targeted at one feature or plan. Read-only, idempotent.
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
string or null |
|
A feature tag or a plan stem/path. Omit it for a project-wide rollup. |
Call status without a target to get a rollup: the features in the vault (name,
document count, latest activity, whether a plan exists, lifecycle status, plan tier, and
plan completion percent), the plans currently in flight (stem, feature, tier, open and
closed step counts, completion percent, and the next open step), and vault-wide totals.
Every response carries a tool_schema_version field so a client can detect a server
upgrade.
Pass a target to trace one plan or feature instead. The response then reports each
plan’s steps in full detail (canonical ID, display path, checked state, the ledger stem
when rows exist, the row count, and the last verify: result), the grounding documents
behind the plan (grouped by document type), and any exec documents that name no step. A
target that resolves to no plan or feature fails the whole call with a protocol error.
status returns no hashes; get a blob_hash from find to guard an edit.
Rollup example, captured against this project’s own vault and cut to one feature and one plan of the eleven and four it returned:
File contents
{
"tool_schema_version": "0.1.73",
"kind": "rollup",
"features": [
{
"name": "docs-site",
"doc_count": 115,
"latest_activity": "2026-09-02",
"has_plan": true,
"status": "In Progress",
"plan_tier": "L3",
"plan_completion_percent": 99.1
}
],
"features_total": 11,
"plans_in_flight": [
{
"stem": "2026-08-28-docs-site-plan",
"feature": "docs-site",
"tier": "L3",
"open_steps": 1,
"closed_steps": 110,
"total_steps": 111,
"completion_percent": 99.1,
"next_open_step": "W09.P18.S109"
}
],
"totals": {
"total_docs": 394,
"total_features": 11,
"counts_by_type": { "adr": 14, "audit": 15, "plan": 11, "exec": 327 },
"orphaned_count": 0,
"dangling_link_count": 0
},
"plans": []
}
status on a feature is written for a person - In Progress, Completed - rather than
as a slug, so match it exactly rather than assuming a lowercase hyphenated form.
counts_by_type above is cut to four of its seven keys.
Trace example:
File contents
{
"kind": "trace",
"trace_kind": "feature",
"target": "search-api",
"plans": [
{
"stem": "2026-07-12-search-api-plan",
"steps": [
{ "canonical_id": "S01", "display_path": "S01", "checked": true, "record_stem": null }
]
}
]
}
checkLink to Tools, check
Run the vault health-check suite and, when asked, apply safe repairs. Not read-only, not destructive, idempotent - a repair never overwrites authored prose, so running it again converges on the same clean state.
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
string or null |
|
Restrict the check to one feature tag, without the |
|
boolean |
|
Apply safe auto-corrections instead of only reporting them. |
The response reports an overall status: "ok" when the total error count is zero,
"failed" otherwise. Warnings alone don’t fail a check. Alongside status, the
response carries fixed (whether a repair ran), total_errors, total_warnings,
total_fixed, a checks array with one summary row per validator (error, warning,
info, and fixed counts, plus a clean flag), and a findings array listing every error
and warning (informational diagnostics are dropped). Each finding names the check that
raised it, the affected path, a message, its severity, and whether it’s fixable.
Set fix to true to apply safe corrections in the same call. The response then
reports fixed: true, and if every issue found was auto-fixable, status returns to
"ok".
Clean example:
File contents
{
"status": "ok",
"fixed": false,
"total_errors": 0,
"total_warnings": 0,
"total_fixed": 0,
"checks": [{ "check": "dangling", "error_count": 0, "warning_count": 0, "info_count": 0, "fixed_count": 0, "clean": true }],
"findings": []
}
Finding example:
File contents
{
"check": "dangling",
"path": ".vault/adr/2026-07-01-search-api-adr.md",
"message": "wiki-link target not found",
"severity": "error",
"fixable": false
}
plan_progressLink to Tools, plan_progress
Mark plan steps checked or unchecked by canonical identifier. There’s no toggle operation - each step change states its target explicitly, which keeps the tool idempotent. Not read-only, not destructive, idempotent.
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
string |
- (required) |
A feature tag or a plan stem/path. |
|
list of step-state changes |
- (required, at least one entry) |
Each entry is |
Plan resolution follows one rule across every tool that accepts a plan parameter: an
exact stem or path match wins first; failing that, a feature tag resolves only when it
matches exactly one plan. A feature with several plans refuses the call and lists the
candidate stems rather than guessing. An unresolvable or ambiguous plan fails the whole
call with a protocol error, as does an empty steps list.
Each step change reports its own outcome:
"updated"- the step changed to the target state"unchanged"- the step was already at the target state, which counts as a success"failed"- an unknown step, an ambiguous ID, or an invalid state, with anerror.message
A failure in one step change doesn’t abort the batch - the aggregate status becomes
"ok", "mixed", or "failed" depending on the outcomes. The plan file is written
once, at the end of the batch, and only if a step changed. That write refreshes the
plan’s modified stamp.
Every response ends with the plan’s post-batch state: total_steps, steps_completed,
completion_percent, and next_open_step (a display path, or null when every step is
closed).
File contents
{
"status": "ok",
"plan": "2026-07-12-search-api-plan",
"items": [{ "step_id": "S01", "state": "checked", "status": "updated" }],
"total_steps": 2,
"steps_completed": 1,
"completion_percent": 50.0,
"next_open_step": "S02"
}
plan_editLink to Tools, plan_edit
Author plan steps: add, insert, edit, or remove step rows. Not read-only, destructive - removing a step retires its ID permanently - and not idempotent.
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
string |
- (required) |
A feature tag or a plan stem/path, resolved the same way as in |
|
list of plan-edit operations |
- (required, at least one entry) |
See the operation fields below. |
Each operation has these fields. Omit a field that does not apply rather than sending
null:
Field |
Type |
Default |
Description |
|---|---|---|---|
|
string |
required |
One of |
|
string or null |
|
The step’s imperative statement. Required for |
|
string or null |
|
The file or area the step touches. Required for |
|
string or null |
|
The |
|
string or null |
|
The |
|
string or null |
|
The |
|
string or null |
|
The target step for |
Operations apply in order against a single parsed plan, so an earlier add in the same
call is visible to a later edit or remove. A failure in one operation doesn’t abort
the batch; the aggregate status can come back "mixed". An empty operations list
fails the whole call with a protocol error.
Identifier guarantees belong to the plan core, not this tool:
addallocates the next free ID in sequence.insertalways allocates a new ID, even when inserting before an existing step. Nothing is ever renumbered.removeretires an ID permanently. Lateraddoperations skip past retired IDs rather than reusing them.
A write guard refuses any write that would retire identifiers the operation didn’t intend to touch.
plan_edit covers step rows only. Restructuring above the step level - tiers, waves,
phases, or the epic frame - is gateway-or-CLI territory and isn’t a first-class
operation here.
Add example:
File contents
{
"status": "ok",
"plan": "2026-07-12-search-api-plan",
"items": [{ "operation": "add", "status": "created", "step_id": "S03" }],
"total_steps": 3,
"steps_completed": 1,
"next_open_step": "S02"
}
logLink to Tools, log
Append one step’s evidence to its plan’s execution ledger. Not read-only, not
destructive (append-only), idempotent: re-logging a row changes nothing. The ledger is
one document per plan, created on the first call; it is the only execution artifact, and
this tool and vaultspec-core vault exec log share one writer.
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
string |
- (required) |
Feature tag, with or without |
|
string |
- (required) |
The parent plan’s stem. |
|
string |
- (required) |
Canonical step id ( |
|
list of strings or null |
|
One |
|
string or null |
|
A check that ran, as |
|
string or null |
|
The persona that closed the step. |
|
list of strings or null |
|
Exception notes only (data loss, skipped work, a scaffold left in code, a persistent failure). |
A malformed row or verify spec, or a feature, plan, or step that does not resolve, fails the whole call with a protocol error and writes nothing.
File contents
{
"path": ".vault/exec/2026-07-12-search-api/2026-07-12-search-api-ledger.md",
"step": "S01",
"rows": 3,
"notes": 0,
"changed": true,
"created": true
}
discover and invokeLink to Tools, discover and invoke
discover and invoke are the gateway pair for everything else. The eight hot tools
cover the verbs a client reaches for constantly; discover and invoke reach the rest
of the CLI - the long tail of verbs that don’t earn a dedicated tool but still need to
run. discover searches that catalog and returns ranked schemas. invoke runs one of
the verbs the catalog names.
discover: search the verb catalogLink to discover and invoke, discover: search the verb catalog
Read-only, idempotent.
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
string |
required |
Verb words or an intent phrase to rank the catalog against. |
|
integer |
|
Maximum number of ranked verbs to return. |
It returns the query echoed back, a count, and a list of ranked verb schemas.
Each verb schema carries:
verb- the space-joined verb path, for example"vault list"description- the verb’s help textscore- a relevance score; results are ranked by relevance, highest firstsupports_json- whether the verb accepts a--jsonflagflags- the verb’s flags, each with a name, whether it takes a value, and its help textarguments- the verb’s positional arguments, each with a name, whether it’s required, and whether it’s variadic
A discover call looks like this:
File contents
{
"query": "list vault documents"
}
And the response:
File contents
{
"query": "list vault documents",
"count": 1,
"verbs": [
{
"verb": "vault list",
"description": "List or filter vault documents.",
"score": 7.5,
"supports_json": true,
"flags": [
{ "name": "--feature", "takes_value": true, "help": "Filter by feature tag." }
],
"arguments": []
}
]
}
invoke: run a cataloged verbLink to discover and invoke, invoke: run a cataloged verb
Not read-only. Destructive, because the catalog reaches mutating verbs. Idempotency depends on the verb invoked.
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
string |
required |
The space-joined verb path returned by |
|
object or null |
|
The verb’s flags as a mapping. A list value repeats the flag; a boolean passes |
|
list of strings or null |
|
The verb’s positional operands, in CLI order. |
|
number |
|
Subprocess wall-clock budget in seconds. |
Three flags are reserved for the server: --target, --json, and --help. Do not pass
them in arguments. The server sets --target to the resolved workspace and adds
--json automatically when the verb supports it, and supplying a reserved flag fails
the call before any process starts.
At the user level, invoke runs the named verb as a subprocess of the installed
vaultspec-core binary against the resolved workspace. When the verb supports --json,
invoke parses the output into the response’s data field; otherwise it returns the
raw standard output as text. Because the catalog mixes mutating and read-only verbs,
invoke carries the destructive-operation annotation unconditionally: the host prompts
for confirmation on every call, not just the ones that change state.
A successful call looks like this:
File contents
{
"verb": "vault add",
"positionals": ["research"],
"arguments": { "feature": "search-api" }
}
File contents
{
"verb": "vault add",
"ok": true,
"exit_code": 0,
"format": "json",
"data": { "path": ".vault/research/2026-07-11-search-api-research.md" },
"stdout": null,
"error": null,
"command": ["vaultspec-core", "--target", "<root>", "vault", "add", "research", "--feature", "search-api", "--json"]
}
Two kinds of failureLink to discover and invoke, Two kinds of failure
invoke distinguishes a failed call from a failed verb. Three checks run before any
process spawns and raise a protocol error: an unknown verb path, a denylisted verb, or
an invalid argument (a reserved or undeclared flag, or a malformed positional). Nothing
runs when one of these fires.
Once the verb runs, its outcome is a per-call result, not a protocol error. The
response’s ok field is true only when the process exits zero and its output parses
successfully. A verb that runs and exits non-zero returns ok: false with a populated
error, and the call itself still succeeds:
File contents
{
"verb": "vault plan status"
}
File contents
{
"verb": "vault plan status",
"ok": false,
"exit_code": 2,
"format": "text",
"data": null,
"stdout": null,
"error": {
"kind": "nonzero_exit",
"exit_code": 2,
"stderr": "Error: missing argument 'PLAN'",
"message": "verb exited with status 2"
},
"command": ["vaultspec-core", "--target", "<root>", "vault", "plan", "status", "--json"]
}
error.kind takes one of three values: nonzero_exit (the process exited with a
non-zero status), json_parse (the process exited zero, but its declared JSON output
didn’t parse), or timeout (the process didn’t finish within the timeout budget).
The denylistLink to discover and invoke, The denylist
A verb never reaches a subprocess if it’s on the denylist. discover excludes
denylisted verbs from its results, and invoke rejects them before spawning. The
denylist covers:
uninstall- tears down the frameworkvaultspec-core spec mcps add,vaultspec-core spec mcps remove,vaultspec-core spec mcps sync, andvaultspec-core spec mcps uninstall- MCP definition or enrollment mutation, owned by thevaultspec-core spec mcpslifecycle (the read-onlyvaultspec-core spec mcps listandvaultspec-core spec mcps statusstay available)vaultspec-core vault feature index- index documents stay uncreatable through the MCP surfacevaultspec-core install,vaultspec-core sync,vaultspec-core spec hooks sync,vaultspec-core spec hooks trust,vaultspec-core spec triggers add,vaultspec-core spec triggers run, andvaultspec-core spec triggers trust- a hook or trigger file declares a shell command, and these are the verbs that write one, approve one, render one into an agent’s config, or fire the lifecycle event that runs them. Every other gateway defence is about argv hygiene and cannot help here, because a well-formed invocation of these verbs is exactly the dangerous one, and because a model’s context can contain text from a cloned repository. Approving either is an operator decision taken at a terminal, so no tool call can stand in for one. The deprecatedvaultspec-core spec hooks addandvaultspec-core spec hooks runaliases delegate to the trigger verbs and carry the same authority, so they are denied too, and leave the denylist when the aliases do. The read-onlylist,show, andstatusverbs of both groups stay available, so an agent can still read and explain what a workspace declares.
Flags the gateway does not passLink to discover and invoke, Flags the gateway does not pass
A verb can be perfectly safe to run and still declare one option that is not safe to
accept from a tool call. --editor is the example: it names a command for the CLI to
execute, which is the point of the flag at a terminal and has no meaning at all through
invoke, where no terminal is attached.
The gateway’s other screens do not catch it. The positional guard sees no operand, and
the flag-name guard sees an option the verb genuinely declares; neither looks at the
value. So --editor is refused by name for every gateway call, whichever verb declares
it, and it is withheld from the schemas discover returns.
Separately, every subprocess invoke spawns is marked as non-interactive in its
environment, and the CLI declines to open an editor when it sees that mark - from the
flag, from .vaultspec/config.toml, from VAULTSPEC_EDITOR, VISUAL or EDITOR, or
from the built-in fallback. The mark is added to the environment the gateway itself
composes, which no tool argument reaches, so a caller can neither set nor clear it. The
edit verbs stay reachable for everything else they do; only the interactive launch is
off.
Server lifetimeLink to Server lifetime
The server’s lifetime is its client connection. It exits when stdin reaches EOF, and a lifetime watchdog backstops the cases where EOF never arrives - on Windows, a client that restarts or dies can leave the server’s stdin pipe held open by inherited handles, which would otherwise leave orphaned server processes running indefinitely.
The watchdog anchors shutdown to the client process itself. On Windows it identifies the
process that created the server’s stdin pipe and exits the moment that process
terminates; when stdin is not a client-created pipe (a console launch, for example) it
watches the server’s ancestor process chain instead, ignoring short-lived launcher
processes. On other platforms a coarse poll exits the server if it is ever orphaned.
When the watchdog fires, one JSON event line ("event": "stdio_watchdog_exit") is
written to stderr before the server exits cleanly. Its reason field says why:
watched_process_exit when the client died, unanchored_orphan when the server reaped
itself as described below.
On Windows the watchdog can also lose track of the client entirely - every launcher in
the chain exits before it looks, and stdin is not a pipe it can trace. Rather than give
up for the rest of the run, it keeps looking on an interval and re-anchors as soon as a
live ancestor appears. Only if repeated checks agree that no live process is left to
serve does it write a "stdio_watchdog_disarmed" event and then exit as an orphan.
Two knobs control it:
VAULTSPEC_STDIO_WATCHDOG- set to0,false,off, ornoto disable the watchdog entirely; the server then exits only on stdin EOF.--parent-pid <PID>- names an explicit client process to watch in addition to the detected one. Useful for hosts that spawn the server through wrappers the detection cannot see through.
A watchdog that cannot start never prevents the server from serving, and an ambiguous reading is never treated as a dead client: the server exits only when a check confirms the client is gone, never when a check is inconclusive.
LoggingLink to Logging
All server logs go to stderr. Stdout is reserved exclusively for the JSON-RPC protocol stream, so nothing but protocol data can be written there without breaking the client connection.
Getting helpLink to Getting help
If the server misbehaves, start with the commands in the Verification section - they diagnose the most common configuration problems without leaving the terminal. For anything they don’t explain, file bugs and questions on the vaultspec-core issue tracker.