Provisioning: what is downloaded and how it is checkedLink to Provisioning: what is downloaded and how it is checked
Examples use the installed-tool form and call vaultspec-rag directly. If
vaultspec-rag is a project dependency, prefix each command with uv run; see
the installation guide for lane selection, including the
standalone tool and no-install routes.
A host installation needs two things that do not come with the Python package: the model files and, for the default backend, a Qdrant server executable. vaultspec-rag fetches both for you. This page is the one place that says what is fetched, from where, how it is checked, where it is kept, which commands do it, and how to run without a network. Settings are listed in the configuration reference; this page links to them rather than repeating their syntax.
On this pageLink to On this page
What is fetchedLink to What is fetched
What |
Default source |
Checked against |
Kept in |
|---|---|---|---|
Qdrant server executable |
|
Two SHA256 digests compiled into vaultspec-rag |
|
Dense, sparse, reranker models |
|
A SHA256 digest for every file, compiled into vaultspec-rag |
The Hugging Face cache, under |
vaultspec-rag does not download PyTorch. vaultspec-rag install writes the CUDA wheel
index into your pyproject.toml, and uv downloads the build from it; see
download sources and
pin the GPU build.
Nothing else is fetched at run time. Indexing, searching, the status commands and the MCP server make no download of their own.
Which commands provisionLink to Which commands provision
Four commands fetch, all through one implementation, so they report the same progress and the same failures.
Command |
Fetches |
Skip or switch off with |
|---|---|---|
|
Missing model files, then the Qdrant server |
|
|
Missing model files, then the Qdrant server |
|
|
Missing model files |
Hub offline mode |
|
The Qdrant server |
|
Every one of them is safe to run again. A model repository already in the cache and a server install that still passes its check cost no download.
install --dry-run and server qdrant install --dry-run preview what would be
fetched. They make no request and write nothing.
What server start does firstLink to Which commands provision, What server start does first
Starting the service is the consent to fetch what it needs. On a host installation
server start works in this order and stops at the first step that fails, before any
service process exists:
It looks for a service that is already running, and reports that one if it finds it.
It judges the environment that would run the service. A client installation is refused, and so is a host whose PyTorch build cannot use a GPU. Nothing has been fetched or written by this point.
It checks that the port is free and that no other service owns this machine.
It fetches any model files that are missing, and checks the ones that are there.
It checks the Qdrant server that would run, and downloads the pinned release if none is installed.
Then it starts the service. The service itself downloads nothing: it checks the Qdrant executable again before every launch, and loads models from the cache only.
A failed step exits 1. With --json the one document carries an error code from the
failure table and the step’s detail.
Exit codes and interruptsLink to Which commands provision, Exit codes and interrupts
install exits 1 with status failed when a provisioning step failed; the report
still lists every step. server warmup exits 1 when a model could not be fetched.
server qdrant install exits 1 when the install failed. All three exit 0 when
there was nothing to do. server start, server warmup and server qdrant install
take --json and then write exactly one document, on every way they can end.
Pressing Ctrl+C during a fetch ends the command with one outcome, like any other
failure, and exit 1. server start reports start_interrupted and says that no
service was started. server warmup and server qdrant install report interrupted.
Files that finished downloading are kept, nothing half-installed is left behind, and
the outcome names the command that continues. One limit applies: while the Qdrant
download is waiting on a source that has gone silent, the interrupt is acted on when
that wait ends, which takes at most 30 seconds.
A client installation provisions nothingLink to A client installation provisions nothing
A client installation sends requests to a service that a host installation runs. It loads no model and runs no server, so it has nothing to fetch, and no command makes it fetch anything:
installreports each provisioning step as skipped.server warmupandserver qdrant installreport that they are not needed and exit0.--archivechanges nothing: the file is not read.server startis refused before it looks at a model or an executable.
The installation guide explains the two roles.
A host that cannot run the service yetLink to A client installation provisions nothing, A host that cannot run the service yet
A host installation whose PyTorch build cannot use a GPU cannot run the service, so
nothing is fetched for it either. This is the ordinary state of a project straight
after vaultspec-rag install has written the PyTorch configuration and before
uv sync has applied it.
install,server warmupandserver qdrant installreport the model files and the Qdrant server as skipped, say why, and say thatserver startfetches them once the environment is ready. They exit0.server startis refused with the reason and the repair.server doctorreports what is missing truthfully, and does not send you to a command that would skip.
install --local-only is still recorded for such a host, so the server start that
follows the sync does not download a server you declined.
The Qdrant serverLink to The Qdrant server
What is installedLink to The Qdrant server, What is installed
One release of the Qdrant server is pinned in each vaultspec-rag release.
vaultspec-rag server qdrant status prints it as Managed version. The release asset
depends on the platform:
Platform |
Release asset |
|---|---|
Windows x64 |
|
macOS, Apple silicon |
|
macOS, Intel |
|
Linux x64 |
|
Linux arm64 |
|
Both Linux builds are static musl builds and do not depend on the host’s C library. A platform that is not in the table has no managed server: the install fails with a message that says so, and you supply your own executable.
The archive is fetched from {base}/v{version}/{asset}, where the base is
https://github.com/qdrant/qdrant/releases/download unless you
set a mirror.
How it is checkedLink to The Qdrant server, How it is checked
Two SHA256 digests for each asset are compiled into vaultspec-rag: one for the archive and one for the executable inside it. No setting, file or mirror can change them. A mirror changes where the bytes come from and never which bytes are accepted.
The archive is hashed and compared before anything is extracted from it.
Only the one executable is taken from the archive. Paths stored in the archive are ignored.
The extracted executable is hashed and compared before it replaces an install.
The installed executable is hashed and compared again before every launch: a start, a restart, and a retry.
The directory the executable is installed in is listed before every launch and again once the server is running. It may hold the executable, its manifest, and the working files of an install in progress. Anything else stops the launch.
A mismatch at any step is a failure. Nothing is installed from an archive that failed, and a previous install is left as it was.
The last step exists because a file beside an executable is not inert. A library the server asks for by name can be loaded from the server’s own directory, so a file with the right name there would run inside a server whose own digest is correct. Nothing can stop a file being written into a directory, so a file put there and removed again between the two listings is not seen.
The manifest written beside the executable records what was installed. It is never what a check trusts: the expected digest always comes from vaultspec-rag itself.
server qdrant status and server doctor run the same check a launch runs, so they
do not report an install as usable when a start would refuse it. In the readiness
report a check that cannot finish in 10 seconds is reported as
qdrant_binary_unchecked, and the install is not reported ready.
Between the check and the launch. On Windows the file is held so that nothing can change it between the check and the launch. On Linux the server is started from the file that was hashed, whatever its path names by then; a rewrite of that same file in place in that instant is detected by a second hash after the launch, not prevented. On macOS the path is confirmed and the file hashed again after the launch; a replacement that is put back before that second look is not seen.
One consequence on Linux: the managed server is started from the open file, so process
listings show it under a number instead of the name qdrant. Find it with
vaultspec-rag server qdrant status, which prints its process ID, or with
pgrep -f bin/qdrant.
Where it is keptLink to The Qdrant server, Where it is kept
The executable and its manifest are in bin/qdrant/<version>/ under the service
directory, which is ~/.vaultspec-rag unless --status-dir or
VAULTSPEC_RAG_STATUS_DIR moves it. A lock file, bin/qdrant/provision.lock, makes
sure only one command installs at a time. A second command run meanwhile says which
process it is waiting for, then finds the install present and downloads nothing.
vaultspec-rag server qdrant clean --yes deletes the installed servers. It deletes
executables only. Index data is kept elsewhere and is not touched. The next
server start downloads the server again.
Replace an install that fails its checkLink to The Qdrant server, Replace an install that fails its check
An install whose executable no longer matches is refused at start with
qdrant_binary_unverified. server qdrant install does not overwrite it unasked,
because something changed that file and you may want to look first. Replace it with:
Command
vaultspec-rag server qdrant install --upgrade
--upgrade leaves an install that passes its check alone. On Windows, stop the service
first: a running server holds its executable, and the install says so.
An install with anything else in its directory is refused at start with
qdrant_install_foreign_files, and the message names what is there. No install
removes it, with or without --upgrade. Delete everything from that directory except
the executable and manifest.json, then start again.
Use your own executableLink to The Qdrant server, Use your own executable
To run a Qdrant server you built or obtained yourself, name it with two environment variables, set together:
VAULTSPEC_RAG_QDRANT_BINARYis the absolute path of the executable. It must be a regular file, not a symbolic link.VAULTSPEC_RAG_QDRANT_BINARY_SHA256is the SHA256 of that file. Print it withGet-FileHash -Algorithm SHA256 <path>on Windows, orsha256sum <path>orshasum -a 256 <path>elsewhere.
The file is hashed and compared with your digest before every launch, exactly as the
managed install is compared with the compiled-in one. Your executable is never copied,
and while both variables are set install and server start download no server: a
file that fails the check stops the command and is never replaced by the pinned
release. server start announces the server as operator-supplied, and
server qdrant status shows its source as operator-supplied (env).
The directory your executable is in is yours and is not checked. Keep it writable only by accounts you trust: a library placed beside the executable can be loaded into it.
Both variables are read from the process environment only. See supplying your own server binary.
A qdrant executable on PATH or in the working directory is never used.
The model filesLink to The model files
What is fetchedLink to The model files, What is fetched
Search uses up to three public models. Each is a repository on the Hugging Face Hub, and its files are downloaded into the Hugging Face cache.
Model |
Default repository |
Fetched |
|---|---|---|
Dense |
|
Always |
Sparse |
|
Unless |
Reranker |
|
Always, whether or not reranking is on |
No account or token is needed. Model selection lists the settings that name other repositories.
How they are checkedLink to The model files, How they are checked
Each default model is pinned to one commit of its repository, and vaultspec-rag carries the SHA256 of every file of that commit. No setting, file or mirror can change those digests. A mirror changes where the bytes come from and never which bytes are accepted.
A fetch ends with the check. Every file of the snapshot is hashed and compared, and a repository is reported as downloaded only when all of them match.
A cached file that fails is downloaded again once. A repository that still fails is a failure,
models_unverified, that names the file.A file in the snapshot that the pinned commit does not have is a failure too. It is never deleted for you: the message gives its path.
Every load begins with the same check. A model is never built from a snapshot that did not pass.
Weights are loaded from the safetensors format only. A model that ships its weights
only as a pickle file is refused, because a pickle file is code from whatever endpoint
served it.
The sparse model builds itself from a source file in its own repository. vaultspec-rag does not ask the model library to import it. It reads that file once, compares it with its committed digest, and runs the bytes it compared, so what runs is what was checked.
A model you name yourself is unpinned. Setting a model to another repository, or a
default model to another commit, leaves vaultspec-rag with no digests for it. It is
checked for its files and for safetensors weights, loaded without running any code
from its repository, and named as unpinned by install, server warmup,
server doctor and the readiness report. The dense and reranker commits are settings;
the sparse model has none, because a commit you could change would select code that
has no committed digest. See model selection.
Between the check and the load. On Windows every file of the snapshot is held open
from before it is hashed until the model has loaded, so nothing can change one in
between, and a file added to the directory meanwhile discards the load. Elsewhere a
process that can write the model cache could replace a file between the hash and the
library’s own read of it. Because weights are safetensors and the one executed source
file is the bytes that were hashed, that window can yield wrong numbers, not run code.
vaultspec-rag server doctor runs the full check without starting the service. It reads
several gigabytes and takes a few seconds.
How they are fetchedLink to The model files, How they are fetched
Each repository that is missing is downloaded by a separate download process, one repository at a time. The command you ran makes no request to the hub itself; it starts that process, watches what it reports, and can always stop it.
The download process asks the hub how large the repository is.
The command compares that size with the free space in the cache, less the files of that repository the cache already holds. A download that cannot fit is refused before its first byte, with the shortfall.
The files are downloaded. A file that finished is kept. A file that was interrupted is removed and fetched whole next time; nothing is resumed part-way.
Only one command fetches into a cache at a time. A second one says which process it is waiting for, then looks at the cache again and downloads only what is still missing.
Loading never downloadsLink to The model files, Loading never downloads
The service does not download a model when it loads one. It reads the
cache, checks the files, and stops with one message when a model is absent or fails
its check. The message names the model and the command that repairs it:
vaultspec-rag server warmup fetches what is missing, and vaultspec-rag server doctor
confirms the cache passes. The fetch that server start runs before it starts the
service is the only download on that path.
Where they are keptLink to The model files, Where they are kept
In the Hugging Face cache: the hub directory under HF_HOME, which defaults to
~/.cache/huggingface. Set HF_HOME before the first download to keep the models
somewhere else. See Hugging Face cache.
Change where downloads come fromLink to Change where downloads come from
Download |
Setting |
Reference |
|---|---|---|
Qdrant server |
|
|
Model files |
|
Both sources must be https. The Qdrant download follows a redirect only to https
and only to a host on the allowed list. A mirror that redirects to its own storage
host needs that host listed.
A mirror with a private certificate authority. Certificate verification is never
skipped. For the Qdrant download, add the authority to the system trust store or name
its bundle in SSL_CERT_FILE. For the model hub, set SSL_CERT_FILE to a PEM bundle
that includes the authority, or SSL_CERT_DIR to a directory of such certificates.
Run without a networkLink to Run without a network
The Qdrant server from a local fileLink to Run without a network, The Qdrant server from a local file
On a host that cannot reach a release source, install the server from a copy of the official release archive:
On a machine with network access, download the asset for the target platform from the table under what is installed.
server qdrant install --dry-runon the target host prints the exact asset name, version and digest.Copy the file to the host. Its name does not matter; it is identified by its digest.
Install from it:
Command
vaultspec-rag server qdrant install --archive <file>
The file goes through the same archive and executable checks as a download. No request
is made, and the file is read where it lies: it is not copied, moved or removed. Keep
it outside ~/.vaultspec-rag/bin/qdrant. The report names the source it used. Add
--upgrade to replace an install that fails its check.
The models from a copied cacheLink to Run without a network, The models from a copied cache
On a machine with network access, run
vaultspec-rag server warmupwith the same model settings the offline host will use.Copy the
hubdirectory of that machine’s Hugging Face cache into the same place underHF_HOMEon the offline host.Set
HF_HUB_OFFLINE=1in the environment that runs vaultspec-rag there.
With the offline switch set, nothing is requested from the hub. A model that is
missing from the cache is then a failure, models_offline, that names the switch; it
is never a download attempt. TRANSFORMERS_OFFLINE=1 has the same effect.
No server at allLink to Run without a network, No server at all
--local-only on install and server start selects the on-disk store, which needs
no Qdrant server. See storage backends.
How long a fetch can takeLink to How long a fetch can take
No fetch waits without limit. These are the bounds, so you can tell a slow link from a command that is stuck.
Fetch |
Limit |
|---|---|
Qdrant server |
15 minutes for the whole download, every retry included |
Qdrant server |
30 seconds with nothing arriving ends one attempt; a transient failure is tried 3 times in all |
Qdrant server |
A response that declares or sends more than 256 MiB is refused |
Qdrant server |
A second command waits up to 16 minutes for the first to finish installing |
Model files |
10 seconds with nothing arriving ends one attempt at a file; 6 attempts in a row with no data fail the file, about a minute |
Model files |
10 seconds for the hub to answer how large a repository is |
Model files |
Fewer than 1 MiB in any 10 minutes after the first byte stops the download |
Model files |
A download that receives nothing at all is stopped after 11 minutes |
Model files |
A command waits up to 1 hour in all for another one that is fetching into the same cache |
Model files |
4 hours for the whole fetch, every repository and every wait included |
The first two model limits are the Hugging Face Hub client’s own. Raise them with
HF_HUB_DOWNLOAD_TIMEOUT and HF_HUB_ETAG_TIMEOUT on a link that pauses for longer.
The last is a setting, VAULTSPEC_RAG_MODEL_FETCH_DEADLINE_SECONDS; raise it for a
link that needs longer. A fetch it stops keeps every file that finished, so running
the command again continues. The others are vaultspec-rag’s and are not settings. They exist because a hub that
keeps sending a few bytes at a time never trips a read timeout. One mebibyte in ten
minutes is about 14 kbit/s, so no link anyone would wait on is caught by it. Stopping
a download adds at most about 11 seconds: the progress is looked at every half second,
and the download process is given 10 seconds to be confirmed gone.
A refusal that asking again will not change is not retried: a missing file, a certificate that does not verify, a redirect off the allowed hosts, a digest mismatch.
Both fetches check free space before the first byte and fail with the shortfall when the volume cannot hold the download.
When provisioning failsLink to When provisioning fails
Each failure ends the command with one outcome that names the cause and what to do.
server start --json reports the code as error; install --json reports it as the
step’s code.
Model filesLink to When provisioning fails, Model files
What happened |
Code |
What to do |
|---|---|---|
A model is missing and the hub is in offline mode |
|
Unset |
The hub cannot be reached, or stopped answering |
|
Check network access to the hub; raise |
The hub has no such repository or revision |
|
Correct the model setting that names it |
The cache volume cannot hold the download |
|
Free space on that volume, or set |
The hub’s certificate is not trusted |
|
Set |
Too little arrived for too long and the download was stopped |
|
Copy a complete cache and set |
The download process ended without a result |
|
Run the command again; if it ends the same way, report the detail it printed |
Another process was still fetching after an hour |
|
Wait for that process to finish, then run the command again |
A snapshot fails its check: a file differs, is extra, or the weights are pickle only |
|
The detail names the file. Point |
The time allowed for the whole fetch ran out |
|
Run the command again to continue, or raise |
A file or directory of the model cache could not be used |
|
Make the cache writable by the account running the command, or set |
|
|
Reinstall vaultspec-rag |
When several repositories fail for different reasons the code is
models_fetch_failed and the detail says which repository had which. Every repository
is attempted before the command reports, so one failure does not hide the next.
Qdrant serverLink to When provisioning fails, Qdrant server
What happened |
Code |
What to do |
|---|---|---|
No server is installed and downloading is switched off |
|
Run |
The install failed: see the causes below |
|
The detail names the cause and its remedy |
The installed executable is not the pinned release |
|
|
The executable cannot be read |
|
Close whatever holds it, or correct its permissions, then try again |
Something that is not a file is at the install’s name |
|
Remove it and run |
The install’s directory holds something no install wrote |
|
Delete everything there except the executable and |
Your own executable’s settings do not name a usable file |
|
Correct |
Your own executable does not match its declared digest |
|
Correct the path or the digest, or unset both |
Causes of qdrant_provision_failed, each named in the detail with what to do:
The source cannot be reached, is failing or rate-limiting, or is too slow for the 15 minute limit. The detail also names the route that needs no network:
server qdrant install --archive <file>.The source’s certificate is not trusted, the file is not there, or the source refused the request.
A redirect led to a host that is not allowed, or the redirects never end at the file.
The response is larger than the 256 MiB cap, so the source is not serving the release asset.
The archive or the executable inside it does not match its digest.
The volume cannot hold the download, or filled while it was being written.
The installed executable is in use, on Windows, because the service is running.
A local archive given with
--archiveis not the pinned one.The platform has no managed server.
Another command was still installing after 16 minutes.
The model files are fetched before the Qdrant server, so a start that cannot get its models is not also sent to download a server it cannot use yet.
A separate failure belongs to the environment and not to any fetch: server start
reports service_env_no_gpu for a client installation and for a host whose PyTorch
build cannot use a GPU, with the reason and the repair.
After a failureLink to When provisioning fails, After a failure
There is nothing to clean up by hand, and running the same command again is always the next step once the cause is fixed.
A failed Qdrant install removes its working files. A previous install is untouched. Working files left by a run that was killed are removed by the next install.
A failed model fetch keeps every file that finished and removes the one that did not. The next run downloads only what is missing.
Where to go nextLink to Where to go next
Installation guide - choose a host or client installation and set it up.
Configuration reference - every setting named on this page.
Storage backends - the managed server and the local-only store.
CLI reference - every option of the commands above.