Replies: 1 comment
|
Following up now that 1.3.0 has shipped external executors, since that covers most of this and it is worth closing the loop. The executor seam matches the constraints in the original post: off by default, explicit The piece that has not landed is the one I flagged as carrying most of the value and none of the privacy surface: a different local model. Is that seam still under consideration, or is the intended answer now to use an executor? Either is a reasonable outcome. A line in the docs would settle it, because the current wording reads as a deliberate pin rather than a decision that was made. Two notes since the original post: The GPU argument for the local seam is stronger than when I wrote this. #749 fixed WSL2 accelerator detection, and the release factory at HEAD still targets Built-in provisioning is currently unavailable on Linux x64 regardless of model choice, because no semantic asset keys are published in the release metadata. Filed as #910 rather than here, since it is a release-plumbing bug and not part of this idea. |
Uh oh!
There was an error while loading. Please reload this page.
The ask: let the semantic embedding backend be selectable, off by default, with the current local e5-small staying the default and the recommendation.
Not cloud history sync, and not a change to any default. The offline contract reads as deliberate:
cloud.modereturns a "no longer supported" error,offline_default.rspins staleCTX_CLOUD_*variables to hostile values and asserts nothing dials out, and the privacy contract indocs/product-contract.mdis unambiguous. That should stay as it is. What I am after is the seam, not any particular backend behind it.Why
It is a permanent cost, not a setup step.
run_pending_core_semantic_catch_upruns on every daemon iteration, bound to the current Core generation, with its own retry and backoff. Every new lexical generation means more embedding, so for daily agent use the backend is a permanent property of the install rather than a one-time choice.The default model is a quality ceiling.
intfloat/multilingual-e5-smallis 33M parameters, 384 dimensions, 512-token window. Fast, multilingual, runs anywhere. But retrieval quality is the point of enabling semantic search, and there is no way to trade resources for it, even locally.It occupies the machine indefinitely. Embedding runs continuously at a throttled 25% duty cycle because it has to share. Someone who would rather spend an API call than a core, or point at a GPU box they already run, cannot say so.
What I would most like
A different local model, no network path.
voyage-4-nanois Apache-2.0 on Hugging Face, 340M parameters, 32K context, with local development named as an intended use on the model card. Roughly ten times e5-small while staying on the machine, so better retrieval no longer has to mean sending history anywhere.Then a self-hosted endpoint.
huggingface/text-embeddings-inferenceis Apache-2.0, and Ollama is an existing precedent. Nothing leaves the machine. For me it would also be the practical way to use a GPU ctx cannot currently detect, without ctx having to solve GPU detection first.Hosted APIs last, with the caveats below.
Prior art
ChunkHound indexes source code for coding agents, describes itself as "Local-first" in its README, and makes the provider a config choice:
The local path stays first-class and the docs name the zero-egress option explicitly. ChunkHound is also indexing source code, which is at least as sensitive as agent transcripts.
This may be less invasive than it looks
The sidecar already treats the model as a variable:
semantic_model_contract_descriptor()folds model id, dimensions, pooling, normalization, prefixes, and max sequence length into one descriptorsource_contract_fingerprint_with_authority, withcomplete_model_descriptor_participates_in_semantic_contract_identityasserting itThe flat-F32 store looks dimension-generic too:
SegmentDescriptorcarriesdimensions: u32,vector_stride(contract.dimensions)computes stride at runtime,validate_vectorchecks width against the contract, and there is aMAX_DIMENSIONSbound rather than a hardcoded 384.SEMANTIC_DIMENSIONSlooks like the pin rather than the format.If that reading is right, a second backend is mostly a trait behind
TextEmbeddingplus config plumbing. Please correct me if I have misread it.On hosted APIs
Voyage's ToS (2026-05-27) grants a perpetual, irrevocable license to train on submitted content unless you opt out, and:
The opt-out is prospective only, needs an org admin with a payment method on file, and is one-way in the dashboard. So the free tier and the privacy-respecting configuration are mutually exclusive.
Paying from the first token, my corpus is roughly 585M embeddable tokens by sampling, which at
voyage-4-literates is about $12 initially and near a dollar a month ongoing.OpenAI does not train on API data by default; Jina states it never does, though that sits on a sales FAQ rather than binding terms. One practical detail: Voyage's 4 series emits 256, 512, 1024, or 2048 dimensions, so no hosted provider will match 384.
If there is appetite
search.semanticis nowconfig.tomlctx statusnames the active backend, worth doing regardless since a CPU fallback is already invisibleThe counter-argument
"Your history never leaves your machine" is a clean promise, and "unless you turn on a thing" is weaker and needs explaining every time. That is a real cost and it is your call. The narrow version, a different local model with no network path, avoids it entirely, and I would be glad if that were the only part that landed.
All reactions