ProximaDB as the Code Context Graph (CCG) Backend¶
Status: Embedded backend and local Tier-A/Tier-B boundary implemented behind a
per-repo flag; SQLite remains default (tracked as TD-11, TD-12, TD-13 in
../tech-stack.md). Proxima columnar/service mode is WIP.
Date: 2026-08-12
Implementation status (2026-08-05)¶
The embedded ProximaDB backend is implemented and parity-verified at the adapter
level; SQLite stays the default and nothing flips automatically.
victor-codegraphexclusively owns stable symbol identity and emits the full
correlatedgraph/{repo}/node/{symbol_oid}recordoid; storage persists it
unchanged.victor/storage/proxima_runtime.pyprovides optional-dependency
detection, repository/graph naming, theProximaEmbeddingMode(memory/cold)
encoding of the RustEmbeddingMode, and embedded bootstrap.victor/storage/vector_stores/proximadb_provider.py— now a real
EmbeddingProvideroverproximadb_sdk's embedded API with an in-process
embedding model (in-RAM fp32 =EmbeddingMode::Memory). Documents are keyed by
theiroid, so the always-emptyembedding_refbridge is unnecessary.victor/storage/graph/proxima_store.py—ProximaGraphStoreimplements
GraphStoreProtocoloverproximadb_sdk.graph.ProximaDBGraph
(upsert_nodes/edges,get_neighbors,search_symbols,find_nodes,
multi_hop_traverse_parallel, …). Each Tier-A symbol is durably written as one
ProximaRecord envelope containing its graph properties, vector, and staleness
markers under oneoid;embedding_refis dropped. The indexing pipeline no
longer performs a vector write followed by a graph-metadata write.victor/storage/graph/cpg_fragments.py— the Tier-B boundary is live. The
Proxima adapter routesstatementnodes and every CFG/CDG/DDG edge away from
ORION into a durable SQLite fragment index keyed by file, scope, statement
type, and edge endpoint. Default scans and unfiltered traversals remain Tier A;
explicit CCG filters and statement/scope lookups drill into Tier B. File/repo
deletion, restart persistence, bounded edge iteration, and per-tier stats are
contract-tested. This local indexed representation establishes the replaceable
boundary; Proxima PAX/columnar fragments remain the service-mode follow-up.- Selection:
create_graph_store("proxima", …), or per-repo via a
<project>/.victor/graph_backendmarker honored bycreate_graph_store("auto", …)
(defaultsqlite).impact_analysisand the hybrid graph query tool resolve
the backend through"auto". - Parity:
tests/unit/storage/graph/test_proxima_store_parity.pydrives the real
ProximaDBGraphagainst an in-memory fake client and asserts impact_analysis
(forward/backward) and hybrid seed→expand match SQLite — runs without the
server binary.tests/integration/storage/graph/test_proxima_embedded_parity.py
repeats it against a real embedded instance, skipping when the binary is absent. - WIP / gated: the multi-tenant service path (
server_url=,
EmbeddingMode::Cold/SQ8) is marked WIP — gated on ProximaDB TD-127 (secondary
indexes) + TD-130/131 (graph bulk-load + REST v2 hybrid). As of 2026-06-22 the
engine-side gate is now satisfied on ProximaDBdevelop: TD-127/128 merged
(PR #215,40c08076), TD-130 graph bulk-load merged (PR #220,967f15db), and
REST v2 hybrid is live (/api/v2/hybrid/search+/strategiesin the SDK). The
service-mode WIP guard and the SQLite default stay in place until a live
bench (not just parity) passes. Arrow Flight bulk-load and ORION native
centrality (steps below) remain pending on that measurement. - Live parity is verified (2026-08-05). An earlier revision of this document
claimed live parity could not be measured because noproximadb-serverbinary
existed here. That was stale: a release binary was available, and both live
suites now pass against a real embedded instance —impact_analysis
(forward/backward) and hybrid seed→expand match SQLite, and a vector resolves to
its graph node byoid. Getting there required two fixes, because a live run
exercised seams the fake-client unit tests cannot reach: - Embedded transport. ProximaDB #264 made embedded mode portless (UDS
sockets, no TCP port), sorest_urldegrades to the host-header sentinel
http://localhost.ProximaRepoConnectionbuilt a plain TCP client from it,
so every ORION call failed withECONNREFUSEDwhile ProximaRecord writes
still succeeded through the SDK's own UDS-aware client. That split transport
meant the authoritative record committed and its projection never landed —
the exact skew the atomic-record boundary exists to prevent. Fixed by
plumbing the socket through (needsProximaDBClient(uds_path=…), proximaDB
PR #1447). - Stale parity fixture. The embedded parity test injected a pre-built
graph/client, bypassing the connection that owns the record collection, so it
failed before any assertion. It now drives the production bootstrap. - Local source dependency gate: Victor's development virtualenv resolves the
pure-Pythonproximadb0.2.2 SDK directly from../proximaDB/clients/python.
Do not pin a newer PyPI version until ProximaDB publishes it. The native
proximadb_embeddedwheel could not be rebuilt because ProximaDB #1021 moved
the PyO3 bindings intocrates/binding/proximadb-embeddedwithout updating the
Python build config: maturin still targeted the root manifest (whosepython
feature is now an empty stub and whose lib isrlib-only), requested the
removedpylibfeature, and no crate declared acdylib. Fixed in proximaDB
PR #1448. Note this wheel is a separate artifact from the pure-Python SDK
Victor actually uses at runtime — Victor's embedded mode spawns a
proximadb-serversubprocess and never imports the native module, so the wheel
gates the PyPI release, not Victor's local verification.
Atomic record boundary and failure model¶
“Atomic” has one precise meaning here: the authoritative write for one symbol is
one /api/v2/collections/{collection}/records/batch ProximaRecord containing
id, vector, and the complete props map. ProximaDB commits that envelope on
its canonical record/WAL path. ORION does not currently accept this rich public
record shape as a graph mutation, so its node is a rebuildable traversal
projection applied only after the record commit succeeds.
- A new or changed symbol first writes a pending record with full graph props,
has_embedding=false, and a 384-dimensional zero placeholder. The placeholder
is necessary because the current v2 public record contract requires a vector;
semantic search always addshas_embedding=true, so pending records cannot
enter retrieval. - Embedding completion replaces that same record once with the real vector,
has_embedding=true, andcontent_version. There is no subsequent metadata
mutation. - A record failure prevents the ORION projection from being created or changed.
A projection failure leaves a correct committed record and is retriable.
Logical edge enumeration reads the record authority exclusively: it never
unions ORION into the result and fails closed if the record scan is unreadable.
Otherwise a lagging projection could resurrect a canonically deleted edge.
Cached enumerations are revalidated against ProximaDB's server-owned
content_revision_token. The token combines the monotonic content revision
with a server-incarnation epoch, so a restart cannot validate stale data via
revision-number reuse. Process-local mutation generations remain a fast path,
not the cross-process freshness authority. - Metadata changes preserve the committed vector and replace the complete record.
If Victor cannot prove it has the vector needed for replacement, it fails
closed instead of overwriting it with a placeholder. - File/repository deletion removes the unified record before deleting its ORION
projection. The legacyclear_embeddingsargument cannot split the modalities
and is retained only as a compatibility input.
This boundary eliminates graph-only and vector-only authoritative states. It
does not claim a distributed transaction between the record WAL and ORION; that
would require ProximaDB to expose a graph projection sourced transactionally from
the rich ProximaRecord itself.
Why¶
Victor's durable code memory — the thing that lets the agent answer "who calls X / blast radius /
what is semantically near this" without re-reading files — lives today in two embedded stores:
- SQLite (
.victor/project.db, ~2.4 GB):graph_node/graph_edge/graph_module_metric/
graph_node_fts— a statement-level Code Property Graph (CFG/CDG/DDG + CALLS/IMPORTS/INHERITS). - LanceDB (
.victor/embeddings/embeddings.lance): 384-dBAAI/bge-small-en-v1.5vectors at
symbol granularity (measured 77,902 vectors; function 65,622 + class 12,280), with the symbol
snippet co-stored.
These are hand-joined: graph_node.embedding_ref is meant to bridge them but is unpopulated, and a
watch daemon keeps both in sync on file change. The abstraction to swap them already exists — a
GraphStoreProtocol (sqlite/memory/duckdb-stub) and an EmbeddingProvider protocol with a
proximadb_provider.py referencing ProximaDB's SST (vector) + ORION (graph) engines.
The opportunity: collapse the two stores into one authoritative ProximaDB record collection where a
code symbol is one durable entity — complete graph properties plus an HNSW vector — addressed by a
single oid, with ORION as its traversal projection. This removes authoritative dual-write skew, makes
the embedding and staleness update one record replacement, and gives the agent native graph algorithms
(impact analysis, centrality, hybrid seed→expand) instead of hand-rolled Python.
Measured shape (one real repo, 3,659 files)¶
| value | |
|---|---|
| graph nodes | 1.26M (94% statement) |
| symbol nodes (module/class/function/method) | 79,744 |
| edges | 2.97M |
| Tier-A cross-fn edges (CALLS/IMPORTS/INHERITS/CONTAINS/…) | 96,538 (CALLS 61% cross-file) |
| Tier-B intra-fn edges (DDG/CFG/CDG) | 2,875,449 (DDG measured 100% intra-file) |
| embeddings (Lance) | 77,902 rows / 68,612 distinct @ 384-d ≈ 100 MB f32 / 25 MB SQ8 + 5.5 MB snippet |
Current correlation reality (measured): the two stores are disjoint — SQLite graph_node has
signature/docstring/embedding_ref 0% populated (pure topology + file/line); code snippet +
vector are LanceDB-only, keyed symbol:{file}:{name}; the graph hex node_id and the Lance id
namespace do not intersect (correlation is implicit by (file, symbol_name), the embedding_ref
bridge is empty). Only ~5.4% of nodes (≈86% of symbols) carry a vector. The ProximaDB one-oid
record makes embedding optional-per-node (NF² props), removing the always-empty bridge column.
Design — three tiers, one oid per symbol¶
A code symbol becomes one authoritative ProximaDB record with the victor-codegraph-emitted
oid = graph/{repo}/node/{symbol_oid}, carrying its
properties, its embedding, and a ref to its intra-procedural detail. The same oid
keys its vector index entry and derived ORION node, so vector hit → graph node is
identity, not a join.
- Tier A — semantic graph (HOT, in-RAM): ~80K symbol nodes + ~96K cross-fn edges + per-node 384-d
vector → ORION graph + co-indexed vector. Drivesimpact_analysis(forward/backward k-hop), call paths,
and hybrid semantic-seed → expand. ~120 MB f32 / ~35 MB SQ8 — fits memory. - Tier B — intra-procedural CPG (COLD fragments): statements + DDG/CFG/CDG
are routed to the durable fragment store and fetched on dataflow drill-down.
The local implementation is SQLite indexed by file/scope; the service target
is PAX/columnar fragments per function. Tier B is never included in default
global scans, so the real size driver stays off the hot graph. - Tier C — relational facts:
code_file/code_import/code_module_metric/code_file_mtime
served from the same records; point-reads/upserts on the re-index hot path.
What changes in Victor¶
victor/storage/vector_stores/proximadb_provider.py→ make real (currently emerging); use ProximaDB's
EmbeddingMode::Memoryfor the embedded/local case so semantic BFS scores neighbors inline.victor/storage/graph/→ add aProximaGraphStoreimplementingGraphStoreProtocol
(upsert_nodes/edges,get_neighbors,search_symbols,multi_hop_traverse_parallel) over the
ProximaDB graph/hybrid API. The watch daemon's incremental path becomes idempotentinsert_proxima_records
upserts; initial load uses Arrow Flight bulk.victor/core/graph_rag/retrieval.py(MultiHopRetriever) andvictor/framework/search/hybrid.py
(RRF) → can delegate to ProximaDB's nativeGraphHybridQuery(VectorFirst fusion) instead of hand-rolled
seed→expand + fuse.graph_module_metric(pagerank/betweenness/coupling/instability/hotspot) → can be computed by ORION's
native centrality/community algorithms rather than in Python.
Embedded vs service¶
- Embedded (local single-repo): one
proximadb-serversubprocess owns the
repo's local writable roots; clients share that owner over UDS. ProximaDB
rejects a second server for any overlapping local authority. SDK startup
accepts/healthonly when its process identity matches the child it spawned,
so a pre-existing server cannot be mistaken for a successful launch.
EmbeddingMode::Memoryremains the Tier-A vector mode. The separate native
PyO3 package is not the runtime used by this adapter. - Service (multi-tenant, via anvaiops): collection
{tenant}_{repo}_codegraph,graph_id=repo,
branch_id=git branch,EmbeddingMode::Cold(SQ8) to bound RAM. See the ProximaDB design spec
proximaDB/docs/12-design/CODE_GRAPH_CORRELATED_SUBSTRATE_2026_06_22.adocand the anvaiops ADR.
ProximaDB-side enabling work (not Victor's)¶
The ProximaDB engine asks are filed there as TD-127..TD-134 (OLTP secondary indexes by name/file,
IN-list pushdown, ON CONFLICT DO UPDATE, graph edge bulk-load, graph REST v2 + co-planned hybrid,
Tier-B PAX fragment contract, optional transactional multi-modal write, code-embedding KEU meter).
Migration / verification (when picked up)¶
- ✅ Stand up
ProximaGraphStorebehind the existingGraphStoreProtocol; keep SQLite as the default. - ✅ Parity test on a fixture repo:
impact_analysis(forward/backward)and hybrid seed→expand match
the SQLite store on known symbols — adapter-level always-on, and verified
live against a real embedded instance on 2026-08-05 (previously unmeasured). - ✅ Enforce the local Tier-A/Tier-B storage boundary with durable, on-demand
CPG fragments and restart/routing/deletion contract tests. - ✅ Replace separate Tier-A vector and graph-metadata mutations with one
authoritative ProximaRecord replacement; keep ORION explicitly rebuildable. - ⏳ Replace the local Tier-B representation with Proxima PAX/columnar
fragments and verify per-function drill-down parity in service mode. - 🟡 Bench footprint + k-hop latency — measured 2026-08-06 via
scripts/benchmark_graph_backends.py; see "Measured backend comparison"
below. Arrow Flight bulk-load is not benched and cannot be: Victor has no
Arrow Flight code path (ingest is RESTinsert_records+batch_create_nodes),
so that clause is a feature to build, not a measurement to take. - ⏳ Flip the default provider per-repo once parity holds (per-repo
.victor/graph_backendflag exists; SQLite stays default).
Measured backend comparison (2026-08-06)¶
Superseded — the SQLite footprints below are inflated by a measurement bug.
The harness summed the project directory before SQLite's write-ahead log was
checkpointed. Victor holds a process-global connection, soproject.db-wal
survivedstore.close()and was counted, then deleted moments later at
interpreter exit. The error is large and one-sided: ProximaDB measured
identically before and after, so every ratio here flatters ProximaDB.
Corrected repo-scale figure: SQLite 196.8 MB as measured vs 125 MB at rest.
The "SQLite costs ~33 KB/node with embeddings" result is the same artifact —
at rest it is ~3.8 KB/node. Fixed in_checkpoint_sqlite_wals; see the
2026-08-19 section for numbers taken with the fix in place.
Same corpus, same GraphIndexingPipeline, both backends holding verified
identical data (the bench suppresses the ratio outright if node counts differ,
because a footprint comparison across differing data is worse than no number).
Tier-A only (enable_ccg=False).
| corpus | embeddings | nodes | SQLite | ProximaDB | ratio | SQLite B/node | Proxima B/node |
|---|---|---|---|---|---|---|---|
| 10 files | off | 251 | 1.4 MB | 309 KB | 4.58× | 5,777 | 1,261 |
| 87 files | off | 1,328 | 3.0 MB | 1.1 MB | 2.80× | 2,361 | 843 |
| 10 files | on | 251 | 8.1 MB | 987 KB | 8.38× | 33,729 | 4,026 |
| 87 files | on | 1,328 | 42.8 MB | 4.3 MB | 9.88× | 33,827 | 3,424 |
Read the ratios above with the caveat below. Every row runs
enable_ccg=False, which is not whatvictor initdoes. In the default
CCG-on configuration the footprint advantage disappears — see
"End-to-end reality check". The vector-efficiency result is real but applies
to the symbol tier, which statement-level CPG dwarfs in practice.
Four findings from the Tier-A slice:
- The footprint win is real and comes from vectors. Graph-only, the ratio
shrinks with scale (4.58× → 2.80×). With embeddings it grows (8.38× →
9.88×), and SQLite's cost stays pinned near 33 KB/node while Proxima's falls.
This is the design's premise — the SQLite+LanceDB pair stores vectors
inefficiently — confirmed by measurement rather than projection. It also means
a graph-only benchmark understates the case, and extrapolating any single
corpus size overstates or understates depending on which mode it ran in. - Ingest reverses under embeddings. Proxima is ~3× faster to index without
them (3.5 s vs 10.4 s) and ~3.5× slower with them (40.3 s vs 11.5 s). The
embedding path replaces one full ProximaRecord per symbol, so vector
completion pays a whole-record write; SQLite updates a column. Worth
attention before recommending Proxima for large first-time indexes. - SQLite's ~10× traversal advantage does not matter at these magnitudes.
k-hop p50 is 0.06 ms vs 0.65 ms — a large ratio on a negligible absolute
number. Graph reads happen in tool calls inside an agent turn whose LLM
round-trip is measured in seconds; even 100 graph queries per turn costs
~65 ms on Proxima, well under 1% of the turn. Retrieval latency is not a
differentiator between these backends and should not be weighted as one. - Footprint is the decisive axis, and worktrees are why. Victor development
routinely runs many linked worktrees at once, each carrying its own
.victor/project.dband embeddings. Per-worktree indexing is where a ~10×
reduction stops being a nice-to-have: the SQLite+Lance pair does not scale
across concurrent worktrees; Proxima does. - Ingest is the one real regression. ~3.5× slower with embeddings, paid on
first index and session start — the moments a developer actually waits.
Extrapolating the 87-file result linearly to Victor's ~1,452 source files
suggests roughly 12 min (Proxima) vs 3.4 min (SQLite). That is an
order-of-magnitude estimate, not a measurement.
End-to-end reality check (victor init, 2026-08-06)¶
The table above measures a slice, not the product. Running the real CLI —
victor init --no-deep --no-interactive --force, CCG on (the default), identical
70-file corpora, backend chosen by the .victor/graph_backend marker:
| SQLite | ProximaDB | |
|---|---|---|
| init wall time | 22.59 s | 10.71 s |
| nodes | 27,265 | 25,937 |
| edges | 75,626 | 74,204 |
| footprint | 52 MB | 54 MB |
Footprint is a wash, not ~10×. The Tier-A bench disables CCG; init enables it,
and statement-level CFG/CDG/DDG then dominates the graph. The Tier-A/Tier-B
boundary routes all of that to a local SQLite fragment store, so the Proxima
configuration is mostly SQLite by volume:
50 MB .victor/proximadb/cpg_fragments.sqlite3 <- Tier-B SQLite
1.1 MB .victor/proximadb/data <- actual ProximaDB storage
ProximaDB holds 1.1 MB of the 54 MB. The vector-efficiency result stays true but
governs only the symbol tier. Until Tier-B moves to Proxima PAX/columnar
fragments (checklist step 5), a footprint argument for this backend is measuring
the wrong thing.
What the default configuration does show is ingest: ProximaDB is 2.1× faster
end to end (10.71 s vs 22.59 s). That, not disk, is the adoption argument
today — which makes the write-amplification gap (ProximaDB issue #1479, no
partial/vector-only record update) the thing worth pushing upstream, since it is
what stops that lead from being larger.
Blocking gap for step 7: the backends do not hold identical graphs through
init — Proxima is short 1,328 nodes and 1,422 edges. Flipping any default before
that has a root cause would ship a silently lossier index.
Two bugs had to be fixed before this comparison could run at all, both silent:
victor init hardcoded the SQLite backend while the read paths honored the
marker (split-brain), and ProximaGraphStore.stats() omitted top-level
nodes/edges, so init raised KeyError('nodes'), printed only
! CCG indexing skipped: 'nodes', and reported success having built no index.
Not yet measured: hybrid seed→expand latency under load, SQ8 cold mode, and
behaviour at the 3,659-file / 2.4 GB scale the original figure came from.
Repo-scale measurement and the ingest blocker (2026-08-19)¶
The 2026-08-06 comparison above measured corpora of 10–87 files. This run
measured the whole repository, and the conclusion changes at scale: the
blocker is not traversal latency, it is ingest and recovery.
Both backends indexed the same corpus through the same GraphIndexingPipeline
and ended at byte-identical graph size — 93,457 nodes / 177,031 edges —
so nothing below is a data-difference artifact. Embeddings off (Tier-A only).
| SQLite | ProximaDB | ratio | |
|---|---|---|---|
| index time, before the fix below | 447 s | 3,269 s | 7.3× |
| index time, current | 395 s | 1,748 s | 4.4× |
| footprint | 196.8 MB | 473.4 MB | 2.4× larger |
The ingest gap was superlinear, and the cause was on the client. The same
comparison on a 33,205-node corpus gave 379 s vs 331 s — only 1.14×. A ratio
that grows with the graph is the signature of per-item work scaling with graph
size, and here it was upsert_edges preloading graph.get_all_edges() to
dedup: the SDK implements that as per-node outgoing-edge scans, ~93k HTTP
requests, re-paid on cache invalidation. Removing it (below) cut ProximaDB
ingest 47% with no server change, and the ratio fell 7.3× → 4.4×.
What remains is a constant-factor ~4.4× on ingest — a far less alarming
shape than the quadratic it replaced, and one nothing here has yet attributed.
Candidates are the HTTP/UDS boundary and per-batch JSON, the canonical record
write accompanying each edge batch, and WAL volume. bulk_insert_edges already
emits validate/wal/mempool/index/csr timings, so the per-phase trace is cheap
to obtain; it should be taken before anyone optimizes.
A correction worth recording: this section first attributed the ingest gap
to ORION rebuilding its CSR per edge insert. That was wrong. batch_create_edges
already routes through bulk_insert_edges, which batches index creation and
defers to a threshold-triggered compaction, so the per-edge rebuild was never on
the ingest path. It is real on the replay path, which is why recovery
behaves as described below; anvai-labs/proximaDB#1673 was rescoped accordingly
and anvai-labs/proximaDB#1678 fixes it.
Traversal latency is not quoted here on purpose. k-hop p50 measured
SQLite 0.33 ms vs Proxima 1.76 ms in one run and SQLite 4.77 ms vs Proxima
1.50 ms in the next — SQLite moved 14× between runs while Proxima barely moved.
The harness samples 75 traversals with seeds re-chosen per run and no cache
control or repeats, so it cannot support a claim in either direction yet.
Recovery inherited the same defect — fixed and measured. Reopening the
data directory replays ~77k batched operations through the same per-edge
rebuild. anvai-labs/proximaDB#1678 defers the CSR rebuild for the duration of
replay and commits it once at the end. Timed on a 467 MB / 93,457-node /
177,030-edge directory, each arm on its own copy:
| binary | load at start | time to serving |
|---|---|---|
| pre-fix | 16.85 | 824.5 s |
| pre-fix | 11.67 | 655.0 s |
| post-fix | 4.13 | 171.1 s |
| post-fix | 4.66 | 176.7 s |
The post-fix figure is stable across runs; the pre-fix figure is load-sensitive,
so the honest speedup is ~3–4× rather than the 4.8× the extreme pair
suggests. Repo-scale startup is now under three minutes.
Correction. Earlier revisions of this section said the graph "never
reachedserving" and was "unusable regardless of query speed". That was
wrong: pre-fix recovery completes at 824.5 s, and the original run was
killed at 813 s — about eleven seconds early. The distinction matters,
because "cannot be reopened" and "takes fourteen minutes to reopen" justify
different decisions. The quadratic behaviour was real and the fix is real;
the catastrophic framing was not.
Recommendation: keep SQLite as the default backend for repo-scale graphs.
(proximaDB#1673/#1678 have since landed and recovery is fixed; the
recommendation now rests on ingest cost and the absence of a storage
advantage, not on recovery.) Traversal latency is not the reason and never was —
a millisecond inside an agent turn whose LLM round-trip is measured in seconds
is noise, as the 2026-08-06 analysis already argued, and the current numbers are
too unstable to quote anyway. Recovery is no longer the blocker either — it is
under three minutes since proximaDB#1678, validated above. What is left is the
ingest cost and the absence of any offsetting advantage. Ingest and recovery are the
gating axes.
What this run fixed on the Victor side¶
Two defects found while measuring, both fixed here:
upsert_edgesno longer preloads every edge id. It called
graph.get_all_edges(), which the SDK implements as per-node outgoing-edge
scans — one HTTP request per node, ~93k round trips, re-paid on any cache
invalidation. It existed to work around ORION's create-only batch aborting at
the first duplicate and silently discarding the rest; proximaDB#1647 replaced
that with per-edge admission ("remainder of the batch applies"), so the scan
is obsolete. A free in-process set still prevents re-sending within one run.- The 30 s startup ceiling is gone.
start_embedded_dbhardcoded it, and
since proximaDB#1668 an explicit caller timeout deliberately overrides the
SDK's progress-aware wait — so Victor's constant became the binding limit.
Startup cost scales with the data directory (replay and rebuild happen before
the listener binds), so a fixed budget is the wrong shape: it passes on toy
repos and fails on real ones. The default now delegates to the phase-watching
wait, which gives up on a stall rather than a clock.
Neither fix can show its full benefit while the server is quadratic; re-measure
ingest once #1673 is fixed.
Method notes (so these numbers can be re-run)¶
- Reproduce with
scripts/benchmark_graph_backends.py --corpus . --backends sqlite,proxima --proxima-binary /path/to/proximadb-server; the report includes
post-close reopen time and recovered node/edge parity. Measure retrieval
quality withscripts/eval_graph_retrieval_value.py. - The SDK under measurement must be a clean checkout — the benchmark refuses
to measure otherwise, because a number that cannot be attributed to a commit
cannot be reproduced. - Timing runs need a quiet machine. Recall does not: it is unaffected by load,
so quality numbers from a busy machine are still sound while latency is not. - Two plausible causes were checked and rejected before landing on the CSR
rebuild: WAL replay re-appending to the WAL (theis_replaying()gate works;
the surrounding debug lines print even when the append is suppressed), and
replay dropping edges on ordering (exactly 1 failed frame out of ~77k).
Total-cost comparison with embeddings (2026-08-19)¶
Every figure in the section above is graph-tier only — those runs had
embeddings off, so the SQLite arm never wrote LanceDB and the ProximaDB arm
never wrote vectors. That excludes exactly the tier the backend was adopted
for, so it cannot support a storage conclusion in either direction.
With embeddings on, both arms measured after the footprint fix
(_checkpoint_sqlite_wals), victor/storage corpus, 1,350 nodes / 2,258 edges:
| SQLite + LanceDB | ProximaDB | ratio | |
|---|---|---|---|
| index time | 15.78 s | 6.33 s | ProximaDB 2.5× faster |
| footprint at rest | 4.9 MB | 7.0 MB | ProximaDB 1.4× larger |
| bytes/node | 3,829 | 5,404 |
The ingest result above does not survive scale — see the repo-scale table
below, which supersedes it. It is kept here because the contrast is the
point: a 1,350-node corpus reported ProximaDB 2.5× faster, and the same
comparison at 93,458 nodes reports it 3.6× slower. Small-corpus graph
benchmarks in this repo have now inverted a conclusion twice, and should not
be used to decide anything.
Repo scale, embeddings on (the number that counts)¶
Full repository, both arms identical at 93,458 nodes / 177,032 edges,
footprint fix in place and verified against at-rest bytes (reported
395.7 / 437.1 MB vs measured 395 / 437 MB):
| SQLite + LanceDB | ProximaDB | ratio | |
|---|---|---|---|
| index time | 522 s | 1,858 s | ProximaDB 3.6× slower |
| footprint at rest | 395.7 MB | 437.1 MB | ProximaDB 1.10× larger |
| bytes/node | 4,439 | 4,904 |
Storage is a wash, and that is the headline. 1.10× is parity, not an
advantage. The 9.88× ProximaDB win previously reported was the WAL artifact;
the 1.4× penalty from the small corpus was corpus size. With vectors included
at real scale, neither backend has a meaningful disk edge — which removes the
premise ProximaDB was adopted on. What remains is a backend that costs ~3.6×
on ingest, matches on disk, and today cannot reopen a repo-scale graph.
One anomaly, recorded rather than smoothed. ProximaDB measured 473 MB
graph-only but 437 MB with embeddings added — smaller with strictly more
data. That points at variable unreclaimed WAL in its data directory
(consistent with the retention work scoped but never built), so its footprint
carries run-to-run noise SQLite's does not. Treat storage as "parity ±10%"
until that is understood.
Timing caveat: machine load ran 13–30 during this run, so 522 s and
1,858 s are not clean absolute numbers. The 3.6× ratio is consistent with the
4.4× measured graph-only on an idle machine, so the direction holds even if
the seconds do not.
Vector read cost remains unmeasured on both sides; nothing here compares
LanceDB similarity search against ProximaDB's vector engine, so no claim is
made about it. That is the last unmeasured quadrant of the total-cost
picture.
Vector read cost (2026-08-20) — the last quadrant¶
Every prior section measured writes, storage, recovery or graph traversal.
Vector reads were never compared, and they were the one axis where the
co-located design should win: in ProximaDB a vector hit is a graph node,
while the SQLite arm searches LanceDB and then joins back into SQLite to turn
ids into nodes.
Same corpus indexed into both backends, 1,350 symbol vectors at 384-d, 50
query vectors sampled with a fixed seed and handed identically to both arms,
top_k = 10. Reproduce with scripts/eval_vector_read.py <workdir>.
| LanceDB (in-process) | ProximaDB | |
|---|---|---|
| search p50 | 4.07 ms | 35.46 ms |
| search p95 | 16.9 ms | 38.8 ms |
| + join to graph nodes | 4.23 ms | not needed |
| recall@10 vs exact cosine | 0.996 | 0.990 |
ProximaDB is ~8.4× slower at equivalent recall. Both backends return
essentially what exact brute-force cosine returns, so this is not a
speed/accuracy trade.
The co-location advantage is real and small. The join the SQLite arm
cannot avoid — LanceDB ids back into graph nodes — costs 0.16 ms
(4.23 vs 4.07). That is the entire prize the single-store design was meant to
collect on this axis. ProximaDB's search costs 31 ms more than LanceDB's.
Transport does not explain it. The obvious suspect is the HTTP/UDS hop to
the embedded server, but the same store over the same transport in the same
run answers k-hop traversal at 0.68 ms p50. Round-trip overhead is
sub-millisecond, so ~31 ms of the 35 ms is search work. Filed upstream as
anvai-labs/proximaDB#1690.
Two measurement hazards the harness avoids, both of which would have produced
a flattering-but-wrong number: LanceDBProvider.search_similar takes a
string and embeds internally while semantic_search takes a vector,
so timing those against each other charges one arm for embedding generation;
and latency without a correctness check is meaningless, because an ANN index
can always be fast by returning the wrong neighbours.
Scale caveat, and it is not a formality: ratios in this comparison have moved
with corpus size before — graph ingest measured 1.14× at 33k nodes and 4.4× at
93k. This is 1,350 vectors. The gap is an order of magnitude and reproducible
across runs, but it should be re-taken at repo scale before it is treated as
settled.
Where this leaves the backend decision¶
With reads measured, every axis has now been measured at least once:
| axis | result |
|---|---|
| ingest (repo scale, embeddings) | ProximaDB 3.6× slower |
| storage at rest (repo scale) | parity, 1.10× |
| recovery | fixed by proximaDB#1678 — 171 s, no longer a differentiator |
| graph traversal | too unstable to quote; harness needs seed control and repeats |
| vector read | ProximaDB ~8.4× slower at equal recall |
The recommendation is unchanged — SQLite remains the default for repo-scale
graphs — but it no longer rests on a defect. It rests on there being no
measured advantage anywhere, and a real cost on the two axes that dominate a
first index and a query.
The in-process path, measured at last (2026-08-23)¶
Everything above measured the subprocess ProximaDB — a managed
proximadb-server child reached over UDS+HTTP+JSON. That is the availability
fallback, not the product CODE_GRAPH_CORRELATED_SUBSTRATE_2026_06_22.adoc
specifies, which is a PyO3 in-process engine. The wheel became buildable
(anvai-labs/proximaDB#1696) and then correct (#1701 typed properties, #1718
durability), so the comparison the vision actually asks for is finally
possible.
Two defects had to be fixed before any number meant anything:
- The in-process path persisted nothing. Ingest reported 177,046 edges,
graph_statsagreed,flush()/checkpoint()/close()all returned
success — and the graph WAL on disk was 0 bytes, so a fresh process saw
edges=0. Both halves of durability (flush_walon shutdown,
recover_all_graphson startup) were gated on the network server
object, which port-free embedded never constructs. A transport choice
silently decided whether data survived. Fixed in proximaDB#1718. - Typed graph properties were stringified at the boundary, so
line: 42
persisted as"42"and the engine's numeric index could never populate.
Fixed in proximaDB#1701.
Write-only comparison, identical pre-parsed payload¶
The earlier repo-scale figures (404 s SQLite / 1,748 s ProximaDB) include
parsing the repository, which both arms share. To isolate the graph
store itself, the same 93,469 nodes / 177,046 edges are replayed from an
existing project.db into each backend:
| SQLite (raw) | ProximaDB in-process | |
|---|---|---|
| durable write | 0.92 s | 3.14 s (2.73 s ingest + 0.41 s close) |
| footprint at rest | 49.7 MB | 49.0 MB |
| traversal p50 | 0.021 ms | 0.018 ms |
| traversal p95 | 0.119 ms | 0.049 ms |
| reopen → queryable | n/a (always open) | 4.05 s, 177,046 edges recovered |
In-process ProximaDB is at rough parity with a raw SQLite graph store:
~3.4× slower on bulk write (2.2 s absolute across a whole repository index),
identical footprint, and equal-or-better traversal — notably better at p95.
That is a different world from the subprocess path, whose graph-write cost
was ~1,344 s over the shared parse baseline. Moving from subprocess to
in-process improves ProximaDB's own graph-write path by roughly 400×, and
it is what makes the backend viable at all.
It does not, however, beat SQLite on the graph tier. The vision's claim
that the embedded engine replaces SQLite+LanceDB has to be carried by the
vector tier and by queries the stitch cannot express in one call
(graph+vector fusion, impact analysis, branch/time-travel) — not by graph
write throughput, where parity is the honest result.
Caveats, stated because they bound the claim:
- The SQLite arm here is a simplified schema (five columns, three
indexes) built for a write-cost comparison, so it is a lower bound on the
product's real SQLite cost; Victor's actual schema does more work. - Durability here means across a clean
close(). Per-batch fsync is a
separate decision with its own throughput trade (proximaDB#1695 documented
the real contract rather than changing it silently). - Graph tier only, embeddings off, matching the arms it is compared against.
- A second in-process instance on the same directory still blocks forever on
exclusive.lock(proximaDB#1715), so reopen must happen in a fresh
process — which is what an embedded application does anyway.
Recommendation¶
Unchanged for today — SQLite remains Victor's default — but the reason has
narrowed to something actionable rather than damning. ProximaDB in-process is
now a credible backend on the graph tier (parity, not penalty), so the
decision should be re-taken when (a) a wheel is published so consumers can
install it without a local Rust build, and (b) the vector tier is measured
in-process, where the single-store design has its only remaining structural
advantage.