Regression suite
reflect has to stay provably stable while its recall features and storage backends evolve. The suite answers two questions: did the Postgres backend change what recall returns, and do the 57 ported recall features (the 4.1.0 “ports”) still behave? It is built in layers, from fast deterministic unit tests to heavy full-stack golden runs. Only the cheap layers gate pull requests.
┌───────────────────────┐ ┌──────────────────────┐ ┌───────────────────────┐│ Unit + contract tests │ │ Behavioral proofs │ │ Golden snapshot diff ││ tests/, CI gate │ │ 60 files, real engine│ │ R@5 vs baseline.json ││ slim install, fast │ │ manual / non-blocking│ │ manual (workflow_disp)│└───────────────────────┘ └──────────────────────┘ └───────────────────────┘┌───────────────────────┐ ┌──────────────────────┐ ┌───────────────────────┐│ Backend parity │ │ e2e (Playwright) │ │ Plugin contract ││ local vs Postgres │ │ reflect serve UI │ │ manifests + assets ││ needs a database │ │ CI gate │ │ CI gate │└───────────────────────┘ └──────────────────────┘ └───────────────────────┘What runs where
Section titled “What runs where”| Layer | Location | What it proves | Needs | Runs in |
|---|---|---|---|---|
| Unit tests | tests/test_*.py, tests/postgres/ (no-DB files) |
Frontmatter schema, write flow, metrics, model daemon, fleet import, domain boost, recency, issues pipeline, reflect serve curation, Postgres SQL builders and models |
[dev] install only |
CI job test, Python 3.11 and 3.12 (gate) |
| Fleet guards | test_fleet_importer.py, test_domain_boost.py, test_fleet_context_format.py, test_recency_norm.py |
Fleet importer and recall ranking inputs stay deterministic | [dev] install |
CI job fleet-guards (gate) |
| Fleet proofs F1 to F3 | tests/eval/behavioral/proofs/proof_F*.py |
Fleet ingest isolation, domain boost ranking, quarantine enforcement against the real engine | [graph] + model |
Same job, non-blocking (skips or fails on slim) |
| e2e | tests/e2e/ (Playwright) |
Every memory browser flow against a fixture KB | Node 20, Chromium | CI job e2e (gate) |
| Plugin contract | scripts/check_plugin_contract.py, manifest JSON checks |
Manifests are valid JSON with name and version; every ${CLAUDE_PLUGIN_ROOT} path exists (and shell scripts are executable); files external consumers rely on (for example the statusline’s plugin/scripts/reflect_timeline.sh) are still at their contracted paths |
stdlib | CI job plugin-manifest (gate) |
| Behavioral proofs | tests/eval/behavioral/proofs/ |
Each ported feature has an observable, deterministic invariant on a real-engine KB | [graph], embedding model, full-stack reflect |
Manual |
| Golden snapshot diff | tests/eval/snapshot_diff.py, tests/eval/results/baseline.json |
Ranking quality on 20 golden queries has not regressed | [graph] + model |
CI job golden-diff, workflow_dispatch only, never blocks a PR |
| Backend parity | tests/postgres/nanographrag/ |
Local and Postgres backends return identical evidence | Live Postgres with pgvector | Manual (skips cleanly without a database) |
| Plugin tests | plugin/tests/, plugin/adapters/tests/ |
Drain, hooks, recall scripts, cascade, adapters | [dev] install; some tests shell out to uv and the recall script, so they are slow on a cold machine |
Not run by any CI workflow (see gaps) |
| LOCOMO benchmark | tests/eval/locomo/ |
Answer quality, not regression | Heavy, costs money | Manual, see benchmarks |
At the time of writing, the CI gate command passes locally on a slim install with 278 passed, 18 skipped (the skips are database-backed and graph-stack tests that need extra dependencies).
Run it
Section titled “Run it”# CI gate, slim install (no torch, no models)uv venv --python 3.12 && uv pip install -e '.[dev]'uv run pytest -q tests --ignore=tests/eval
# Plugin-side tests (not in CI)uv run pytest -q plugin/tests plugin/adapters/tests
# Packaging contractpython3 scripts/check_plugin_contract.py
# Memory browser e2e (starts reflect serve against a fixture copy)cd tests/e2e && npm install && npx playwright install chromium && npx playwright testHeavy layers need the full stack. Use a venv with the graph extra and point the harness at its reflect:
uv venv .venv-eval --python 3.12VIRTUAL_ENV=$PWD/.venv-eval uv pip install -e '.[dev,graph]'export RECALL_EVAL_BIN_DIR=$PWD/.venv-eval/bin # so recall.py and the harness resolve the same CLI
# the behavioral proofs, one pytest process per file, writes a verdict matrixpython3 tests/eval/behavioral/run_proofs.py # allpython3 tests/eval/behavioral/run_proofs.py R7 R8 # only those ports
# golden snapshot gate (exit 1 on regression, exit 0 and SKIP when the env cannot run)python3 tests/eval/snapshot_diff.pypython3 tests/eval/snapshot_diff.py --update-baseline # only after an intentional ranking change
# backend parity, needs Postgres with the pgvector extensionDATABASE_URL=postgresql://... uv run pytest -m integration tests/postgresThe 13 lookup types
Section titled “The 13 lookup types”Recall combines several lookups. Only the ones that read nano-graphrag’s storage can be affected by the Postgres backend.
| # | Lookup | Layer | Backend-coupled | Source |
|---|---|---|---|---|
| 1 | Vector / semantic (naive) | nano-graphrag | yes (local vector store vs pgvector) | src/reflect_kb/cli/graph_engine.py |
| 2 | Graph-local (entity neighborhood) | nano-graphrag | yes | graph_engine.py |
| 3 | Graph-global (community reports) | nano-graphrag | yes | graph_engine.py |
| 4 | BM25 / lexical (QMD) | QMD | no (own index) | plugin/skills/recall/scripts/recall.py |
| 5 | Typed-link graph (R1) | nano-graphrag + recall | yes | src/reflect_kb/cli/graph_links.py |
| 6 | Cross-encoder rerank (R2) | recall | no | src/reflect_kb/recall/cross_encoder.py |
| 7 | Embed + MMR diversity (R3) | recall | no (own embed call) | reflect embed in src/reflect_kb/cli/learnings_cli.py |
| 8 | Temporal (R5, R6) | recall | no | recall.py |
| 9 | Entity / alias lookup | nano-graphrag | yes | src/reflect_kb/cli/entity_store.py |
| 10 | Corpus saved-filter (M7) | recall | no (frontmatter scan) | src/reflect_kb/recall/corpus.py |
| 11 | RRF fusion + recency / confidence / tag rerank | recall | no | recall.py |
| 12 | Staged 3-layer recall (M1) | recall | no | plugin/skills/recall/scripts/recall_stages.py |
| 13 | Per-project sharding / global scope (R15, R16) | recall | no | recall.py |
The backend swap is inert unless both REFLECT_PG_DSN and REFLECT_WORKSPACE_ID are set (the generic DATABASE_URL does not trigger it). Of the 57 ported features, exactly one (R1, the graph arm) routes through nano-graphrag’s storage; the other 56 are recall-layer and backend-agnostic. tests/postgres/test_recall_backend_independence.py pins that by scanning the recall script for references to the Postgres backend (REFLECT_PG, reflect_kb.postgres, the Pg storage classes, pgvector, psycopg) and failing if any appear. It needs no database.
Behavioral proofs
Section titled “Behavioral proofs”Each proof file under tests/eval/behavioral/proofs/ follows one pattern: seed specific learnings into a hermetic real-engine KB, run recall.py the way SessionStart does, and assert an observable invariant on ranking, inclusion, exclusion or metadata. No LLM takes part in the assertion, so the seeds and flags fully determine the result.
There are 60 proof files. 57 are the recall-upgrade ports, plus 3 fleet proofs:
| Family | Files | Theme |
|---|---|---|
R |
17 | Retrieval: R1 to R16 and R20 (graph arm, rerank, MMR, token budget, temporal, OOD gate, bounded boosts, fuzzy cache, tiered inject, per-arm thresholds, sharding, project affinity, skills index) |
S |
10 | Storage structuring (structured fields, typed links, numeric confidence, provenance, belief revision, history, chunk-hash dedup) |
SG |
8 | Signals (contradiction, git events, idle sweep, test outcomes, loop detection, knowledge gaps, todo completion, permission replies) |
M |
8 | claude-mem style safeguards (staged recall, writer breaker, quota abort, modes, commit verification, private-tag strip, corpus Q&A, token economics) |
A |
6 | agentmemory style (pinned slots, bitemporal edges, TTL forget, followup diagnostic, synthetic compression, branch isolation) |
C |
5 | Consolidation (semantic dedup, auto-consolidation, graph maintenance, lifecycle events, export/import) |
O |
3 | Open-domain (observations layer, conventions doc, persona fields) |
F |
3 | Fleet (ingest isolation, domain boost, quarantine) |
The feature behind each family is described on the retrieval features page and the 4.1.0 entry of the changelog.
run_proofs.py runs each file in its own pytest process (a crash in one cannot poison the rest), takes the verdict from the pytest return code (0 pass, 1 test failure, anything else error), writes tests/eval/behavioral/results/matrix.json and prints a markdown table. The matrix.json committed in the repo covers only 11 ports, so it is a sample from an earlier run, not a current full result.
Golden snapshot diff
Section titled “Golden snapshot diff”tests/eval/snapshot_diff.py builds a hermetic KB from tests/eval/fixtures/corpus/, scores the 20 queries in tests/eval/fixtures/golden_queries.yaml, and compares the result with the committed tests/eval/results/baseline.json.
| Query class | Count | Purpose |
|---|---|---|
exact |
8 | Unique terminology; vector and BM25 should hit directly |
graph |
5 | Best answer is one entity hop away from the lexical match (exercises R1) |
temporal |
3 | Current convention must outrank the superseded one |
ood |
4 | Nothing relevant exists; returned results count as noise |
The gate fails (exit 1) when either:
- overall recall at 5 drops by more than 0.05 against the baseline, or
- an
exactquery that had a relevant document in the baseline top 5 no longer has that document in its top 5.
The committed baseline records recall at 5 of 1.0, MRR of 0.9375 and a noise rate of 0.2 across the 20 queries (the four ood queries all return results, because recall.py leaves the OOD gate off unless --min-overlap or REFLECT_RECALL_MIN_OVERLAP is set). If the environment cannot run the real engine or no baseline exists, the script prints a SKIP and exits 0 instead of failing. The same check is available as a pytest test (test_no_recall_regression) that skips under the same conditions.
Backend parity (Postgres)
Section titled “Backend parity (Postgres)”tests/postgres/nanographrag/test_backend_parity.py seeds an identical corpus into the local-default backend and the Postgres backend (same pinned embedding, canned LLM extraction), runs naive, local and global queries, and asserts that the evidence set is identical on both. A second test asserts the local backend writes a .graphml file while the Postgres backend writes none. test_backend_parity_realmodel.py repeats the parity check with the real all-mpnet-base-v2 embedding model and skips unless sentence-transformers is installed.
The rest of tests/postgres/ covers the adapters themselves:
| File | Proves |
|---|---|
test_server_is_dumb.py |
Storage adapters do no LLM or embedding work (source scan, no database) |
test_sql_builders.py, test_models.py, test_normalize.py |
workspace_id is always the first bound parameter and values are never interpolated; tenant scope is mandatory; dedupe hashing is stable (no database) |
test_integration_store.py |
Insert and full-text search, idempotent ingestion, graph neighborhood, tenant isolation, RLS fail-closed (live database) |
nanographrag/test_pg_storage_conformance.py |
Adapters satisfy nano-graphrag’s storage contracts across two isolated instances |
nanographrag/test_cross_machine_graphrag.py |
A real GraphRAG pipeline inserts on “machine A” and answers local and naive queries on a fresh “machine B” from Postgres |
nanographrag/test_ng_rls.py |
Row-level security fails closed and isolates workspaces |
Database tests read DATABASE_URL or REFLECT_TEST_DATABASE_URL, apply the two migrations in supabase/migrations/, and skip with a clear message when no database, psycopg or pgvector is available.
Known gaps
Section titled “Known gaps”plugin/tests/andplugin/adapters/tests/are not run by CI. The CI test job runspytest -q tests --ignore=tests/eval, and no other workflow invokes the plugin test directories. That includes the drain regression tests behind the fixes in 5.2.1 to 5.2.5, so run them locally before release.- Behavioral proofs are not a PR gate. Only F1 to F3 run in CI, and non-blocking. The other 57 rely on a manual run with the full stack.
- The golden diff is manual (
workflow_dispatch) because it installs about 2 GB of torch plus a model download. - Ruff and mypy are advisory. Both run in CI with
continue-on-errorwhile a style and typing backlog is cleared.
Adding coverage
Section titled “Adding coverage”- A new recall feature should ship a proof named
proof_<PORT>_<slug>.pyundertests/eval/behavioral/proofs/;run_proofs.pyderives the port id from the filename. - A fix to the drain or hooks should include a test under
plugin/tests/that drives the real script against a stubclaude, as the 5.2.x tests do. - If an intentional change shifts golden-query ranking, regenerate the baseline with
--update-baselineand commit it with an explanation.