Skip to content

Retrieval styles

A map of the retrieval strategies in reflect, by style rather than by feature id. For each: what it does, where the code is, what has been measured, and whether it is live. The per-switch defaults and costs are on Retrieval features; how the stages chain together on one query is on Recall pipeline. The only quality numbers come from the LOCOMO pilot and the small recall eval; where a style has no measurement, this page says so.

Status values: live (runs by default), opt-in (shipped, off by default), helper (code exists, not called by recall()), dropped (tried, removed or abandoned), not built (planned or considered, no code).

query
│
┌────────┬────────┴────────┬──────────┐ 4 arms, parallel thread pool
▼ ▼ ▼ ▼
vector BM25 graph temporal
(naive) (QMD) (local) (date window)
└────────┴────────┬────────┴──────────┘
▼
per-arm floor (R12, off) ──▶ RRF fusion (k=60)
▼
cross-encoder rerank × bounded boosts ──▶ MMR ──▶ OOD gate (R7) ──▶ budget
Style What it adds Status Measured impact
Dense vector (GraphRAG naive) semantic match live baseline arm
Lexical BM25 (QMD) exact tokens, error strings live if qmd installed no isolated measurement
Graph neighbourhood (local) one-hop connected notes live no isolated measurement
Graph communities (global) cluster-level context helper, placeholder reports none
Typed-edge filtering filter hops by link type helper none
Hybrid fusion (RRF) merge arms without score calibration live none isolated
Cross-encoder rerank meaning-level reorder live LOCOMO fix B (bundled)
Bounded boosts recency, confidence, tags, proof, project live none isolated
MMR diversity de-duplicate near-twins live none isolated
Temporal date-window arm, query date parsing live LOCOMO temporal 0.80 (not isolated)
Entity-aware entities at index time, not query time index-time only none
Query expansion (HyDE) answer-shaped query opt-in LOCOMO +0.06 overall
Embedder swap stronger vectors opt-in LOCOMO fix B
Gates and floors return nothing when nothing fits R7 off in recall.py; R12 off LOCOMO regression at 0.15
Caches skip repeat work live none
Scoping and sharding project and branch isolation live where shards exist none
Staged, tiered, persona, slots cut tokens, not rank mixed none
Shared Postgres store same vectors on every machine opt-in parity tests only
Not built or declined

What. nano-graphrag’s naive mode: embed the query, nearest-neighbour search over the chunk vectors. It is the default primary arm (DEFAULT_MODE = "naive").

Where. LearningsGraphEngine.search(mode=...) in src/reflect_kb/cli/graph_engine.py; recall calls it as reflect search QUERY --mode naive --format json. Default embedder all-mpnet-base-v2 (768-d, CPU), set in model_daemon.py and overridable with REFLECT_EMBED_MODEL. Since 5.2.0 the model lives in a persistent unix-socket daemon, so a warm embed or rerank round trip is about 0.2 s instead of 7.5 s cold (from the 5.2.0 changelog).

Status. Live. In the committed recall eval (tests/eval/results/baseline.json) every attributed top-5 slot came from the vector arm (graphrag_only 48, qmd_only 0, both 0, neither 8); QMD most likely was not installed in that environment (inference).

What. QMD’s qmd search, a BM25 index over the same markdown files. It catches literal tokens (a file name, an error string) that embeddings blur.

Where. fetch_qmd in plugin/skills/recall/scripts/recall.py. It is active exactly when qmd is on $PATH; a missing or failing qmd returns no results and recall continues. reflect reindex runs qmd update and qmd embed only if qmd is installed.

Why search and not query. The commit that introduced it (2026-04-23, e80992567) records the choice: qmd query runs LLM query expansion plus reranking and took about 25 s, while qmd search is BM25 only at about 0.45 s. Recall wanted the cheap complementary signal, and does its own reranking. QMD’s own vector and LLM-rerank modes are therefore not used by recall (the TUI learnings screen does use qmd query; see TUI learnings plugin).

Status. Live when installed. No quality number isolates BM25: the LOCOMO calibration run recorded a BM25 sample of 0 (arm_thresholds_bge.txt), consistent with QMD being absent there (inference).

What. nano-graphrag’s local mode walks the entity neighbourhood in the stored graph, so a note linked to a lexical or vector hit can surface even when it shares no words with the query (R1).

Where. Same search call with --mode local, added as a second arm in recall() unless the primary mode is already local. Entities and relationships come from .entities.yaml sidecars written next to each note (entity_store.py); the graph is rebuilt by reflect reindex.

How the graph is built. nano-graphrag normally calls an LLM to extract entities. reflect replaces that call with a passthrough in graph_engine._llm_complete that feeds the sidecar entities straight in, so indexing makes no model calls. The graspologic dependency is replaced by a small networkx shim (graspologic_shim.py) to avoid its numba/llvmlite chain.

Status. Live (RECALL_GRAPH_ARM, default on). Behavioural proof proof_R1_graph_expansion_arm.py shows the two-hop example; there is no ablation of its quality.

What. global mode retrieves community-level reports, meant for broad “what patterns exist” questions.

Status: helper, with a caveat. recall.py --mode global and reflect search --mode global work, but recall never uses them by default. Because indexing uses the passthrough LLM, the community reports are a fixed placeholder (“A group of related technical concepts and patterns”, rating 5.0) rather than real summaries, so global retrieval carries no per-community content. This is read from the code (graph_engine._llm_complete), not measured. The Postgres backend tests do cover global parity between backends, which tests storage, not usefulness.

What. Each sidecar relationship has a type (for example solves, caused_by, requires). The type survives into the graph as a [type] prefix on the edge description, and graph_links.py recovers it from the stored GraphML with the standard library only, so slim builds can still query by link type (“what enabled this fix?”).

Where. src/reflect_kb/cli/graph_links.py, exposed as LearningsGraphEngine.get_typed_edges.

Status. Helper. The module docstring describes filtering the R1 arm by link type, but recall.py does not import it, so the live graph arm does not filter by type today.

What. Reciprocal-rank fusion over the arm result lists: score = sum over arms of 1 / (60 + rank), de-duplicated by note id or chunk hash. Ranks, not scores, are fused because BM25, cosine and graph scores are not comparable.

Where. rrf_fuse and RRF_K = 60 in recall.py; the four arms run in a ThreadPoolExecutor(max_workers=4) and a failed arm returns an empty list.

History. The first retrieval plan (April 2026, Retrieval roadmap) proposed a weighted blend, qmd_score * 0.6 + graph_degree * 0.4. What shipped is RRF over BM25 and vector (commit e80992567), later widened to graph and temporal arms.

Status. Live. Not ablated against a weighted blend.

What. After fusion, the top 20 candidates are scored jointly with the query by a local cross-encoder (default cross-encoder/ms-marco-MiniLM-L-6-v2, REFLECT_CE_MODEL to swap). The score becomes the primary sort key; the boosts multiply it (R2).

Where. src/reflect_kb/recall/cross_encoder.py and the reflect rerank subcommand; RECALL_CROSS_ENCODER=0 disables. Slim builds without sentence-transformers fall back to formula-only ordering.

Measured. Only bundled: LOCOMO fix B swapped both the embedder and the reranker (to bge) and gained +0.01 overall, with single-hop 0.70 to 0.80.

Status. Live.

What. Secondary signals multiply the rerank score but each is clamped, so they break ties and cannot override a decisive relevance gap (R8, R16, F3, SG3). Defaults in recall.py: confidence, recency and tag alpha 0.2 each (about +/-10%), proof count 0.1, project affinity 0.2, domain 0.2, authority 0.1, speculative down-rank 0.2. Recency is a linear decay over 365 days with a floor of 0.1, and neutral 0.5 when the date is missing.

History. The first recall (v3.1.0) used confidence x recency (60-day half-life) x tag overlap. The R8 commit replaced the exponential decay, which crushed old notes to near zero, with the bounded linear form.

Status. Live. No ablation.

What. Maximal Marginal Relevance on the final top-k: keep the best hit, then pick each next note to maximise lambda * relevance - (1 - lambda) * max similarity to those already picked, lambda 0.7 (RECALL_MMR_LAMBDA). It stops three near-identical notes filling three slots. Similarity uses the same embedding space as the index (reflect embed).

Where. mmr_select in recall.py; RECALL_MMR=0 or --no-mmr disables. It runs concurrently with the cross-encoder, so added latency is the larger of the two, not the sum.

Status. Live. Proof: proof_R3_mmr_diversity.py. No quality ablation.

What. Two parts. R6 parses natural-language dates in the query (“last week”, “since 2026-01-01”, “in March”) into a range with a confidence. R5, only when a range was found, adds an arm that scans note timestamps for the window and feeds RRF. Date-free queries get nothing from this arm. An older design also considered bitemporal edge validity (A2); the filter exists as a helper and recall() does not call it.

Where. plugin/skills/recall/scripts/temporal_extraction.py, fetch_temporal in recall.py; RECALL_TEMPORAL, RECALL_TEMPORAL_ARM, RECALL_BITEMPORAL_EDGES.

Measured. LOCOMO temporal questions scored 0.80 in the best configuration, but LOCOMO queries are phrased as questions about stored note dates and the benchmark did not isolate the arm.

Status. Live (A2 helper).

reflect is entity-aware at index time only. Each note can carry an .entities.yaml sidecar (entities with types and descriptions, typed relationships); auto_extract_entities falls back to heuristics (backtick terms, frontmatter) when the writer produced none. These feed the graph that the local arm walks. There is no reflect-side entity detection on the query: recall does not parse entity names out of the question. nano-graphrag’s own local mode matches the query embedding against entity descriptions and expands from there, which is the only query-time entity step. A fuzzy entity-by-name-or-alias lookup exists in the Postgres MemoryStore (MemoryStore.lookup_entities in postgres/store.py), but recall does not call it.

Status. Index-time live; query-time entity linking not built.

What. One headless claude -p call writes a single invented answer sentence, which is appended to the query so retrieval embeds answer-shaped text. It bridges the vocabulary gap between a question and the statement that answers it.

Where. _hyde_expand in recall.py, gated by REFLECT_RECALL_HYDE=1; model from REFLECT_DRAIN_MODEL (default sonnet), 60 s timeout, any failure falls back to the raw query. No new API key. Costs about $0.02 per recall plus latency (REPORT.md).

Measured. The largest single gain in the LOCOMO pilot: 0.74 to 0.80 overall, multi-hop 0.70 to 0.80, open-domain 0.60 to 0.70, adversarial 0.80 to 0.90 (Opus judge, n=50).

Status. Opt-in. The hooks run recall.py as a child process and inherit the environment, so setting REFLECT_RECALL_HYDE=1 in the harness env turns it on for every recall, including SessionStart, adding a model call to each.

What. REFLECT_EMBED_MODEL (for example BAAI/bge-base-en-v1.5) and REFLECT_CE_MODEL replace the default models. The embedding dimension is derived from the loaded model, but vectors are dimension-specific, so a swap needs reflect reindex --force.

Measured. LOCOMO fix B: single-hop 0.70 to 0.80, open-domain 0.50 to 0.60, +0.01 overall. Kept as a knob, not made the default.

Status. Opt-in.

Gate What it does Status
OOD gate (R7) if the top hit covers too few query terms, inject nothing --min-overlap default 0 (off) in recall.py; SessionStart hook uses 0.2
Per-arm floors (R12) each arm drops sub-floor candidates before fusion off (all 0); calibrated suggestions 0.1 vector, 0.15 BM25, 0 graph, 0.05 temporal; recall.py --calibrate-thresholds
Token budget (R4) pack notes until a token budget, not a count --max-tokens, REFLECT_RECALL_MAX_TOKENS

Dropped result. In LOCOMO the abstention gate (R7 at 0.15) over-suppressed: 27 of 50 answers became “NOT MENTIONED” against 12 of 50 without it, and overall J fell to 0.44. It was dropped for that benchmark; the shipped SessionStart default is 0.2 with a different goal (do not inject junk into a new repo). A gentler value was never tuned.

Exact cache (1 h, per query), fuzzy cache (R9, Jaccard similarity of 0.85 or more against recent queries, RECALL_FUZZY_CACHE), and a follow-up diagnostic (A4). They save repeated work and change no ranking; the LOCOMO harness passes --no-cache so they cannot affect scores.

Per-project shards (R15) and branch sub-shards (A6) keep recall to the current project and branch by default; --global and --all-branches union them; a project-affinity boost (R16) softly prefers the current project in global mode. An empty shard falls back to the global KB (5.0.2 fix). Live where shards exist.

These reduce tokens rather than change ranking.

  • Staged recall (M1): recall_stages.py index / timeline / hydrate returns a token-capped index first, then full bodies only for chosen ids.
  • Tiered inject (R10, R11): at SessionStart, look up curated skills first; a strong, fresh hit short-circuits the raw recall. Off unless REFLECT_TIERED_INJECT=1.
  • Persona fields (O3) and memory slots (A1): direct answers and pinned text that bypass search.
  • Corpus Q&A (M7): snapshot a filtered set of notes and answer questions over it with no retrieval models.

Details and defaults: Retrieval features.

An opt-in backend (REFLECT_PG_DSN plus REFLECT_WORKSPACE_ID) swaps the local vector and graph storage for tenant-scoped Postgres with pgvector, so every machine queries the same index. Only the vector and graph arms touch it; BM25, rerank, MMR, temporal and the boosts are backend-independent (pinned by tests/postgres/test_recall_backend_independence.py). Evidence is parity tests (the local and Postgres backends return the same evidence set), not a quality comparison. See Shared Postgres backend.

Idea Outcome
Weighted blend of BM25 and graph degree replaced by RRF before it shipped
QMD hybrid query (LLM expansion plus rerank) inside recall declined for latency (about 25 s vs 0.45 s)
Real LLM community summaries for global mode not built; reports are placeholders
Query-time entity linking not built
Helpfulness-feedback reranking (roadmap phase 6) not built; recall only logs to an append-only JSONL and the A4 follow-up diagnostic exists
Retrieval exposed as an MCP tool open question in the 2026-04-20 handover; not built
SQLite single-file vector backend considered in the Postgres work, explicitly not built
Abstention gate at 0.15 for LOCOMO tried and dropped

Code locations and defaults above were read from this repo at main (48968dd) and spot-checked: RRF_K = 60, DEFAULT_MODE = "naive", CE_CANDIDATES = 20, MMR lambda 0.7, alphas, RECENCY_WINDOW_DAYS = 365, the passthrough LLM stub, HyDE gating. The --calibrate-thresholds entry point is a flag on recall.py, not a reflect subcommand. The 60 proof files in tests/eval/behavioral/proofs/ demonstrate behaviour with each knob on and off; they do not measure retrieval quality.