Skip to content

Recall pipeline

Recall turns “what is the agent about to do” into a few lines of prior learnings in its context. It is one Python script (recall.py) wrapped by three hooks. This page follows one recall from trigger to injected text and gives the defaults as they are in the code.

reflect recall pipeline Stages of a recall from hook trigger to injected context: persona lookup, scope, cache, four parallel arms (vector, BM25, graph, temporal), RRF fusion, cross-encoder and embeddings, bounded-boost scoring, filters, OOD gate, MMR, token budget, injection. 4 arms in parallel (4 threads); a failed arm returns [] SessionStart hook query: project, branch, commits limit 3, 1500 chars, overlap 0.2 UserPromptSubmit hook query: the prompt (min 12 chars) fetch 9, keep 3 new, 1500 chars SubagentStart hook query: agent type, cwd, task limit 3, 1500 chars, 5 s cap recall.py QUERY (hook subprocess timeout 30 s) 0. persona lookup (O3) aggregate question: answer from project_persona, then stop 1. scope (R15, A6) branch or project shard, else pooled ~/.learnings 2. cache (R9) exact sha1, fuzzy Jaccard 0.85 1 h TTL; hit skips arms + CE vector reflect search --mode naive BM25 qmd search -c learnings graph (R1) reflect search --mode local temporal (R5) date-window scan, only on a date phrase 3. RRF fusion, k = 60 (per-arm floors first, off by default) 4a. cross-encoder (R2) top 20 via reflect rerank MiniLM-L-6-v2 logits 4b. embeddings (R3) top 20 via reflect embed all-mpnet-base-v2 vectors 5. score = sigmoid(CE logit) x bounded boosts confidence, recency, tags, proof, project, domain, authority, speculative 6. filters quarantine, --confidence 7. OOD gate (R7) best of top 5 below min-overlap: nothing 8. MMR (R3) lambda 0.7, top LIMIT 9. budget (R4) --max-tokens (0 = off), then --max-chars 10. inject as additionalContext SessionStart adds a token-economics footer UserPromptSubmit dedupes per session any failure or empty result: empty context, exit 0

Standalone SVG (follows your OS light/dark setting).

Everything after the arms is local: fusion, scoring, gates and MMR run in recall.py; the cross-encoder and embeddings are two small subprocess calls to the reflect engine (served by a warm model daemon when available). Every stage past the primary vector arm is a booster, not a blocker: on any failure it returns nothing and the pipeline continues with what it has.

Source: plugin/skills/recall/scripts/recall.py, function recall().

Four callers run recall.py. Each builds its own query and passes its own limits.

Caller Fires on Query Limit and size Other flags
session_start_recall.py SessionStart project + branch + commit tags (see below) --limit 3, --max-chars 1500 --min-overlap 0.2, --max-tokens 0, --no-gap-log, --no-followup
user_prompt_submit_recall.py UserPromptSubmit the prompt text fetches --limit 9, keeps 3 not-yet-injected, --max-chars 3000, then fits up to 3 whole learnings into 1500 chars (a learning that does not fit is dropped, not cut, and is not marked injected) passes --session-id; no min-overlap, no gap/followup suppression
subagent_start_recall.py SubagentStart subagent <type> | cwd <cwd> | agent_id <id> | <task prompt> --limit 3, --max-chars 1500 --no-gap-log, --no-followup, 5 s timeout
/reflect:recall skill you type it your query --limit 10, --max-chars 2000 all flags available

Hook subprocesses time out at 30 s (REFLECT_RECALL_TIMEOUT); the subagent hook at 5 s (REFLECT_SUBAGENT_RECALL_TIMEOUT). The SessionStart and UserPromptSubmit hooks set HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1 (via setdefault) because a cold model load that hits the network blows the timeout; models must already be cached under ~/.reflect/models/.

Env overrides for the callers: REFLECT_RECALL_LIMIT and REFLECT_RECALL_MAX_CHARS (explicit recall defaults), REFLECT_RECALL_MIN_OVERLAP and REFLECT_RECALL_MAX_TOKENS (SessionStart), REFLECT_SUBAGENT_RECALL_LIMIT, REFLECT_SUBAGENT_RECALL_MAX_CHARS, REFLECT_SUBAGENT_CONTEXT (replaces the recall result with a fixed string).

The hooks are wired in plugin/.claude-plugin/plugin.json for Claude Code, and in codex-hooks.json and copilot-hooks.json for the other harnesses. See Hooks reference.

pretooluse_context.py is not recall. It matches deterministic policy rules from ~/.reflect/policy-rules.jsonl (or REFLECT_POLICY_FILE) and never queries the KB.

build_query() in session_start_recall.py:

  1. Project name: basename of git remote get-url origin (minus .git), else the cwd basename.
  2. Branch, unless it is main, master or detached: slashes, underscores and hyphens become spaces (feat/foo-bar becomes foo bar).
  3. Up to 3 tokens from the last 5 commit subjects (3+ chars, conventional-commit words and filler dropped, ranked by frequency).
  4. Up to 3 more tokens from the last 5 records of ~/.reflect/commits.jsonl (written by the post-commit hook), if present.
  5. Words are lowercase-deduplicated in order. The commit tokens are also passed as --tags, which feeds the tag boost.

So the query is a bag of keywords, not a sentence. That suits the BM25 arm and the embedding arm equally.

If the query looks like an aggregate question (“what testing style do we use?”) and the project has a high-confidence field in the project_persona table, recall() returns that single line and stops. Closed-domain queries or a miss fall through. Any DB problem is swallowed.

resolve_kb_root() picks the KB root, first match wins:

  1. $GLOBAL_LEARNINGS_PATH already set: use it untouched (eval and test harness contract).
  2. --global or RECALL_GLOBAL=1: the pooled KB, ~/.learnings (override root with RECALL_LEARNINGS_ROOT).
  3. Current project plus current non-trunk branch: ~/.learnings/shards/<project>/branches/<branch>/.
  4. Current project, trunk or detached: ~/.learnings/shards/<project>/.
  5. Anything else, or a shard with no documents/*.md and no vector index: the pooled KB.

--all-branches / RECALL_ALL_BRANCHES=1 skips the branch level. Branch comes from RECALL_BRANCH, else git rev-parse --abbrev-ref HEAD; slashes become __.

Cache files live in ~/.reflect/recall_cache/ ($REFLECT_STATE_DIR).

  • Exact tier: key is sha1(version | query | mode | fetched_limit), with the shard folded in. fetched_limit = max(limit * 2, 10), so limit 1 and limit 3 share an entry (SessionStart relies on this: its probe and its main recall hit the same file).
  • Fuzzy tier: if the exact lookup misses, the stopword-filtered token set of the query is compared (Jaccard) against an index of up to 200 recent entries; best match at or above 0.85 wins. Queries with fewer than 2 meaningful tokens never fuzzy-match.
  • Validity: 1 h TTL (--cache-ttl), and invalid once ~/.learnings/nano_graphrag_cache is newer than the file.
  • What a hit skips: the arms, the cross-encoder and the embeddings. Scoring, filters, OOD gate, MMR and budget are re-applied on every call, so tags, --confidence and --limit can differ between calls sharing a fetch.

Run in a 4-worker thread pool. Each returns at most fetched_limit learnings, or [] on any error.

Arm Command Timeout Runs when
Vector reflect search QUERY --mode naive --format json 60 s always (the primary arm)
BM25 qmd search QUERY -c learnings 10 s qmd is on $PATH
Graph (R1) reflect search QUERY --mode local --format json 60 s RECALL_GRAPH_ARM not 0 and mode is not already local
Temporal (R5) scan documents/*.md for notes timestamped inside a parsed date window n/a, capped at 5000 files RECALL_TEMPORAL_ARM not 0 and the query contains a date phrase

Only if the vector arm fails and no other arm returned anything does recall() give up with an error (silent, shown only with REFLECT_RECALL_DEBUG=1).

Note: reflect search accepts --limit but the engine call (LearningsGraphEngine.search) does not pass it on, so the vector and graph arms return whatever the engine’s default QueryParam yields. The limit is enforced later, after MMR.

Date phrases are extracted by a stdlib regex pass in temporal_extraction.py: yesterday, N days ago, last week/month/year, last <weekday>, last 3 days, ISO dates, in march, march 2024, last sprint (14 days), recently, and before/until/since/after <anchor> modifiers. A phrase with no resolvable date (“before the rewrite”) returns nothing.

Reciprocal rank fusion over the four lists: score(doc) = sum of 1 / (60 + rank) across arms (RRF_K = 60). Documents are matched by frontmatter id (else a hash of the first 256 chars). No score normalisation is needed, which is why it suits arms whose native scores are incomparable.

Before fusion, each arm may drop candidates below its own query-term-coverage floor (R12). All four floors default to 0, which is off. See Retrieval features.

For the top 20 fused candidates, two subprocess calls run concurrently, so added latency is the slower of the two:

  • reflect rerank QUERY: cross-encoder/ms-marco-MiniLM-L-6-v2 scores each (query, first 2000 chars) pair. Output is a raw logit per candidate.
  • reflect embed QUERY: unit-normalised all-mpnet-base-v2 vectors for the query and the same 20 candidates, used by MMR.

Both need sentence-transformers (the [graph] extra of reflect-kb). On a slim install each returns available: false and recall silently continues without it.

score = sigmoid(ce_logit) x confidence x recency x tags x proof
x project x domain x authority x speculative

Each factor is 1 + alpha * (norm - 0.5), with norm clamped to 0..1, so it stays inside [1 - alpha/2, 1 + alpha/2] and the neutral value (norm 0.5) is exactly 1.0. Candidates past the top 20 get a cross-encoder component of 1e-6, so they sort below every scored one. Without cross-encoder scores the boost product is the whole score.

Factor Env Default alpha Swing Norm
confidence RECALL_CONFIDENCE_ALPHA 0.2 +/- 10% (confidence_num - 0.3) / 0.6; HIGH 0.9, MEDIUM 0.6, LOW 0.3
recency RECALL_RECENCY_ALPHA 0.2 +/- 10% linear decay over 365 days, floor 0.1, 0.5 if undated (see note)
tags RECALL_TAG_ALPHA 0.2 +/- 10% fraction of query tags the note carries; 0.5 with no query tags
proof count RECALL_PROOF_ALPHA 0.1 +/- 5% 0.5 + ln(proof_count)/10
project RECALL_PROJECT_ALPHA 0.2 up to +10% 1.0 for same project, else neutral
domain RECALL_DOMAIN_ALPHA 0.2 up to +10% 1.0 only if --domain-hint equals the note’s domain
authority RECALL_AUTHORITY_ALPHA 0.1 +/- 5% law/promoted 1.0, advisory 0.5, archived 0.0
speculative RECALL_SPECULATIVE_ALPHA 0.2 down to -10% 0.0 if tagged speculative, else neutral

Alphas are clamped to 0..2; a malformed value falls back to the default. Recency reads the <!-- archived: ISO --> header in the note body (written by the ingest archive step), not frontmatter; a note without that header counts as undated and gets the neutral 0.5. The project boost is skipped (neutral) when the corpus is already a single-project shard.

Before any of this (after fusion, before rerank), filter_superseded() drops notes retired by frontmatter (superseded_by set, status superseded or archived), notes whose id sits in archived/ or documents/.forgotten/, and notes whose ledger row has is_latest = 0. Matching is on note id (never name) and content hash, and only retired notes whose file name or ledger row links to a candidate are opened, so the cost does not grow with the retired backlog. REFLECT_RECALL_INCLUDE_SUPERSEDED=1 turns it off. A missing ledger is fine; ledger ids differ from note ids, so a ledger-only retirement with no artifact_path cannot always be linked to a file.

Applied in this order after scoring:

  1. Quarantine: fleet-imported notes with quarantine set are dropped unless --include-quarantined (implied by --format fleet-context).
  2. Confidence: --confidence HIGH|MEDIUM|LOW|ANY (hooks use ANY).
  3. OOD gate (R7): if the best of the top 5 notes covers less than --min-overlap of the query’s content terms, the result is emptied and flagged ood_gated. Coverage is the fraction of the query’s stopword-filtered tokens (3+ chars) found in the note text. 0 disables it. Defaults: recall.py 0.0, SessionStart 0.2.
  4. MMR (R3): keep the top note, then repeatedly pick argmax(lambda * rel - (1 - lambda) * max_sim_to_selected), with lambda = 0.7 and rel the rerank score divided by the window max. Similarity is cosine between the reflect embed vectors. Notes outside the 20-note window only fill leftover slots. With no embeddings, this is plain [:limit].
  5. Token budget (R4): with --max-tokens N > 0, keep notes in order until the estimate (characters / 4) would exceed N. The first note is always kept. Default 0 (off).

render_markdown() builds one block. --max-chars is a budget on the rendered entries; when the next entry would exceed it the block ends with - _(...N more truncated)_.

## Prior learnings relevant to `billing-svc payments proto` - 2 learnings, ~320 tok injected, est ~3880 tok saved
- **[lrn-0f3a9c]** Regenerate gRPC clients after editing payments.proto - ⚒ D:3000 → R:180 (-94%)
How to apply: Run `make proto-gen` after any .proto edit; CI gates on it.
- **[lrn-7b21de]** Pin the retry budget per client, not per call - ⚖ D:1200 → R:140 (-88%)
How to apply: Set it in the client constructor so retries cannot compound.
memory: 2 learnings, ~320 tok injected, est ~3880 tok saved

The sample is invented, and the real output separates the header counts and the economics with a long dash instead of the hyphen shown here. Per row: the glyph comes from the active mode’s learning_types (work_emoji), D is estimated discovery cost, R is the read cost of the stored note, and the percentage is the saving. The final memory: footer is added by the SessionStart hook only. RECALL_ECONOMICS=0 removes the economics, byte-for-byte the pre-economics format. With --field rule each hit is just that one frontmatter field.

The block is handed back as hook JSON:

Harness Envelope
Claude Code, Codex {"hookSpecificOutput": {"hookEventName": "SessionStart" or "UserPromptSubmit", "additionalContext": "..."}}
Copilot (REFLECT_HARNESS=copilot) {"additionalContext": "..."} (the exact shape Copilot expects is not confirmed in the source)

On UserPromptSubmit the hook also dedupes: ids found in the output ([lrn-...]) are compared with ~/.reflect/session-injected/<session_id>.json, already-seen notes are dropped, and the new ids are saved. SessionStart does not record its ids there, so a note injected at boot can appear again on the first matching prompt. Notes whose id does not start with lrn- are not deduped.

Situation Result
SessionStart with cwd equal to $HOME empty context
SessionStart without uv or recall.py only the slots and conventions blocks, if enabled
UserPromptSubmit with a prompt under 12 characters no recall at all
Best hit covers less than --min-overlap of the query terms empty (on by default at SessionStart only)
Persona hit one persona line, no learnings
reflect CLI not on $PATH empty, silent
Vector arm fails and no other arm has hits empty, silent (REFLECT_RECALL_DEBUG=1 prints the reason)
Recall subprocess exceeds its timeout empty
Any uncaught hook exception empty context, exit 0, breadcrumb in ~/.reflect/last-event.json
UserPromptSubmit and every result already injected this session empty

Hooks never block the session: they always exit 0 with valid JSON.

An empty final result from a real prompt (UserPromptSubmit or /reflect:recall) is appended to ~/.reflect/knowledge-gaps.jsonl as a knowledge gap (RECALL_GAP_LOG=0 disables). SessionStart and SubagentStart suppress this because their queries are synthetic. Infrastructure errors are not logged as gaps.

SessionStart does more than one recall:

cwd == $HOME? ──yes──▶ empty
│ no
▼
Tier 0 slots (REFLECT_SLOTS, opt-in) + conventions pointer (REFLECT_TIERED_INJECT, opt-in)
prepended to whatever follows, never suppress it
▼
Tier 1 skills index (REFLECT_TIERED_INJECT, opt-in): strong hit ──▶ inject skills, stop
▼
probe recall.py --limit 1 --confidence HIGH --format json
fresh (archived within 30 days) and score above 0.8 ──▶ inject that one note, stop
▼
Tier 2 recall.py --limit 3 (the full pipeline above)

Two details worth knowing, both from session_start_recall.py:

  • The probe (R11) is not behind a flag. It always runs first. Its “score” is derived from the note’s confidence tier (HIGH 1.0, MEDIUM 0.7, LOW 0.4), and --confidence HIGH already limits it to HIGH notes, so in practice a HIGH-confidence top hit whose <!-- archived --> header is at most 30 days old short-circuits the boot. A note with no such header has unknown age and never qualifies. The probe warms the cache, so the Tier 2 call that follows reuses the fetch.
  • The skills tier and slots are opt-in env flags (REFLECT_TIERED_INJECT, REFLECT_SLOTS), both off by default. See Retrieval features.

The code comments give these figures; treat them as orders of magnitude, not guarantees.

Cost Source
Cold load of the embedding model and cross-encoder in a fresh process: roughly 11 to 16 s hook comments in session_start_recall.py
Cross-encoder on 20 candidates once loaded: about 50 ms on CPU cross_encoder.py docstring
In-process fallback holds about 3.5 GB RSS with torch cross_encoder.py comment
Cache hit: skips the arms, cross-encoder and embeddings recall()

To avoid paying the cold load per recall, reflect_kb/model_daemon.py keeps both models warm behind a unix socket, auto-spawned on first use and shared by every KB on the box. It exits after REFLECT_IDLE_TIMEOUT seconds idle (default 1800). REFLECT_NO_DAEMON=1 forces in-process loading. REFLECT_EMBED_MODEL and REFLECT_CE_MODEL swap the models; changing the embedding model requires a reindex because similarity must live in the index’s space.

Related state, all under ~/.reflect/ unless REFLECT_STATE_DIR is set: recall_cache/, recall_log.jsonl (one line per recall: query, mode, count, cache tier, economics), knowledge-gaps.jsonl, recent-searches.json (follow-up diagnostic), session-injected/.