Recall pipeline
Recall turns “what is the agent about to do” into a few lines of prior learnings in its context. It is one Python script (recall.py) wrapped by three hooks. This page follows one recall from trigger to injected text and gives the defaults as they are in the code.
The pipeline at a glance
Section titled “The pipeline at a glance”Standalone SVG (follows your OS light/dark setting).
Everything after the arms is local: fusion, scoring, gates and MMR run in recall.py; the cross-encoder and embeddings are two small subprocess calls to the reflect engine (served by a warm model daemon when available). Every stage past the primary vector arm is a booster, not a blocker: on any failure it returns nothing and the pipeline continues with what it has.
Source: plugin/skills/recall/scripts/recall.py, function recall().
Entry points
Section titled “Entry points”Four callers run recall.py. Each builds its own query and passes its own limits.
| Caller | Fires on | Query | Limit and size | Other flags |
|---|---|---|---|---|
session_start_recall.py |
SessionStart |
project + branch + commit tags (see below) | --limit 3, --max-chars 1500 |
--min-overlap 0.2, --max-tokens 0, --no-gap-log, --no-followup |
user_prompt_submit_recall.py |
UserPromptSubmit |
the prompt text | fetches --limit 9, keeps 3 not-yet-injected, --max-chars 3000, then fits up to 3 whole learnings into 1500 chars (a learning that does not fit is dropped, not cut, and is not marked injected) |
passes --session-id; no min-overlap, no gap/followup suppression |
subagent_start_recall.py |
SubagentStart |
subagent <type> | cwd <cwd> | agent_id <id> | <task prompt> |
--limit 3, --max-chars 1500 |
--no-gap-log, --no-followup, 5 s timeout |
/reflect:recall skill |
you type it | your query | --limit 10, --max-chars 2000 |
all flags available |
Hook subprocesses time out at 30 s (REFLECT_RECALL_TIMEOUT); the subagent hook at 5 s (REFLECT_SUBAGENT_RECALL_TIMEOUT). The SessionStart and UserPromptSubmit hooks set HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1 (via setdefault) because a cold model load that hits the network blows the timeout; models must already be cached under ~/.reflect/models/.
Env overrides for the callers: REFLECT_RECALL_LIMIT and REFLECT_RECALL_MAX_CHARS (explicit recall defaults), REFLECT_RECALL_MIN_OVERLAP and REFLECT_RECALL_MAX_TOKENS (SessionStart), REFLECT_SUBAGENT_RECALL_LIMIT, REFLECT_SUBAGENT_RECALL_MAX_CHARS, REFLECT_SUBAGENT_CONTEXT (replaces the recall result with a fixed string).
The hooks are wired in plugin/.claude-plugin/plugin.json for Claude Code, and in codex-hooks.json and copilot-hooks.json for the other harnesses. See Hooks reference.
pretooluse_context.py is not recall. It matches deterministic policy rules from ~/.reflect/policy-rules.jsonl (or REFLECT_POLICY_FILE) and never queries the KB.
How the SessionStart query is built
Section titled “How the SessionStart query is built”build_query() in session_start_recall.py:
- Project name: basename of
git remote get-url origin(minus.git), else the cwd basename. - Branch, unless it is
main,masteror detached: slashes, underscores and hyphens become spaces (feat/foo-barbecomesfoo bar). - Up to 3 tokens from the last 5 commit subjects (3+ chars, conventional-commit words and filler dropped, ranked by frequency).
- Up to 3 more tokens from the last 5 records of
~/.reflect/commits.jsonl(written by the post-commit hook), if present. - Words are lowercase-deduplicated in order. The commit tokens are also passed as
--tags, which feeds the tag boost.
So the query is a bag of keywords, not a sentence. That suits the BM25 arm and the embedding arm equally.
Stage by stage
Section titled “Stage by stage”0. Persona short-circuit (O3)
Section titled “0. Persona short-circuit (O3)”If the query looks like an aggregate question (“what testing style do we use?”) and the project has a high-confidence field in the project_persona table, recall() returns that single line and stops. Closed-domain queries or a miss fall through. Any DB problem is swallowed.
1. Scope: which KB gets read (R15, A6)
Section titled “1. Scope: which KB gets read (R15, A6)”resolve_kb_root() picks the KB root, first match wins:
$GLOBAL_LEARNINGS_PATHalready set: use it untouched (eval and test harness contract).--globalorRECALL_GLOBAL=1: the pooled KB,~/.learnings(override root withRECALL_LEARNINGS_ROOT).- Current project plus current non-trunk branch:
~/.learnings/shards/<project>/branches/<branch>/. - Current project, trunk or detached:
~/.learnings/shards/<project>/. - Anything else, or a shard with no
documents/*.mdand no vector index: the pooled KB.
--all-branches / RECALL_ALL_BRANCHES=1 skips the branch level. Branch comes from RECALL_BRANCH, else git rev-parse --abbrev-ref HEAD; slashes become __.
2. Cache (R9)
Section titled “2. Cache (R9)”Cache files live in ~/.reflect/recall_cache/ ($REFLECT_STATE_DIR).
- Exact tier: key is
sha1(version | query | mode | fetched_limit), with the shard folded in.fetched_limit = max(limit * 2, 10), so limit 1 and limit 3 share an entry (SessionStart relies on this: its probe and its main recall hit the same file). - Fuzzy tier: if the exact lookup misses, the stopword-filtered token set of the query is compared (Jaccard) against an index of up to 200 recent entries; best match at or above 0.85 wins. Queries with fewer than 2 meaningful tokens never fuzzy-match.
- Validity: 1 h TTL (
--cache-ttl), and invalid once~/.learnings/nano_graphrag_cacheis newer than the file. - What a hit skips: the arms, the cross-encoder and the embeddings. Scoring, filters, OOD gate, MMR and budget are re-applied on every call, so tags,
--confidenceand--limitcan differ between calls sharing a fetch.
3. Candidate arms
Section titled “3. Candidate arms”Run in a 4-worker thread pool. Each returns at most fetched_limit learnings, or [] on any error.
| Arm | Command | Timeout | Runs when |
|---|---|---|---|
| Vector | reflect search QUERY --mode naive --format json |
60 s | always (the primary arm) |
| BM25 | qmd search QUERY -c learnings |
10 s | qmd is on $PATH |
| Graph (R1) | reflect search QUERY --mode local --format json |
60 s | RECALL_GRAPH_ARM not 0 and mode is not already local |
| Temporal (R5) | scan documents/*.md for notes timestamped inside a parsed date window |
n/a, capped at 5000 files | RECALL_TEMPORAL_ARM not 0 and the query contains a date phrase |
Only if the vector arm fails and no other arm returned anything does recall() give up with an error (silent, shown only with REFLECT_RECALL_DEBUG=1).
Note: reflect search accepts --limit but the engine call (LearningsGraphEngine.search) does not pass it on, so the vector and graph arms return whatever the engine’s default QueryParam yields. The limit is enforced later, after MMR.
Date phrases are extracted by a stdlib regex pass in temporal_extraction.py: yesterday, N days ago, last week/month/year, last <weekday>, last 3 days, ISO dates, in march, march 2024, last sprint (14 days), recently, and before/until/since/after <anchor> modifiers. A phrase with no resolvable date (“before the rewrite”) returns nothing.
4. Fusion (RRF)
Section titled “4. Fusion (RRF)”Reciprocal rank fusion over the four lists: score(doc) = sum of 1 / (60 + rank) across arms (RRF_K = 60). Documents are matched by frontmatter id (else a hash of the first 256 chars). No score normalisation is needed, which is why it suits arms whose native scores are incomparable.
Before fusion, each arm may drop candidates below its own query-term-coverage floor (R12). All four floors default to 0, which is off. See Retrieval features.
5. Cross-encoder and embeddings (R2, R3)
Section titled “5. Cross-encoder and embeddings (R2, R3)”For the top 20 fused candidates, two subprocess calls run concurrently, so added latency is the slower of the two:
reflect rerank QUERY:cross-encoder/ms-marco-MiniLM-L-6-v2scores each (query, first 2000 chars) pair. Output is a raw logit per candidate.reflect embed QUERY: unit-normalisedall-mpnet-base-v2vectors for the query and the same 20 candidates, used by MMR.
Both need sentence-transformers (the [graph] extra of reflect-kb). On a slim install each returns available: false and recall silently continues without it.
6. Scoring
Section titled “6. Scoring”score = sigmoid(ce_logit) x confidence x recency x tags x proof x project x domain x authority x speculativeEach factor is 1 + alpha * (norm - 0.5), with norm clamped to 0..1, so it stays inside [1 - alpha/2, 1 + alpha/2] and the neutral value (norm 0.5) is exactly 1.0. Candidates past the top 20 get a cross-encoder component of 1e-6, so they sort below every scored one. Without cross-encoder scores the boost product is the whole score.
| Factor | Env | Default alpha | Swing | Norm |
|---|---|---|---|---|
| confidence | RECALL_CONFIDENCE_ALPHA |
0.2 | +/- 10% | (confidence_num - 0.3) / 0.6; HIGH 0.9, MEDIUM 0.6, LOW 0.3 |
| recency | RECALL_RECENCY_ALPHA |
0.2 | +/- 10% | linear decay over 365 days, floor 0.1, 0.5 if undated (see note) |
| tags | RECALL_TAG_ALPHA |
0.2 | +/- 10% | fraction of query tags the note carries; 0.5 with no query tags |
| proof count | RECALL_PROOF_ALPHA |
0.1 | +/- 5% | 0.5 + ln(proof_count)/10 |
| project | RECALL_PROJECT_ALPHA |
0.2 | up to +10% | 1.0 for same project, else neutral |
| domain | RECALL_DOMAIN_ALPHA |
0.2 | up to +10% | 1.0 only if --domain-hint equals the note’s domain |
| authority | RECALL_AUTHORITY_ALPHA |
0.1 | +/- 5% | law/promoted 1.0, advisory 0.5, archived 0.0 |
| speculative | RECALL_SPECULATIVE_ALPHA |
0.2 | down to -10% | 0.0 if tagged speculative, else neutral |
Alphas are clamped to 0..2; a malformed value falls back to the default. Recency reads the <!-- archived: ISO --> header in the note body (written by the ingest archive step), not frontmatter; a note without that header counts as undated and gets the neutral 0.5. The project boost is skipped (neutral) when the corpus is already a single-project shard.
7. Filters, gate, MMR, budget
Section titled “7. Filters, gate, MMR, budget”Before any of this (after fusion, before rerank), filter_superseded() drops notes retired by frontmatter (superseded_by set, status superseded or archived), notes whose id sits in archived/ or documents/.forgotten/, and notes whose ledger row has is_latest = 0. Matching is on note id (never name) and content hash, and only retired notes whose file name or ledger row links to a candidate are opened, so the cost does not grow with the retired backlog. REFLECT_RECALL_INCLUDE_SUPERSEDED=1 turns it off. A missing ledger is fine; ledger ids differ from note ids, so a ledger-only retirement with no artifact_path cannot always be linked to a file.
Applied in this order after scoring:
- Quarantine: fleet-imported notes with
quarantineset are dropped unless--include-quarantined(implied by--format fleet-context). - Confidence:
--confidence HIGH|MEDIUM|LOW|ANY(hooks useANY). - OOD gate (R7): if the best of the top 5 notes covers less than
--min-overlapof the query’s content terms, the result is emptied and flaggedood_gated. Coverage is the fraction of the query’s stopword-filtered tokens (3+ chars) found in the note text.0disables it. Defaults:recall.py0.0, SessionStart 0.2. - MMR (R3): keep the top note, then repeatedly pick
argmax(lambda * rel - (1 - lambda) * max_sim_to_selected), withlambda = 0.7andrelthe rerank score divided by the window max. Similarity is cosine between thereflect embedvectors. Notes outside the 20-note window only fill leftover slots. With no embeddings, this is plain[:limit]. - Token budget (R4): with
--max-tokens N > 0, keep notes in order until the estimate (characters / 4) would exceed N. The first note is always kept. Default 0 (off).
8. Render and inject
Section titled “8. Render and inject”render_markdown() builds one block. --max-chars is a budget on the rendered entries; when the next entry would exceed it the block ends with - _(...N more truncated)_.
## Prior learnings relevant to `billing-svc payments proto` - 2 learnings, ~320 tok injected, est ~3880 tok saved- **[lrn-0f3a9c]** Regenerate gRPC clients after editing payments.proto - ⚒ D:3000 → R:180 (-94%) How to apply: Run `make proto-gen` after any .proto edit; CI gates on it.- **[lrn-7b21de]** Pin the retry budget per client, not per call - ⚖ D:1200 → R:140 (-88%) How to apply: Set it in the client constructor so retries cannot compound.
memory: 2 learnings, ~320 tok injected, est ~3880 tok savedThe sample is invented, and the real output separates the header counts and the economics with a long dash instead of the hyphen shown here. Per row: the glyph comes from the active mode’s learning_types (work_emoji), D is estimated discovery cost, R is the read cost of the stored note, and the percentage is the saving. The final memory: footer is added by the SessionStart hook only. RECALL_ECONOMICS=0 removes the economics, byte-for-byte the pre-economics format. With --field rule each hit is just that one frontmatter field.
The block is handed back as hook JSON:
| Harness | Envelope |
|---|---|
| Claude Code, Codex | {"hookSpecificOutput": {"hookEventName": "SessionStart" or "UserPromptSubmit", "additionalContext": "..."}} |
Copilot (REFLECT_HARNESS=copilot) |
{"additionalContext": "..."} (the exact shape Copilot expects is not confirmed in the source) |
On UserPromptSubmit the hook also dedupes: ids found in the output ([lrn-...]) are compared with ~/.reflect/session-injected/<session_id>.json, already-seen notes are dropped, and the new ids are saved. SessionStart does not record its ids there, so a note injected at boot can appear again on the first matching prompt. Notes whose id does not start with lrn- are not deduped.
What is skipped, and when
Section titled “What is skipped, and when”| Situation | Result |
|---|---|
SessionStart with cwd equal to $HOME |
empty context |
SessionStart without uv or recall.py |
only the slots and conventions blocks, if enabled |
| UserPromptSubmit with a prompt under 12 characters | no recall at all |
Best hit covers less than --min-overlap of the query terms |
empty (on by default at SessionStart only) |
| Persona hit | one persona line, no learnings |
reflect CLI not on $PATH |
empty, silent |
| Vector arm fails and no other arm has hits | empty, silent (REFLECT_RECALL_DEBUG=1 prints the reason) |
| Recall subprocess exceeds its timeout | empty |
| Any uncaught hook exception | empty context, exit 0, breadcrumb in ~/.reflect/last-event.json |
| UserPromptSubmit and every result already injected this session | empty |
Hooks never block the session: they always exit 0 with valid JSON.
An empty final result from a real prompt (UserPromptSubmit or /reflect:recall) is appended to ~/.reflect/knowledge-gaps.jsonl as a knowledge gap (RECALL_GAP_LOG=0 disables). SessionStart and SubagentStart suppress this because their queries are synthetic. Infrastructure errors are not logged as gaps.
SessionStart is tiered
Section titled “SessionStart is tiered”SessionStart does more than one recall:
cwd == $HOME? ──yes──▶ empty │ no ▼Tier 0 slots (REFLECT_SLOTS, opt-in) + conventions pointer (REFLECT_TIERED_INJECT, opt-in) prepended to whatever follows, never suppress it ▼Tier 1 skills index (REFLECT_TIERED_INJECT, opt-in): strong hit ──▶ inject skills, stop ▼probe recall.py --limit 1 --confidence HIGH --format json fresh (archived within 30 days) and score above 0.8 ──▶ inject that one note, stop ▼Tier 2 recall.py --limit 3 (the full pipeline above)Two details worth knowing, both from session_start_recall.py:
- The probe (R11) is not behind a flag. It always runs first. Its “score” is derived from the note’s confidence tier (HIGH 1.0, MEDIUM 0.7, LOW 0.4), and
--confidence HIGHalready limits it to HIGH notes, so in practice a HIGH-confidence top hit whose<!-- archived -->header is at most 30 days old short-circuits the boot. A note with no such header has unknown age and never qualifies. The probe warms the cache, so the Tier 2 call that follows reuses the fetch. - The skills tier and slots are opt-in env flags (
REFLECT_TIERED_INJECT,REFLECT_SLOTS), both off by default. See Retrieval features.
Where recall spends time
Section titled “Where recall spends time”The code comments give these figures; treat them as orders of magnitude, not guarantees.
| Cost | Source |
|---|---|
| Cold load of the embedding model and cross-encoder in a fresh process: roughly 11 to 16 s | hook comments in session_start_recall.py |
| Cross-encoder on 20 candidates once loaded: about 50 ms on CPU | cross_encoder.py docstring |
| In-process fallback holds about 3.5 GB RSS with torch | cross_encoder.py comment |
| Cache hit: skips the arms, cross-encoder and embeddings | recall() |
To avoid paying the cold load per recall, reflect_kb/model_daemon.py keeps both models warm behind a unix socket, auto-spawned on first use and shared by every KB on the box. It exits after REFLECT_IDLE_TIMEOUT seconds idle (default 1800). REFLECT_NO_DAEMON=1 forces in-process loading. REFLECT_EMBED_MODEL and REFLECT_CE_MODEL swap the models; changing the embedding model requires a reindex because similarity must live in the index’s space.
Related state, all under ~/.reflect/ unless REFLECT_STATE_DIR is set: recall_cache/, recall_log.jsonl (one line per recall: query, mode, count, cache tier, economics), knowledge-gaps.jsonl, recent-searches.json (follow-up diagnostic), session-injected/.
- Retrieval features: every optional feature with default, key and cost.
- Index and storage: what the arms read.
- Configuration reference: the TOML and env surface.
- Troubleshooting: empty recall, slow recall.