Architecture
reflect is a loop with four stages and one rule: hooks never call a model. Hooks record cheap facts and queue transcripts. A detached drain turns the queue into learnings. The engine indexes them. Recall injects them at the start of the next session or prompt.
The loop
Section titled “The loop” harness session ┌────────────────────┐ │ hooks (no LLM) │ write notes directly ┌───────────────────────┐ │ PostToolUse │ ──────────────────────────▶ │ ~/.learnings/ │ │ UserPromptSubmit │ (test fix, todo done, │ documents/*.md │◀─ source of truth │ Notification │ mini-learning, perm.) └──────────┬────────────┘ │ │ │ reflect add / reindex │ PreCompact, Stop, │ enqueue transcript path ┌──────────▼────────────┐ │ SessionEnd, ... │ ─────────┐ │ derived index │ └─────────▲──────────┘ │ │ qmd (BM25) │ │ additionalContext ▼ │ nano-graphrag files │ │ ┌──────────────┐ │ or shared Postgres │ │ │ ~/.reflect/ │ └──────────┬────────────┘ │ │ pending_ │ SessionStart │ │ │ reflections │ ─▶ detached drain │ │ │ .jsonl │ gate, slice, │ │ └──────────────┘ one claude -p ────┘ writes notes │ + sidecars, reindexes └────────────────── recall (hybrid search, rerank, budget) ◀──────┘| Stage | Page | One line |
|---|---|---|
| Capture | Capture | Hooks write deterministic notes and queue transcripts. $0, no model. |
| Drain | Drain | A detached script gates, slices, and runs one model call per transcript. Capped on turns, time, tokens, and daily count. |
| Index | Index and storage | reflect add and reflect reindex build the lexical and graph indexes from the markdown. |
| Recall | Recall pipeline | Fan out to the arms, fuse, rerank, gate, pack to a budget, inject. |
Components
Section titled “Components”| Component | Where | Runs | Calls a model | Owns |
|---|---|---|---|---|
| Lifecycle hooks | plugin/hooks/*.py, plugin/skills/recall/hooks/*.py |
In the harness process, per event | No | Queue, armed watchers, hook-written notes under ~/.reflect/ and the KB |
| Drain | plugin/hooks/reflect-drain-bg.sh, plugin/scripts/reflect_cascade.py, plugin/scripts/drain_extract.py |
Detached, started by SessionStart |
Yes, claude -p, one call per transcript |
Lock, debounce, retry and cost ledgers, poison file |
| Engine | src/reflect_kb/ (reflect CLI) |
On demand | No LLM. Local embedding and rerank models | Index build and search, reflect serve, the model daemon |
| Recall | plugin/skills/recall/scripts/recall.py |
Called by recall hooks and /reflect:recall |
No | Recall cache and log under ~/.reflect/ |
| Ledger | plugin/scripts/reflect_db.py |
Imported by scripts | No | ~/.reflect/reflect.db (SQLite) |
Design rules the code enforces:
- Hooks are silent-fail. Every hook wraps its body, writes a breadcrumb on error, and exits 0. A broken hook cannot break the session.
- Producers and the consumer are separate.
PreCompact,Stop,SessionEnd, andSubagentStoponly append to the queue. The detachedSessionStartdrain is the only thing that runs/reflectfrom hooks. - No extra API key. Capture shells out to the
claudeCLI. Embeddings and reranking run on a local model (all-mpnet-base-v2and a MiniLM cross-encoder by default). - Markdown is the source of truth. Indexes and the SQLite ledger are derived or auxiliary. A lost index is rebuilt with
reflect reindex.
One session, end to end
Section titled “One session, end to end”Harness coverage
Section titled “Harness coverage”Same scripts, three wiring files. Event names differ in case between harnesses.
| Harness | Wiring file | Events wired |
|---|---|---|
| Claude Code | .claude-plugin/plugin.json |
13: SessionStart, UserPromptSubmit, Notification, PreToolUse, PermissionRequest, PostToolUse, PostToolUseFailure, Stop, PostCompact, SubagentStart, SubagentStop, SessionEnd, PreCompact |
| Codex CLI (0.129+) | plugin/codex-hooks.json |
10: as Claude minus Notification, PostToolUseFailure, SessionEnd |
| GitHub Copilot | plugin/copilot-hooks.json |
13: camelCase equivalents (sessionStart, userPromptSubmitted, agentStop for Stop, and so on), no postCompact, plus errorOccurred |
Per-event behaviour is in Hooks reference.
Two storage modes
Section titled “Two storage modes”The markdown notes are always local. Only the derived vector and graph store moves.
| Local (default) | Shared (Postgres) | |
|---|---|---|
| Derived store | nano_graphrag_cache/ under the KB root, plus qmd’s own index |
Four tables in the reflect_memory schema (ng_kv, ng_graph_nodes, ng_graph_edges, ng_vectors) |
| Enable | Nothing | Install the [postgres] extra, apply supabase/migrations/0001_*.sql and 0002_*.sql, set REFLECT_PG_DSN and REFLECT_WORKSPACE_ID |
| Across machines | Sync the notes (git), run reflect reindex on each machine |
Every machine reads the same store |
Both env vars must be set for the Postgres store to switch on. REFLECT_PG_DSN is the trigger, not the generic DATABASE_URL. nano-graphrag itself is unchanged; it is handed Postgres storage classes. The database does no LLM or embedding work, and tenancy is the workspace_id column enforced by row-level security. Details in Index and storage.
Local model daemon
Section titled “Local model daemon”Each reflect search, embed, or rerank used to cold-boot torch (about 3.5 GB RSS). Since 5.2.0 a unix-socket daemon loads the models once and serves every CLI call. It auto-spawns on first use, exits after REFLECT_IDLE_TIMEOUT seconds idle (default 1800), and one daemon serves every KB for a given user, model pair, and TMPDIR. If it cannot start, calls fall back to in-process loading, with a lock that allows one concurrent load. REFLECT_NO_DAEMON=1 disables it.
Where state lives
Section titled “Where state lives”| Path | Purpose | Detail |
|---|---|---|
~/.learnings/ (or $GLOBAL_LEARNINGS_PATH) |
KB root: notes, sidecars, derived graph cache | Index and storage |
~/.reflect/ (or $REFLECT_STATE_DIR) |
Queue, armed watchers, drain ledgers, logs, reflect.db |
Capture, Drain |
$REFLECT_STATE_DIR/reflect.toml or ~/.reflect/reflect.toml |
User config over the bundled plugin/reflect.toml |
Configuration |