Quickstart
Two tracks. Steps 1 to 4 prove the engine works on its own (no harness needed). Steps 5 and 6 wire it into Claude Code so recall and capture happen automatically. Other harnesses: see Codex, Copilot, Hermes.
Prerequisites
Section titled “Prerequisites”| Need | Why |
|---|---|
| Python 3.11+ | requires-python = ">=3.11" |
uv |
Installs the CLI, and every plugin hook runs as uv run --script |
claude CLI |
The background drain calls claude -p to write learnings (steps 5 and 6 only) |
qmd (optional) |
Adds a BM25 arm to recall. Without it, recall uses the graph and vector arm only |
1. Install the engine
Section titled “1. Install the engine”uv tool install --upgrade --torch-backend cpu \ 'git+https://github.com/stevengonsalvez/ainb-reflect-memory.git[graph]'reflect --versionreflect --version prints reflect, version 0.3.0 (the engine version, not the plugin’s). --torch-backend cpu skips roughly 4 GB of CUDA wheels; drop it on a GPU box. If your uv rejects the flag, use UV_TORCH_BACKEND=cpu uv tool install ... instead.
2. Create the knowledge base
Section titled “2. Create the knowledge base”reflect initCreates ~/.learnings/ with documents/ and nano_graphrag_cache/, runs git init, and writes a .gitignore that excludes the cache. Set GLOBAL_LEARNINGS_PATH first to use another location.
3. Add a learning
Section titled “3. Add a learning”A learning is a markdown file with YAML frontmatter. title, category and key_insight are required; confidence and tags are optional but used for ranking.
cat > bun-not-node.md <<'EOF'---title: "Use Bun, not Node, in the billing-api repo"category: toolingkey_insight: "billing-api runs on Bun; node commands fail on the lockfile and test runner."confidence: hightags: [bun, node, billing-api]---
## ProblemThe agent ran `npm install` and `node --test` in billing-api and corrupted the lockfile.
## FixUse `bun install` and `bun test`. The repo has a bun.lockb and no package-lock.json.EOF
reflect add ./bun-not-node.mdreflect add copies the note to ~/.learnings/documents/<slug>-<hash>.md, auto-generates an entity sidecar (heuristic, no LLM), and inserts it into the graph index. The first run downloads the embedding model, so allow a few minutes. If qmd is installed, add also runs a qmd sync, which can take up to two minutes on large KBs.
Expected tail of the output:
Indexed into graphAdded: /home/you/.learnings/documents/use-bun-not-node-in-the-billingapi-repo-a44a62.mdTitle: Use Bun, not Node, in the billing-api repoCategory: toolingEntities: 8, Relationships: 64. Recall it
Section titled “4. Recall it”reflect search "which package manager for billing-api"The result panel contains the note, found by meaning rather than exact words. Useful flags: --mode naive|local|global (default naive), --limit, --format rich|json|simple. See the CLI reference.
This is the raw engine query. What a harness injects is the fuller recall pipeline (fusion, rerank, gating): recall pipeline.
5. Install the plugin (Claude Code)
Section titled “5. Install the plugin (Claude Code)”claude plugin marketplace add stevengonsalvez/ainb-reflect-memoryclaude plugin install reflect@ainb-reflect-memoryStart a new Claude Code session and run:
/reflect:recall bun billing-apiOutput is a markdown block headed Prior learnings relevant to ... with one line per hit, showing your note. From now on, the SessionStart and UserPromptSubmit hooks run the same recall automatically and inject up to 3 learnings (capped at 1500 characters) before the agent acts.
6. Capture a real learning
Section titled “6. Capture a real learning”Correct the agent in a session (“no, we use bun here”), then end the session or let it compact. What happens next:
Stop,SessionEndorPreCompactruns a $0 gate over the transcript and, if it has signal, appends it to~/.reflect/pending_reflections.jsonl. No LLM is called.- The next
SessionStartlaunches the drain in the background (debounced to once per 10 minutes). It slices the transcript and runsclaude -pon Sonnet to extract learnings. - Written learnings are indexed automatically (the drain runs
reflect reindexafter a successful batch).
To capture immediately instead of waiting, run /reflect inside the session.
Check it worked:
reflect stats # document, entity and relationship countstail ~/.reflect/drain.logIn Claude Code, /reflect:status shows pending reviews, sidecar coverage and GraphRAG health, and /reflect:cost shows what the drain has spent.
Where things live
Section titled “Where things live”| Path | Contents |
|---|---|
~/.learnings/documents/ |
Learning notes and .entities.yaml sidecars (source of truth) |
~/.learnings/nano_graphrag_cache/ |
Local vector and graph index (derived) |
~/.cache/qmd/index.sqlite |
BM25 index, only if qmd is installed (derived) |
~/.reflect/ |
Queue (pending_reflections.jsonl), drain.log, errors.json, config overrides |
- Overview: how the loop fits together.
- Per-harness install: hooks, adapters and options in full.
- Memory browser:
reflect serveto browse and curate what was captured.