Skip to content

Comparison

Eight coding-agent memory systems on four axes. Short version: tools that need no second API key tend to hoard noise; tools that curate well make you run a server and pay a second provider; and none of the others capture corrections, tests, git events or skill changes as typed signals.

For the underlying problem, see problem and fit.

# Axis Question
1 Selective or everything Does it store only signal-bearing moments, or every tool call?
2 Separate key Does memory need its own LLM or embedding provider and subscription?
3 Infrastructure Does it need an always-on server (Postgres, Redis, a daemon, a vector DB)?
4 Coding signals Are corrections, test outcomes, git events and skill upgrades captured as typed events, or hoped for from prose?
Tool What it stores Separate LLM or embed key Runs as Coding signals first-class License
reflect Selective learnings; markdown is the source of truth No separate key; the drain calls claude -p, embeddings are local Local files; optional shared Postgres Yes: corrections, tests, tool loops, git, todos, permissions, contradictions, skill refresh MIT
Hindsight LLM-extracted facts and mental models Yes for writes (see note) Local daemon, Docker (FastAPI and Postgres), or cloud No: extracted from prose MIT
Mem0 LLM-extracted facts Yes: OpenAI by default at ingest Library, or Docker (Postgres and Neo4j) No: hooks, but no correction, git or test capture Apache-2.0 and SaaS
ByteRover Curated markdown tree Curation makes its own LLM calls Local files and a node daemon; optional cloud sync No: curation is agent-directed, not passive Elastic 2.0 (not OSI open source)
claude-mem Every tool call, compressed to observations No: reuses Claude auth and local embeddings Always-on Bun daemon; optional ChromaDB No: probabilistic extraction Apache-2.0
agentmemory Every tool call, verbatim Optional; value degrades without it Always-on Rust daemon No: raw events, little structure without an LLM Apache-2.0
Honcho User and peer models Yes when self-hosted; cloud is metered Postgres, Redis and a deriver worker; or cloud No: built for end-user personalization AGPL-3.0
OpenViking Tiered LLM-extracted memories and skills Yes: OpenAI or Volcengine by default Rust, Go and C++ server, always on No: git, test and skill capture not first-class AGPL-3.0

Note on Hindsight: its retain step needs an LLM. At evaluation time its “reuse your Claude subscription” loopback was documented as personal-use only under Anthropic’s terms, so it could not be a shipped default for a tool other people install.

Pattern: the only other tool that matches reflect on “no extra key” (claude-mem) stores everything and grows noisy. The tools that curate well (Hindsight, OpenViking, ByteRover) lose on key or infrastructure. Typed signal capture is a column of “no”.

Scores are 1 to 5, assigned by reflect’s author and weighted for one workflow: capturing corrections into behavior, cheaply, with no infrastructure. Treat the weights as an opinion, not a measurement.

Criterion (weight) reflect Hindsight claude-mem ByteRover Mem0 agentmemory Honcho OpenViking
Signal capture (0.25) 5 2 2 2 2 2 1 2
No extra key (0.22) 5 2 5 3 2 3 2 2
Local-first (0.20) 5 3 3 4 2 2 2 2
Selective, low noise (0.16) 4 4 2 4 3 1 2 4
Retrieval quality (0.12) 4 5 3 3 4 3 4 4
License and portability (0.05) 5 5 4 2 4 4 2 2
Weighted total 4.72 3.03 3.08 3.06 2.50 2.28 1.99 2.56

Memory looks cheap if you only price retrieval. The real cost is the write side (LLM extraction and consolidation) plus the read side (context injection).

System Write side Read side Extra spend
reflect Gated and queued; one claude -p extraction per queued transcript (Sonnet by default), local embeddings Local vector, graph and optional BM25; no LLM unless HyDE is on No second key; drain runs consume your Claude quota (see /reflect:cost)
claude-mem Haiku on your Claude auth, per tool-call batch Local FTS and ONNX vectors No second key, but high write volume
Hindsight LLM extraction on retain Vector and graph; synthesis has a cost Separate key or metered cloud
Mem0 OpenAI extraction at ingest Vector and BM25 Separate OpenAI key, plus a paid tier for graph memory
OpenViking VLM extraction plus tier summaries Vector, recursive Separate provider key
Honcho Deriver LLM per batch Context reads are cheap; dialectic queries are metered One to three keys self-hosted, or per-query
agentmemory Optional LLM compression, off by default Local embeddings and rank fusion None if the LLM is off, but value degrades
ByteRover Curation uses its own LLM calls BM25 with LLM fallback Own key or a paid plan

reflect 4.1.0 on LOCOMO, a long-term conversational memory benchmark. This is a small pilot: 50 stratified questions (10 per category) from one conversation (conv-26), answered and written by Sonnet, graded by an Opus reference judge. Retrieval runs reflect’s real engine; the dialogue-to-note extraction is a LOCOMO-specific adapter. Configurations were kept or dropped by their score on these same 50 questions, and per-cell noise is about 0.1.

Config (Opus judge) single-hop multi-hop temporal open-domain adversarial overall
reflect 4.1.0, tuned recall and extraction 0.70 0.75 0.80 0.50 0.90 0.73
plus retrieval fixes 0.80 0.80 0.80 0.70 0.90 0.80

The fixes are two env-gated, no-new-key options, both off by default: a stronger local embedder (REFLECT_EMBED_MODEL=BAAI/bge-base-en-v1.5) and HyDE query expansion (REFLECT_RECALL_HYDE=1, which uses reflect’s own claude -p).

LOCOMO positioning of reflect against other memory systems

reflect lands mid-field, on par with Memobase and Zep and above Mem0. Newer systems (ByteRover, Honcho, Hindsight) score higher but are self-reported on their own harnesses. Judges and harnesses differ across the field, so read this as directional placement, not a ranking. Methodology, tuning steps and judge calibration: benchmarks.

When memory is wrong, can you see it, edit it, delete it and trace its origin?

  • reflect: learnings are plain markdown you open and edit in your editor, or curate through the memory browser. Notes carry provenance such as the source session. Volatile bookkeeping lives in a local SQLite database, keeping git diffs clean. Optional per-row TTL (forget_after) and contradiction detection retire stale beliefs.
  • Most others: facts live in Postgres or vector stores, inspected through an API or SQL. claude-mem ships a web viewer, but observations are database rows. ByteRover is the exception, with an editable markdown tree.

In most systems a correction is a sentence in a transcript that an extraction LLM may or may not turn into a fact. reflect routes it as typed signals at hook time: contradiction, git commit and revert, test pass and fail, tool loop, todo completion, permission reply, idle sweep, and knowledge gaps from empty recalls. When a revision changes a learning that a skill depends on, it queues an automatic skill refresh.

The aim is to capture what changed, not only what was said, and feed it back as the next session’s behavior. Hook-level detail: hooks reference and capture.

Verdict: build and port, with eyes open. reflect ported retrieval ideas from Hindsight and others (graph expansion, RRF fusion, cross-encoder rerank, temporal handling). The cost is owning that code: retrieval is the least differentiated and highest-maintenance part of reflect, and it trails the best systems on benchmarks.

If you would adopt instead: four pilot tests

Section titled “If you would adopt instead: four pilot tests”
  1. Zero extra credentials. Install on a fresh machine with only the agent’s existing auth. Do capture and recall work with no new key or account? (reflect needs the claude CLI for capture; claude-mem passes; most others fail.)
  2. No always-on server. Reboot. Does memory work with nothing running in the background? (reflect passes; daemon-based tools fail.)
  3. Correction to behavior. Correct the agent once. Next session, does the rule resurface without manual curation?
  4. Noise over 30 days. Run daily for a month. Is recall still sharp, or flooded with stale and duplicate entries? (Fire-hose tools degrade here.)
  1. Quickstart: run reflect on your machine.
  2. OKF versus reflect: design note comparing Google’s Open Knowledge Format (OKF) with reflect.
  3. Recall pipeline: what reflect does at retrieval time.