Forensic memory
The first version wrote its own memories. It made 2,366 of them, and the handmade ones were the only clean ones. I deleted it and kept the search.
Problem
Every agent session on this machine writes a transcript to disk. Those transcripts answer real questions. How did we fix this error before? Which session made that decision? Where did this number come from? Without an index, the answers exist and cannot be found.
The common build does more than index. An extractor reads each finished session, decides what matters, and writes memories into a long-term store. The agent then treats that store as truth. I built that version first: a compiler on the session stop hook, and a poller that did the same across agents.
The question that matters is not whether a machine can write memories. It is whether a machine-written memory can be a record. I measured mine, and the answer was no.
System
In May I counted the extractor’s output: 2,366 open assertion rows, sludge that no session ever read again. The captures I wrote by hand in the same database stayed clean. That difference was the verdict, not a tuning problem. I deleted the extractor and its poller the same day, and I kept the search.
What survived is deliberately simple. One local server indexes every transcript into one SQLite database with full-text search. The transcripts come from four sources: my Claude profile (CLI), my employer’s work profile (also CLI), the Claude desktop app, and Codex. Results rank on text match, then on recency, and tool output is down-weighted because it is one third of the index. I refused embeddings and semantic search, and I recorded the reasons: the plain index already answers the real query shape. Subtraction first.
The survivor is still held to a measurement. A gold set of 15 real forensic queries runs against the live server, and the session each query must surface is pinned by hand. The baseline is 15 of 15 in the top five, with a mean reciprocal rank of 1.0. The eval exits non-zero on a failure, so it gates every change to the ranking code. The down-weight on tool output holds its default on that evidence: the sweep without it loses a case.
The instrument needed its own fix. Each search from an agent session is ingested into the corpus within seconds, and the eval’s own queries are no exception. Without a frozen date, the eval matched its own echo. Eleven weeks later, a test query’s top hit was the session that ran the test, minutes old. Every gold case now scores against the corpus at a frozen date, and the date moves only when I examine every case again. When a system records you, the instrument is part of the system.
The rule that keeps this system in its place is written at the top of the repo: forensic lookup only, never load-bearing. When a search surfaces something that matters, I promote it into the vault by hand — refine and cite, never copy. The archive points at what happened. Only the vault holds what I know.
What runs today
The server runs on this machine now, under a keep-alive service, and every agent session reaches it with one curl call. The corpus holds every agent conversation from the four sources, and a watcher syncs a new transcript within minutes of its last line. The work boundary is active: a work session pins every recall call to the work account, and personal context never enters it.
The artifact
The panels show the deletion that made the system small, the gate that lets it stay, and one real question answered from the archive.
CLAUDE.md §History — what the extractor earned- **2026-05-16** — OMS auto-extraction dismantled: the Stop-hook compiler + cross-agent poller generated 2,366 un-closed assertion "sludge" rows while hand-made captures stayed clean. … docs/optimization-backlog.md — considered and argued against| 11 | REJECTED | Shrink FTS index (`detail=`/contentless) || 12 | REJECTED | Backups for conversations.db || 13 | REJECTED | New endpoints / embeddings / semantic search | FTS5 + recency + noise down-weighting already answers the "how did we solve X" query shape well (verified live). Embeddings add an index to maintain and a model dependency for marginal recall on a corpus this size. Subtraction first.
The dismantling and the refusals are written where the code lives, each with its reason. The 2,366 rows are the reason the system is small.
$ uv run python evals/run_eval.pyfts-punctuation-fix rank 1embeddings-rejected rank 1oms-sludge rank 1codex-ingestion rank 1twenty-crm-url rank 1launchd-kickstart rank 1wikijs-railway rank 1backup-rotation rank 1thread-viewer rank 1ghl-cli-hyphen rank 1memory-reminder-hyphen rank 1zclaude-alias rank 1openui-dashboard rank 1recency-project-scoped rank 1codex-origin-scoped rank 1server-py-punct passenv-local-punct passhub-ops-punct passagentmail-overconstrained MISS [L3 FTS recall] (known-hard, not scored) ranked: hit@5 15/15 hit@10 15/15 mrr 1.0 smoke: 3/3 passEXIT: 0 the sweep — the case for the tool-output down-weightconfig hit@5 hit@10 mrrbaseline 15/15 15/15 1.0pure-bm25 15/15 15/15 1.0no-tool-demotion 15/15 15/15 0.897no-glm-demotion 15/15 15/15 1.0heavy-recency 15/15 15/15 1.0half-life-7d 15/15 15/15 1.0half-life-90d 15/15 15/15 0.967
Fifteen pinned queries, each at rank 1, and one documented miss kept out of the score on purpose. The exit code gates every change to the ranking code. The sweep is the evidence for the down-weight: without it, the rank decreases. One gold case, oms-sludge, is the deletion in the panel beside this one.
$ curl -s "http://127.0.0.1:8025/api/search?q=propagation+og.png&project=mattbastar-mattbastar-com&limit=2"{ "results": [ { "timestamp": "2026-08-13T20:59:37.647+00:00", "kind": "message", "content": "Resolved on its own — that was propagation lag in the seconds after upload. The apex now serves `/og.png` as `image/png` at 57,667 bytes, byte-identical to the local file. …", "session_id": "4800d551-b71d-4db0-9295-b78516a3cc03", "origin": "personal", "bm25": -18.9243, "recency_score": 0.9036, "score": 27.4748, … }, { "timestamp": "2026-08-13T21:01:45.544+00:00", "content": "**The share card is live.** Verified on the apex rather than assumed from the upload: …", "score": 12.4362, … } ], "query": "propagation og.png", "count": 2, "tool_result_weight": 0.3, …}
The question is three days old: what was the og.png trap at deploy? Rank 1 is the session that resolved it, scored by match and recency with tool output down-weighted. The result points at what happened. What mattered from that session was promoted into the repo docs by hand — the valve in one query.
I captured all three panels live on 2026-08-16: the repo files at commit d37327a, the eval and its sweep against the live server, and one search round-trip. The query is staged — chosen for this shot, and about this site’s own build. Everything returned is real; long lines are re-wrapped to fit the frame and elisions are marked. No panel shows client, employer, or third-party conversation content.