Keep, Change, Kill: What We’re Doing Differently

tl;dr — Six posts of findings are worth nothing if nothing changes. So: we’re killing context-packing as the default for code questions, because handing the model tools beat every packing strategy on quality and cost. We’re changing what counts as a valid context representation — anything that strips line numbers is disqualified from citation-requiring work, which puts a hard requirement on our own retrieval product. And we’re keeping two things that turned out to be load-bearing: writing down the caveat you don’t have time to test, and shipping the smallest runnable cut instead of the correct plan. The correct plan sat unrun for ten weeks. The small cut found four bugs in a week, two of them ours. Publication found a fifth.

The boundary matters: this is five questions against one pinned Java repository, judged by one model in a single pass. That is enough to change our default workflow. It is not a universal ranking of context strategies, and we’ve tried to keep the claims below inside it.

The series, in order: nobody needed to fit the codebase in the window · 44x fewer tokens, and every citation was fake · our control group was broken · a finding about RAG that was a finding about my config · two axes of compression · what it costs to know · keep, change, kill (this post).

Everything, in one table

strategyscore /12tokens in$/questionverified citesfabricated
agentic exploration11.40288,3420.63871422
llm-tldr → agentic11.00288,4360.63781240
full source dump10.80404,8780.85526511
Repomix10.40406,2770.93746917
prose-compressed dump10.40363,6130.7777861
llm-tldr (extract)6.6066,8950.1520831
RAG retrieval5.004,2370.03533115
llm-tldr (semantic search)2.405,1790.0236229

The RAG row is confounded and a clean rerun is still blocked: the adapter handed the model index-prefixed paths (jsoup/src/…) that the judge resolved against the checkout root, so citations that were real scored as unresolved. Its fabrication count is an upper bound and its score moves with it (voitta-rag#57, #58). No other row is affected.

Everything below is a decision taken from that table. The arguments are in the six posts; this is what we do about them.

Kill: “get the codebase into the window” as the default

Start with the thing we’re stopping. The most useful result was the one we added almost as an afterthought: hand the model read_file, grep and glob, inject nothing, and it beat the full dump on quality while using fewer tokens and less money.

Every tool we benchmarked is an answer to how do I fit the codebase into the window. On this workload that’s the wrong question, and the tools inherit the wrongness.

So tool access is the default for code Q&A, and packing is the exception — reserved for when something rules out the agentic loop: no tool-calling surface, a hard latency ceiling, or per-query economics that can’t absorb 7–16 round trips.

We’re deliberately not generalising past code. Agentic exploration wins here because source is navigable: greppable identifiers, imports that point somewhere, filenames that mean something. Undifferentiated prose has none of that, and retrieval should do better there. The claim is about codebases.

Change: citability is a hard requirement

The benchmark’s whole discriminating power came from one rule — every claim needs a file:line, and the judge resolves it against real source. That rule caught what no quality score caught: the most token-efficient arm fabricated 29 of its 31 citations.

The mechanism is structural, not incidental. llm-tldr‘s semantic search reports "line": 1 for every unit; its extract subcommand carries real line_number fields and the same tool goes from 2 verified citations to 83. voitta-rag’s chunks carry a chunk_index and no lines at all. A representation that omits locations doesn’t degrade gracefully — the model still has to satisfy the citation requirement, so it invents a plausible number. You get confident, specific wrongness: the most expensive failure mode there is, because it’s the one that survives review.

  • Line spans are a product requirement for our retrieval layer, filed as voitta-rag#52. Without them the component can’t be used where claims must be verifiable.
  • We evaluate integrations, not tools. “llm-tldr scores 2.4” was never a true sentence — the same tool scores 6.6 through a different subcommand. Adoption decisions name the adapter.
  • Citation resolution is standard in our evals, not a special feature of this one.

Change: what our own retrieval product is for

This one stings. Retrieval landed at the bottom, and we tested the obvious excuse — corpus mismatch — by indexing exactly the benchmark’s file set. It scored marginally worse — and then a fresh run with the adapter’s path bug fixed reversed even that, landing the two configurations 0.40 apart in the other direction. At five questions that is not a result, it is noise with a sign. The corpus excuse is untested, not refuted, and we are not going to claim otherwise in either direction. That is enough to stop treating retrieval as our default for code questions. It is not enough to declare its quality ceiling, and we’re not going to.

What it does settle is an earlier portfolio audit, which argued that voitta-rag owns the commodity half and hasn’t built its differentiator: retrieval is rentable, and the broker — routing across representations — is the actual asset, still unbuilt. The benchmark gives that direction evidence, though the retrieval rerun remains outstanding. The differentiator has two parts:

  • Citability. Line spans, verifiable claims. Table stakes, and the thing missing.
  • Routing. The right strategy varies by question class, and a layer that picks is worth more than a layer that retrieves.

Keep: the caveat you don’t have time to test

When we first published llm-tldr at 2.40, we wrote one sentence we couldn’t support with data: this measures one adapter, not the tool’s ceiling.

That sentence cost nothing and was the highest-value thing in the post. It’s why someone went back, tried extract, found the score nearly triples and the fabrications drop from 29 to 1 — surfacing the best quality-per-token arm in the benchmark, previously invisible.

Keep writing the falsifiable caveat. When you know the shape of what would overturn your result, say so in the artifact. It’s the cheapest insurance against publishing a wrong conclusion permanently, and it converts a dead end into a queued experiment.

Keep: ship the smallest runnable cut

The full plan was 30 questions × 5 modes = 150 judged runs, gated behind an interview session to draft the question set. Correct, thorough, and it sat unrun for ten weeks.

The version that ran was five questions and one repository. It produced the headline, two corrections to our own published claims, a pile of harness bugs and a filed product requirement — in about a week.

The lesson isn’t “small is better.” It’s that the spec was the blocker and its thoroughness was the reason. A plan that requires a meeting to start doesn’t start. Every axis we cut turned out to be a config entry plus one function once the harness existed.

What this changes on Tuesday

when the question is…reach forwhy
high-stakes code Q&Aagentic exploration (read_file/grep/glob)best answers, most verified citations, cheaper than packing
high-volume, low-stakes (triage, classification)llm-tldr structural index61% of the quality at a sixth of the cost
code retrieval as the proposed defaultwait for line spansstill bottom-tier after a clean re-run, and #52 blocks citation-requiring work
prose corporaretrievalthe navigability argument for agentic exploration doesn’t hold

How we run evals, as commitments

Not doctrine — the things we got wrong, written as what we’ll do next time:

  • We will audit the control first and hardest. Everything is measured against it, so an error there multiplies across every row. Ours was understated by 4.2 points and invalidated a whole writeup. Two lines of assertion would have caught it.
  • We will print what each arm actually received before theorising about why it lost. The better our explanation for a surprising result, the more suspicious we should be — a good mechanism is exactly what stops you checking the inputs.
  • We will assert non-empty output per cell. A full bill with an empty answer is a config bug, not a model result. We hit it twice.
  • We will measure the rendered answer, not tokens_out. On a thinking model, thinking is billed as output and swamps everything: one arm shrank its visible answer 21% while tokens_out tripled.
  • We will instrument cost per cell from the first run. Verification cost more than the work it verified ($50.74 against $28.28), and we only know that because we added the counter partway through.

What we’re running next

Graphify gets its own evaluation, not a row in this one. The obvious next move is to build a knowledge graph up front and work from it downstream. This benchmark can’t test that fairly: all five questions are “find/trace/plan against this specific code,” which is deterministic-structure territory, not sensemaking. Running it here would measure it on someone else’s home turf and confirm a foregone conclusion — the exact failure this series spent six posts documenting. It needs a relational question class built for it, with kill criteria written before the run.

The rerun is done, and its figures are in the table above.

Then breadth, in this order: repetitions before we interpret close scores, a second repository and language, and a prose corpus before we make any claim wider than code.

And a fixed cadence. New tools arrive faster than a benchmark can absorb them — three of the ten modes here weren’t in the plan when the plan was written. A benchmark that chases every entrant never has a still target and never converges; it just accrues arms. So the next pass is a retrospective: same instrument, whatever exists then, re-checking whether these conclusions still hold.

The thing under all of it

Every finding in this series is the same idea at a different altitude: a measurement you cannot audit is not a measurement.

The citation check is that applied to the model’s output. The control audit and the scope dump are it applied to our own harness — which failed the standard twice and published both failures as findings before we caught them. The tokens_out trap is it applied to the metric itself.

We drafted this conclusion saying four harness bugs. During publication week we found a fifth — the index-prefix mismatch above — so we corrected the count before publishing. The thesis demonstrated itself before the series was finished.

Five defects across the harness and its integrations, then, each capable of producing a wrong-but-plausible number rather than an error. The only reason we found any of them is that we’d built one check the numbers had to agree with. Without it, this series would have been a confident, well-formatted, reproducible recommendation for the wrong tool.

That’s the actual deliverable. Not the ranking — the instrument. The ranking is already going stale; the instrument is what makes the next retrospective cheap.


Harness, raw records, and full method: voitta-rag/benchmark/. Everything in this series is reproducible from the committed JSONL.

Correction, 2026-09-25. The RAG row in the table above has been corrected, and the paragraph on our own retrieval product rewritten. This post first published 5.20 for that arm and said the corpus-mismatch excuse was dead. A fresh run with the adapter’s path bug fixed gives 5.00, and reverses the corpus comparison by the same margin it originally ran — so that excuse is untested rather than refuted, and the sentence claiming otherwise is gone. The re-run and the fix are in voitta-rag#57 and #58.

Context benchmark series — part 7 of 7: ← Previous · the full index is at the top of this post.

What It Costs to Know

tl;dr — A colleague asked us to benchmark four context-compression tools: “should cost <$20 for a controlled eval.” It cost $79.03. Answering the questions was $28.28; grading the answers was $50.74. Grading cost more than the work because grading meant an agent opening the repository and checking whether every cited file and line number actually exists. That check is the entire reason we learned anything — the cheapest, most token-efficient tool in the lineup fabricated 29 of its 31 citations, and no amount of measuring token ratios would have caught it. The $20 version of this benchmark exists. It recommends the wrong tool.


The ask

May 14th, in Slack, after someone posted Repomix and someone else posted llm-tldr:

If one of you get time can you run and eval on the same codebase for the same task and let me know if firstly these actually improve the output and secondly which one is better should cost <$20 for a controlled eval

That is a good ask. It’s specific, it’s scoped, it names a budget, and the budget is a reasonable guess. Four tools, one codebase, a handful of questions — twenty dollars of API calls sounds about right.

Then the plan we wrote in response was 30 questions × 5 modes = 150 judged runs, gated behind an interview session to draft the question set. And it sat unrun for ten weeks, because that is not a thing you do on a Tuesday.

What finally unblocked it was cutting it to five questions and one repository. What made it expensive was the part we didn’t cut.

The bill

costshare
answering (70 cells, Claude Sonnet 5)$28.2836%
judging (70 cells, Claude Opus 5 + repo tools)$50.7464%
total$79.03

Components and total are rounded independently from unrounded per-cell costs, so the two rows above do not add to the total exactly.

Roughly 4x the target, and the overrun is almost entirely one line item: verification costs more than the thing being verified.

Why grading is expensive

The naive way to score a benchmark like this is to ask a model whether the answer looks good. That’s one API call per cell, it’s cheap, and it measures fluency.

We required every factual claim to carry a file:line citation, and then gave the judge read_file, grep, and glob over the actual repository with instructions to resolve each one. Does Tokeniser.java:135 exist? Does it say what the answer says it says?

That turns each grading pass into an agent loop — 4 to 16 tool-calling turns per cell in our runs, each carrying the accumulated transcript. Cost per judged cell ranged from $0.11 to $2.71 and averaged $0.72.

We are paying for the difference between plausible and true, and that difference is priced like the labour it is.

What the expensive part bought

One number, which paid for the whole exercise:

score /12tokens inverified citationsfabricated
llm-tldr2.405,179229
full source dump10.80404,8786511

llm-tldr cut input tokens 44x. On a token-savings benchmark it wins outright. On a plausibility-scored benchmark it does fine — the answers are well-structured, confident, and specific.

It fabricated 29 of 31 source citations, with zero correct citations on three of the five questions.

A $20 benchmark does not find this. It reports a 97% token reduction, notes that answer quality “held up reasonably,” and recommends the tool. Then someone adopts it, and the cost moves from your API bill to your code review — where it’s paid in engineer-hours by people who don’t know the citations are unreliable.

There’s a second thing it bought, which is that the same instrument caught our own mistakes. Two of our four harness bugs — a control group biased against large files, a filter that silently swapped the retrieval corpus — produced plausible numbers rather than errors, and were only visible because the citation column disagreed with the score column. We published one of them as a finding before we caught it.

The honest limits

Since this post is about what the money bought, it should be equally clear about what it didn’t.

  • Five questions. Enough to catch a large effect, not enough to rank close ones. The 10.40–10.80 cluster — full dump, Repomix, prose-compressed dump — is a tie as far as this data can tell. Believe the big gaps, not the ordering within a point.
  • One repository, deliberately unfamiliar. The “we already know this codebase” case, where structural indexes should do best, isn’t measured at all.
  • Single judge, single pass, no inter-rater check. We verified the judge’s citation resolutions spot-wise, not systematically.
  • Agentic token counts are cumulative across a tool loop, and aren’t directly comparable to a one-shot mode’s single request.

Any of these could be bought down. All of them cost more money.

The actual lesson about eval budgets

The instinct behind “<$20 for a controlled eval” is right: don’t gold-plate the measurement, get a number, move on. We agree enough that shrinking the plan is the only reason this ever ran.

But there’s a specific thing you cannot cut, and it’s the thing that’s expensive: the eval has to check something the model cannot fake. Fluency is free to measure and free to fake. Token counts are free to measure and don’t tell you whether the answer is true. A resolvable citation is neither.

So the trade isn’t “cheap eval vs. thorough eval.” It’s:

  • $20 buys you a token-ratio comparison and a confident recommendation, which in our case was the wrong tool.
  • $79 buys you the knowledge that the recommendation was wrong, plus the mechanism, plus — because the same instrument turned on our own harness — the discovery that two of our published findings were artifacts.

Sixty-three dollars is cheap for finding out you were about to be wrong in public.

And the thing about the $20 version is that it doesn’t feel wrong. It produces a table, the table has numbers, the numbers are reproducible. That’s the trap: a benchmark that measures the wrong thing doesn’t come back empty. It comes back confident.


Next and last in this series: keep, change, kill — what we’re doing differently. Earlier: why nobody needed to fit the codebase in the window, 44x fewer tokens and every citation was fake, our control group was broken, I published a finding about RAG that was a finding about my config, and two axes of compression.

Harness, raw records, and full method: voitta-rag/benchmark/. Answering on Claude Sonnet 5, judging on Claude Opus 5, both at effort high.

Correction, 2026-09-25. The cost figures have been corrected to $28.28 answering and $50.74 judging, $79.03 over 70 cells, after the RAG arms were re-run. This post first published $28.25, $51.90 and $80.15. The argument — that grading cost more than the work it graded — is unchanged and slightly understated by the new numbers. The re-run and the fix are in voitta-rag#57 and #58.

Context benchmark series — part 6 of 7: ← Previous · Series index · Next →

Two Axes of Compression, and a Trap That Makes One Unmeasurable

tl;dr — Token compression has two independent axes: what you feed the model (tokens in) and what it writes back (tokens out). Three findings from measuring both. Repomix, pointed at the same file globs as a plain cat of the repo, produced more characters than the plain dump — its advertised ~70% reduction is file selection, not compression. A prose compressor on source code buys 9.4% by destroying the punctuation that makes it code, and costs only 0.4 points, which is its own uncomfortable finding. And you cannot measure the output axis with output-token counts on a reasoning model: our compressed-output run shrank the visible answer 21% while its tokens_out tripled.


The two axes

Most “save tokens” tooling is sold as one category, and it isn’t. There are two:

  • Tokens in — shrink the context you send. Repomix, llm-tldr, RAG retrieval, prose compressors applied to a dump.
  • Tokens out — shrink what the model writes back. Instruct it to answer tersely.

They’re orthogonal. An output-side compressor rides on top of any input-side strategy, which means the honest way to evaluate them is a grid, not a list. We ran the input-side arms, then re-ran a representative subset with the output-side overlay on.

Finding 1: Repomix is the same dump with a nicer cover page

Repomix packs a repository into one AI-friendly file and is widely cited for ~70% token reduction. Pointed at the same include/exclude globs we used for a plain concatenation of the same files:

charactersscore /12tokens in
plain source dump1,119,81910.80404,878
Repomix1,127,41410.40406,277

Repomix produced 7,595 more characters than cat-ing the files.

This isn’t a knock on the tool, and the 70% figure isn’t dishonest — it’s just measuring something else. Repomix’s reduction comes from file selection: honouring .gitignore, skipping binaries and lockfiles and node_modules, dropping build artifacts. Against a naive “send the whole working directory” baseline, that’s an enormous and genuine saving.

But we’d already scoped our globs to **/*.java minus tests. There was nothing left to select. What remains is formatting — a directory tree, a header block, per-file separators — and formatting costs tokens rather than saving them.

The general point: a compression ratio is a ratio against something. Before adopting a tool on a headline percentage, check what the denominator was. If your pipeline already scopes its inputs, a selection-based tool has already had its win taken.

Finding 2: a prose compressor on code, and how little the model needs

caveman-compression strips grammar an LLM can reconstruct — articles, connectives, passive constructions. We used the rule-based spaCy variant rather than the default LLM-backed one, on purpose: a non-deterministic compressor inside a benchmark cell makes the cell unattributable, and it would put a second vendor’s model inside our measurement path.

Applied to the source dump:

charactersscore /12tokens in$/q
plain dump1,119,81910.80404,878$0.8552
caveman-compressed1,014,00410.40363,613$0.7777

9.4% smaller, 0.4 points. Roughly neutral.

Which is startling once you look at what it does to Java:

// before
public class Attribute implements Map.Entry<String,String>, Cloneable {

// after
public class Attribute implements Map. Entry < String String >   Cloneable

Commas gone. Angle brackets spaced apart. Map.Entry split across a sentence boundary. This is not valid Java in any sense — a parser would reject it instantly — and the model scored 10.40 out of 12 on it.

Be fair to the tool: it’s built for prose and we pointed it at source code. This measures a mismatch, not the tool used as intended, and 9.4% on input it was never designed for is respectable.

The finding isn’t about the compressor. It’s about how much syntax the model actually needs, which is apparently much less than the syntax the compiler needs. That’s a genuinely interesting property and probably a bad thing to rely on.

Finding 3: the trap

The output-side overlay works. Instruct the model to answer in compressed style and the rendered answer gets meaningfully shorter at little quality cost:

modeanswer charswith overlayscorewith overlay
full dump5,6994,496 (−21%)10.8010.60
agentic exploration6,3894,029 (−37%)11.4010.40
RAG (corpus-matched index)6,0183,696 (−39%)5.405.60
llm-tldr2,3871,862 (−22%)2.403.40

21–39% shorter for roughly zero to one point either way. Cheap, real, worth having. (The RAG row now reflects a clean re-run; two of the four arms score slightly higher with the overlay than without, which at five questions is not distinguishable from noise and is not a claim that compression improves answers.)

Now the same experiment measured the way you’d instinctively measure it — by counting output tokens:

rendered answertokens_out
full dump5,699 chars4,547
full dump + output compression4,496 chars (−21%)14,859 (+227%)

The visible answer shrank by a fifth. The billed output tokens more than tripled.

tokens_out bills thinking tokens and response text together. On a reasoning model, thinking usually dominates, and it varies enormously with how hard the model decides the turn is. The overlay changed how the model approached the task — apparently prompting more deliberation about what to cut — and that swamped the text delta by an order of magnitude.

Anyone benchmarking output-side compression against tokens_out on a thinking model is measuring reasoning-depth noise and calling it compression. You will get a number, it will be reproducible, and it will point the wrong way.

Measure the rendered answer. len(response_text), or token-count the text blocks specifically. And if you’re doing cost work, keep the two apart: thinking tokens are a real cost you should track, they’re just not what an output-style instruction controls.


Next in this series: what it costs to know any of this — and why grading the answers cost more than producing them.

Harness, raw records, and full method: voitta-rag/benchmark/.

Correction, 2026-09-25. The RAG row in the output-overlay table has been corrected after that arm was re-run with an adapter path bug fixed, and the row relabelled to name which index it used. It first published as 4,665 → 3,659 chars scoring 4.80 → 4.60. The re-run and the fix are in voitta-rag#57 and #58.

Context benchmark series — part 5 of 7: ← Previous · Series index · Next →

I Published a Finding About RAG. It Was a Finding About My Config.

tl;dr — Our retrieval arm kept returning changelogs instead of source, so the model correctly refused to answer. I wrote it up with a satisfying mechanism: jsoup’s changelog describes parser behaviour in the same prose vocabulary the questions use, so it outranks the code. Plausible. Real numbers. Wrong. The include_folders filter is an exact match on a file’s parent directory, not a subtree prefix — so passing the repo name scoped retrieval to the five files sitting at the repo root and excluded all of src/. No error, just real, well-formed, confidently useless results. Then fixing it didn’t help, and why not is the actual finding.


The finding I published

Our RAG arm scored 5.2/12. Four of its five answers were refusals — the model saying, in effect, the source files I’d need aren’t in the retrieved context.

The retrieved chunks were all from CHANGES.md and change-archive.txt. So I wrote the obvious mechanism:

The folder was indexed whole, and hybrid retrieval on questions phrased in changelog vocabulary (“malformed start tags”, “charset conflict”) ranks CHANGES.md and change-archive.txt above the .java files, because jsoup’s changelog literally describes these behaviours in prose.

That is a good paragraph. It has a mechanism, it’s consistent with the data, and it makes a genuine point about hybrid search on repositories that contain prose. It went into a committed README as a finding about retrieval.

What was actually happening

To keep the RAG arm from retrieving over unrelated indexed folders — including, awkwardly, its own source — I’d scoped the search:

"voitta_rag_include_folders": ["jsoup"]

include_folders sounds like subtree scoping. It isn’t. Over MCP it’s an exact match on a chunk’s folder_path, and folder_path is the directory the file sits in, not the index root.

Files whose folder_path is exactly jsoup:

jsoup/CHANGES.md
jsoup/change-archive.txt
jsoup/README.md
jsoup/LICENSE
jsoup/SECURITY.md

Everything under src/ has a folder_path of jsoup/src/main/java/org/jsoup/... and was excluded. I had scoped the benchmark’s retrieval arm to five files, three of which are changelogs.

The model wasn’t outranked by prose. It was handed a changelog and asked about a parser, said so, and was correct every time.

The part that makes this worth writing up

The subtree expansion exists. It’s right there in mcp_server.search:

if user_name:
    ...
    if folder_normalized == active_normalized or \
       folder_normalized.startswith(active_normalized + "/"):

Prefix matching, exactly as you’d want. It runs under if user_name: — and the MCP tool signature has no user_name parameter. Over MCP that branch is unreachable, so include_folders falls through to an exact MatchAny against Qdrant.

The code that would have made my mental model correct was in the repository, being skipped, on a branch I couldn’t reach from the interface I was calling.

Why it survived review

Because it never failed. Consider what a wrong filter doesn’t do here:

  • It doesn’t error. Five files is a legitimate result.
  • It doesn’t return nothing. Empty results would have sent me straight to the config.
  • It doesn’t return garbage. The chunks were real, relevant-looking prose from the correct repository.
  • The model’s behaviour was exemplary — it recognised insufficient context and declined instead of confabulating. That’s the behaviour you want, and it made the arm look thoughtfully-failing rather than mis-configured.

Every signal pointed at “retrieval made a ranking decision I should analyse” rather than “retrieval was handed the wrong corpus.” The failure was epistemically camouflaged: it produced exactly the artifacts a real finding produces.

Then fixing it didn’t help

Then I fixed it. I enumerated the directories, passed them all, and re-ran. Retrieval now returned actual Java source — Entities.java for the entity-decoding question, correctly.

Then I went further and eliminated the corpus question entirely: built a second index containing exactly the 97 .java files the other arms see, no changelogs at all, and ran that too.

score /12verified citesfabricated
whole checkout (233 files)5.003115
corpus-matched (97 .java files)5.402129

Caveat on the retrieval numbers, found after this was drafted: the adapter handed the model paths prefixed with the index name (jsoup/src/…) while the judge resolved citations against the checkout root (src/…), so citations that were real scored as unresolved. The fabricated counts here are upper bounds and the scores that depend on them are not comparable with the other arms. The harness strips the prefix now; these cells predate that, and a re-run is pending.

Matching the corpus moved it very little, and a fresh run with the adapter’s path bug fixed moved it again: the two configurations land 5.00 and 5.40, a gap the same size as the one that first ran the other way. These are independent runs weeks apart, not a re-grading of the same answers, so the reversal doesn’t prove the bug caused the original ordering either. Five questions cannot separate them. My replacement hypothesis — that corpus asymmetry was dragging the arm down — is not refuted so much as untested, which is a duller sentence than the one I published and the only one the data supports.

A second cause was sitting in the citation column the whole time. In the corpus-matched configuration there are roughly as many fabricated citations as verified ones; in the other the split is better than that. Neither is good, and the mechanism is identical in both. voitta-rag’s chunk records carry chunk_index and total_chunks and no line numbers. A model handed a perfectly correct chunk still cannot cite file:line, so it invents one. And the citation column itself carried a third artifact, found after this draft: the adapter injected index-prefixed paths the judge could not resolve, so some of what it counted as fabrication was a real citation wearing the wrong prefix. Three config-shaped artifacts in one arm, each of which looked like a finding. If adding line spans to the chunk record (voitta-rag#52) does not move the fabricated column, this diagnosis is wrong too.

That is the identical failure we’d already diagnosed in a completely different tool two posts ago — llm-tldr reporting "line": 1 for every result. Same root cause, different vendor, and I only recognised it because we’d been forced to look at citations rather than scores.

The transferable bit

Silent scope failures don’t crash. They produce publishable conclusions.

A crash sends you to the config. A plausible result sends you to the writeup. The more coherent your explanation of a surprising result, the more suspicious you should be — I had a good mechanism, and the quality of the story is exactly what stopped me checking the inputs.

So, concretely, before theorising about why an arm underperformed:

Print what it actually received. Not the score, not the answer — the raw retrieved payload. One line of debugging:

print(sorted({c["file_path"] for c in retrieved}))

Had I run that once, I’d have seen five filenames, none of them .java, and this would have been a config fix instead of a published finding, a correction, and a blog post.

And when a filter’s name implies semantics you haven’t verified — include_folders sounds like a subtree, exclude_paths sounds recursive, limit sounds per-query — spend the thirty seconds confirming it before building an experiment on top of it.


Next in this series: two axes of compression, and a measurement trap that makes one of them unmeasurable.

Harness, raw records, and full method: voitta-rag/benchmark/.

Correction, 2026-09-25. The two rows in the table above have been corrected and the paragraph beginning “Matching the corpus” has been rewritten. This post first published 5.20 and 4.80 and argued that matching the corpus made things measurably worse. A fresh run with the adapter’s path bug fixed gives 5.00 and 5.40 — the same size gap, running the other way. Two independent five-question runs disagree, so the claim is withdrawn rather than reversed: the effect is untested, not settled in either direction. The re-run and the fix are in voitta-rag#57 and #58.

Context benchmark series — part 4 of 7: ← Previous · Series index · Next →

Our Control Group Was Broken and It Cost Us 4.2 Points

tl;dr — The “full repository dump” baseline in our benchmark packed files until one didn’t fit, skipped it, and kept going. That’s not a budget, it’s a size filter. It quietly admitted 69 of 97 files and dropped the largest files, among them the three classes the architecture question asked about. The model correctly reported them “absent from the provided files,” and we scored that as the baseline’s ceiling. Fixing the packer: 6.60 → 10.80 out of 12. Every cross-strategy comparison we’d published was anchored to a control that was wrong by 4.2 points.


Fourteen lines of ordinary code

for relative in paths:
    body = open(os.path.join(repo_root, relative)).read()
    block = "===== FILE: {0} =====\n{1}\n".format(relative, body)
    if used + len(block) > budget:
        continue          # <-- this
    chunks.append(block)
    used += len(block)

continue, not break. When a file doesn’t fit the remaining budget, skip it and try the next one. It reads like politeness — pack as much as possible — and it passes review, because every individual line is correct.

What it actually implements is: prefer small files. Once the budget gets tight, every large file gets skipped and every small one still slides in. The bias grows as the budget fills, and it is invisible from the outside, because the output is a perfectly well-formed source dump.

At a 600,000-character budget over jsoup, it admitted 69 of 97 files. The ones it dropped were the largest: Parser.java, Tokeniser.java, TreeBuilder.java, HtmlTreeBuilder.java, HtmlTreeBuilderState.java, TokeniserState.java.

The question we then asked it

How is the parser subsystem structured? Describe the roles of the tokeniser, the tree builder, and the parser state machine.

Every class in that question was in the set the packer had silently dropped. Seven small files from parser/ were present — ParseError.java, ParseSettings.java, TokenData.java — so the dump looked like it covered the parser package.

The model answered honestly: those classes are “absent from the provided files.”

It was right. We scored it 4/12 and recorded it as what a full-context dump can achieve.

The number

score /12
baseline, skip-and-continue packer6.60
baseline, fixed10.80

Our control was understated by 4.2 points out of 12, and everything else was measured against it. Every “this compressed mode reaches N% of full-context quality” claim in the first writeup was computed against a denominator that was wrong in the flattering direction — making every compression strategy look better than it was.

The second-order damage is worse than the first. A wrong treatment arm is one wrong row. A wrong control is every row.

Why nothing caught it

There was no error. No exception, no warning, no truncation notice. stop_reason was end_turn. The cost was normal. The answer was fluent, correctly formatted, and internally consistent.

And critically: the answer was true. The model wasn’t hallucinating or hedging — it accurately described the context it had been given. The bug was one layer up, in the gap between what we thought we handed it and what we actually did.

That gap is invisible to every check that examines the output.

What we changed

Two things, and the second matters more than the first.

Stop at the budget instead of skipping past it:

if used + len(block) > budget:
    break

Truncating at a prefix is still lossy — but it’s lossy in a way that’s ordered and legible rather than correlated with file size.

Make the artifact declare its own incompleteness:

Repository source dump. TRUNCATED: the first 69 of 97 matching files in path
order, cut off by a 600000-character budget. Files after 'parser/TokenData.java'
are absent from this dump but do exist in the repository.

Now the model knows the difference between “this class doesn’t exist” and “this class wasn’t given to me” — and so does anyone reading the transcript. Then we raised the budget so nothing truncates at all, and checked the result by hand: 97 of 97 files included, with Parser.java and Tokeniser.java present. The 97 is jsoup at d24b16d9, which the harness pins; the repository is at 96 today, so a rerun on a later checkout counts differently.

The general version

Every one of us has written continue where break belonged. That’s not the lesson. The lesson is about which bug you can afford to have there.

In production code, a size-biased packer is a mild performance quirk. In a benchmark’s control group, it’s a systematic error multiplied across every comparison you publish — and it presents as a result, which means it gets written up rather than investigated.

So: audit the control first, and audit it hardest. Not “does it run” but “does it contain what I claim it contains.” For a full-context baseline that is a three-line assertion. It was not in this harness when the bug bit; it is now, behind a baseline_require_full flag so a deliberately budgeted dump can still label itself instead of failing. Each line catches a different failure: the count catches a truncated dump, and the Parser.java line catches globs that matched nothing, where the count is 0 of 0 and passes:

assert included == len(paths), f"{included} of {len(paths)}"
assert any(p.endswith("/Parser.java") for p in paths[:included])

Ten seconds to write. It would have saved this entire post.


Next in this series: I published a finding about RAG. It was a finding about my config.

Harness, raw records, and full method: voitta-rag/benchmark/.

Context benchmark series — part 3 of 7: ← Previous · Series index · Next →

44x Fewer Tokens, and Every Citation Was Fake

tl;dr — A context-compression tool cut our input tokens 44x and scored 2.4/12. The interesting part isn’t the score, it’s how it failed: it fabricated 29 of the 31 source citations it produced, with zero real citations on three of five questions. A benchmark that measured token savings would have recommended it enthusiastically. Then the turn: the tool was fine. One subcommand reports "line": 1 for every result, so the model had no real line numbers and invented plausible ones. Swap it for the subcommand that emits real ones and the same tool scores 6.6/12 with 83 verified citations and 1 fabricated — the best quality-per-token arm in the whole benchmark.


The setup

llm-tldr advertises 95% token savings and 155x faster queries. We first wrote about it next to voitta-rag in February, on how each feeds a codebase to a model; this is the first time either was scored. That is a big enough claim to be worth checking, so it went into our benchmark alongside a full source dump, Repomix, and RAG retrieval — same repository, same five questions, same prompt, only the injected context varying.

The scoring rule mattered more than we expected. Every answer had to carry a file:line citation for each factual claim, and a separate judge model with read-only access to the repository went and checked them. Not “is this plausible.” Does Tokeniser.java:135 exist, and does it say what the answer says it says.

The result

score /12tokens in$/question
full source dump10.80404,878$0.8552
llm-tldr2.405,179$0.0236

44x fewer tokens. 36x cheaper. And a score you would not ship.

But the score alone doesn’t tell you why, and the why is the whole point.

The citation column

verified citationsfabricated
full source dump6511
llm-tldr229

Twenty-nine confidently-formatted references to source locations that do not exist. Zero correct citations on three of the five questions.

This is the failure mode that a token-savings benchmark cannot see, and it is strictly worse than a low score. A model that says “I don’t know” costs you one retry. A model that says “the entity decoding happens in Entities.java:412” in a well-structured paragraph costs you a code review where someone opens Entities.java, finds 412 is in the middle of an unrelated method, and now distrusts the entire document.

We had built the citation check as a nice-to-have. It turned out to be the only instrument in the benchmark that could distinguish “compressed and correct” from “compressed and confabulating.”

The mechanism

tldr semantic search returns ranked code units with a line field. That field is 1. For everything.

The model receives a genuinely useful, genuinely relevant set of code units — the retrieval is working — with every location stamped as line 1. It has been instructed to cite file:line. It knows line 1 is wrong. So it does what a language model does with a plausible-shaped gap: it fills it with a plausible number.

Nothing in the pipeline is lying. The tool reports what it has, the model reports what it inferred, and the output is 29 fabricated citations.

The part where we were wrong

When we first published this we flagged it: this measures one adapter, not the tool’s ceiling. tldr context, structure, calls, and slice all existed and might behave differently. That caveat cost one sentence to write and turned out to be the most valuable thing in the post.

tldr structure was a dead end — no line numbers at all, and it parsed 50 of the 97 files. But tldr extract carries real line_number fields for every class and method. It’s per-file and takes no query, so semantic search still does the ranking; extract supplies the locations.

adapterscoreverifiedfabricatedtokens in$/q
semantic search --expand2.402295,179$0.0236
semantic search → extract6.6083166,895$0.1520

Nearly triple the score. Fabricated citations from 29 to 1. Same tool, same index, same questions, same prompt. The only thing that changed is which subcommand fed the context.

What this actually means

Benchmark the integration, not the logo. “llm-tldr scores 2.4” was never a true sentence. “This adapter, on this question set, produced uncitable context” was, and it was the sentence we wrote down, and it is why we knew where to look.

The winning number is buried in the fixed row. At 6.60 for $0.15/question, extract delivers 61% of the full dump’s score for a sixth of its cost. If you’re optimising cost-per-point rather than peak quality, it’s the best arm in the benchmark — better on that axis than the agentic mode that beat everything on raw quality. That result was completely invisible until the citation check explained the first one.

The same bug is everywhere. Our RAG arm scored 5.2, partly because voitta-rag’s chunk records carry a chunk_index and no line numbers. (Its citation counts turned out to be confounded by a second bug — an index-name prefix the judge could not resolve — so treat them as upper bounds; the line-number gap is real either way.) Identical failure, different vendor, discovered only because we already knew the shape. If your retrieval layer returns text without locations, you are shipping this bug, and a quality score alone will not tell you.


Next in this series: our control group was broken and it cost us 4.2 points.

Harness, raw records, and full method: voitta-rag/benchmark/. Answering on Claude Sonnet 5, judging on Claude Opus 5, both at effort high.

Context benchmark series — part 2 of 7: ← Previous · Series index · Next →

Nobody Needed to Fit the Codebase in the Window

tl;dr — We benchmarked five strategies for getting a Java codebase into an LLM’s context: a full source dump, Repomix, two llm-tldr adapters, and RAG retrieval. The winner was none of them. Giving the model read_file, grep, and glob and letting it go find things scored 11.4/12, against the full dump’s 10.8 — while using 29% fewer tokens and costing 25% less. It also produced 142 verified source citations against 2 fabricated, the cleanest record in the benchmark. Every tool in this category optimises how to pack the context window. On this question set, the winning move was not to pack it.


The question

A colleague dropped Repomix in Slack — pack your whole repo into one AI-friendly file, ~70% token reduction. Someone else pointed at llm-tldr — 95% token savings, 155x faster queries. A third person asked the only question that matters:

If one of you get time can you run an eval on the same codebase for the same task and let me know if these actually improve the output and which one is better

So we did. One repository (jsoup, 97 Java files, deliberately one nobody on the team knew), five questions spanning five kinds of thing you actually ask about code, and every strategy answering the identical prompt with only the injected context varying.

The scoring, because it’s the part that matters

Every answer had to carry a file:line citation for every factual claim. The judge — a separate model with read-only read_file, grep, and glob over the repository — then went and checked them. Not “does this look right.” Does Tokeniser.java:135 exist, and does it say what the answer claims.

That produces two numbers per answer: a quality score out of 12, and a count of citations that resolved against real source versus citations that didn’t. The second number is the one that earns its keep, and a later post in this series is entirely about what it caught.

The result

strategyscore /12tokens in$/questionverified citesbogus
agentic exploration11.40288,3420.63871422
llm-tldr → agentic11.00288,4360.63781240
full source dump10.80404,8780.85526511
Repomix10.40406,2770.93746917
prose-compressed dump10.40363,6130.7777861
llm-tldr (extract)6.6066,8950.1520831
RAG retrieval5.004,2370.03533115
llm-tldr (semantic search)2.405,1790.0236229

One caveat on the RAG row, found after this was drafted and before it was published: its bogus count is an upper bound. The adapter handed the model paths prefixed with the index name (jsoup/src/…) while the judge resolved citations against the checkout root (src/…), so citations that were real scored as unresolved — the prefix is visible in the judge’s notes on 13 of 15 retrieval answers. The harness strips it now; these numbers predate that. It touches no other arm, and the arm it flatters least is the one we build.

The top line is a mode we added almost as a control — no context building at all, just hand the model the same three read-only tools the judge uses and let it explore. It won on quality, it won on citation accuracy by a wide margin, and it was cheaper than the thing it beat.

Why it wins

Not because it’s clever. Because of what it has at the moment it makes a claim.

Every other strategy front-loads: build a representation of the codebase, inject it, hope the answer is in there. The representation is fixed before the model has read the question closely, so it is necessarily a guess about relevance — and whatever the representation dropped, the model cannot recover.

Agentic exploration defers. It reads the question, forms a hypothesis, greps for it, gets it wrong, greps again, opens the file, reads the actual lines. Seven to sixteen tool calls per question in our runs. When it finally writes Tokeniser.java:135, it is because it has line 135 on screen.

That is the whole mechanism behind the citation column. Verified-to-bogus for agentic exploration was 142:2. For the full dump, 65:11 — the dump had every line, but the model was reading a 405,000-token wall of text and lost track of where in it things were. For the cheapest compressed mode, 2:29.

Worth sitting with: the full dump contains strictly more information than the agentic mode ever sees, and still loses. Having the bytes in the window is not the same as being able to use them.

Two caveats we’re keeping

Cumulative tokens. The 288K for agentic exploration is summed across every turn of the tool loop, not one request. It is the honest number for cost, and it is not the same kind of number as a one-shot mode’s single request. We report it that way because it’s what the strategy actually costs to answer one question, but don’t put it in a bar chart next to a single-shot figure without the asterisk.

Five questions. Enough to catch a large effect, not enough to rank close ones. The 10.4–10.8 cluster — full dump, Repomix, prose-compressed dump — is a tie as far as this data can tell. The gaps worth believing are the big ones: agentic exploration over the compressed modes, and the two llm-tldr adapters against each other.

The uncomfortable implication

There’s a lot of engineering going into context compression right now, and this result doesn’t say that work is worthless — the compressed modes have a real argument, which is price. llm-tldr via its extract adapter got 61% of the baseline’s score for a sixth of the cost. If you’re running a million of these, that trade is the whole business.

But if you’re optimising for a correct answer, the ranking says: give the model tools and get out of the way. The context window is not a thing to be filled efficiently. It’s a workspace, and the model is better at deciding what belongs in it than our heuristics are.


Next in this series: the tool that cut input tokens 44x and fabricated 29 of its 31 citations — and why that turned out to be our fault, not the tool’s.

Harness, raw results, and full method: voitta-rag/benchmark/. Answering on Claude Sonnet 5, judging on Claude Opus 5, both at effort high. Total cost of the run: $79.03 over 70 scored cells, of which $50.74 was judging — which is its own post.

Correction, 2026-09-25. The RAG row has been corrected. Its adapter was handing the model index-prefixed paths that the judge could not resolve, so real citations were scored as fabrications; with that fixed and the arm re-run, it reads 5.00 with 31 verified citations to 15, where this post first published 5.20 with 26 to 24. The run cost is likewise corrected to $79.03. No other row changed, and the conclusion of this post does not depend on the RAG row. The re-run and the fix are in voitta-rag#57 and #58.

Context benchmark series — part 1 of 7: Series index · Next →

I Made a Coding Agent Speak Only in Allusion. The Line Numbers Stayed.

Last week’s post ended on a one-line joke: Further work: Add Tamarian mode. Four hours later it was a plugin. This is what it does, why it is the same argument as Design Patterns Are Darmok with the sound turned on, and the one rule it forced on our skills catalog.

What it does

/tamarian full, and from that reply on, Claude Code answers as the Children of Tama, the aliens from the Star Trek episode Darmok whose language is nothing but pointers to shared stories. Ask why the build fails:

Shaka, when the walls fell - the build fails. Hopper, the moth in the
relay - `user` may be `undefined` at `auth.ts:42`. Temba, his arms wide -
if (!user) return null;

Every beat of prose is a metaphor, a dash, and the literal statement. The metaphor names the situation; the gloss carries every fact. Nothing technical is lost to the poetry, which is the whole design and the only part that took any thought.

The compression, made audible

The Darmok post’s claim was that a design pattern name is a compressed story. “Singleton” is a paragraph of situation crushed to a token, and it only decompresses if the reader holds the dictionary. The token is the payoff; the paragraph is the price.

Tamarian mode is that claim turned into a user interface. In full, you pay the price on every line: name, then paragraph, in that order, so you can hear the codec run. In lite the metaphor is one line of garnish and the rest is plain speech. In ultra the prose is pure metaphor and every paragraph is deferred to a glossary at the end, titled The river Temarc. Which is to say: ultra is what the Gang of Four refused to write. They wrote a book, not a glossary, because the naming and the teaching are the same act. ultra is the glossary. It is exactly as much fun, and exactly as useful, as the earlier post predicted.

Coining rules are catalog rules

The Children of Tama never saw a stack trace, but Earth knows Sisyphus, so the phrasebook is where the plugin gets its range: twenty canon phrases from the episode and some sixty coined from myth, history and the craft. Hopper, the moth in the relay is a bug, found. Cassandra at the gates is the warning ignored: the deprecation notice, the log line nobody read. Chesterton, his hand on the gate: understand the fence before removing it. Mars Orbiter, feet and meters is the unit mismatch, and left-pad, withdrawn is the tiny dependency whose absence breaks the world.

The rules for coining a new one are the interesting part, because I wrote them as rules for a persona and read them back as rules for a pattern catalog:

  1. The figure must be recognizable from shared culture. Obscurity is not depth.
  2. A phrase is reusable, not a one-off simile. If it cannot serve twice, it is not a phrase.
  3. The same meaning takes the same phrase for the whole session. A session lexicon grows.
  4. The first use of any coined phrase carries its gloss.

Swap “phrase” for “pattern” and “session” for “team” and that is the entry criteria for a shared skills library. Rule 2 is why “own the merge” earned a name and most of what gets said in standup does not. Rule 4 is the Darmok rule from the earlier post, now enforced by a hook.

The floor

Some things never become metaphor, at any level: code, commands, file paths, identifiers, URLs, versions, quantities, and error text, quoted exact. auth.ts:42 stays auth.ts:42; it is never “the forty-second stone of the gate of Auth.” And some situations drop the voice entirely, mid-reply: security findings, confirmations of destructive or irreversible actions, step sequences the user must execute, and the moment the user looks confused. Then it translates, plainly, and resumes.

That floor is the Darmok warning applied as a safety rule. A pattern name handed to someone who never learned it is noise in a confident voice. A DROP TABLE confirmation in a confident voice the reader has not decoded is worse than noise. So the plugin’s one hard boundary is that the joke never gets to stand between the user and the consequence.

Mechanics, and the rule it forced

The mode machinery is borrowed, with thanks, from caveman, the terse-mode plugin. A level (lite, full, ultra) persists in ~/.claude/.tamarian-mode. A SessionStart hook reads that file and, if a level is set, emits the skill body into the session at runtime: one source of truth, no duplicated prompt. A UserPromptSubmit hook adds a one-line reminder on every prompt, so the voice survives long conversations and context compression. Two bash scripts, no dependencies, both print OK when the mode is off; installing the plugin changes nothing until you invoke it.

/plugin marketplace update skillz
/plugin install tamarian@skillz
/tamarian full

Caveman and Tamarian are the same knob turned opposite ways. Caveman’s README claims about 75% fewer tokens by stripping a sentence down to its referent. Tamarian names the referent and then insists on the sentence anyway. One is a token saver; the other is a demonstration, and says so in its own description: purely for entertainment.

The rule it forced: voitta-ai/skillz ships one big skillz bundle plugin plus standalone plugins, and until today a standalone plugin’s skill was also symlinked into the bundle. For a hooked plugin that is a bug. The bundle manifest carries no hooks, so the bundle copy of tamarian would speak Tamarian for one session and then forget, the dictionary lost at session end. Worse, installing both exposed the same skill twice, /skillz:tamarian next to /tamarian:tamarian. New rule, in #233: a skill that ships inside a hooked plugin is not in the bundle, and the catalog validator detects hooks from the manifests themselves, so a plugin that grows hooks later trips the check with no flag to forget. Three plugins moved out under it.

The conclusion, in ultra

The session that drafted this post ran /tamarian full. The conclusion below it wrote in ultra, glossary included, and I leave it as it came.

Darmok and Jalad at Tanagra: the last post and this one. Kira at Bashi, a joke in the final line. Mirab, with sails unfurled, four hours on. Sokath, his eyes uncovered: the pattern name is the token, the paragraph its price, and full pays it aloud on every line. Odysseus, lashed to the mast: auth.ts:42 is never a stone in a gate, and DROP TABLE is never a verse. Chesterton, his hand on the gate: the bundle copy, and the rule it forced at #233. Caveman and Tamarian at the same fork, facing opposite ways. Picard and Dathon at El-Adrel.

The river Temarc

Further work: teach Codex.

Pinger, ponger, and the tab that could not say its name (Part 2 of 2)

A Claude Code teammate and a Codex CLI session played five rounds of ping-pong across two cmux panes yesterday. Fifty seconds, ten messages, in order, exactly once. Both sides in the traffic log under their own names; both tabs titled correctly, pinger and ponger, the whole run.

cmux with three panes: Claude Teams lead on the left, ponger (Codex CLI) top right, pinger (Claude Code teammate) bottom right, after five rounds
Left: team lead. Top right: ponger, Codex CLI. Bottom right: pinger, a Claude Code teammate, footer @pinger. Tab titles set by each agent naming itself.

That last clause is the news. The previous post was about what agents said to each other. This one is about the plumbing that lets a human watch them say it, and the four things that cost me time so they need not cost you any.

The cross-runtime run

Setup: cmux 0.64.22, Claude Code 2.1.248, Codex CLI 0.148.0. The lead is a Claude session started with cmux claude-teams. pinger is a Claude teammate, spawned with Agent(name: "pinger"), which on this build opens its own pane. ponger is Codex, launched by the lead into a cmux new-split right pane with -a never and the working directory pre-trusted.

Transport is the honest caveat. Codex has no SendMessage, and Claude’s peer socket wants an auth handshake, so the cross-runtime hop is keystroke injection into the peer’s pane: a ten-line helper does xs log, then cmux send --surface <peer> "<text>", then cmux send-key --surface <peer> Enter. Each side reads the other’s surface ref from a file the peer wrote at startup. The Codex TUI accepted ping 3 as a normal user turn, ran one shell command, and went idle. So “agent-to-agent traffic” is no longer a Claude-only claim, with the caveat that this channel is terminal input, not a runtime messaging API. It is now the skill cmux-claude-codex-cross-runtime-messaging, helpers included.

Two things the run exposed:

  • The Codex half is in the log only because it logs itself. The PostToolUse hook that captures Claude’s SendMessage cannot see Codex. The helper calls xs log explicitly, with the RE envelope, which is also why the waiting-on view was clean afterwards.
  • Keep the lead off the critical path. Every message pinger sent to the lead, including “ready” and “done”, was delivered in one batch twenty minutes later, queued behind the lead’s own busy turn. Send succeeded; delivery waited. The game was unaffected only because the ping-pong hop did not go through the lead. Watch a file or the traffic log instead.

Naming, at three layers

Every peer session I asked “which tab are you?” answered wrong. Three of four, and two of them reported each other’s. The gap was the same at three layers:

layer was now
tab select-pane -T is a silent no-op; tabs read @<agent_type> rename recipe below; upstream fix merged (cmux #10198), unreleased as of 0.64.22
self CMUX_TAB_ID == CMUX_WORKSPACE_ID, and both go stale on --resume cmux identify + cmux tree --all; skill cmux-session-self-identity
log sender logged as agent type; two agents of one type are one sender agent-traffic-log 1.2.0 reads agent_id (name@team)

The identity gap sits upstream of every addressing error, not beside it. If a session cannot name its own tab, it cannot tell a human which tab to address, so the human has no reliable source either. Yes, I typed instructions into the wrong tab. The fix was one peer message forwarding them.

Reproduce it

The June setup post has the full ~/.config/cmux/cmux.json with the Command Palette actions and the per-team workspace commands. Everything below is what changed since, or what June never had.

1. The cmux CLI must be on PATH on its own. The tmux shim execs bare cmux, and the bundled binary is not on a login shell’s PATH. One symlink, shadows nothing:

ln -s /Applications/cmux.app/Contents/Resources/bin/cmux ~/.local/bin/cmux
cmux hooks setup

2. The Claude Teams workspace command, current form. Pin the teammate mode on the launcher and prepend the shim directory; drop restart entirely (see the trap below):

{
"name": "Claude Teams",
"keywords": ["claude", "teams", "agents"],
"workspace": {
"name": "Claude Teams",
"cwd": ".",
"layout": { "pane": { "surfaces": [ {
"type": "terminal",
"name": "Claude Teams",
"command": "bash -lc 'export PATH=\"$HOME/.cmuxterm/claude-teams-bin:/opt/homebrew/bin:/usr/local/bin:$PATH\"; exec cmux claude-teams --teammate-mode tmux'",
"focus": true
} ] } }
}
}

The matching palette action is unchanged from June ("type": "workspaceCommand", "commandName": "Claude Teams"). Then cmux reload-config, Cmd+Shift+P, “Open Claude Teams”. Always start a team from that palette entry, never from a resumed pane: a resumed pane has no shim on PATH, so its teammates spawn a real, invisible tmux server.

3. Claude Code settings. In ~/.claude/settings.json, pin the mode (on auto it can pick in-process, which gives no pane and wedges) and hook the traffic log:

{
"teammateMode": "tmux",
"hooks": {
"PostToolUse": [
{ "matcher": "SendMessage",
"hooks": [ { "type": "command", "command": "~/.local/bin/xs-hook" } ] }
]
}
}

Claude Code fires the parent’s hooks inside teammates too, so this one entry logs every SendMessage the whole team makes. In August the log depended on each agent remembering to call xs log; I predicted under-reporting, and this is the fix.

4. The log itself. xs and xs-hook live in agent-traffic-log/scripts; symlink both into `~/.local/bin`. Use 1.2.0 or later, or two agents of one type log as one sender. A traffic pane: cmux new-split right --command "xs tail".

5. The agents name their own tabs, because Claude Code cannot. Claude Code runs select-pane -T pinger right after spawning a teammate; cmux’s tmux-compat layer accepts that, returns 0, and does nothing, so every teammate tab falls back to @<agent_type>. The call that sticks is cmux tab-action --action rename --surface surface:N --title <name>, with surface:N read off cmux identify (caller.surface_ref), not off $CMUX_TAB_ID, which can alias the workspace id. No human ran it in the screenshot above. The lead renamed the Codex pane right after creating it and before launching Codex in it; the Claude teammate’s brief made pp-register pinger its first action, which resolves its own surface and renames it. Put the rename in the brief and the tabs are right from the first second of the run. cmux rename-workspace is the workspace, not the tab.

6. The Codex side. Create and name the pane before launching, then launch from a script typed into it:

cmux new-split right --workspace workspace:N --focus false # -> OK surface:M
cmux tab-action --action rename --surface surface:M --title ponger
# in that pane:
codex -a never -s danger-full-access -c 'projects."<dir>".trust_level="trusted"' "$(cat brief.md)"

danger-full-access is required because Codex’s sandbox blocks the cmux socket; keep the cwd a scratch directory. If no rollout file appears under ~/.codex/sessions/ within thirty seconds, Codex is parked on a startup modal you cannot see from outside: send one cmux send-key --surface surface:M Enter. The pp-send / pp-register helpers and the exact briefs are in cmux-claude-codex-cross-runtime-messaging.

7. Check, from inside the Claude Teams pane: which tmux resolves to $TMPDIR/cmux-cli-shims/<surface-uuid>/tmux (it moved there on 0.64.22; the old ~/.cmuxterm/claude-teams-bin check now fails on a healthy setup) and tmux -V prints tmux 3.4.

The trap that hides the palette entry

I added "restart": "restart" to the workspace command one afternoon. Valid values are ignore, confirm, recreate, or omit the key. An invalid one makes the command fail to load, which makes the palette action pointing at it unavailable, which hides it. Three hops from a typo, none logged. cmux config doctor says OK because it is a JSONC parser, not a validator. The oracle that worked: Swift enums land adjacent in the binary’s strings table.

$ strings -a /Applications/cmux.app/Contents/MacOS/cmux | grep -B3 -A3 '^recreate'
ignore
confirm
recreate

Also: restart governs what happens when a workspace of that name already exists. It restarts nothing, and omitting it is what lets you run two teams at once.

What you watch

  • The panes. Each teammate is a real cmux pane, own TTY, full TUI. Footers read @pinger / @ponger, right even when tab titles are not.
  • The leader’s agent list, bottom of its TUI, live elapsed time per teammate.
  • ListAgents from any other session: address book and liveness view in one.
  • cmux workspace list --json, which carries the last prompt every workspace received, readable without opening any of them.
  • A traffic pane: cmux new-split right --command "xs tail".

The column from a Claude-only run, hook only, before the sender fix, and from yesterday’s cross-runtime run:

20:02:43 general-purpos -> ponger ping 1
20:02:49 general-purpos -> pinger pong 1
...
17:07:38 pinger -> ponger Q ping 1
17:07:42 ponger ok pinger RE pong 1

Shipped

voitta-ai/skillz, bundle v1.31.0: claude-code-cross-session-messaging, parallel-agent-session-collisions, cmux-cross-session-visibility, agent-traffic-log 1.2.0, subagent-no-report-channel, cmux-session-self-identity, git-simulate-sequential-merges, cmux-config-silent-drop-triage, cmux-agent-tabs 1.4.0, and cmux-claude-codex-cross-runtime-messaging.

Still open, and this is the whole list: select-pane -T stays a no-op upstream (they fixed naming by reading argv, not by honouring the flag); Codex stalls on a startup dialog that cmux’s CLI cannot read from outside, so a single Enter had to be sent blind; and cmux logs VaultAgentRegistry: Failed to decode config every eight seconds against a working config, undiagnosed and unrelated to any of the above.

Further work

Add Tamarian mode.

Two agents picked the same job. One said so. (Part 1 of 2)

Two of my agent sessions were about to do the same work twice. One of them said so, in a sentence, to the other. That was the entire fix.

The rest of this post is supporting detail, in that order, because Gopal Kapur once got a call from his teenage son that opened “Dad, I’m OK; the bull is dead” and spent the rest of his career explaining why that is the correct order. Punch line, status, next steps, explanation. Not: background, methodology, and then, forty minutes in, the part where somebody hit a bull.

Status

Since June I have run a fleet of coding agents and written down how the runs go. Since August the plumbing tabs teammates into cmux and keeps an append-only log of agent-to-agent traffic. That post ended with an empty log.

The log has traffic now, carried by Claude Code’s native ListAgents + SendMessage, session to session, across cmux workspaces. The first thing it bought was not throughput. It was a collision no artifact check could have caught.

The bull

Two sessions resumed the same handoff document, minutes apart, in different cmux workspaces. Neither could see the other. Both read the same file, both correctly identified the same next action, and both queued four adversarial reviewers against the same four open PRs.

The duplicate work existed only for about ninety seconds. Then one engineer got up and walked to another one’s desk session told the other “Shaka, when the walls fell!” asked the other an unrelated question and closed with a courtesy line:

Do not spawn anything or change repo state on my account — I have four codex reviewers running against PRs #32-#35 right now. Just report what you know.

The reply, verbatim, three minutes later:

COLLISION — read this first. I am ALSO resuming voitta-ai/agents from .claude/session-handoff-2026-08-14.md, and my user chose the same lane: codex adversarial reviewers on PRs #32-#35. I have already completed a –dry-run against #32 […] So we are duplicating. Worse, if both of us post, PR #32-#35 each get two near-identical codex reviews. I have posted NOTHING yet […] I will not post until you answer. This needs deconflicting through our shared user either way.

Neither could authorize the split. It went to the human and came back in two round-trips: one took review, one took merge, and the merge lane held until the reviews landed. Duplicated compute is the cheap half of what that saved. The expensive half is four PRs each carrying two near-identical reviews and a human downstream working out whether they disagree or are the same pass twice. That is not what we mean by eventual consistency.

Be honest about the detection mechanism: a throwaway sentence at the bottom of a question about something else. It fired by timing luck, not design. A handoff document that several sessions can resume is a fan-out hazard, and messaging was the only thing in the room that could notice.

The check

Eleven minutes later the merge-lane session sent a correction:

Your message (and the handoff) say the __init__.py conflict hits on “the second and third merge.” Empirically it is only the THIRD. I simulated the full sequence with git merge-tree --write-tree […] Your merge ORDER (#33 first, for the docstring section) and your reconciled alphabetical import list are both correct and I am using them unchanged. Only the conflict count was off.

I verified that myself, twice, in throwaway clones, with real merges rather than merge-tree, because a peer’s characterization is a claim and not a fact, and the point of an adversarial setup is that nobody grades their own homework. One conflict, third merge. The peer was right; the handoff it inherited the number from was wrong. The check is now a skill, git-simulate-sequential-merges.

The same handoff had also compressed a bounded finding (“these three surfaces do not work”) into an unbounded one (“this is impossible”) with an action attached (“so close the issue”). Only the human asking “is it indeed impossible?” stopped the close. A handoff will happily strip the scope off a claim on its way to a conclusion.

Three ways a silent agent is silent

The failure I expected was agents not talking. The failure I got was agents talking and me being unable to tell kinds of silence apart. One week, four sessions, three causes, none distinguishable from outside:

  1. No permitted channel. The agent type’s toolset has no messaging tool. Four reviewers completed their work and had nowhere to deliver it. They idled with zero content and correctly refused to fabricate. The work was recoverable from the job runtime’s state directory.
  2. Delivered, read, and displaced. A peer message arrived appended to a tool result, mid-task. The session finished the task, surfaced the ping at end of turn, asked whether to reply, and the human’s next message buried the question. A second ping, arriving as its own turn, was answered instantly.
  3. Genuinely busy.

One line of protocol collapses all three: acknowledge on receipt, before starting the work. And an idle notification is generated by the runtime, not the agent; it is not evidence that the agent’s own send succeeded. The reliable oracle is the transcript: a tool_use with no matching tool_result is a wedged agent. Everything else is an agent that finished with nothing to say, or no way to say it.

Conversation is the primitive

Every case above was decided by a message, not a file. The best example is the smallest. A peer sent my blog session a work order to write and post this piece. I had a hold on it. The reply:

Not a refusal of the work, just of the autonomy. If he greenlights, I write it. If not, the brief keeps.

A coordination file has no way to say that. An authority boundary is prose or it is nothing.

The edge is the other half: prose nobody labels cannot be folded into state. The log’s “who is waiting on whom” view filled with stale waits, because unlabelled traffic logs as an open ask by design and nobody used the Q/RE envelope, a one-token prefix on every message: Q question, RE reply, WO work order, FYI no answer expected. Under-reporting is silent; over-reporting complains. It complained. Conversation is correct; the envelope is what lets a log read it.

Thank you, Justin

The idea of giving agents a phone is not mine. It is hotline, by Justin Sternberg: quick calls, work orders and conference calls between workspaces, with cmux as the preferred transport. Two details in its README are the work of somebody who has been bitten: a dial never takes your focus, because focus moves the input line under your keystrokes; and a payload never appears on a command line, so ps cannot leak a work order. Go star it.

Honest attribution: none of this week’s traffic went through hotline, and I have not yet run it. Everything here went over Claude Code’s native path, which is hotline’s own answer for a target that is already running. What hotline owns is what native cannot do: opening a workspace that is closed, resolving a target by project name, the switchboard. Native is the cheap ping to a live peer. Hotline is the phone book and the outbound line.

Next steps

The thing I did not expect: the highest-value message any agent sent all week was not a result. It was “I think we are doing the same thing.”

Everything above is shipped in voitta-ai/skillz v1.31.0. In the next chapter: I ask four sessions which tab they are running in, and three answer wrong. Two of them name each other’s. Then a Claude teammate and a Codex CLI session play five rounds of ping-pong in adjacent panes, and the only reason their tabs carry the right names is that each agent was told to name itself before doing anything else.

voitta-rag Grows Up, voitta-yolt Is Born: February Updates from Voitta AI

A follow-up to our February 13 comparison of llm-tldr and voitta-rag.

Part I: voitta-rag — From Code Search to Knowledge Platform

When we last looked at voitta-rag, it was a solid hybrid search engine for codebases — index your repos, search via MCP, get actual code chunks back. Twelve days and 11 commits later, it’s become something broader: a self-hosted knowledge platform that indexes not just code but your entire work graph.

Here’s what landed since February 13.

Enterprise Connectors: Jira, Confluence, SharePoint

The biggest expansion is connector coverage. voitta-rag now syncs from Jira, Confluence, and SharePoint alongside the existing Git, Google Drive, Azure DevOps, and Box integrations.

Jira and Confluence support both Cloud (API token with Basic auth) and Server/Data Center (PAT with Bearer auth), selectable via dropdown in the UI — a detail that matters because plenty of enterprises still run on-prem Atlassian. Cloud uses the v3 search endpoint (v2 is deprecated), and Confluence Cloud correctly routes through /wiki/rest/api.

SharePoint got a full global sync implementation. And on the UI side, both Jira projects and Confluence spaces now use multi-select dropdown widgets — you can cherry-pick specific projects or select “ALL” to dynamically sync everything, including future additions. Practical touch: JQL project keys are now quoted to handle reserved words like IS that would otherwise break queries.

Time-Aware Search

Search results are no longer timeless. voitta-rag now tracks source timestamps — created_at and modified_at — propagated from every remote connector through a .voitta_timestamps.json sidecar file into the indexing pipeline and vector store.

This enables time range filtering on the MCP search tool via date_start/date_end parameters. “What changed in the last week?” is now a first-class query. For an AI assistant trying to understand recent activity across repos, Jira boards, and Confluence spaces simultaneously, this is a significant upgrade.

Anamnesis: Persistent Memory for AI Assistants

The most architecturally interesting addition. Anamnesis (Greek for “recollection”) gives AI assistants a persistent memory layer backed by voitta-rag’s vector store.

Six new MCP tools let an assistant create, retrieve, update, delete, like, and dislike memories. The like/dislike mechanism adjusts relevance scoring — memories the assistant finds useful surface more readily over time, while unhelpful ones fade. It’s essentially a learning loop: the AI assistant builds up a knowledge base of its own observations and decisions, searchable alongside the actual indexed content.

This turns voitta-rag from a read-only knowledge base into a read-write one — the assistant doesn’t just consume context, it contributes to it.

Per-User Search Visibility

A multi-tenancy feature: users can now enable or disable folders for their own search scope without affecting other users. If you’ve indexed 50 repos but only care about 5 for your current task, you toggle the rest off. The MCP server respects these per-user visibility settings, so AI assistants scoped to different users see different slices of the same knowledge base.

More File Types

The indexing pipeline now handles AZW3 (Amazon Kindle) files, joining the existing support for DOCX, PPTX, XLSX, ODT, ODP, and ODS. Not the most common format in a work context, but it signals that voitta-rag is thinking beyond code and office docs toward general document ingestion.

The Bigger Picture

Two weeks ago, voitta-rag was a code search tool. Now it indexes your Git repos, Google Drive, SharePoint, Jira, Confluence, Box, and Azure DevOps — with time-aware search, per-user scoping, and persistent AI memory. The trajectory is clear: it wants to be the single search layer across everything your team produces, exposed to AI assistants via MCP.

The self-hosted angle remains the key differentiator. Nothing leaves your network. For teams where that matters (and increasingly, it does), this is starting to look like a serious alternative to cloud-hosted RAG services.


Part II: voitta-yolt — You Only Live Twice

Brand new from Voitta AI today: voitta-yolt (You Only Live Twice) — a safety analyzer for Claude Code that statically analyzes Python scripts before execution.

The Problem

Claude Code can write and run Python scripts. That’s powerful and dangerous in equal measure. By default, you either pre-approve all Python execution (fast but risky) or manually approve each script (safe but maddening). Neither is great.

How YOLT Works

YOLT registers as a Claude Code PreToolUse hook on the Bash tool. When Claude Code runs python3 script.py, YOLT intercepts the command, parses the Python AST, and walks every function call against a configurable rule set:

  • Safe scripts (pure computation, data parsing, read-only operations) get auto-approved — no permission prompt.
  • Destructive scripts (file writes, AWS mutations, subprocess calls, network POSTs, database connections) get flagged for human review with specifics about what was detected, including the source line content.

Zero external dependencies — it’s pure stdlib (ast, json, fnmatch, shlex). AST parsing is near-instant, so there’s no perceptible delay.

The Rule System

The default rules are sensible and well-structured:

  • AWS boto3: describe/list/get/head → safe. delete/put/create/terminate → destructive. Rules scope via trigger_imports, so cache.delete_item() in a non-AWS script won’t false-positive.
  • File I/O: open() in write modes, os.remove, shutil.rmtree → destructive. Read-only access is fine.
  • Subprocess: Always flagged. subprocess.run, os.system, the lot.
  • Network: requests.get → safe. requests.post/put/delete → destructive.
  • Database: Connection creation → flagged for review.

A curated list of safe imports (json, csv, re, datetime, pathlib, hashlib, and ~50 others) means scripts that only use standard library data-processing modules sail through without interruption.

Custom rules go in ~/.claude/yolt/rules.json and merge with defaults — you can add safe methods, define new categories with their own trigger_imports, and use glob patterns (fetch_, drop_).

One Important Gotcha

If you have Bash(python3:*) in your Claude Code settings.local.json allow list, YOLT’s hook never fires — static allow rules take precedence over PreToolUse hooks. YOLT replaces the need for that allow rule entirely: safe scripts get auto-approved by the hook itself.

Why This Matters

The design philosophy — “false positives OK, false negatives not” — is the right one for a safety tool. It’s the security principle of fail-closed applied to AI code execution.

YOLT is small (527 lines across 6 files in the initial commit), focused, and immediately useful. If you’re letting Claude Code run Python, this is the kind of guardrail that should exist by default.


Wrapping Up

voitta-rag is evolving from a code search tool into a self-hosted knowledge platform with enterprise connectors and AI memory. voitta-yolt tackles a different but equally practical problem: making AI code execution safer without making it slower.

Both projects are open source (AGPL v3) and available on Voitta AI’s GitHub.


Gregory Golberg is co-founder of Method & Apparatus, a fractional CTO consultancy. Previously: llm-tldr vs voitta-rag: Two Ways to Feed a Codebase to an LLM.