Two Axes of Compression, and a Trap That Makes One Unmeasurable

tl;dr — Token compression has two independent axes: what you feed the model (tokens in) and what it writes back (tokens out). Three findings from measuring both. Repomix, pointed at the same file globs as a plain cat of the repo, produced more characters than the plain dump — its advertised ~70% reduction is file selection, not compression. A prose compressor on source code buys 9.4% by destroying the punctuation that makes it code, and costs only 0.4 points, which is its own uncomfortable finding. And you cannot measure the output axis with output-token counts on a reasoning model: our compressed-output run shrank the visible answer 21% while its tokens_out tripled.


The two axes

Most “save tokens” tooling is sold as one category, and it isn’t. There are two:

  • Tokens in — shrink the context you send. Repomix, llm-tldr, RAG retrieval, prose compressors applied to a dump.
  • Tokens out — shrink what the model writes back. Instruct it to answer tersely.

They’re orthogonal. An output-side compressor rides on top of any input-side strategy, which means the honest way to evaluate them is a grid, not a list. We ran the input-side arms, then re-ran a representative subset with the output-side overlay on.

Finding 1: Repomix is the same dump with a nicer cover page

Repomix packs a repository into one AI-friendly file and is widely cited for ~70% token reduction. Pointed at the same include/exclude globs we used for a plain concatenation of the same files:

charactersscore /12tokens in
plain source dump1,119,81910.80404,878
Repomix1,127,41410.40406,277

Repomix produced 7,595 more characters than cat-ing the files.

This isn’t a knock on the tool, and the 70% figure isn’t dishonest — it’s just measuring something else. Repomix’s reduction comes from file selection: honouring .gitignore, skipping binaries and lockfiles and node_modules, dropping build artifacts. Against a naive “send the whole working directory” baseline, that’s an enormous and genuine saving.

But we’d already scoped our globs to **/*.java minus tests. There was nothing left to select. What remains is formatting — a directory tree, a header block, per-file separators — and formatting costs tokens rather than saving them.

The general point: a compression ratio is a ratio against something. Before adopting a tool on a headline percentage, check what the denominator was. If your pipeline already scopes its inputs, a selection-based tool has already had its win taken.

Finding 2: a prose compressor on code, and how little the model needs

caveman-compression strips grammar an LLM can reconstruct — articles, connectives, passive constructions. We used the rule-based spaCy variant rather than the default LLM-backed one, on purpose: a non-deterministic compressor inside a benchmark cell makes the cell unattributable, and it would put a second vendor’s model inside our measurement path.

Applied to the source dump:

charactersscore /12tokens in$/q
plain dump1,119,81910.80404,878$0.8552
caveman-compressed1,014,00410.40363,613$0.7777

9.4% smaller, 0.4 points. Roughly neutral.

Which is startling once you look at what it does to Java:

// before
public class Attribute implements Map.Entry<String,String>, Cloneable {

// after
public class Attribute implements Map. Entry < String String >   Cloneable

Commas gone. Angle brackets spaced apart. Map.Entry split across a sentence boundary. This is not valid Java in any sense — a parser would reject it instantly — and the model scored 10.40 out of 12 on it.

Be fair to the tool: it’s built for prose and we pointed it at source code. This measures a mismatch, not the tool used as intended, and 9.4% on input it was never designed for is respectable.

The finding isn’t about the compressor. It’s about how much syntax the model actually needs, which is apparently much less than the syntax the compiler needs. That’s a genuinely interesting property and probably a bad thing to rely on.

Finding 3: the trap

The output-side overlay works. Instruct the model to answer in compressed style and the rendered answer gets meaningfully shorter at little quality cost:

modeanswer charswith overlayscorewith overlay
full dump5,6994,496 (−21%)10.8010.60
agentic exploration6,3894,029 (−37%)11.4010.40
RAG (corpus-matched index)6,0183,696 (−39%)5.405.60
llm-tldr2,3871,862 (−22%)2.403.40

21–39% shorter for roughly zero to one point either way. Cheap, real, worth having. (The RAG row now reflects a clean re-run; two of the four arms score slightly higher with the overlay than without, which at five questions is not distinguishable from noise and is not a claim that compression improves answers.)

Now the same experiment measured the way you’d instinctively measure it — by counting output tokens:

rendered answertokens_out
full dump5,699 chars4,547
full dump + output compression4,496 chars (−21%)14,859 (+227%)

The visible answer shrank by a fifth. The billed output tokens more than tripled.

tokens_out bills thinking tokens and response text together. On a reasoning model, thinking usually dominates, and it varies enormously with how hard the model decides the turn is. The overlay changed how the model approached the task — apparently prompting more deliberation about what to cut — and that swamped the text delta by an order of magnitude.

Anyone benchmarking output-side compression against tokens_out on a thinking model is measuring reasoning-depth noise and calling it compression. You will get a number, it will be reproducible, and it will point the wrong way.

Measure the rendered answer. len(response_text), or token-count the text blocks specifically. And if you’re doing cost work, keep the two apart: thinking tokens are a real cost you should track, they’re just not what an output-style instruction controls.


Next in this series: what it costs to know any of this — and why grading the answers cost more than producing them.

Harness, raw records, and full method: voitta-rag/benchmark/.

Correction, 2026-09-25. The RAG row in the output-overlay table has been corrected after that arm was re-run with an adapter path bug fixed, and the row relabelled to name which index it used. It first published as 4,665 → 3,659 chars scoring 4.80 → 4.60. The re-run and the fix are in voitta-rag#57 and #58.

Context benchmark series — part 5 of 7: ← Previous · Series index · Next →

Nobody Needed to Fit the Codebase in the Window

tl;dr — We benchmarked five strategies for getting a Java codebase into an LLM’s context: a full source dump, Repomix, two llm-tldr adapters, and RAG retrieval. The winner was none of them. Giving the model read_file, grep, and glob and letting it go find things scored 11.4/12, against the full dump’s 10.8 — while using 29% fewer tokens and costing 25% less. It also produced 142 verified source citations against 2 fabricated, the cleanest record in the benchmark. Every tool in this category optimises how to pack the context window. On this question set, the winning move was not to pack it.


The question

A colleague dropped Repomix in Slack — pack your whole repo into one AI-friendly file, ~70% token reduction. Someone else pointed at llm-tldr — 95% token savings, 155x faster queries. A third person asked the only question that matters:

If one of you get time can you run an eval on the same codebase for the same task and let me know if these actually improve the output and which one is better

So we did. One repository (jsoup, 97 Java files, deliberately one nobody on the team knew), five questions spanning five kinds of thing you actually ask about code, and every strategy answering the identical prompt with only the injected context varying.

The scoring, because it’s the part that matters

Every answer had to carry a file:line citation for every factual claim. The judge — a separate model with read-only read_file, grep, and glob over the repository — then went and checked them. Not “does this look right.” Does Tokeniser.java:135 exist, and does it say what the answer claims.

That produces two numbers per answer: a quality score out of 12, and a count of citations that resolved against real source versus citations that didn’t. The second number is the one that earns its keep, and a later post in this series is entirely about what it caught.

The result

strategyscore /12tokens in$/questionverified citesbogus
agentic exploration11.40288,3420.63871422
llm-tldr → agentic11.00288,4360.63781240
full source dump10.80404,8780.85526511
Repomix10.40406,2770.93746917
prose-compressed dump10.40363,6130.7777861
llm-tldr (extract)6.6066,8950.1520831
RAG retrieval5.004,2370.03533115
llm-tldr (semantic search)2.405,1790.0236229

One caveat on the RAG row, found after this was drafted and before it was published: its bogus count is an upper bound. The adapter handed the model paths prefixed with the index name (jsoup/src/…) while the judge resolved citations against the checkout root (src/…), so citations that were real scored as unresolved — the prefix is visible in the judge’s notes on 13 of 15 retrieval answers. The harness strips it now; these numbers predate that. It touches no other arm, and the arm it flatters least is the one we build.

The top line is a mode we added almost as a control — no context building at all, just hand the model the same three read-only tools the judge uses and let it explore. It won on quality, it won on citation accuracy by a wide margin, and it was cheaper than the thing it beat.

Why it wins

Not because it’s clever. Because of what it has at the moment it makes a claim.

Every other strategy front-loads: build a representation of the codebase, inject it, hope the answer is in there. The representation is fixed before the model has read the question closely, so it is necessarily a guess about relevance — and whatever the representation dropped, the model cannot recover.

Agentic exploration defers. It reads the question, forms a hypothesis, greps for it, gets it wrong, greps again, opens the file, reads the actual lines. Seven to sixteen tool calls per question in our runs. When it finally writes Tokeniser.java:135, it is because it has line 135 on screen.

That is the whole mechanism behind the citation column. Verified-to-bogus for agentic exploration was 142:2. For the full dump, 65:11 — the dump had every line, but the model was reading a 405,000-token wall of text and lost track of where in it things were. For the cheapest compressed mode, 2:29.

Worth sitting with: the full dump contains strictly more information than the agentic mode ever sees, and still loses. Having the bytes in the window is not the same as being able to use them.

Two caveats we’re keeping

Cumulative tokens. The 288K for agentic exploration is summed across every turn of the tool loop, not one request. It is the honest number for cost, and it is not the same kind of number as a one-shot mode’s single request. We report it that way because it’s what the strategy actually costs to answer one question, but don’t put it in a bar chart next to a single-shot figure without the asterisk.

Five questions. Enough to catch a large effect, not enough to rank close ones. The 10.4–10.8 cluster — full dump, Repomix, prose-compressed dump — is a tie as far as this data can tell. The gaps worth believing are the big ones: agentic exploration over the compressed modes, and the two llm-tldr adapters against each other.

The uncomfortable implication

There’s a lot of engineering going into context compression right now, and this result doesn’t say that work is worthless — the compressed modes have a real argument, which is price. llm-tldr via its extract adapter got 61% of the baseline’s score for a sixth of the cost. If you’re running a million of these, that trade is the whole business.

But if you’re optimising for a correct answer, the ranking says: give the model tools and get out of the way. The context window is not a thing to be filled efficiently. It’s a workspace, and the model is better at deciding what belongs in it than our heuristics are.


Next in this series: the tool that cut input tokens 44x and fabricated 29 of its 31 citations — and why that turned out to be our fault, not the tool’s.

Harness, raw results, and full method: voitta-rag/benchmark/. Answering on Claude Sonnet 5, judging on Claude Opus 5, both at effort high. Total cost of the run: $79.03 over 70 scored cells, of which $50.74 was judging — which is its own post.

Correction, 2026-09-25. The RAG row has been corrected. Its adapter was handing the model index-prefixed paths that the judge could not resolve, so real citations were scored as fabrications; with that fixed and the arm re-run, it reads 5.00 with 31 verified citations to 15, where this post first published 5.20 with 26 to 24. The run cost is likewise corrected to $79.03. No other row changed, and the conclusion of this post does not depend on the RAG row. The re-run and the fix are in voitta-rag#57 and #58.

Context benchmark series — part 1 of 7: Series index · Next →