Reference, read on 2026-09-30
Token-saving tools for coding agents: claims and third-party studies
What 45 projects claim, and what 8 public benchmarks and 15 arXiv preprints measured, with a link to every source.
- tools catalogued
- 45
- public benchmarks
- 8
- arXiv papers
- 15
Disclosure. Florian Bruniaux, who maintains this site, is a core contributor to RTK, one of the tools below. This page does not recommend a tool. It sets out what each project claims and what third parties measured, with links to every source, and reports RTK's results as measured, including the unfavourable ones. Tools enter the catalogue through a fixed rule stated there, not by preference.
Catalogue of token-saving tools
45 projects, each with its own claim and any third-party result found. Open a row for the full claim, every result and the limits.
How tools are listed
A project is listed when it is public, not archived, aimed at coding agents, and has at least 100 GitHub stars, or was measured by one of the benchmarks below, or is already covered by the guide. Stars, license and last push come from the GitHub API; claims were read in each project's README or product page on 2026-09-30. The list is sorted by name by default. GitHub stars measure attention, not quality. Each result comes from one study: see Claims and studies before reading a number.
-
Caveman Output style 33.2% of input tokens through its proxy, 54 pinned runs 6 third-party results 109k GitHub stars
Rule-file skill that makes the agent answer in terse caveman prose, now bundled with a local proxy that shrinks tool output the agent reads and a middleware library for application code.
- Project claim
-
-33.2%
Measured on: provider-reported input tokens through its proxy over a pinned 54-run Claude Code suite (885,793 to 591,673); for the skill alone the README reports 50% fewer output tokens at the median against an 'Answer concisely.' control Claim source β - Repository description
-
Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- Pinned 54-run Claude Code suite: 6 cases x 3 runs x 3 arms, every answer checked against an exact oracle, 18 of 18 answer checks passed, case-clustered 95% interval 14.6% to 48.5%. Per case: CSV outlier -55.1%, log needle -50.2%, YAML drift -46.2%, test output -27.8%, deployment JSON -26.4%, dashboard HTML +9.9%. In the same suite Headroom's wrap saved 6.7% and failed 3 of 18 checks. Skill-only eval snapshot: 50% fewer output tokens at the median on ten dev questions versus a plain 'Answer concisely.' control (length only, not correctness).
- Limits and context
- The 75% figure no longer appears in the README, and no '65% of output tokens' headline was found either. The string 65% appears only in the title of the JetBrains blog post, 'Speaking to AI Agents like Cavemen Saves 65% of Tokens. We Test.', while JetBrains's measured result is 8.5% fewer output tokens (about 10% cost, quality flat, sign test p = 0.82). Adobe Research paper (CAVEWOMAN, arXiv 2606.24083) is cited for 1.4 to 2.4x realized cost cut, up to 3x. README says raw harness artifacts are not in the checkout, so the 54-run result is 'a pinned report, not a public reproduction', and the rules add input tokens on every call. The HTML row is left red at +9.9%. Licence is split: skill MIT, engine, proxy and runtime BSL-1.1 (source-available, converts to Apache-2.0 on 2030-06-21 or four years after each release); hosting for third parties needs a commercial licence. CLI telemetry (anonymous counts) is on by default.
Project page βThird-party results (6)
- Dasein Labs β Total cost β19% (19% cheaper) 58 resolved.
- Stet.sh β Workload cost, run 1 then run 2 +9% then β12% (9% more expensive, then 12% cheaper)
- JetBrains (Caveman) β Output tokens β8.5% (8.5% fewer output tokens) Forced activation, so a ceiling. Quality 8 better, 10 worse, 64 equal.
- THOL β End-to-end cost, 7 long tasks β13.6% (13.6% cheaper) Interval 30.2% cheaper to 3.6% more expensive.
- Tura β Modeled cost, one task, 2 runs β3.9% (3.9% cheaper) Score 86.5% vs 78.9% baseline.
- Codepointer β Share of replayed spend saved β0.4% (0.4% cheaper) Replay, not a live run.
-
Claude Context Code index and search about 40% of tokens at equal retrieval quality No third-party result 13k GitHub stars
MCP server that indexes a codebase with AST chunking into a Milvus or Zilliz Cloud vector database using hybrid BM25 plus dense embeddings and returns relevant snippets for natural-language queries.
- Project claim
-
Claude Context MCP achieves ~40% token reduction under the condition of equivalent retrieval quality.
Measured on: tokens in the project's controlled retrieval evaluation (evaluation/ directory) at equivalent retrieval quality; task and model details are not in the README Claim source β - Project's own benchmark
- Controlled evaluation with an efficiency chart referenced in the README; method lives in the evaluation directory, not reviewed for this page.
- Limits and context
- Conflict of interest: Zilliz sells the vector database it uses. The quick start requires a vector database (README's default path is a free Zilliz Cloud account and API key) plus an embedding provider key (OpenAI by default; VoyageAI, Ollama and Gemini also supported). README FAQ links a fully local deployment answer, not verified. Code is sent to the embedding provider unless a local one is used. MIT licence.
Third-party results (0)
None found on 2026-09-30.
- Requires a vector database sold by the publisher
-
claude-mem Memory about 10x fewer tokens for memory retrieval No third-party result 95k GitHub stars
Captures tool-use observations through lifecycle hooks, summarises them into a local SQLite plus Chroma store, and injects prior context at session start with a 3-layer search, timeline and get_observations MCP workflow.
- Project claim
-
~10x token savings by filtering before fetching details
Measured on: tokens for memory retrieval: compact index results at ~50-100 tokens each versus ~500-1,000 tokens per fully fetched observation; not a measure of whole-session cost Claim source β - Limits and context
- The installer prompts sign-in to a hosted 'claude-mem observer' (free up to 14 days, then falls back to the user's Anthropic plan unless subscribed) and a CMEM Pro cloud sync; local providers can be chosen with --provider. Memory generation itself consumes model tokens. README ends with a section on CMEM, a crypto token 'created by a 3rd party but officially embraced' by the author. Apache-2.0 licence.
Project page βThird-party results (0)
None found on 2026-09-30.
-
claude-token-efficient Output style 63% fewer words over four prompts 1 third-party result 6.1k GitHub stars
A single CLAUDE.md rule file (with profiles) that tells Claude to skip preamble and sycophancy, prefer targeted edits, read files once and keep answers terse.
- Project claim
-
63%
Measured on: words in answers over four prompts (465 to 170), single run each; the README calls it a directional indicator Claim source β - Project's own benchmark
- Reproducible token benchmark (2026-06, benchmark/, N=5): output-token reduction with the current minimal CLAUDE.md is about 4% (haiku), 12% (sonnet), 7% (opus). Head-to-head on the external testing-claude-agent harness: v8 config $0.935 total versus $1.131 for C-structured (-17.4%).
- Limits and context
- README states the 63% is 'achievable with a stricter rules profile' while the current file gives 4-12%, and that current-model baselines already show 0% preamble, sycophancy and smart quotes, so those rules 'carry input cost without changing output'. Also states the file adds input tokens on every turn and is 'a net token increase' on low-output exchanges. The 17.4% cost figure comes from an external harness (adam-s/testing-claude-agent, issue 1) run by the project against its own config. MIT licence.
Third-party results (1)
- THOL β End-to-end cost, 7 long tasks β11.6% (11.6% cheaper) Interval 24.2% cheaper to 1.4% more expensive.
- Headline superseded by its own benchmark
-
Code Context Engine (CCE) Code index and search 94% of retrieval tokens against full-file reads No third-party result 425 GitHub stars
Local MCP server that indexes code into semantic chunks with hybrid vector and BM25 search plus a call graph, compresses retrieved chunks to signatures, records decisions across sessions, and logs savings.
- Project claim
-
94% token savings, reproducibly benchmarked.
Measured on: retrieval tokens per query versus reading the full content of every touched file (83,681 to 4,927 tokens/query on FastAPI, 20 queries); not against what Claude Code does with grep and partial reads Claim source β - Repository description
-
Save 94% on AI coding tokens.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- FastAPI: 94% retrieval savings, Recall@10 0.90, 20 queries. Other repos: Django 93% (R 0.95), Express 94% (R 1.00), chi 76% (R 0.67), fiber 93% (R 0.07).
- Limits and context
- README's own baseline note: 'The 94% number is measured against full-file reads, not against what Claude Code actually does... the real-world savings compared to normal Claude Code behavior will be lower than 94%.' Fiber recall of 0.07 shows retrieval can miss on monorepos. Also ships output compression rules with stated levels of ~30%/~65%/~75% (README table) and ~25%/~70%/~80% (FAQ table) that are estimates, not in the headline. MIT licence.
Project page βThird-party results (0)
None found on 2026-09-30.
-
CodeGraph Code index and search 44% cheaper on 7 repos, one question each 1 third-party result 72k GitHub stars
Local Rust-kernel knowledge graph of symbols, call edges and dependencies exposed through MCP tools such as codegraph_explore, updated incrementally on file save.
- Project claim
-
88% fewer tool calls Β· 53% faster Β· 62% fewer tokens Β· 44% cheaper Β· file reads cut to zero on all seven repos.
Measured on: average across 7 open-source repos, one architecture question per repo, headless Claude Code on Opus 4.8 with versus without CodeGraph, median of 4 runs per arm (re-measured 2026-08-05) Claim source β - Project's own benchmark
- 7 repos in 7 languages; per-repo cost reduction ranges from about even (Gin) and 13% (Django) to 78% (Excalidraw); harness blocks the codegraph CLI in both arms (0 of 28 without-arm runs contaminated).
- Limits and context
- README states a counter-effect: in multi-turn sessions CodeGraph's responses leave about 80% more retrieval context resident at session end (VS Code: 67k tokens against 18k), so 'on that axis CodeGraph costs more, not less'. One question per repo, so results reflect discovery-heavy questions. MIT licence.
Project page βThird-party results (1)
- THOL β End-to-end cost, 7 long tasks β7.6% (7.6% cheaper) Interval 21.6% cheaper to 8.7% more expensive. Opt-in tool.
-
codesight Code index and search 7x to 12x smaller map than an estimated exploration cost No third-party result 1.4k GitHub stars
One-command scanner that generates a compiled map of routes, models, components and dependencies for a repo (plus optional wiki articles and MCP tools) for the agent to read instead of exploring files.
- Project claim
-
7x-12x base scan; 60-131x with targeted wiki queries
Measured on: output map tokens versus an estimated exploration cost computed by formula (routes x 400 + models x 300 + components x 250, plus a revisit multiplier), on three private SaaS projects and several open-source repos Claim source β - Project's own benchmark
- Three SaaS projects: 11.7x, 7.2x and 11.4x base reduction (46,020, 26,130 and 47,450 estimated exploration tokens versus 3,936, 3,629 and 4,162 map tokens); README quotes 'Average combined reduction: 91x' with the wiki layer.
- Limits and context
- The baseline is an estimate by multiplier, not a measured agent run; only output size is measured (chars / 4). Savings compare a one-time map to per-session exploration and do not include agent quality. MIT licence.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Context Mode MCP tools and gateways 98% of tool-output bytes kept out of context 1 third-party result 24k GitHub stars
MCP server plus hooks that run tool work in a sandbox so only the printed result enters context, index session events in SQLite FTS5 for retrieval after compaction, and route agents to these tools.
- Project claim
-
315 KB becomes 5.4 KB. 98% reduction.
Measured on: bytes of raw tool output kept out of the context window (not tokens, not the bill), on the vendor's scenarios Claim source β - Repository description
-
Sandboxes tool output (98% reduction), persists session memory, and enforces routing across 17 platforms via MCP + hooks.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- Table of scenarios, raw vs context bytes: Playwright snapshot 56.2 KB to 299 B (99%), GitHub issues x20 58.9 KB to 1.1 KB (98%), access log 500 requests 45.1 KB to 155 B, repo research subagent 986 KB to 62 KB (94%); a 21-scenario BENCHMARK.md is linked (not reviewed for this page). No model, task success or dollar cost.
- Limits and context
- The per-platform table shows '~98% saved' with hook enforcement and '~60% saved' without hooks (routing by instruction file only, '~60% compliance' on Gemini CLI and Codex-style instruction files). Cursor accepts hook context but does not surface it to the model (known limitation). Savings depend on the agent choosing the sandbox tools; it changes the workflow ('Think in Code'). Elastic License 2.0 (source-available, not OSI open source; no hosted-service offering allowed).
Project page βThird-party results (1)
- Stet.sh β Workload cost, run 1 then run 2 +72% then +33% (72% more expensive, then 33% more expensive)
-
Dynamic Context Pruning (DCP) Compaction and pruning No quantified claim No third-party result 4.3k GitHub stars
OpenCode plugin that exposes a compress tool so the model replaces closed conversation ranges with summaries, plus automatic deduplication of repeated tool calls and purging of errored tool inputs.
- Project claim
- No quantified claim.
- Limits and context
- README 'Project Status': development has slowed because new context-management work moved to Sleev (sleev.ai), a local proxy that supports Claude Code, Codex, OpenCode, Pi and Hermes; new features land in Sleev first and the README recommends it for new users. README states the cache trade-off itself: pruning 'changes messages, which invalidates cached prefixes' and 'can increase cache misses'. Licence AGPL-3.0-or-later.
-
Edgee API proxies and context compression up to 70% of the token bill, measurement not stated 1 third-party result 137 GitHub stars
Rust CLI that points coding agents at Edgee's hosted agent gateway, which routes requests, compresses tool outputs and meters usage per developer, repo and model.
- Project claim
-
Together that's up to 70% off your token bill
Measured on: not stated (the README combines routing to cheaper models, compression and visibility under one 'up to' figure) Claim source β - Project's own benchmark
- compression-lab (Edgee, published 2026-09-30), SWE-bench Lite against vanilla Claude Code, 300 paired comparisons per layer, layers measured one at a time: brevity -30.6% cost, tool result trimming -6.3%, tool surface reduction -16.7% (aggregate). Endurance run: 26.5 instructions instead of 21, total session cost +19.6%, cost per instruction -5.1%. Edgee's docs give 15-20% in production and 50% on SWE-bench Lite as a ceiling.
- Limits and context
- Hosted gateway: this repository ships the CLI; routing, compression and metering run on Edgee's service. The compression-lab README says SWE-bench was augmented with MCP requests to test tool surface reduction, that layers were not run combined, that cost is computed from session logs at list prices. No resolution rate is published. The README credits RTK as the inspiration for its trimming engine.
Third-party results (1)
- THOL β End-to-end cost, 7 long tasks β14.7% (14.7% cheaper) Interval 37.1% cheaper to 11.1% more expensive. THOL flags no token accounting for this tool.
- Hosted gateway
-
Fermat's Last Token (Quotient Labs) API proxies and context compression 47% cheaper per session, on the vendor's SessionBench 1 third-party result
Closed, paid wrapper that runs Claude Code through a local proxy gateway, replaces the native file tools with Search and Edit MCP facades and compresses context.
- Project claim
-
47% cheaper on average.
Measured on: per-session agent token cost, averaged over K=5 runs per benchmark, versus vanilla Claude Code on the vendor's SessionBench (SWE-Atlas refactoring tickets), claude-opus-5; range shown on the page is 34-51% Claim source β - Project's own benchmark
- SessionBench: 15 paired sessions across three repos (Go, TypeScript, Python) on real SWE-Atlas refactoring tasks; the page says Fermat was cheaper in every run and matched or beat vanilla on quality in 13 of 15 runs.
- Limits and context
- Pricing model on the page: 'Pay just 10% of what you saved'. Requires an entitled Quotient Labs account; the Dasein bench says outsiders cannot reproduce the Fermat arm without a paid account and that no self-host recipe exists. The vendor figure and the bench figure use different task sets (SWE-Atlas vs SWE-bench Verified), models (claude-opus-5 vs claude-sonnet-4-6) and metrics (token cost vs cache-aware dollars per solved task). Page version read: v0.1.10. No public repository found.
Third-party results (1)
- Dasein Labs β Total cost β22% (22% cheaper) Run on 2026-09-21, not paired with the other arms. 55 resolved.
- Paid, closed
-
Governor Compaction and pruning 45.5% of tokens on the project's fixtures No third-party result 134 GitHub stars
Claude Code plugin with hooks that filter repetitive tool output, compact responses, compress memory files such as CLAUDE.md, and log telemetry; other agents get the same behaviour as a prompt-based rules file.
- Project claim
-
45.5%
Measured on: average token savings in the project's V2 Sonnet fixture table, where Caveman scores 69.1% Claim source β - Project's own benchmark
- V2 Sonnet fixture run: Governor 45.5% average token savings, valid-context loss 0.00, decision preserved 100.0%; Caveman 69.1%, valid-context loss 0.14, decision preserved 87.5%. Separate multi-turn pilot: output tokens 10,997 to 10,113 (-8.0%), cost $0.5169 to $0.4933 (-4.6%), described as 'a narrow pilot, not a universal claim'.
- Limits and context
- README concedes 'Caveman still wins on raw compression'. Tool-filter signal checks: 64.0% blocked on a noisy pytest log, 90.9% on a Burp-style MCP payload, 0.0% on a large Read of source code. For non-Claude agents the filtering is done by the model following a rules file, not by hooks.
Project page βThird-party results (0)
None found on 2026-09-30.
-
graphify Code index and search No quantified claim 2 third-party results 123k GitHub stars
Builds a queryable knowledge graph of code (local tree-sitter AST, no LLM) and of docs, PDFs, images and video (model-based semantic pass), with query, path and explain commands over graph.json.
- Project claim
- No quantified claim.
- Project's own benchmark
- Memory-retrieval benchmarks on the same harness and model: LOCOMO (n=300) recall@10 0.497 versus mem0 0.048 and supermemory 0.149; QA accuracy 45.3% versus supermemory 49.7% and mem0 27.3%; LongMemEval-S (n=50) 76%, tied with dense RAG. Graph build uses 0 LLM credits for code.
- Limits and context
- No token-saving figure in the README; benchmarks measure retrieval quality, not tokens. README promotes a hosted platform (graphify.com, 14-day free trial) as the always-on version. Apache-2.0 licence per badge. Docs and media extraction sends content to the assistant's model or a configured API key.
Project page βThird-party results (2)
- Marmelab β Cost on the large transformation +1% (1% more expensive)
- THOL β End-to-end cost, 7 long tasks β3.3% (3.3% cheaper) THOL reports that the agent never used the tool, so the gap is not a tool effect.
-
grepai Code index and search No quantified claim No third-party result 1.9k GitHub stars
Go CLI and MCP server for semantic code search using vector embeddings from Ollama, LM Studio or OpenAI, with a file watcher and call-graph tracing.
- Project claim
-
Drastically reduces AI agent input tokens by providing relevant context instead of raw search results.
Measured on: not stated (no figure in the README) Claim source β - Limits and context
- No quantified claim in the README; it links a blog for benchmarks (not reviewed for this page) and quotes user testimonials, including a Reddit post titled 'I reduced Claude Code input tokens by...' (title truncated in the URL slug). Requires an embedding provider; 100% local only with Ollama or LM Studio. MIT licence.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Greplica Memory 40% to 50% of tokens on showcased planning cases No third-party result 436 GitHub stars
Builds a queryable graph of components, flows and claims from repo structure and past agent session transcripts in local SQLite, which the agent queries with greplica graph context before exploring.
- Project claim
-
In several showcased planning cases, Greplica cut token usage by 40-50%. In the strongest run we measured, Greplica used 75.0% fewer tokens and finished about 38% faster.
Measured on: total tokens on held-out planning tasks built from SWE-chat sessions, baseline run with no Greplica context versus a run allowed to query greplica graph context; three showcased cases only Claim source β - Repository description
-
Saves ~50% tokens and ~30% time while planning
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- Three showcased SWE-chat cases: Gemini Voyager 1,925,152 to 480,988 tokens (75.0% fewer), IPTVnator 1,592,080 to 881,541 (44.6%), Marin Harbor 2,992,034 to 2,193,023 (26.7%).
- Limits and context
- The showcased rows are selected cases, and README says 'several showcased' cases reached 40-50%, not the full set. Local mode is free and offline; managed shared memory needs a login to a hosted server (memory.autoloops.ai) operated by the vendor, with GitHub App and OIDC integration. MIT licence.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Headroom API proxies and context compression 20% fewer tokens for coding agents, per the repository description 4 third-party results 74k GitHub stars
Local compression layer (library, proxy, agent wrapper and MCP server) that routes JSON, code and text through dedicated compressors, keeps originals retrievable, and can trim output tokens.
- Project claim
-
20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers.
Measured on: tokens, per the repository description; the README's four seeded scenarios show 21% to 57% Claim source β - Project's own benchmark
- Accuracy suite (python -m headroom.evals suite --tier 1), N=100 each: GSM8K 0.870 vs 0.870, TruthfulQA 0.530 vs 0.560, SQuAD v2 97% at 19% compression, BFCL 97% at 32% compression; README notes a +/-0.03 delta is inside the confidence interval. Latency 0.21 ms p50 at 10K tokens.
- Limits and context
- README says savings scale with repetition: repeated JSON arrays and log lines clear 90% in bench_latency.py, prose and already-dense output compress very little; run headroom savings on your own traffic. Output-token reduction is opt-in (HEADROOM_OUTPUT_SHAPER=1) and reported as an estimate with a confidence interval (example shown: 31.7%, 95% CI 27.7% to 35.7%). The '60-95% fewer tokens' headline that the Dasein bench quotes from github.com/chopratejas/headroom is not in the current README; the repo now lives under headroomlabs-ai. Note: the Dasein bench runs Headroom in token mode behind a gateway and is sponsored by a competitor (Parsec's vendor); THOL is maintained by Tokenade's author. Apache-2.0. Compression runs locally, no content sent for compression.
Project page βThird-party results (4)
- Dasein Labs β Total cost +44% (44% more expensive) 58 resolved.
- Marmelab β Cost per operation, cache mode then token mode β4% then +32% (4% cheaper, then 32% more expensive) One task, three runs. The authors call -4% noise.
- THOL β End-to-end cost, 7 long tasks +52.8% (52.8% more expensive) Interval 10.2% to 120.2% more expensive, entirely above zero.
- Codepointer β Share of replayed spend saved β2.8% (2.8% cheaper) Replay, not a live run. Median 54% cut on the grep and diff output it touches.
-
Honey (I Shrunk the AI) Output style 24% cheaper on 23 single-call tasks, Opus 5 No third-party result 310 GitHub stars
Skill that combines minimal-code (YAGNI ladder) and terse-prose rules with a denser ESON format for agent-to-agent handoffs, with auto-chosen intensity and safety carve-outs.
- Project claim
-
β24%
Measured on: cost per task on Claude Opus 5, 23 tasks Γ 3 runs, single-call generation (p<0.001) Claim source β - Repository description
-
cuts AI coding-agent token usage and LLM API costs [...] (β53%, lossless in benchmarks) with no loss of quality.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- Vendor bench/ suite, 23 tasks, 3 runs each. On Opus 4.8, whole suite: LOC -43%, output -29%, judge win/loss/tie 8/11/2 (p=0.648, a tie). Caveman -28% LOC and -22% output, Ponytail -33% LOC and -7% output on the same suite. On GPT-5.5 output -20% (p=0.004) but cost +14% (not significant).
- Limits and context
- README caveats itself: 'The dollar saving is unproven at this sample size' (-21% on Opus 4.8 at p=0.104), output delta on user-facing tasks is a tie (-7%, p=0.673), and Honey loses the code judge 2/11/1 (p=0.022) on Opus. The benchmark is single-call, not an agent loop; an end-to-end Cline harness exists but no result was read. Published by GreenPT, which is also a listed Ponytail sponsor.
Project page βThird-party results (0)
None found on 2026-09-30.
-
ICM (Infinite Context Memory) Memory 44% of input context in a later session, one project No third-party result 580 GitHub stars
Single-binary MCP memory server backed by SQLite, FTS5 and sqlite-vec with decaying episodic memories, permanent knowledge graphs ('memoirs') and feedback records, shared by every configured agent.
- Project claim
-
-44%
Measured on: input context in session 3 of a multi-session run on one 12-file Rust project, haiku, 3 runs averaged (74.7k to 41.6k) Claim source β - Project's own benchmark
- Agent benchmark: 10 sessions, haiku, 3 runs, real claude -p calls; session 2 turns -29%, cost -17%; session 3 turns -40%, cost -22%. Recall benchmark (fictional document, 10 questions): average score 5% without ICM vs 68% with. LongMemEval oracle variant: 100.0% retrieval, 82.0% answer accuracy with Sonnet.
- Limits and context
- Published by the rtk-ai organization, the same organization as RTK. README banner: 'Project status: experimental', pre-1.0, breaking changes possible in any minor release, and the maintainer says his focus is RTK so ICM updates are 'best-effort cadence'. Apache-2.0 licence. Single project and single small task for the token figures.
Third-party results (0)
None found on 2026-09-30.
- Experimental
- Same organization as RTK
-
LeanCTX Memory 99.4% of input-side cost on a replayed session 1 third-party result 3.8k GitHub stars
Local binary and MCP server that caches file reads (cached re-read costs about 13 tokens), compresses shell output with 85+ patterns, runs an optional prompt-cache-safe proxy, and keeps session memory across chats.
- Project claim
-
99.4% input-side saving on cache-priced rails
Measured on: input-side cost on a deterministic 72-turn replayed session priced per model (lean-ctx benchmark dual-arm, digest f5ed145e61ce3689), not a live coding task Claim source β - Project's own benchmark
- Self-verify replay of a 72-turn session with the proxy rail; fixed per-session footprint of about 3.0K tokens measured in CI. A model-free A/B gate checks that the JSON crusher keeps gold answers in its fixtures.
- Limits and context
- README says the earlier per-read-mode compression table (map and signatures over 50 files) is 'withdrawn' along with other historical figures, and that savings 'depend on the workload'. The CI replay is called a 'mechanism gate', 'not evidence that compression preserves answer quality', and a powered quality study is still open (issue 1905). Also positions itself as an 'AI Value Gate' with governance features. Apache-2.0 licence.
Third-party results (1)
- THOL β End-to-end cost, 7 long tasks +10.3% (10.3% more expensive) Interval 12.6% cheaper to 46.4% more expensive.
- Earlier figures withdrawn by the project
-
llmtrim API proxies and context compression 66% of round-trip cost over 112 A/B cases No third-party result 240 GitHub stars
Local MITM proxy (Rust, also CLI, MCP and library) that trims LLM API requests (logs, tool schemas, repeated JSON) and can steer output terseness, without touching the cached prefix.
- Project claim
-
β31% input Β· β74% output Β· β66% round-trip cost Β· 112 live A/B cases
Measured on: input tokens (71,031 to 49,062), output tokens (25,843 to 6,628) and round-trip cost ($0.0365 to $0.0126) over 112 A/B cases where each request is sent original and compressed Claim source β - Repository description
-
-31% input / -74% output, measured live.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- 112 live A/B cases, each sent twice and scored, billed at real rates; answer quality 78.9% original vs 82.2% compressed. Named-benchmark suites (GSM8K n=12 plus three others n=20) run on qwen3-next-80b at a conservative preset. Head-to-head tables against RTK, Caveman and others in crates/llmtrim-cli/bench (not reviewed for this page).
- Limits and context
- Stated limits: Anthropic and Gemini token counts are approximate (OpenAI exact); output savings are not measured live, the status 'saved' figure is input-side; the default is quality-gated, not lossless (use the safe preset for byte-faithful). The dollar saving depends on the model's output/input price ratio (-66% here, projected -57% at GPT-4o rates, -59% at Claude Sonnet rates). Requires a private name-constrained CA in ~/.llmtrim and shell env changes. The README says compression 'cannot raise your bill or break a request; worst case is zero savings' (not tested for this page). MPL-2.0.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Magic Compact Compaction and pruning No quantified claim No third-party result 173 GitHub stars
Replaces old assistant turns with per-turn summaries, keeps user messages verbatim, and prunes bulky tool input and output into a cache retrievable through a read_omitted_content tool.
- Project claim
- No quantified claim.
- Limits and context
- README states 'Magic Compact development has been paused in favor of Operator Memory, its successor' and that Magic Compact 'remains fully functional'. Claims qualitative advantages ('Maximum token savings', 'Lossless quality') with no figure. On Claude Code it creates a new compacted session to resume, and /magic-stats is not implemented there.
-
maki Alternative agents 165 tokens per turn saved on reads, in the author's usage No third-party result 1.1k GitHub stars
Rust terminal coding agent with a tree-sitter index tool (file skeletons with line ranges before reads), a sandboxed code_execution tool for filtering data outside context, tiered subagent models and optional RTK use.
- Project claim
-
For my usage it adds 59 tok/turn but saves 224 tok/turn on read calls, saving 165 tok/turn.
Measured on: tokens per turn on the author's own usage: the index tool adds 59 and saves 224 on read calls (net 165); a personal measurement, no task set Claim source β - Project's own benchmark
- FrontierHarness Eval result for Maki 0.5.5: 70% pass rate and $2.06 per pass against the FrontierHarness Eval baselines (figures read from the image alt text in the README; report.zip linked from the release, not reviewed for this page).
- Limits and context
- Useful data point on RTK dilution stated by the author: with rtk installed maki saves about 50% of bash output tokens, but 'bash is just 12% of total token usage' (so about 6% overall) while reads are 65% of total usage, which is why he prefers the index tool; he plans his own bash output filtering. Also states over 90% of maki's code was written by maki, guided by humans. The README's benchmark chart is an image, so the baseline agents and models are not in the text. License not stated in the README.
Project page βThird-party results (0)
None found on 2026-09-30.
-
mcp2cli MCP tools and gateways 96% to 99% of tool-schema tokens per turn No third-party result 2.4k GitHub stars
Python CLI that turns an MCP server, OpenAPI spec or GraphQL endpoint into shell commands at runtime, so the agent discovers tools on demand instead of holding all tool schemas in context.
- Project claim
-
Save 96β99% of the tokens wasted on tool schemas every turn.
Measured on: tokens of MCP/OpenAPI tool schemas resident in context each turn (not total session tokens, not the bill) Claim source β - Project's own benchmark
- README points to a token-savings test (tests/test_token_savings.py) and a blog writeup at orangecountyai.com (not reviewed for this page). Example in README: --list on 96 tools costs about 1,400 tokens; top 10 names only about 20 tokens.
- Limits and context
- The README figure is about schema overhead only; it does not quantify the cost of discovery and help calls the agent makes instead (--list, --search, --help), nor task success. Skill install is separate (npx skills add). TOON output is described as 40-60% fewer tokens than JSON for large uniform arrays (unmeasured in README). MIT.
Project page βThird-party results (0)
None found on 2026-09-30.
-
PandaFilter Shell output filters 82% of command output on its test fixtures No third-party result 100 GitHub stars
Rust CLI (panda) that hooks agent tool calls, compresses command output with per-command handlers and an on-device BERT summarizer, diffs file re-reads and saves a session digest before compaction.
- Project claim
-
β82%
Measured on: tokens in command outputs, summed over the 47 command types of the project's test fixtures (81,882 to 14,347) Claim source β - Project's own benchmark
- handler_benchmarks.rs unit-test fixtures, one row per command (pip install 1,787 to 9 tokens, git diff 6,370 to 861, tsc 2,598 to 1,320, stylelint 1,100 to 845). Number of fixtures per command and the tokenizer are not stated in the README; no agent task or model.
- Limits and context
- Range across rows is -23% (stylelint) to -99%. README also claims 'Both save 60β95% of re-read tokens' for file delta re-reads (v1.3.0) with no measurement shown. First run downloads a ~90 MB BERT model (all-MiniLM-L6-v2) from HuggingFace. Adaptive router is opt-in (use_router = true). Internal inconsistency: the agents section says 7 agents while the install section lists 8. MIT. Repo name is homebrew-pandafilter, project name PandaFilter (README links AssafWoo/PandaFilter).
Project page βThird-party results (0)
None found on 2026-09-30.
-
Parsec (Dasein) API proxies and context compression 44% cheaper per solved task, in its maker's benchmark No third-party result
Hosted service from Dasein that curates the coding agent's working context (native Anthropic Messages proxy plus a turn-0 brief hook and a stop-decision hook) so the model carries less context.
- Project claim
-
Parsec, our first product, makes Claude Code and other frontier coding agents do more with less: Cheaper 44%, More successful 10%, Less input needed 58%
Measured on: not stated on the page; the bench README defines -44% as cost per solved task versus baseline on 100 SWE-bench Verified tasks (the bench reports 62 solved vs 57, total cost -39%, input tokens -54%) Claim source β - Project's own benchmark
- Dasein Code-Compression Bench, run 2026-07-04: headless Claude Code, claude-sonnet-4-6, 100 SWE-bench Verified tasks, official Docker grader, cache-aware pricing, one run per arm. Parsec 62/100 solved, $1.45 per solved task (-44% vs baseline), total $89.65 (-39%), input 144.8M (-54%), wall clock 10.8 h (-25%).
- Limits and context
- Dasein makes Parsec: the site says 'Parsec, our first product', the bench README says 'Sponsored and operated by Dasein' and FACT-VS-FICTION.md says 'Benchmark sponsored and operated by Dasein; the methodology is identical for every arm, including Parsec's'. No public repo: the bench's Parsec arm is 'a thin over-the-wire client to a hosted compression service; this repo contains no vendor internals' and needs a DASEIN_API_KEY, so outsiders cannot reproduce that arm. 3 of 100 Parsec runs ended by context limit. The page's '10% more successful' is not the same as the bench's 62 vs 57 (+8.8%); the page gives no denominator for its three figures. Single run per arm, no confidence intervals in the README.
Third-party results (0)
None found on 2026-09-30.
- Hosted service
- Made by the benchmark sponsor
-
Ponytail Output style about 20% cheaper on 12 tasks with Haiku 4.5 3 third-party results 149k GitHub stars
Skill that makes the agent write the smallest working solution by walking a ladder (does it need to exist, reuse, stdlib, native feature, one line) while keeping validation, security and accessibility.
- Project claim
-
~54% less code (up to 94%) Β· ~20% cheaper Β· ~27% faster Β· 100% safe
Measured on: lines of code in the git diff on 12 feature tasks in one FastAPI plus React repo, headless Claude Code on Haiku 4.5, n=4, against the same agent with no skill; cost -20%, tokens -22%, time -27% Claim source β - Project's own benchmark
- Agentic benchmark (2026-06-18): 12 tasks, Haiku 4.5, n=4. Versus baseline, ponytail LOC -54%, tokens -22%, cost -20%, time -27%, safety 100%; a caveman control LOC -20%, tokens +7%, cost +3%; a 'YAGNI + one-liners' prompt LOC -33%, cost -21%, safety 95%.
- Limits and context
- README says an earlier single-shot benchmark reported 80-94% less code and that against a fair agentic baseline that is 'the per-task ceiling, not the average'; issue 126 pointed out the baseline padded its answers. Savings are in code written, not context read, and the cut is 'near zero on code that is already minimal'. README notes a terse reasoning model (GPT-5.5) can cost more. Author-run benchmark. GreenPT is listed as a sponsor and also publishes the competing Honey skill. MIT licence.
Third-party results (3)
- Stet.sh β Workload cost, run 1 then run 2 +20% then β2% (20% more expensive, then 2% cheaper) Tests 1 win, 4 losses, 15 ties.
- THOL β End-to-end cost, 7 long tasks β3.5% (3.5% cheaper) Interval 18.1% cheaper to 12.4% more expensive.
- Tura β Modeled cost, one task, 2 runs β8.9% (8.9% cheaper) Score 80.8% vs 78.9% baseline. The two runs spread by 51.7% of their mean.
- Earlier figure reframed as a ceiling
-
pxpipe API proxies and context compression 59% to 70% of the end-to-end bill on Claude Code traffic No third-party result 7.5k GitHub stars
Local proxy that renders bulky request context (system prompt, tool docs, old history, large tool results) as dense PNG images so the model reads it through the vision channel at fewer tokens.
- Project claim
-
At current Fable list prices that lands as a ~59β70% lower end-to-end bill
Measured on: the whole bill on Claude Code traffic: all requests including uncompressed small ones, cache writes and reads, and output tokens (59% on a 13,709-request snapshot, ~70% on an 8,904-compressed-request trace) Claim source β - Project's own benchmark
- SWE-bench Lite pilot 10/10 both arms at -65% request size; SWE-bench Pro 14/19 with pxpipe vs 15/19 without at -60% (small n, one split re-resolved 3/3); a model quality matrix (arithmetic, gist, dense hex) per model, model named Fable 5 by the README.
- Limits and context
- Self-described as lossy: exact 12-char hex recall 13/15 on Fable 5, 0/15 on Sol, misses are silent confabulations; byte-exact values must stay text. One reported real-world failure (a name recalled wrongly from imaged history). Loses money on sparse prose; a profitability gate images only where it wins. Savings depend on the client re-sending uncached bulk (Claude Code typically 60-70%). Compressed-only figures run higher (~72-74%) and are quoted separately. Effective-context benefits are stated as unproven. Savings are measured per request against a free count_tokens counterfactual logged locally. README says most commits were written by agent sessions running behind pxpipe. MIT. Windows is community-supported.
Project page βThird-party results (0)
None found on 2026-09-30.
-
QMD (Query Markup Documents) Code index and search No quantified claim No third-party result 30k GitHub stars
On-device search engine for markdown notes, transcripts and docs that combines BM25, vector search and an LLM re-ranker through node-llama-cpp with GGUF models, usable by agents.
- Project claim
- No quantified claim.
- Limits and context
- No token-saving claim; it is a local document search tool for agentic flows, not a coding-context reducer. Downloads local GGUF models.
Project page βThird-party results (0)
None found on 2026-09-30.
-
routatic-proxy Model routing No quantified claim No third-party result 976 GitHub stars
Go CLI proxy that accepts Claude Code's Anthropic-format requests, converts them to provider formats and routes them to OpenCode Go, OpenCode Zen, AWS Bedrock, OpenRouter or Anthropic with scenario routing and fallback chains.
- Project claim
- No quantified claim.
- Limits and context
- No measured saving in the README. The only price statement is that OpenCode Go costs $5/month (then $10/month), which describes a provider's subscription, not a measured reduction. Also formerly named oc-go-cc (kept as compatibility alias). AGPL-3.0. Cost effect depends entirely on which models the user routes to; quality of substituted models is not evaluated in the README.
Project page βThird-party results (0)
None found on 2026-09-30.
-
RTK (Rust Token Killer) Shell output filters up to 90% of the bash output the agent reads 6 third-party results 82k GitHub stars
Single Rust binary that rewrites Bash commands through a hook and compresses their output (grouping, truncation, deduplication) before the agent reads it.
- Project claim
-
RTK cuts up to 90% of the bash output your agent reads. That is what RTK measures, and it is not the same as cutting your bill by 90%.
Measured on: bytes of bash (shell command) output read by the agent, estimated as bytes / 4 (RTK ships no tokenizer) Claim source β - Repository description
-
CLI proxy that reduces LLM token consumption by 60-90% on common dev commands.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- RTK AI Labs re-ran SkillsBench with RTK v0.45.0 (2026-08-07): -4.8% cost on 13 dev tasks (permutation p = 0.305), task-level results from -32% to +14%. It reports that RTK rewrote about one third of Bash calls, covering about 20% of tool-output characters, and puts the compression ceiling at about 3% to 4% of the bill. Published by RTK's maintainers.
- Limits and context
- The README states its own caveat: bash output is one contributor to input tokens, input tokens are only part of the bill, so the reduction dilutes at every step; percentages are reliable but absolute token numbers are approximate. The hook only runs on Bash tool calls: Claude Code built-in Read, Grep and Glob bypass it (README 'Important' note). Per-command figures in the README (for example pytest -90%, cargo build -80%) are reductions in command output, not session-level results. Disclosure: the README core-team list names Florian Bruniaux, who maintains this site, as a core contributor. Apache-2.0. Telemetry is opt-in. README version check text says 'rtk 0.28.2' while THOL tested v0.42.3, so the README copy is stale on that point.
Third-party results (6)
- Dasein Labs β Total cost +13% (13% more expensive) 54 resolved. Best cache hit rate (97.3%) but 6,131 agent steps vs 5,325.
- Stet.sh β Workload cost, run 1 then run 2 +13% then β9% (13% more expensive, then 9% cheaper)
- JetBrains (RTK) β Median cost per task, low effort then high effort +7.6% then +0.1% (7.6% more expensive, then 0.1% more expensive) Low effort p=0.004; high effort p=0.99. Quality unchanged.
- THOL β End-to-end cost, 7 long tasks +7.1% (7.1% more expensive) Interval 7.2% cheaper to 26.0% more expensive.
- Tura β Modeled cost, one task, 2 runs +7.2% (7.2% more expensive) Score 76.9% vs 78.9% baseline, 44% more agent rounds. The author says the data do not identify a plugin effect.
- Codepointer β Share of replayed spend saved β0.5% (0.5% cheaper) Replay, not a live run. 33% to 99% cut on the shell output it recognizes.
Maintainer response
- Guide author is a core contributor
-
Semble Code index and search about 99% of tokens against grep plus full-file reads No third-party result 6.2k GitHub stars
CPU-only code search library, CLI and MCP server that chunks code with tree-sitter and fuses static Model2Vec embeddings with BM25 to return ranked snippets for natural-language queries.
- Project claim
-
Uses ~99% fewer tokens than grep+read
Measured on: tokens needed to reach a given recall level in the project's retrieval benchmark (~1,250 queries over 63 repositories in 19 languages), compared with a grep plus full-file-read baseline; the savings command uses full matched files as baseline Claim source β - Repository description
-
Uses 99% fewer tokens than grep+read
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- Retrieval benchmark: NDCG@10 of 0.854, indexing ~380x faster and querying ~17x faster than CodeRankEmbed; semble hits 97% recall at 2k tokens while grep+read needs a 100k context to reach 85%.
- Limits and context
- README's own savings formula is (file chars - snippet chars) / 4 and it calls the full-file baseline 'conservative'; it measures retrieval token cost, not end-to-end agent sessions. MIT licence. Downloads a model from Hugging Face on first use.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Serena Code index and search No quantified claim No third-party result 30k GitHub stars
MCP server giving agents IDE-like symbol-level retrieval, refactoring and editing over language servers (40+ languages) or a paid JetBrains plugin, plus a memory system.
- Project claim
- No quantified claim.
- Project's own benchmark
- README quotes agent self-evaluations (Opus 4.6 in Claude Code, GPT 5.4 in Codex CLI and Copilot CLI) from an evaluation prompt of ~20 routine tasks; these are qualitative statements, with a fuller evaluation on its docs site (not reviewed for this page).
- Limits and context
- Says it makes agents 'faster, more efficiently and more reliably' but gives no token or cost figure. JetBrains backend is a paid plugin with a free trial. Licence: SolidLSP MIT, rest GPL-3.0-or-later, and distributions combining both are subject to the GPL; contributions require a CLA. README warns not to install via MCP or plugin marketplaces because their commands are outdated.
Project page βThird-party results (0)
None found on 2026-09-30.
-
shunt (Spotify Portal AI plugins) Delegation 82% to 94% of Claude-side tokens on large reads No third-party result 2.3k GitHub stars
Claude Code plugin whose hooks block large file reads and redirect them to a cheaper worker model (AiKA modes bulk-reader and code-writer) called through the Portal CLI actions registry.
- Project claim
-
A Claude Code plugin that shunts I/O-heavy work to AiKA modes, saving 82-94% of tokens on large file reads and boilerplate generation.
Measured on: Claude-side tokens for reading large files (the corpus goes to the worker model and never enters Claude's context) on a 162K-line Java monorepo; mean bulk-read savings 90% Claim source β - Project's own benchmark
- Four scenarios on a 162K-line Java monorepo: single large file 33,684 to 5,737 tokens (82%), source plus test pair 75,990 to 4,148 (94%), multi-file cross-service 16,221 to 821 (94%), code-write 40,614 tokens plus generation vs 833 lines to disk (no percentage). Number of runs, worker model and quality of the worker's answers are not stated.
- Limits and context
- Requires a Spotify Portal instance with AiKA enabled, the portal plugin and jq, so it is not usable standalone. Figures count Claude's context tokens; the worker model's cost is not netted in. Hook default blocks Read on files over 350 lines (SHUNT_MIN_LINES). Stated limits: no enforcement for code-writer, request must fit in ARG_MAX (SHUNT_MAX_PAYLOAD_BYTES), 180 s timeout per invocation, each call is one-shot with no follow-up state. README says not to delegate debugging, editing, small files or architectural decisions. The repo README says Claude Code only for now.
Project page βThird-party results (0)
None found on 2026-09-30.
-
snip Shell output filters 60% to 90% of tokens in shell command output No third-party result 456 GitHub stars
Go CLI that sits between the agent and the shell and filters command output through declarative YAML pipelines, passing unmatched commands through unchanged.
- Project claim
-
snip - Reduce LLM Token Usage by 60-90%
Measured on: not stated; the supporting table measures tokens in individual shell command outputs (before and after filtering) Claim source β - Repository description
-
CLI proxy that reduces LLM token usage by 60-90%.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- Per-command table on the author's repository: cargo test 591 to 5 tokens (99.2%), go test ./... 275 to 8 (97.1%), git log 371 to 53 (85.7%), git status 112 to 16 (85.7%), git diff 355 to 66 (81.4%); plus one real Claude Code session report (128 commands, 2.3M tokens saved, average 99.8%). No model, task success or session cost stated.
- Limits and context
- The 60-90% headline is not tied to a stated measurement in the README; the sample session shows 99.8%, driven by go test output (top three rows are go test runs). Savings are on command output only. Filters are YAML data, extensible. The README documents a caveat for the compact_path step.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Stacklit Code index and search about 4,000 tokens of index for a 108,000-line repo No third-party result 106 GitHub stars
Go binary that scans a repo with tree-sitter into a committed stacklit.json module index (plus a Mermaid dependency diagram and optional ~250-token map injected into CLAUDE.md) with a 7-tool MCP server.
- Project claim
-
108,000 lines of code. 4,000 tokens of index.
Measured on: size of the generated index on FastAPI (108,075 lines, 4,142 index tokens); the README's 'Without stacklit: ~400,000 tokens' versus 'With stacklit: ~4,000 tokens' comparison is an illustrative estimate of an agent reading 8-12 files Claim source β - Project's own benchmark
- Index sizes measured on four projects: Express.js 3,765 tokens (21,346 lines), FastAPI 4,142 (108,075), Gin 3,361 (23,829), Axum 14,371 (43,997).
- Limits and context
- README measures index size, not agent session cost; the 400,000-token baseline is not sourced to a run. The optional --summary flag calls the Claude API, otherwise parsing is local. MIT licence.
Project page βThird-party results (0)
None found on 2026-09-30.
-
tilth Code index and search No quantified claim No third-party result 352 GitHub stars
Rust binary and MCP server that parses code with tree-sitter (no index) to return outlines, definitions before usages, callers, file dependencies, function-level diffs and a combined grok view of a symbol.
- Project claim
- No quantified claim.
- Limits and context
- README section 'A note on benchmarks': the earlier cost benchmark (v0.5.0, early-2026 models) 'no longer describes the tool' and was retired, so 'tilth makes no cost claim'; harness kept under the benchmark-archive tag. States limits: syntax only, callers matched by name, no type resolution. MIT licence.
Third-party results (0)
None found on 2026-09-30.
- No cost claim (retired benchmark)
-
Token Optimizer Compaction and pruning about 28% of the author's 30-day tokens, against a modelled baseline No third-party result 2.5k GitHub stars
Claude Code plugin that audits context waste and applies hooks for delta reads, structure-map skeletons, tool-output compression, checkpoint and restore across compaction, and model routing, with a local dashboard.
- Project claim
-
About ~$1,396 in API-equivalent value for the month, and 59.6M tokens never sent to the model, roughly 28% of the workload (150.8M spent, about 210.4M without Token Optimizer).
Measured on: total tokens over the author's own last 30 days (1,698 sessions, snapshot 2026-09-14), against a modelled no-plugin counterfactual; most of the dollar value is 'Repeat reads avoided (~$1,124)', labelled estimated Claim source β - Project's own benchmark
- README says compression claims are tested against real sessions and an 87-fixture suite, with methodology in BENCHMARK.md; per-feature savings listed as Delta Mode ~20% on re-reads, Structure Map ~30% (up to 99% per file), Bash Compression ~10%, Search Compression ~15%.
- Limits and context
- README separates metered savings (~$126) from estimated ones and says lean-output nudges are 'Estimated 10-15%; not counterfactually measured'. Licence is PolyForm Noncommercial (badge), so commercial use is not covered. The README also asserts that command-output compressors like RTK cover 15-25% of context; that figure is the project's own claim.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Token Optimizer MCP MCP tools and gateways 0.926x median cost per task on THOL tasks, run by the vendor No third-party result 538 GitHub stars
MCP server with agent hooks that caches and diffs file reads, can deny large full-file reads in enforce mode, keeps a per-project knowledge graph, and reports before/actual-return savings.
- Project claim
-
median cost 0.926x, median turns 0.671x
Measured on: end-to-end agent session cost and turns per task versus a no-proxy control on THOL (16 real tasks), one run per task; the README states the aggregate interval still spans 1.0 Claim source β - Project's own benchmark
- Favourable and unfavourable THOL modes, both from its README. Default 'assist' mode: score 0.971 in 14.4 turns vs control 0.969 in 16.2. 'enforce' mode: score 0.960 in 20.3 turns and 1.471x control cost per task (median across 17 tasks; cheaper on only 2 of them). Compression proof table vs a reimplementation of a competitor's design: code-search 97.6% vs 92.1%, sre-debugging 98.4% vs 92.2%, issue-triage 97.3% vs 72.8%, codebase-exploration 46.0% vs 47.4% (published as a loss). Whole-session replay: 0.953x at 10 turns, 0.915x at 20 turns (4 sessions).
- Limits and context
- Self-stated caveats: one run per task; retention units directly visible in the text sent: 345 for this tool vs 1,890 for the competitor design on the competitor's fixtures (reduction reached partly by eliding harder into a spill, recoverable at the cost of a turn); the competitor arm is the vendor's own reimplementation, not the competitor's binary. Two THOL figures use different task counts (16 tasks in the headline block, 17 in the enforce line). Default posture is assist, enforce is not default and is the unfavourable mode. Telemetry opt-in. MIT.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Token Savior Code index and search 80% of active tokens per task, reported and unverified No third-party result 1.2k GitHub stars
Python MCP server that indexes code by symbol for structural navigation, adds a SQLite-backed persistent memory engine, and optionally rewrites or compacts Bash output through hooks.
- Project claim
-
97.9% on tsbench at -80% tokens.
Measured on: active tokens per task (17,221 plain Claude Code versus 3,395 with Token Savior) over 96 tasks on Claude Opus 4.7, May 2026; score 188/192 versus 141/180 Claim source β - Repository description
-
MCP server that gets Claude to 97.9% (188/192) on a real coding benchmark at -80% active tokens and -83% wall time, vs 78.3% plain.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- tsbench, 96 tasks on a generated toy repo with planted traps; also wall time 110.6 s to 18.9 s per task. README says the benchmark repository is not public and returns 404, so the figures are 'as reported' and 'unverified'.
- Limits and context
- Withdrawn claim: a re-measurement published 2026-08-09 was withdrawn 2026-08-10 because across 143 sessions exactly one called a Token Savior tool (MCP deferred-tool loading hid the tools), so it 'did not measure this server at all'; headline figures stand only 'as reported and unverified' and re-measurement is 'open work'. README also says the PostToolUse Bash compactors 'do not shrink the current turn'; only the PreToolUse rewriter does. Its integration table lists RTK as 'repo currently unreachable' (not verified here). MIT licence.
-
Token-Goat Compaction and pruning 40% to 90% measurement not stated No third-party result 130 GitHub stars
Hook layer that blocks or shrinks repeat file, skill and web reads, strips noisy command output, downsizes screenshots, and injects a structured manifest of edited files before compaction.
- Project claim
-
Reduces AI token use/costs by 40-90%, and improves its focus.
Measured on: not stated; the README also says 'Sessions drop 40-90%+ in cost' and lists 85% smaller reads, 97.4% image compression and 80 to 97% smaller pytest output as per-operation figures Claim source β - Limits and context
- Range is self-reported with no method in the README front matter; the '1.1 Gt tokens saved' banner is a cumulative counter, not a controlled result. Licence is PolyForm Noncommercial. Badge says macOS untested. Built and maintained by one person per the README.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Tokenade Shell output filters 38.9% cheaper on THOL long sessions, a leaderboard run by its author No third-party result 8 GitHub stars
Closed-source local binary (npm wrapper) that compacts command output, wraps MCP servers, dedups re-reads, adds code and web search tools, injects a terse-output style and redacts secrets.
- Project claim
-
The #1 tool to cut your AI agent's token bill.
Measured on: not stated in the sentence; the README ties it to its first place on THOL, a leaderboard maintained by Tokenade's author (38.9% cheaper on 7 long tasks) Claim source β - Project's own benchmark
- The THOL leaderboard, run by Tokenade's author (see notes). Tokenade adoption column reads 5/70 runs (agent called the tool in 5 of 70 long-session runs).
- Limits and context
- CONFLICT OF INTEREST, confirmed at source: the THOL page says it 'is maintained by the author of one of the measured tools (tokenade)' and that the tool at the top of the table is maintained by the benchmark's author; it also says small gaps between neighbouring rows are not a ranking. Tokenade itself is freemium: requires a free account (10M tokens/month), paid plans above; the repo holds only the launcher and installer, the binary is downloaded from downloads.tokenade.net with SHA-256 verification. Only aggregate counters are sent (README). The same THOL table shows other tools near zero, RTK at -7.1%, lean-ctx at -10.3% and Headroom at -52.8%.
Third-party results (0)
None found on 2026-09-30.
- Closed source
- Made by the benchmark maintainer
-
tokf Shell output filters 60% to 90% of CLI output tokens, estimated No third-party result 199 GitHub stars
Rust CLI that runs a command, applies a TOML filter to its output and emits a reduced version, installed as a hook for coding agents.
- Project claim
-
reduce LLM context consumption from CLI commands by 60β90%
Measured on: not stated; tokens in CLI command output, estimated as bytes / 3.5 and labelled est. Claim source β - Project's own benchmark
- Before/after examples only (cargo test 61 lines to 1 line, git push 8 lines to 1 line). The README also documents a tokenizer calibration: bytes-per-token divisor measured against cl100k across its test corpus (raw output 3.67, filtered output 2.98, combined 3.53), hence the 3.5 constant.
- Limits and context
- Strong self-stated caveats: every token figure is an estimate ('est.'), cl100k is not Claude's tokenizer, the corpus is weighted toward cargo, git, npm and docker. Percentages are reliable for filters that delete lines (cargo/check est. 84.4% vs real 84.6%) but not for filters that rewrite (docker/ps est. 0.0% vs real -300.0%, git/push est. 33.3% vs real -50.0%); README says to treat reduction percentages on rewriting filters as indicative, not as a claim. Changing the divisor from 4 to 3.5 created a step change in stored absolute counts. MIT.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Tura Alternative agents 77.5% fewer tokens than Codex CLI, Direct mode No third-party result 646 GitHub stars
Open-source agent runtime harness (Rust, TUI, GUI, gateway) that exposes one macro command_run tool executing a multi-step command graph in one model turn, with task-scoped context and CLI-driven compaction.
- Project claim
-
Tura: 16.7% better performance, 77.5% fewer tokens.
Measured on: aggregate tokens versus Codex CLI across 20 DeepSWE v1.1 tasks run three times per agent (180 sessions), GPT-5.6 Sol; 77.5% fewer tokens is the Direct mode, 16.7 percentage points higher success is the Balanced mode Claim source β - Repository description
-
Build agent that uses 80% less token and delivers better results.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- DeepSWE v1.1, 20 tasks x 3 runs: Direct 65.0% verifier success vs Codex CLI 63.3% with 77.5% fewer aggregate tokens and 69.1% fewer turns; Balanced 80.0% success (16.7 points above Codex CLI) with 31.1% fewer tokens and 35.8% fewer turns. Also 5 rewrite tasks (Balanced 389/472 vs Codex CLI 351/472 over 10 sessions each) and 2 design tasks reviewed separately. Benchmark repo Tura-AI/benchmark is run by the vendor.
- Limits and context
- The headline mixes two modes: 77.5% fewer tokens (Direct) and +16.7 points (Balanced) do not occur together; the title says '16.7% better performance' while the body says 16.7 percentage points. The comparison is against Codex CLI, not Claude Code. README states there is no ablation proving command_run alone causes the lower turns and tokens, and that published results do not establish equivalent quality for other providers (Anthropic, Gemini, OpenAI-compatible, local); see docs/KNOWN_ISSUES.md and ROADMAP.md. Compaction claim (2.6 vs an estimated 5.4 rounds to resume) compares explicit Tura events with an estimate for Codex inferred from input-token drops. AGPL-3.0-or-later. Targets not stated as a list; provider setup is per-user.
Project page βThird-party results (0)
None found on 2026-09-30.
-
Weave Router Model routing 40% to 70% of cost, per the repository description; the README states no figure No third-party result 5.4k GitHub stars
Go proxy for Anthropic, OpenAI and Gemini APIs that picks a model per request (per action) with an in-process embedding cluster scorer, with BYOK keys and OTLP traces.
- Project claim
- No quantified claim.
- Repository description
-
Cut costs 40-70% with just an endpoint change.
The short "About" text on GitHub, which can differ from the README. - Project's own benchmark
- bench/README.md reproduces Codex-harness runs, k=2 attempts: SWE-Atlas QnA router /beta vs GPT-5.6 Sol 47.6% vs 53.2% task-mean pass, $262 router-billed vs $445 list; vs Luna 60.1% vs 46.4% but $692 billed vs $127 list; vs Astra (max) 55.6% vs 57.7%, $573 vs $1,250 list. Terminal-Bench 4.0 vs Sol 24.6% vs 30.3% at $399 vs $633 list.
- Limits and context
- Cost columns mix bases: router-billed cost is computed by the router at catalog rates with cache writes at 1.25x, while the direct arms' list cost cannot see cache writes; bench README states this. Against Luna the router cost more ($692 vs $127) while scoring higher. The router arm ran against the vendor's staging router with a named routing package. Renamed and relicensed: repository is now weave-os/router (formerly workweave/router, npm alias @workweave/router kept); the Apache-2.0 release line starts at router-v0.2.24, earlier tags keep the licenses they shipped with (README text). The npx quickstart points tools at the hosted Weave Router; prompts stay local only in the self-hosted stack. Benchmarks are Codex only, Claude Code harnesses out of scope.
-
WOZCODE (Woz) MCP tools and gateways 25% to 50% cheaper measurement not stated 1 third-party result
Paid Claude Code plugin that replaces built-in file tools with AST-aware Search, Edit and SQL tools and routes read-only exploration to a Haiku subagent.
- Project claim
-
80% on TerminalBench 2.0, 5β10Γ faster, 25β50% cheaper
Measured on: not stated on the page meta description; per the Dasein bench's FACT-VS-FICTION.md (third party, not checked at Woz's source) the cost figure is derived from live API usage fields with an undisclosed task mix, and the speed figure is a self-labelled estimate Claim source β - Project's own benchmark
- 80% TerminalBench 2.0; per the Dasein bench, corroborated on the official leaderboard for WOZCODE on Claude Opus 4.7, against a plain Claude Code entry measured on a different model (not checked at source).
- Limits and context
- Requires a Woz account (/woz-login). The plugin replaces Claude Code's commit and PR attribution with its own co-author line unless an attribution entry already exists in settings.json (plugin README). The bench (Dasein) reports input tokens -35% but slower runs (17.8 h vs 14.4 h baseline). Plugin repo has no license text in the README.
Third-party results (1)
- Dasein Labs β Total cost β13% (13% cheaper) 55 resolved, wall-clock +23%.
- Paid, closed
No tool matches these filters.
Also found, below the inclusion bar (37 projects, claims not reviewed)
- V-Songbird/hush
- MaxForAI/Tokenless
- evilayman/opencode-raven
- NodeNestor/claude-rolling-context
- AshishKumar4/better-compact
- giuliastro/HarnessTrim
- howardpen9/kimi-code-mcp
- AgusRdz/chop
- PsYcGoD/sage
- rahadiana/opencode-ultrapress
- Rani367/Skarn
- HandyS11/DotnetTokenKiller
- supermodeltools/cli
- mertcanaltin/composto
- marjoballabani/hypergrep
- yttrium400/reducethemtokens
- phlx0/tokenmiser
- Digital-Threads/token-pilot
- Calluking/ContextSniper
- DexopT/MCE
- Dukeabaddon/Gate-MCP
- AttemorySystem/attemory
- sliday/tamp
- agiwhitelist/tokdiet
- Aditya-Tripuraneni/CacheLane
- Aimaghsoodi/foveance
- PromptForcePrime/prefex
- Neolambo/glyph-compress
- flaviomartil/token-skein
- ypollak2/llm-router
- zx1160763849-hash/codex-cost-router-skills
- rayline-ai/rayline
- TheColliery/CoalTipple
- Adam-Duchemann/agent-model-routing
- 0p9b/TLDR
- GabrielBarberini/laconic
- AkashAi7/stenographer-mode
How to read the results
- Two quantities
- Vendor claims measure one layer: shell output, output tokens, or a single content type. Benchmarks measure the whole task, so the two numbers rarely describe the same quantity.
- Dated models
- The benchmarks used claude-sonnet-4-6, claude-sonnet-5, GPT-5.6 Sol and an unnamed Claude Sonnet; the Codepointer replay applied Opus 4.8 prices. None tested Claude Sonnet 5.5, Claude Opus 5.5 or the GPT-6 family, so read these results as dated evidence about the mechanism, not as current rankings.
- Who publishes
- Three benchmarks come from a vendor with a stake: Dasein makes Parsec, the THOL maintainer makes Tokenade, and the Tura author builds a competing agent. Each says so, and the benchmark table below shows it.
- Sign convention
- Negative means cheaper or fewer tokens than the tool-free baseline. Positive means more expensive or more tokens.
Results by tool
Every third-party result from the 8 benchmarks, one row per tool. Rows are sorted by name, not by outcome. THOL publishes the opposite sign convention; its figures are converted here.
- Cost change
- Token-count change
- Left of zero: cheaper or fewer tokens than the tool-free baseline
- Right of zero: more expensive or more tokens
-
Caveman 6 results cost β19% to +9%output tokens β8.5%
- Dasein Labs β Total cost β19% (19% cheaper; cost) 58 resolved.
- Stet.sh β Workload cost, run 1 then run 2 +9% then β12% (9% more expensive, then 12% cheaper; cost)
- JetBrains (Caveman) β Output tokens β8.5% (8.5% fewer output tokens; output tokens) Forced activation, so a ceiling. Quality 8 better, 10 worse, 64 equal.
- THOL β End-to-end cost, 7 long tasks β13.6% (13.6% cheaper; cost) Interval 30.2% cheaper to 3.6% more expensive.
- Tura β Modeled cost, one task, 2 runs β3.9% (3.9% cheaper; cost) Score 86.5% vs 78.9% baseline.
- Codepointer β Share of replayed spend saved β0.4% (0.4% cheaper; cost) Replay, not a live run.
-
claude-token-efficient 1 result cost β11.6%
- THOL β End-to-end cost, 7 long tasks β11.6% (11.6% cheaper; cost) Interval 24.2% cheaper to 1.4% more expensive.
-
code-review-graph 1 result cost β8.5%
- THOL β End-to-end cost, 7 long tasks β8.5% (8.5% cheaper; cost) THOL reports that the agent never used the tool, so the gap is not a tool effect.
-
CodeGraph 1 result cost β7.6%
- THOL β End-to-end cost, 7 long tasks β7.6% (7.6% cheaper; cost) Interval 21.6% cheaper to 8.7% more expensive. Opt-in tool.
-
Context Mode 1 result cost +33% to +72%
- Stet.sh β Workload cost, run 1 then run 2 +72% then +33% (72% more expensive, then 33% more expensive; cost)
-
Edgee 1 result cost β14.7%
- THOL β End-to-end cost, 7 long tasks β14.7% (14.7% cheaper; cost) Interval 37.1% cheaper to 11.1% more expensive. THOL flags no token accounting for this tool.
-
Fermat 1 result cost β22%
- Dasein Labs β Total cost β22% (22% cheaper; cost) Run on 2026-09-21, not paired with the other arms. 55 resolved.
-
Graphify 2 results cost β3.3% to +1%
- Marmelab β Cost on the large transformation +1% (1% more expensive; cost)
- THOL β End-to-end cost, 7 long tasks β3.3% (3.3% cheaper; cost) THOL reports that the agent never used the tool, so the gap is not a tool effect.
-
Headroom 4 results cost β4% to +52.8%
- Dasein Labs β Total cost +44% (44% more expensive; cost) 58 resolved.
- Marmelab β Cost per operation, cache mode then token mode β4% then +32% (4% cheaper, then 32% more expensive; cost) One task, three runs. The authors call -4% noise.
- THOL β End-to-end cost, 7 long tasks +52.8% (52.8% more expensive; cost) Interval 10.2% to 120.2% more expensive, entirely above zero.
- Codepointer β Share of replayed spend saved β2.8% (2.8% cheaper; cost) Replay, not a live run. Median 54% cut on the grep and diff output it touches.
-
lean-ctx 1 result cost +10.3%
- THOL β End-to-end cost, 7 long tasks +10.3% (10.3% more expensive; cost) Interval 12.6% cheaper to 46.4% more expensive.
-
LSP (code navigation) 1 result cost β13%
- Marmelab β Cost on the large transformation β13% (13% cheaper; cost)
-
Ponytail 3 results cost β8.9% to +20%
- Stet.sh β Workload cost, run 1 then run 2 +20% then β2% (20% more expensive, then 2% cheaper; cost) Tests 1 win, 4 losses, 15 ties.
- THOL β End-to-end cost, 7 long tasks β3.5% (3.5% cheaper; cost) Interval 18.1% cheaper to 12.4% more expensive.
- Tura β Modeled cost, one task, 2 runs β8.9% (8.9% cheaper; cost) Score 80.8% vs 78.9% baseline. The two runs spread by 51.7% of their mean.
-
RTK 6 results cost β9% to +13%
- Dasein Labs β Total cost +13% (13% more expensive; cost) 54 resolved. Best cache hit rate (97.3%) but 6,131 agent steps vs 5,325.
- Stet.sh β Workload cost, run 1 then run 2 +13% then β9% (13% more expensive, then 9% cheaper; cost)
- JetBrains (RTK) β Median cost per task, low effort then high effort +7.6% then +0.1% (7.6% more expensive, then 0.1% more expensive; cost) Low effort p=0.004; high effort p=0.99. Quality unchanged.
- THOL β End-to-end cost, 7 long tasks +7.1% (7.1% more expensive; cost) Interval 7.2% cheaper to 26.0% more expensive.
- Tura β Modeled cost, one task, 2 runs +7.2% (7.2% more expensive; cost) Score 76.9% vs 78.9% baseline, 44% more agent rounds. The author says the data do not identify a plugin effect.
- Codepointer β Share of replayed spend saved β0.5% (0.5% cheaper; cost) Replay, not a live run. 33% to 99% cut on the shell output it recognizes.
-
squeez 1 result cost β6.9%
- THOL β End-to-end cost, 7 long tasks β6.9% (6.9% cheaper; cost) Interval 26.1% cheaper to 18.4% more expensive.
-
Woz 1 result cost β13%
- Dasein Labs β Total cost β13% (13% cheaper; cost) 55 resolved, wall-clock +23%.
Dashed line: Switch to GPT-5.6 Terra xhigh, Stet.sh, workload cost, run 1 then run 2: β49% then β49% (49% cheaper, then 49% cheaper). A model change, not a tool: the only repeated cost drop. Drawn for scale, not as a tool row.
Data table: every result as text
| Tool | Study | Metric | Result | Unit |
|---|---|---|---|---|
| Caveman | Dasein Labs | Total cost | β19% (19% cheaper) | cost |
| Caveman | Stet.sh | Workload cost, run 1 then run 2 | +9% then β12% (9% more expensive, then 12% cheaper) | cost |
| Caveman | JetBrains (Caveman) | Output tokens | β8.5% (8.5% fewer output tokens) | output tokens |
| Caveman | THOL | End-to-end cost, 7 long tasks | β13.6% (13.6% cheaper) | cost |
| Caveman | Tura | Modeled cost, one task, 2 runs | β3.9% (3.9% cheaper) | cost |
| Caveman | Codepointer | Share of replayed spend saved | β0.4% (0.4% cheaper) | cost |
| claude-token-efficient | THOL | End-to-end cost, 7 long tasks | β11.6% (11.6% cheaper) | cost |
| code-review-graph | THOL | End-to-end cost, 7 long tasks | β8.5% (8.5% cheaper) | cost |
| CodeGraph | THOL | End-to-end cost, 7 long tasks | β7.6% (7.6% cheaper) | cost |
| Context Mode | Stet.sh | Workload cost, run 1 then run 2 | +72% then +33% (72% more expensive, then 33% more expensive) | cost |
| Edgee | THOL | End-to-end cost, 7 long tasks | β14.7% (14.7% cheaper) | cost |
| Fermat | Dasein Labs | Total cost | β22% (22% cheaper) | cost |
| Graphify | Marmelab | Cost on the large transformation | +1% (1% more expensive) | cost |
| Graphify | THOL | End-to-end cost, 7 long tasks | β3.3% (3.3% cheaper) | cost |
| Headroom | Dasein Labs | Total cost | +44% (44% more expensive) | cost |
| Headroom | Marmelab | Cost per operation, cache mode then token mode | β4% then +32% (4% cheaper, then 32% more expensive) | cost |
| Headroom | THOL | End-to-end cost, 7 long tasks | +52.8% (52.8% more expensive) | cost |
| Headroom | Codepointer | Share of replayed spend saved | β2.8% (2.8% cheaper) | cost |
| lean-ctx | THOL | End-to-end cost, 7 long tasks | +10.3% (10.3% more expensive) | cost |
| LSP (code navigation) | Marmelab | Cost on the large transformation | β13% (13% cheaper) | cost |
| Ponytail | Stet.sh | Workload cost, run 1 then run 2 | +20% then β2% (20% more expensive, then 2% cheaper) | cost |
| Ponytail | THOL | End-to-end cost, 7 long tasks | β3.5% (3.5% cheaper) | cost |
| Ponytail | Tura | Modeled cost, one task, 2 runs | β8.9% (8.9% cheaper) | cost |
| RTK | Dasein Labs | Total cost | +13% (13% more expensive) | cost |
| RTK | Stet.sh | Workload cost, run 1 then run 2 | +13% then β9% (13% more expensive, then 9% cheaper) | cost |
| RTK | JetBrains (RTK) | Median cost per task, low effort then high effort | +7.6% then +0.1% (7.6% more expensive, then 0.1% more expensive) | cost |
| RTK | THOL | End-to-end cost, 7 long tasks | +7.1% (7.1% more expensive) | cost |
| RTK | Tura | Modeled cost, one task, 2 runs | +7.2% (7.2% more expensive) | cost |
| RTK | Codepointer | Share of replayed spend saved | β0.5% (0.5% cheaper) | cost |
| squeez | THOL | End-to-end cost, 7 long tasks | β6.9% (6.9% cheaper) | cost |
| Switch to GPT-5.6 Terra xhigh | Stet.sh | Workload cost, run 1 then run 2 | β49% then β49% (49% cheaper, then 49% cheaper) | cost |
| Woz | Dasein Labs | Total cost | β13% (13% cheaper) | cost |
Where tool tokens come from
Why compressing one tool's output moves the total by less than the output figure suggests.
How each benchmark was run
Eight checks per benchmark, plus who published it and what they stand to gain.
Read on 2026-09-30. "Not stated" means the publisher's page did not say, not that the answer is no. The tally counts the checks answered yes; it says nothing about the size of any result.
| Benchmark | Paired against a baseline | Task success measured | Cache reported separately | Dollar cost | Versions named | More than 10 tasks | Repeated runs | Code published |
|---|---|---|---|---|---|---|---|---|
| Code-Compression Bench β Yes on 6 of 8 checks | Yes | Yes | Yes | Yes | No | Yes | No | Yes |
| Publisher's interest: Sponsored and operated by Dasein Labs, which makes Parsec, one of the arms. The Parsec arm also injects a turn-0 brief and a stop decision, beyond compression. Method: 100 SWE-bench Verified tasks, headless Claude Code via the Python Agent SDK, claude-sonnet-4-6, official Docker grader, one run per arm, cache-aware cost. | ||||||||
| I Tested Six Ways to Save $$$ on 5.6 Sol β Yes on 4 of 8 checks | Yes | Yes | Not stated | Yes | No | No | Yes | Not stated |
| Publisher's interest: The author builds Stet.sh, the evaluation tool used. No compared tool is theirs. Method: 10 merged changes from one production repo, Codex CLI with GPT-5.6 Sol at medium effort as baseline, 7 arms, 2 repetitions, 140 runs. | ||||||||
| Cutting the Coding Agent Bill β Yes on 3 of 8 checks | Yes | No | Not stated | Yes | No | No | Yes | No |
| Publisher's interest: Marmelab builds Atomic CRM Builder, the workload. No compared tool is theirs. Method: Atomic CRM Builder tasks, 1 to 4 tasks per tool, 3 to 4 runs per configuration. Harness, model and tool versions not stated. | ||||||||
| RTK and Claude Code token savings β Yes on 7 of 8 checks | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Not stated |
| Publisher's interest: No affiliation with RTK stated. Method: 86 SkillsBench tasks, 425 billed trials, Harbor 0.18, Claude Code 2.1.201 headless, rtk 0.43.0, claude-sonnet-5, low and high effort, per-task medians with a Wilcoxon test. RTK AI Labs published a response on 2026-08-07 that questions the one-run-per-task design and reports -4.8% on 13 dev tasks (p = 0.305). | ||||||||
| Speaking to AI agents like cavemen saves 65% of tokens. We test. β Yes on 5 of 8 checks | Yes | Yes | Not stated | Yes | No | Yes | Yes | Not stated |
| Publisher's interest: No affiliation with Caveman stated. Method: 86 SkillsBench tasks, 82 clean pairs, about 240 billed trials, Harbor 0.17, claude-sonnet-5 at low effort, Caveman forcibly activated (its best case). | ||||||||
| Token-Harness Optimizer Leaderboard β Yes on 7 of 8 checks | Yes | Yes | Not stated | Yes | Yes | Yes | Yes | Yes |
| Publisher's interest: Maintained by the author of Tokenade, which tops the table. The page says so. Method: Headless Claude Code 2.1.206 with claude-sonnet-4-6, 17 tasks, 10 runs per task and tool. The headline keeps the 7 tasks where the tool-free control used at least 200,000 tokens (70 runs per tool) and reports the ratio of end-to-end USD with bootstrap 95% intervals. | ||||||||
| Token-saving plugins: measure cost per completed task β Yes on 7 of 8 checks | Yes | Yes | Yes | Yes | Yes | No | Yes | Yes |
| Publisher's interest: Published by the maintainer of Tura, a coding agent that competes with the harnesses these plugins extend. The author says so and calls the post not an independent review. Method: One task (rewrite the Rust eza tool in Python, 52 assertions), Codex CLI 0.144.1 with gpt-5.6-sol at high reasoning, 2 runs per arm against matched baseline runs. Cost is modeled from token counts at list prices, not billed. | ||||||||
| Cutting LLM token costs with rtk, headroom, and caveman β Yes on 3 of 8 checks | No | No | Yes | Yes | Not stated | Yes | No | Not stated |
| Publisher's interest: No relationship with the three projects stated. Method: Counterfactual replay of 500 sessions sampled from the author's own Claude Code history (614M tokens, $926.31 at list prices). Headroom was applied to the recorded payloads; RTK and Caveman at their published per-stream rates. No live runs, task success not measured. | ||||||||
-
Code-Compression Bench 6 of 8 checks
- Paired against a baseline: Yes
- Task success measured: Yes
- Cache reported separately: Yes
- Dollar cost: Yes
- Versions named: No
- More than 10 tasks: Yes
- Repeated runs: No
- Code published: Yes
Publisher's interest: Sponsored and operated by Dasein Labs, which makes Parsec, one of the arms. The Parsec arm also injects a turn-0 brief and a stop decision, beyond compression.
Method: 100 SWE-bench Verified tasks, headless Claude Code via the Python Agent SDK, claude-sonnet-4-6, official Docker grader, one run per arm, cache-aware cost.
Read the benchmark β -
I Tested Six Ways to Save $$$ on 5.6 Sol 4 of 8 checks
- Paired against a baseline: Yes
- Task success measured: Yes
- Cache reported separately: Not stated
- Dollar cost: Yes
- Versions named: No
- More than 10 tasks: No
- Repeated runs: Yes
- Code published: Not stated
Publisher's interest: The author builds Stet.sh, the evaluation tool used. No compared tool is theirs.
Method: 10 merged changes from one production repo, Codex CLI with GPT-5.6 Sol at medium effort as baseline, 7 arms, 2 repetitions, 140 runs.
Read the benchmark β -
Cutting the Coding Agent Bill 3 of 8 checks
- Paired against a baseline: Yes
- Task success measured: No
- Cache reported separately: Not stated
- Dollar cost: Yes
- Versions named: No
- More than 10 tasks: No
- Repeated runs: Yes
- Code published: No
Publisher's interest: Marmelab builds Atomic CRM Builder, the workload. No compared tool is theirs.
Method: Atomic CRM Builder tasks, 1 to 4 tasks per tool, 3 to 4 runs per configuration. Harness, model and tool versions not stated.
Read the benchmark β -
RTK and Claude Code token savings 7 of 8 checks
- Paired against a baseline: Yes
- Task success measured: Yes
- Cache reported separately: Yes
- Dollar cost: Yes
- Versions named: Yes
- More than 10 tasks: Yes
- Repeated runs: Yes
- Code published: Not stated
Publisher's interest: No affiliation with RTK stated.
Method: 86 SkillsBench tasks, 425 billed trials, Harbor 0.18, Claude Code 2.1.201 headless, rtk 0.43.0, claude-sonnet-5, low and high effort, per-task medians with a Wilcoxon test. RTK AI Labs published a response on 2026-08-07 that questions the one-run-per-task design and reports -4.8% on 13 dev tasks (p = 0.305).
Read the benchmark β -
Speaking to AI agents like cavemen saves 65% of tokens. We test. 5 of 8 checks
- Paired against a baseline: Yes
- Task success measured: Yes
- Cache reported separately: Not stated
- Dollar cost: Yes
- Versions named: No
- More than 10 tasks: Yes
- Repeated runs: Yes
- Code published: Not stated
Publisher's interest: No affiliation with Caveman stated.
Method: 86 SkillsBench tasks, 82 clean pairs, about 240 billed trials, Harbor 0.17, claude-sonnet-5 at low effort, Caveman forcibly activated (its best case).
Read the benchmark β -
Token-Harness Optimizer Leaderboard 7 of 8 checks
- Paired against a baseline: Yes
- Task success measured: Yes
- Cache reported separately: Not stated
- Dollar cost: Yes
- Versions named: Yes
- More than 10 tasks: Yes
- Repeated runs: Yes
- Code published: Yes
Publisher's interest: Maintained by the author of Tokenade, which tops the table. The page says so.
Method: Headless Claude Code 2.1.206 with claude-sonnet-4-6, 17 tasks, 10 runs per task and tool. The headline keeps the 7 tasks where the tool-free control used at least 200,000 tokens (70 runs per tool) and reports the ratio of end-to-end USD with bootstrap 95% intervals.
Read the benchmark β -
Token-saving plugins: measure cost per completed task 7 of 8 checks
- Paired against a baseline: Yes
- Task success measured: Yes
- Cache reported separately: Yes
- Dollar cost: Yes
- Versions named: Yes
- More than 10 tasks: No
- Repeated runs: Yes
- Code published: Yes
Publisher's interest: Published by the maintainer of Tura, a coding agent that competes with the harnesses these plugins extend. The author says so and calls the post not an independent review.
Method: One task (rewrite the Rust eza tool in Python, 52 assertions), Codex CLI 0.144.1 with gpt-5.6-sol at high reasoning, 2 runs per arm against matched baseline runs. Cost is modeled from token counts at list prices, not billed.
Read the benchmark β -
Cutting LLM token costs with rtk, headroom, and caveman 3 of 8 checks
- Paired against a baseline: No
- Task success measured: No
- Cache reported separately: Yes
- Dollar cost: Yes
- Versions named: Not stated
- More than 10 tasks: Yes
- Repeated runs: No
- Code published: Not stated
Publisher's interest: No relationship with the three projects stated.
Method: Counterfactual replay of 500 sessions sampled from the author's own Claude Code history (614M tokens, $926.31 at list prices). Headroom was applied to the recorded payloads; RTK and Caveman at their published per-stream rates. No live runs, task success not measured.
Read the benchmark β
What research papers found
Studies that measure token reduction against cost, success or time, rather than tokens alone.
arXiv preprints, read at their abstract pages on 2026-09-30. Preprints are not peer-reviewed. The first group studies cost without selling a method. In the second group the authors propose the method they measure, which is normal in research but is a conflict to keep in mind.
Studies of cost and behavior (9)
- Token Reduction Is Not Cost Reduction β Cost billed or logged
The largest compression setup cut delivered tool-output tokens by 38.4% and raised billed cost by 6.8%. Token reduction and cost reduction correlated weakly (Pearson r = 0.15).
-
Nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench. On Terminal-Bench with Qwen, policies using about one third of the tokens took 20% to 80% longer than the uncompressed agent.
- CAVEWOMAN: How LLMs Behave Under Linguistic Input and Output Compression β Cost billed or logged
Output compression cut realized cost 1.4x to 2.4x per model on most API models. Input compression raised net cost (about 1.15x on the five-benchmark mean). Single generations, not an agent loop.
- Don't Break the Cache: Prompt Caching for Long-Horizon Agentic Tasks β Cost billed or logged
Prompt caching reduced API costs by 41% to 80% and time to first token by 13% to 31% across providers, over 500 agent sessions.
-
June 2026 traces (3.2M users, 13M sessions, 761M LLM calls): cache hit rates average 90% within a turn and fall to 55% across turn boundaries.
-
Across 2,700 runs with Kimi K3, reducing a full task specification to a bare user story raised token spend by 29.7%.
- When Does Restricting a Coding Agent to execute_code Help? β Cost billed or logged
A code-only tool set was cheaper than or tied with the cheapest tool-rich arm in three of four cells; on SWE-bench with Claude it was 14.4% costlier, not significant. Pass rates tied.
- What Does Context Compression Cost an Agent? β Workload study
Completion stayed flat (80% to 85%, p = 1.0) while retrieval calls rose from 21.0 to 63.9 (p = .002): compression moved work into extra tool calls.
-
On 30 ChatDev tasks with GPT-5, the code review stage used 59.4% of tokens; input tokens were 53.9% of the total.
Methods measured by their own authors (6)
-
Input tokens fell 39.9% to 59.7% and total computational cost 21.1% to 35.9% at the same agent performance.
-
Average total token consumption fell 33.0% on SWE-bench Verified with success close to the uncompressed agent. No dollar figure.
-
23% to 54% token reduction on agent tasks such as SWE-bench Verified, with success rates the authors report as equal or better.
- ContextSniper: AntTrail's Token-Efficient Code Memory β Cost estimated
For OpenClaw on SWE-bench Lite, total tokens fell 51.5% and logged cost 36.4%; resolution rates differ by one task out of 50.
- An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents β Cost estimated
Content compression saved about 2% of the cache-priced prefix per turn; the authors say their single-shot benchmark must not be cited as a cost-saving argument.
-
Compressed agent context to 25.7% of its size while retaining 86.5% of uncompressed single-shot solve quality on 300 SWE-bench Lite instances.
Claims and third-party studies
8 tools where a project claim and third-party results can be read side by side, with every result found for each tool.
Before reading these comparisons. These are not benchmarks run by this site. Each result comes from one published study, and each study is one among several. Results depend on the harness, the model, the effort setting, the task set and the date of the run. A project's claim and a study's result often measure different quantities, such as shell output against total task cost. A single result does not show that a tool works or fails on your workload. Open each source, check what it measured, and test on your own tasks.
-
Caveman
Project claim
50% fewer (median)
of output tokens on 10 dev questions against an 'Answer concisely.' control, skill only. The earlier 65% (v1.9.1 release notes) is no longer in the README
Project source βThird-party results (6)
- Dasein Labs β Total cost (cost) β19% (19% cheaper)
- Stet.sh β Workload cost, run 1 then run 2 (cost) +9% then β12% (9% more expensive, then 12% cheaper)
- JetBrains (Caveman) β Output tokens (output tokens) β8.5% (8.5% fewer output tokens)
- THOL β End-to-end cost, 7 long tasks (cost) β13.6% (13.6% cheaper)
- Tura β Modeled cost, one task, 2 runs (cost) β3.9% (3.9% cheaper)
- Codepointer β Share of replayed spend saved (cost) β0.4% (0.4% cheaper)
-
RTK
Different denominatorProject claim
up to 90% fewer
of bash output the agent reads, per the README, which adds that this is not the same as cutting the bill by 90%. The repository description says 60-90% of LLM token consumption on common dev commands
Project source βThird-party results (6)
- Dasein Labs β Total cost (cost) +13% (13% more expensive)
- Stet.sh β Workload cost, run 1 then run 2 (cost) +13% then β9% (13% more expensive, then 9% cheaper)
- JetBrains (RTK) β Median cost per task, low effort then high effort (cost) +7.6% then +0.1% (7.6% more expensive, then 0.1% more expensive)
- THOL β End-to-end cost, 7 long tasks (cost) +7.1% (7.1% more expensive)
- Tura β Modeled cost, one task, 2 runs (cost) +7.2% (7.2% more expensive)
- Codepointer β Share of replayed spend saved (cost) β0.5% (0.5% cheaper)
The claim and the results use different denominators. Each number answers a different question.
-
Ponytail
Third-party results (3)
- Stet.sh β Workload cost, run 1 then run 2 (cost) +20% then β2% (20% more expensive, then 2% cheaper)
- THOL β End-to-end cost, 7 long tasks (cost) β3.5% (3.5% cheaper)
- Tura β Modeled cost, one task, 2 runs (cost) β8.9% (8.9% cheaper)
-
Headroom
Different denominatorProject claim
20% fewer
of tokens for coding agents, per the repository description (its README's four seeded scenarios show 21% to 57%)
Project source βThird-party results (4)
- Dasein Labs β Total cost (cost) +44% (44% more expensive)
- Marmelab β Cost per operation, cache mode then token mode (cost) β4% then +32% (4% cheaper, then 32% more expensive)
- THOL β End-to-end cost, 7 long tasks (cost) +52.8% (52.8% more expensive)
- Codepointer β Share of replayed spend saved (cost) β2.8% (2.8% cheaper)
The claim and the results use different denominators. Each number answers a different question.
-
Context Mode
Different denominatorProject claim
98% reduction
of bytes of raw tool output kept out of the context window, on the vendor's scenarios
Project source βThird-party results (1)
- Stet.sh β Workload cost, run 1 then run 2 (cost) +72% then +33% (72% more expensive, then 33% more expensive)
The claim and the results use different denominators. Each number answers a different question.
-
CodeGraph
Project claim
44% cheaper
of cost on 7 repos, one architecture question each, Opus 4.8, median of 4 runs
Project source βThird-party results (1)
- THOL β End-to-end cost, 7 long tasks (cost) β7.6% (7.6% cheaper)
-
lean-ctx
Different denominatorProject claim
99.4% saving
of input-side cost on a replayed 72-turn session, not a live task
Project source βThird-party results (1)
- THOL β End-to-end cost, 7 long tasks (cost) +10.3% (10.3% more expensive)
The claim and the results use different denominators. Each number answers a different question.
-
Edgee
Project claim
up to 70% off
of the token bill, combining routing, compression and visibility; no measurement stated
Project source βThird-party results (1)
- THOL β End-to-end cost, 7 long tasks (cost) β14.7% (14.7% cheaper)
Sources and further reading
Examined but not used
- tokbench: a pilot with one bug-fix task, one run per arm, on GLM models through the author's own harness. It measures input tokens, not cost. Its author discloses owning the harness and using lean-ctx daily.
- tokenme: its own repository marks the 92.7% result as withdrawn because the outputs were written by hand.
- Project-run benchmarks: kept in the catalogue under each project, labelled as the project's own benchmark, not counted as third-party measurements.
- Blog posts and aggregator lists that repeat a figure without a primary source: used as leads only.
Keep reading
- Independent benchmarks, in the guide β The written analysis behind this page, with sources.
- FinOps levers map β Where the "tool output compression" lever sits among the other cost levers.
- AI FinOps: the lever map β The guide section that maps each cost lever to its regime and its limits.
- LLM market snapshot β Current model prices, which set the size of every percentage above.
- How to read a token-compression benchmark β Blog article: denominators, paired tasks, cache costs and success criteria, with Tokenade versus RTK as a worked example.