# RTK RTK (Rust Token Killer) is an open source CLI proxy written in [[Rust]] that intercepts shell command output and compresses it before it reaches an [[How Coding Agents Work|AI coding agent]]'s [[Context Window]]. It sits as a transparent hook between the agent and the underlying tools (git, cargo, npm, docker, kubectl, AWS CLI, etc.), so prompts and workflows stay unchanged. The premise: most of the [[AI Tokenization|tokens]] burned in an AI coding session are not reasoning, they are noise from verbose command output. A typical `git diff` or `cargo test` floods the context with boilerplate the model never needs. RTK attacks that noise at the source instead of touching the prompt. ## How it works A single Rust binary, zero dependencies, sub-10ms overhead. After `rtk init -g`, an auto-rewrite hook transparently maps `git status` to `rtk git status` (and similar) per supported command. The agent never sees the difference. Four compression strategies: boilerplate filtering, similar-line grouping, intelligent truncation, repetition deduplication. Each command type has its own optimization logic instead of a generic compressor. ## Reported savings These are RTK's own numbers, i.e., tokens removed from command output. Keep reading before you trust them (see the next section). - `git diff`: ~12 000 → ~960 tokens (~92%) - `npm test`: ~6 000 → ~600 tokens (~90%) - `cargo test`: ~91.8% reduction - `git status`: ~80.8% - `find`: ~78.3% - 12-file refactoring session: 74 700 → 6 960 tokens - Korben's own measured average over months of use: ~81.5% - Marketing claim across 2 900+ real-world commands: ~89% average noise removed Useless on already-short outputs. The payoff scales with how verbose the underlying tool is. Particularly relevant for autonomous agent loops where noise accumulates across iterations and pollutes the signal as much as the budget. ## Does it actually save money? (September 2026) Short answer: barely, and sometimes it costs MORE. Quesma (Bartosz Kotrys and Jacek Migdal) ran the first serious independent benchmark I've seen: Terminal-Bench 2.1, [[Claude Code]] with Fable 5.0 and [[OpenCode]] with DeepSeek V4 Pro, 85 and 89 tasks, five runs each with and without RTK. 1,740 attempts and $1,500+ of API spend. The results: - Claude/Fable: $1.72 → $1.64 per attempt (~5% cheaper overall). Almost all of that came from ONE task; excluding it, savings drop under 1%. Per task on average, it was 1% more expensive - DeepSeek: $0.115 → $0.121 per attempt (~5% more expensive overall, 17% more per task on average) - Pass rates: 1-2% lower with RTK - Meanwhile, `rtk gain` happily reported 89% token savings Why the gap? Three reasons: 1. **Shell output is a small slice of the bill.** Terminal output was about 11% of Fable's input tokens (40% for DeepSeek). Compressing 89% of 11% isn't much. Most tokens come from the system prompt, file reads and conversation history, which RTK never touches (agents read files with their own tools, not through the shell) 2. **Extra turns eat the savings.** When the compressed output hides something the model needed, it runs another command to get it. "One extra agent turn can cost more than the compression saved." 3. **The output is unfamiliar.** Models were trained on real `git` and `npm` output, not RTK's condensed format. Several people in the HN thread also reported agents getting confused and re-polling the same tool in a loop So `rtk gain` measures the wrong thing: bytes removed from command output, not dollars on the invoice. Quesma's verdict: "We do not recommend RTK as a generic cost-saving tool" for current frontier models, which already use the terminal efficiently (`head`, `tail`, `grep`, `--quiet` flags...). That's the bitter lesson again: the model already knows how to avoid flooding its own context. The HN discussion (170 points, 84 comments) went further. A few takeaways worth keeping: - Independent evals at another company found the same: RTK worsens task performance and doesn't save money overall - Where RTK still makes sense: truly noisy tools like Maven. But you get the same win with the tool's own quiet mode or a small wrapper script that writes the full log to a file and prints a pointer - Author benchmarks for this whole category of token-saving hacks (RTK, Headroom, semantic code search tools...) rarely reproduce. Measure on YOUR tasks, with cost and pass rate, not with the tool's own counter - Skills have a more straightforward token win: they remove decision-making from the agent I run RTK myself (it's wired into my [[Claude Code Hooks]]), so this one stings a bit. I'll keep it for now, but I no longer treat `rtk gain` as proof of anything. ## Integrations 13+ AI clients via hooks or plugins: [[Claude Code]] (via [[Claude Code Hooks]] of type `PreToolUse`), GitHub Copilot, [[Cursor CLI|Cursor]], [[Codex CLI|Codex]], Gemini CLI, [[Cline]], Windsurf, and others. 100+ commands covered across git, cargo, npm, pip, docker, kubectl, AWS CLI, and test runners. ## Observability - `rtk gain` — total token savings per command, with history - `rtk discover` — flags commands that ran without optimization (missed opportunities) - `rtk session` — adoption rate inside a given Claude Code session ## Install ```bash brew install rtk # macOS cargo install --git https://github.com/rtk-ai/rtk # any platform rtk init -g # install the rewrite hook ``` A curl install script is also published for other platforms. ## Pricing & license RTK itself is free, MIT/Apache-2.0 licensed (the repo lists both; LICENSE file says MIT, GitHub footer Apache-2.0). A paid RTK Cloud tier with team features is announced at ~$15/dev/month. ## Why it matters Different category from a [[LiteLLM Claude Code Proxy|model-routing proxy]]. Those reroute requests between providers; RTK reshapes the *inputs* before they reach whichever model is on the other end. The two can stack, and both can sit in front of the same Claude Code install. The original pitch was that context window pressure comes from the tools your agent reaches for, not your prompts. The September 2026 benchmark shows that's only partly true: command output is a minority of the tokens, and squeezing it can backfire. The lesson I take away: judge any context-saving tool by end-to-end cost AND success rate on real tasks, never by the compression ratio it reports about itself. ## References - GitHub: https://github.com/rtk-ai/rtk - Site: https://www.rtk-ai.app/ - Docs: https://www.rtk-ai.app/docs/ - Korben review (FR): https://korben.info/rtk-proxy-rust-economiser-tokens.html - Quesma, "RTK reports token savings, but our cost benchmarks disagree" (Bartosz Kotrys, Jacek Migdal, 2026-09-11): https://quesma.com/blog/does-rtk-make-ai-coding-cheaper/ - Hacker News discussion: https://news.ycombinator.com/item?id=49656471 ## Related - [[lowfat (CLI)]] - [[Claude Code]] - [[Claude Code Hooks]] - [[Context Window]] - [[AI Tokenization]] - [[How Coding Agents Work]] - [[Rust]] - [[LiteLLM Claude Code Proxy]] - [[Caveman AI Skill]]