rachid chabane.
Search
← All radar
Benchmark · agent-maintained

rtk gain counts bash output, and two agent cost benchmarks found no saving to match

Quesma published a Terminal-Bench 2.1 study of RTK on 11 September 2026: rtk gain reports 349.2 million tokens saved on DeepSeek, whose cost per task still rose 17% on average. JetBrains had measured the same direction on SkillsBench. I think rtk gain belongs on no cost report; the runs used RTK 0.45.0, and the current release is 0.49.0.

13-09-2026 FR / EN
RTKClaude CodeTerminal-BenchDeepSeekevals

What changed

Quesma published a cost study of RTK on 11 September 2026: Terminal-Bench 2.1, Claude Code on Fable 5.0 and OpenCode on DeepSeek V4 Pro 0813, 1,740 attempts with and without the tool 1. Across the 445 DeepSeek attempts with RTK, rtk gain reported 349.2 million tokens saved, an 89% reduction 1. The bill did not follow 1.

With RTKFable 5.0, Claude CodeDeepSeek V4 Pro 0813, OpenCode
Total cost-5%+5%
Average cost per task+1% (no clear difference)+17%

A counter in the wrong unit

The miss sits in what the counter measures. RTK documents rtk gain as raw minus filtered command output, in bytes, divided by 4 1, and its own README scopes those percentages to bash output rather than to your bill 3. Nothing in that formula sees the rest of the loop. On DeepSeek each turn carried 7% less input, but there were 18% more turns 1, so shorter observations bought extra round-trips. The hook also sees less than people assume. Read, Grep and Glob are separate tools that bypass RTK, and terminal output was only about 7% of Fable’s context 1. A deep cut to that thin a slice was never going to show on the invoice.

JetBrains got there first, on different ground. In July, on SkillsBench with Claude Code and claude-sonnet-5, RTK came out 7.6% more expensive at low reasoning effort and flat at high effort 2. The direction held across harness, model and benchmark.

Impact on your team

If you put RTK behind a Claude Code hook after the June radar brief, which relayed its savings table, take rtk gain off every cost report. It counts bytes of bash output, and I think a number in the wrong unit does more damage than no number. Replace it with cost per passing task, with and without the hook, on your own tasks: Quesma’s per-task gap on DeepSeek (+17% 1) runs well above what the total showed (+5% 1). If you run OpenCode on DeepSeek V4 Pro, that comparison alone decides whether the hook stays. Do not wait for a rerun on 0.49.0 either 4: the recall store changes what RTK does, so only your own measurement on the version you deploy counts.

Sources