What changed
Quesma published a cost study of RTK on 11 September 2026: Terminal-Bench 2.1, Claude Code on Fable 5.0 and OpenCode on DeepSeek V4 Pro 0813, 1,740 attempts with and without the tool 1. Across the 445 DeepSeek attempts with RTK, rtk gain reported 349.2 million tokens saved, an 89% reduction 1. The bill did not follow 1.
| With RTK | Fable 5.0, Claude Code | DeepSeek V4 Pro 0813, OpenCode |
|---|---|---|
| Total cost | -5% | +5% |
| Average cost per task | +1% (no clear difference) | +17% |
A counter in the wrong unit
The miss sits in what the counter measures. RTK documents rtk gain as raw minus filtered command output, in bytes, divided by 4 1, and its own README scopes those percentages to bash output rather than to your bill 3. Nothing in that formula sees the rest of the loop. On DeepSeek each turn carried 7% less input, but there were 18% more turns 1, so shorter observations bought extra round-trips. The hook also sees less than people assume. Read, Grep and Glob are separate tools that bypass RTK, and terminal output was only about 7% of Fable’s context 1. A deep cut to that thin a slice was never going to show on the invoice.
JetBrains got there first, on different ground. In July, on SkillsBench with Claude Code and claude-sonnet-5, RTK came out 7.6% more expensive at low reasoning effort and flat at high effort 2. The direction held across harness, model and benchmark.
Impact on your team
If you put RTK behind a Claude Code hook after the June radar brief, which relayed its savings table, take rtk gain off every cost report. It counts bytes of bash output, and I think a number in the wrong unit does more damage than no number. Replace it with cost per passing task, with and without the hook, on your own tasks: Quesma’s per-task gap on DeepSeek (+17% 1) runs well above what the total showed (+5% 1). If you run OpenCode on DeepSeek V4 Pro, that comparison alone decides whether the hook stays. Do not wait for a rerun on 0.49.0 either 4: the recall store changes what RTK does, so only your own measurement on the version you deploy counts.