A telemetry switch is a behaviour setting for your coding agent
I think a coding agent's telemetry switch is a behaviour setting: in Claude Code it has decided which instructions load and which mode an interactive session starts in. An…
I think a coding agent's telemetry switch is a behaviour setting: in Claude Code it has decided which instructions load and which mode an interactive session starts in. An…
Treat an agent's per-session token cap as a tuned value: past a measurable point extra tokens buy less than fresh samples, and splitting the budget needs a cheap scorer. In an…
Your authorization layer checks the artifact an agent produced and never the reasoning that produced it, so the check that failed was never written. Agent payment protocols such…
The state your agent carries between turns is a procurement decision, and ARC Prize has just published what the readable version of it costs. One model, one benchmark, two of the…
An egress allowlist keyed on the HTTP verb bounds nothing, because the far end decides whether a request counts as a read or as a write. A swarm that was allowed to read the…
Localize the fault, infill the suspect span, and you lose to one more plain sample at the same attempt count [s3]; the delivery step costs more than the signal is worth. I think…
Most security money buys rank order inside a queue that is not draining, and the marginal hour returns more spent on remediation throughput. The note that made me write this puts…
I think the design variable in a coding-agent harness is what the failure record says, and the two things I have shipped both put the wrong text in it. The reflex is to log the…
The multi-agent intervention that paid in measured tokens was the one that changed what each agent can read. I read that as the axis the whole design debate keeps missing.
Your coding agent's gain lands on your team's dashboard, and the bill lands on a shared record that no dashboard owns. The gain is real and measured: in a simulated community…
Early abort for agent runs is a capability you buy at serving time, and it bills you every month in the successful runs your recall target chose to kill.
The check that authorizes a write should read environment state, the one evidence the agent's own account cannot manufacture; a detector arrives after the write.
Enumerating the permitted instances of one operation buys you nothing about the effect that operation was supposed to prevent.
MCP's stateless rewrite is a retrofit, and the migration bill lands on the operators who never had the problem it solves. The specification states the change without hedging: the…
Hugging Face's responders could not analyze their own breach through commercial frontier APIs, which is why refusal behavior no longer belongs in any control set I write. The…
The abstention score a benchmark reports is a rate, and the number that decides how you wire an agent is how many irreversible actions run before it stops. Cyera's enterprise…
The reason your agent picked the wrong tool is almost never that your tool description was badly worded. In one 2026 study, prompt-repair recovers at most 23 percent of failures…
Your CLAUDE.md structure is not why the agent stopped listening. Compliance falls about 5.6% in odds with every additional function the agent generates [s1], so the variable that…
The three mid-2026 results that read like a field-wide fight about prompt-injection defense are not disagreeing with each other, and once you name the layer each one attacked,…
The auto-approve config you shipped this quarter checks a string that bash has not finished rewriting. Your gate and your shell disagree about what the command is, and they…
The analysis everyone cites to prove the harness beats the model puts the harness at about 5.3% of the variation in success and the model at 0.7% [s2], a ranking that dwarfs…
The hard part of training a coding agent is no longer generating a solution, it is verifying one, and that single relocation quietly breaks the reward recipe most teams are still…
Normalize the thinking-token budget and the multi-hop reasoning edge people credit to multi-agent orchestration stops reliably showing up [s1]. That one control changes what the…
The pitch for a code review agent is that it reads every pull request so you do not have to. An empirical study of 13 of them inverts that pitch: 60.2% of agent-only PRs land in…
The Claude Code Action breach needed two things to work, a prompt injection and a broken authorization check, and the second is where your next hour of review pays off. A…
Buying an observability dashboard is not the same as knowing your agent is right, yet that is the trade most teams are quietly making. Two independent 2026 surveys show the field…
A 92.5 on the LoCoMo memory benchmark tells you your agent can retrieve a fact, and almost nothing about whether it will keep serving that fact with full confidence after it…
Ninety-four percent of leaders rate AI-generated code as higher quality than human-authored code at review, and yet 82% of them hit at least one production failure tied to that…
The next model release will not fix your coding agent, because the failures that cost you are not the ones a bigger model repairs. They are harness failures: what the agent can…
MCP made wiring an agent to tools trivial by trusting the server's tool descriptions, and that one default quietly moved the trust boundary onto clients built to forward, not to…
You flip on prompt caching expecting the bill to fall and the latency to drop, and most of the time it does. But the first controlled cross-provider study of caching on…
When an agent stalls halfway through a long task, the reflex is to reach for a bigger context window. That is the wrong knob: on long-horizon software work the binding constraint…
Select an agent for long autonomous work by its pass@1 score and you optimize the wrong number: capability and reliability match at horizon one but diverge as the horizon grows,…
Most teams read a SWE-bench number off a leaderboard and treat it as a property of the model, but for the comparison that actually drives procurement, choosing between adjacent…
Most teams budget training runs to the dollar and treat evaluation as effectively free. Then a single agent sweep on the Holistic Agent Leaderboard runs about $40,000 [s1], and…
Single-shot pass@1 and iterative maintainability measure different things, and on the axis that mirrors real software work, today's strongest coding agents are weak and get…
Break an engineering task into verifiable steps, and let the agent fail early rather than late.
An agent that completes the task but wrecks the state hasn’t succeeded.
Something went wrong. Try again.
The assistant is temporarily unavailable. It will be back on {date}.
Curious about Rachid or this site? Ask me.