MCP went stateless for a problem most of us never had
MCP's stateless rewrite is a retrofit, and the migration bill lands on the operators who never had the problem it solves. The specification states the change without hedging: the…
MCP's stateless rewrite is a retrofit, and the migration bill lands on the operators who never had the problem it solves. The specification states the change without hedging: the…
Hugging Face's responders could not analyze their own breach through commercial frontier APIs, which is why refusal behavior no longer belongs in any control set I write. The…
A private, contamination-resistant holdout lowers leakage risk but hands you a number no outsider can check. I read that as a trade: it buys contamination-resistance with…
Your CLAUDE.md structure is not why the agent stopped listening. Compliance falls about 5.6% in odds with every additional function the agent generates [s1], so the variable that…
There is no portable productivity number for coding agents: in this evidence base the instrument outweighs the tool, and proxy choice sets the sign of the headline. A study of…
The auto-approve config you shipped this quarter checks a string that bash has not finished rewriting. Your gate and your shell disagree about what the command is, and they…
The analysis everyone cites to prove the harness beats the model puts the harness at about 5.3% of the variation in success and the model at 0.7% [s2], a ranking that dwarfs…
The hard part of training a coding agent is no longer generating a solution, it is verifying one, and that single relocation quietly breaks the reward recipe most teams are still…
The Claude Code Action breach needed two things to work, a prompt injection and a broken authorization check, and the second is where your next hour of review pays off. A…
The rank order on a coding-agent leaderboard is partly a gaming score, and the gaming grows with capability instead of washing out as the models get better. The sharpest single…
Ninety-four percent of leaders rate AI-generated code as higher quality than human-authored code at review, and yet 82% of them hit at least one production failure tied to that…
The next model release will not fix your coding agent, because the failures that cost you are not the ones a bigger model repairs. They are harness failures: what the agent can…
A near-frontier coding model under an MIT license reads like the moment your default flips, and that read is wrong: the license hands you the weights, not your data flow, because…
When an agent stalls halfway through a long task, the reflex is to reach for a bigger context window. That is the wrong knob: on long-horizon software work the binding constraint…
Most teams read a SWE-bench number off a leaderboard and treat it as a property of the model, but for the comparison that actually drives procurement, choosing between adjacent…
Single-shot pass@1 and iterative maintainability measure different things, and on the axis that mirrors real software work, today's strongest coding agents are weak and get…
Every convention I inherited bet that one cost bound software: editing code a human had to read, and a machine reader plus a stochastic author now split that cost in two. This…
Structure-aware AST chunking helps code RAG, but the cAST headline overstates how much the chunking rule itself earns. cAST reports a 4.3-point Recall@5 gain on RepoEval and a…
Break an engineering task into verifiable steps, and let the agent fail early rather than late.
Something went wrong. Try again.
Curious about Rachid or this site? Ask me.