MCP went stateless for a problem most of us never had
MCP's stateless rewrite is a retrofit, and the migration bill lands on the operators who never had the problem it solves. The specification states the change without hedging: the…
MCP's stateless rewrite is a retrofit, and the migration bill lands on the operators who never had the problem it solves. The specification states the change without hedging: the…
The +11.3% that vLLM's AFD plugin reports at 64A16F [s1] is a number you cannot carry to your own cluster, and the release does not publish what you would need to earn it.…
Hugging Face's responders could not analyze their own breach through commercial frontier APIs, which is why refusal behavior no longer belongs in any control set I write. The…
The abstention score a benchmark reports is a rate, and the number that decides how you wire an agent is how many irreversible actions run before it stops. Cyera's enterprise…
A private, contamination-resistant holdout lowers leakage risk but hands you a number no outsider can check. I read that as a trade: it buys contamination-resistance with…
Your CLAUDE.md structure is not why the agent stopped listening. Compliance falls about 5.6% in odds with every additional function the agent generates [s1], so the variable that…
The three mid-2026 results that read like a field-wide fight about prompt-injection defense are not disagreeing with each other, and once you name the layer each one attacked,…
There is no portable productivity number for coding agents: in this evidence base the instrument outweighs the tool, and proxy choice sets the sign of the headline. A study of…
The auto-approve config you shipped this quarter checks a string that bash has not finished rewriting. Your gate and your shell disagree about what the command is, and they…
The analysis everyone cites to prove the harness beats the model puts the harness at about 5.3% of the variation in success and the model at 0.7% [s2], a ranking that dwarfs…
The hard part of training a coding agent is no longer generating a solution, it is verifying one, and that single relocation quietly breaks the reward recipe most teams are still…
Normalize the thinking-token budget and the multi-hop reasoning edge people credit to multi-agent orchestration stops reliably showing up [s1]. That one control changes what the…
The pitch for a code review agent is that it reads every pull request so you do not have to. An empirical study of 13 of them inverts that pitch: 60.2% of agent-only PRs land in…
The Claude Code Action breach needed two things to work, a prompt injection and a broken authorization check, and the second is where your next hour of review pays off. A…
Buying an observability dashboard is not the same as knowing your agent is right, yet that is the trade most teams are quietly making. Two independent 2026 surveys show the field…
The rank order on a coding-agent leaderboard is partly a gaming score, and the gaming grows with capability instead of washing out as the models get better. The sharpest single…
A 92.5 on the LoCoMo memory benchmark tells you your agent can retrieve a fact, and almost nothing about whether it will keep serving that fact with full confidence after it…
Ninety-four percent of leaders rate AI-generated code as higher quality than human-authored code at review, and yet 82% of them hit at least one production failure tied to that…
The next model release will not fix your coding agent, because the failures that cost you are not the ones a bigger model repairs. They are harness failures: what the agent can…
MCP made wiring an agent to tools trivial by trusting the server's tool descriptions, and that one default quietly moved the trust boundary onto clients built to forward, not to…
A near-frontier coding model under an MIT license reads like the moment your default flips, and that read is wrong: the license hands you the weights, not your data flow, because…
When an agent stalls halfway through a long task, the reflex is to reach for a bigger context window. That is the wrong knob: on long-horizon software work the binding constraint…
Set an LLM to temperature 0 and run the same prompt a thousand times, and you will not get a thousand identical answers; the standard excuse is that GPU floating-point math is…
Select an agent for long autonomous work by its pass@1 score and you optimize the wrong number: capability and reliability match at horizon one but diverge as the horizon grows,…
Most teams read a SWE-bench number off a leaderboard and treat it as a property of the model, but for the comparison that actually drives procurement, choosing between adjacent…
Most teams budget training runs to the dollar and treat evaluation as effectively free. Then a single agent sweep on the Holistic Agent Leaderboard runs about $40,000 [s1], and…
Single-shot pass@1 and iterative maintainability measure different things, and on the axis that mirrors real software work, today's strongest coding agents are weak and get…
Every convention I inherited bet that one cost bound software: editing code a human had to read, and a machine reader plus a stochastic author now split that cost in two. This…
The reason your agent picked the wrong tool is almost never that your tool description was badly worded. In one 2026 study, prompt-repair recovers at most 23 percent of failures…
Structure-aware AST chunking helps code RAG, but the cAST headline overstates how much the chunking rule itself earns. cAST reports a 4.3-point Recall@5 gain on RepoEval and a…
Break an engineering task into verifiable steps, and let the agent fail early rather than late.
Combine BM25 and vectors without tuning ten weights: rank is enough.
Before publishing, the agent must prove its claims, sources attached.
GPTQ, AWQ, GGUF: what quantization actually costs, measured.
An agent that completes the task but wrecks the state hasn’t succeeded.
vLLM, continuous batching, KV-cache: where the VRAM really goes.
Every public number attached to Kimi K3 describes a rented API, not the weight file that is supposed to land on July 27. Artificial Analysis scores the model 57 on its…
You flip on prompt caching expecting the bill to fall and the latency to drop, and most of the time it does. But the first controlled cross-provider study of caching on…
The context length printed on a model's spec sheet is a marketing ceiling, not an operating budget, so I size a prompt for what the model can actually use, not for what caching…
Something went wrong. Try again.
Curious about Rachid or this site? Ask me.