Faster runners alone cannot absorb the CI load agents create
I think buying faster runners will not keep CI ahead of coding agents: they make each test run cheaper, while agents add runs and add changes someone must review. Linear states…
I think buying faster runners will not keep CI ahead of coding agents: they make each test run cheaper, while agents add runs and add changes someone must review. Linear states…
I do not quote a patch-quality rate until I know the grader that produced it would label the same patch the same way twice. That sounds like pedantry until you try to use one of…
An egress allowlist keyed on the HTTP verb bounds nothing, because the far end decides whether a request counts as a read or as a write. A swarm that was allowed to read the…
Most security money buys rank order inside a queue that is not draining, and the marginal hour returns more spent on remediation throughput. The note that made me write this puts…
I think what decides a repository-scale migration is the instrument that grades it: a behaviour-only grade cannot see a migration that never happened. That is not a claim about…
I think the design variable in a coding-agent harness is what the failure record says, and the two things I have shipped both put the wrong text in it. The reflex is to log the…
The watermark is made out of the sampler's freedom [s17], and every practice that makes my coding agent dependable spends that same freedom. Anthropic states the mechanism…
If you are writing an LLM contribution policy for your repository, the premise you start from decides where the rule will leak. Three toolchain projects published theirs…
Hugging Face's responders could not analyze their own breach through commercial frontier APIs, which is why refusal behavior no longer belongs in any control set I write. The…
There is no portable productivity number for coding agents: in this evidence base the instrument outweighs the tool, and proxy choice sets the sign of the headline. A study of…
Ninety-four percent of leaders rate AI-generated code as higher quality than human-authored code at review, and yet 82% of them hit at least one production failure tied to that…
Set an LLM to temperature 0 and run the same prompt a thousand times, and you will not get a thousand identical answers; the standard excuse is that GPU floating-point math is…
Every convention I inherited bet that one cost bound software: editing code a human had to read, and a machine reader plus a stochastic author now split that cost in two. This…
Before publishing, the agent must prove its claims, sources attached.
Something went wrong. Try again.
The assistant is temporarily unavailable. It will be back on {date}.
Curious about Rachid or this site? Ask me.