<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Rachid Chabane · Articles</title><description>AI engineering blog · agents, RAG, systems</description><link>https://rachid-chabane.com/</link><language>en-US</language><item><title>MCP went stateless for a problem most of us never had</title><link>https://rachid-chabane.com/en/blog/mcp-stateless-core-header-routing/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/mcp-stateless-core-header-routing/</guid><description>MCP&apos;s stateless rewrite is a retrofit, and the migration bill lands on the operators who never had the problem it solves. The specification states the change without hedging: the…</description><pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate></item><item><title>What the vLLM AFD numbers ask of your interconnect</title><link>https://rachid-chabane.com/en/blog/what-vllm-afd-numbers-ask-of-your-interconnect/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/what-vllm-afd-numbers-ask-of-your-interconnect/</guid><description>The +11.3% that vLLM&apos;s AFD plugin reports at 64A16F [s1] is a number you cannot carry to your own cluster, and the release does not publish what you would need to earn it.…</description><pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Your Incident Response Runs on a Vendor&apos;s Content Policy</title><link>https://rachid-chabane.com/en/blog/your-incident-response-runs-on-a-vendors-content-policy/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/your-incident-response-runs-on-a-vendors-content-policy/</guid><description>Hugging Face&apos;s responders could not analyze their own breach through commercial frontier APIs, which is why refusal behavior no longer belongs in any control set I write. The…</description><pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Count the irreversible actions an agent takes before it stops</title><link>https://rachid-chabane.com/en/blog/count-the-irreversible-actions-before-an-agent-stops/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/count-the-irreversible-actions-before-an-agent-stops/</guid><description>The abstention score a benchmark reports is a rate, and the number that decides how you wire an agent is how many irreversible actions run before it stops. Cyera&apos;s enterprise…</description><pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate></item><item><title>A private SWE-bench holdout moves the trust problem to its author</title><link>https://rachid-chabane.com/en/blog/swe-bench-private-holdout-moves-the-trust-problem/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/swe-bench-private-holdout-moves-the-trust-problem/</guid><description>A private, contamination-resistant holdout lowers leakage risk but hands you a number no outsider can check. I read that as a trade: it buys contamination-resistance with…</description><pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Kimi K3 ranks fourth, and nobody has measured the model you would run</title><link>https://rachid-chabane.com/en/blog/kimi-k3-ranks-fourth-nobody-measured-the-weights/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/kimi-k3-ranks-fourth-nobody-measured-the-weights/</guid><description>Every public number attached to Kimi K3 describes a rented API, not the weight file that is supposed to land on July 27. Artificial Analysis scores the model 57 on its…</description><pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Your agent attends to the right tool and still picks the wrong one</title><link>https://rachid-chabane.com/en/blog/your-agent-attends-to-the-right-tool-and-still-picks-wrong/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/your-agent-attends-to-the-right-tool-and-still-picks-wrong/</guid><description>The reason your agent picked the wrong tool is almost never that your tool description was badly worded. In one 2026 study, prompt-repair recovers at most 23 percent of failures…</description><pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Your CLAUDE.md structure is not why the agent stopped listening</title><link>https://rachid-chabane.com/en/blog/stop-editing-claude-md/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/stop-editing-claude-md/</guid><description>Your CLAUDE.md structure is not why the agent stopped listening. Compliance falls about 5.6% in odds with every additional function the agent generates [s1], so the variable that…</description><pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Your Injection Defense Was Never Tested Against an Attacker Who Adapts</title><link>https://rachid-chabane.com/en/blog/injection-defense-layer-split/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/injection-defense-layer-split/</guid><description>The three mid-2026 results that read like a field-wide fight about prompt-injection defense are not disagreeing with each other, and once you name the layer each one attacked,…</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Plus 24 Percent and Minus 19 Percent Are Both True</title><link>https://rachid-chabane.com/en/blog/plus-24-percent-and-minus-19-percent-are-both-true/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/plus-24-percent-and-minus-19-percent-are-both-true/</guid><description>There is no portable productivity number for coding agents: in this evidence base the instrument outweighs the tool, and proxy choice sets the sign of the headline. A study of…</description><pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Your Agent&apos;s Command Denylist Checks a String the Shell Then Rewrites</title><link>https://rachid-chabane.com/en/blog/your-agents-denylist-checks-a-string-the-shell-rewrites/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/your-agents-denylist-checks-a-string-the-shell-rewrites/</guid><description>The auto-approve config you shipped this quarter checks a string that bash has not finished rewriting. Your gate and your shell disagree about what the command is, and they…</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The Harness Beat the Model by 4.6 Points of Almost Nothing</title><link>https://rachid-chabane.com/en/blog/there-is-no-best-coding-agent-only-a-best-one-for-this-workload/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/there-is-no-best-coding-agent-only-a-best-one-for-this-workload/</guid><description>The analysis everyone cites to prove the harness beats the model puts the harness at about 5.3% of the variation in success and the model at 0.7% [s2], a ranking that dwarfs…</description><pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The Reward Is the Weak Link in Your Coding Agent</title><link>https://rachid-chabane.com/en/blog/verification-is-the-new-bottleneck-for-coding-agents/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/verification-is-the-new-bottleneck-for-coding-agents/</guid><description>The hard part of training a coding agent is no longer generating a solution, it is verifying one, and that single relocation quietly breaks the reward recipe most teams are still…</description><pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Most Multi-Agent Reasoning Wins Are a Compute Artifact</title><link>https://rachid-chabane.com/en/blog/multi-agent-wins-are-a-compute-artifact/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/multi-agent-wins-are-a-compute-artifact/</guid><description>Normalize the thinking-token budget and the multi-hop reasoning edge people credit to multi-agent orchestration stops reliably showing up [s1]. That one control changes what the…</description><pubDate>Mon, 13 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Code Review Agents Are Optimizing the Wrong Metric</title><link>https://rachid-chabane.com/en/blog/code-review-agents-are-shipping-noise/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/code-review-agents-are-shipping-noise/</guid><description>The pitch for a code review agent is that it reads every pull request so you do not have to. An empirical study of 13 of them inverts that pitch: 60.2% of agent-only PRs land in…</description><pubDate>Sun, 12 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Where to Spend Your Agent Security Budget</title><link>https://rachid-chabane.com/en/blog/claude-code-action-breach-was-an-auth-bug/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/claude-code-action-breach-was-an-auth-bug/</guid><description>The Claude Code Action breach needed two things to work, a prompt injection and a broken authorization check, and the second is where your next hour of review pays off. A…</description><pubDate>Sat, 11 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Observability Is Not Evaluation</title><link>https://rachid-chabane.com/en/blog/observability-is-not-evaluation/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/observability-is-not-evaluation/</guid><description>Buying an observability dashboard is not the same as knowing your agent is right, yet that is the trade most teams are quietly making. Two independent 2026 surveys show the field…</description><pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The Coding Leaderboard Partly Measures Harness Gaming</title><link>https://rachid-chabane.com/en/blog/coding-leaderboards-measure-harness-gaming/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/coding-leaderboards-measure-harness-gaming/</guid><description>The rank order on a coding-agent leaderboard is partly a gaming score, and the gaming grows with capability instead of washing out as the models get better. The sharpest single…</description><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Your Agent&apos;s Memory Benchmark Grades Recall, Not Trust</title><link>https://rachid-chabane.com/en/blog/agent-memory-benchmarks-grade-recall-not-trust/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/agent-memory-benchmarks-grade-recall-not-trust/</guid><description>A 92.5 on the LoCoMo memory benchmark tells you your agent can retrieve a fact, and almost nothing about whether it will keep serving that fact with full confidence after it…</description><pubDate>Sun, 05 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Agent Debt Is a Verification-Placement Problem, Not a Review Problem</title><link>https://rachid-chabane.com/en/blog/agent-debt-is-a-verification-placement-problem/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/agent-debt-is-a-verification-placement-problem/</guid><description>Ninety-four percent of leaders rate AI-generated code as higher quality than human-authored code at review, and yet 82% of them hit at least one production failure tied to that…</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Coding Agents Fail at the Harness, Not the Model</title><link>https://rachid-chabane.com/en/blog/coding-agents-fail-at-the-harness-not-the-model/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/coding-agents-fail-at-the-harness-not-the-model/</guid><description>The next model release will not fix your coding agent, because the failures that cost you are not the ones a bigger model repairs. They are harness failures: what the agent can…</description><pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate></item><item><title>MCP Won on Simplicity and Moved the Trust Boundary to the Client</title><link>https://rachid-chabane.com/en/blog/mcp-moved-the-trust-boundary-to-the-client/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/mcp-moved-the-trust-boundary-to-the-client/</guid><description>MCP made wiring an agent to tools trivial by trusting the server&apos;s tool descriptions, and that one default quietly moved the trust boundary onto clients built to forward, not to…</description><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Prompt Caching Can Make Your Agent Slower</title><link>https://rachid-chabane.com/en/blog/prompt-caching-can-slow-your-agent-down/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/prompt-caching-can-slow-your-agent-down/</guid><description>You flip on prompt caching expecting the bill to fall and the latency to drop, and most of the time it does. But the first controlled cross-provider study of caching on…</description><pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate></item><item><title>GLM-5.2 closes the coding gap. The catch is where you run it.</title><link>https://rachid-chabane.com/en/blog/glm-5-2-where-you-run-it/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/glm-5-2-where-you-run-it/</guid><description>A near-frontier coding model under an MIT license reads like the moment your default flips, and that read is wrong: the license hands you the weights, not your data flow, because…</description><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Teaching an agent to forget: prune the history, not the window</title><link>https://rachid-chabane.com/en/blog/teaching-an-agent-to-forget/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/teaching-an-agent-to-forget/</guid><description>When an agent stalls halfway through a long task, the reflex is to reach for a bigger context window. That is the wrong knob: on long-horizon software work the binding constraint…</description><pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Temperature zero is not deterministic, and the GPU is not to blame</title><link>https://rachid-chabane.com/en/blog/temperature-zero-is-not-deterministic/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/temperature-zero-is-not-deterministic/</guid><description>Set an LLM to temperature 0 and run the same prompt a thousand times, and you will not get a thousand identical answers; the standard excuse is that GPU floating-point math is…</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Your agent&apos;s pass@1 score is blind to long-horizon reliability</title><link>https://rachid-chabane.com/en/blog/pass-at-1-is-blind-to-long-horizon-reliability/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/pass-at-1-is-blind-to-long-horizon-reliability/</guid><description>Select an agent for long autonomous work by its pass@1 score and you optimize the wrong number: capability and reliability match at horizon one but diverge as the horizon grows,…</description><pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate></item><item><title>The scaffold, not the model, decides your SWE-bench score</title><link>https://rachid-chabane.com/en/blog/scaffold-not-model-swe-bench-score/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/scaffold-not-model-swe-bench-score/</guid><description>Most teams read a SWE-bench number off a leaderboard and treat it as a property of the model, but for the comparison that actually drives procurement, choosing between adjacent…</description><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Evaluation is the new compute bottleneck</title><link>https://rachid-chabane.com/en/blog/evaluation-is-the-new-compute-bottleneck/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/evaluation-is-the-new-compute-bottleneck/</guid><description>Most teams budget training runs to the dollar and treat evaluation as effectively free. Then a single agent sweep on the Holistic Agent Leaderboard runs about $40,000 [s1], and…</description><pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Agents Erode the Code They Keep Editing</title><link>https://rachid-chabane.com/en/blog/agents-erode-the-code-they-keep-editing/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/agents-erode-the-code-they-keep-editing/</guid><description>Single-shot pass@1 and iterative maintainability measure different things, and on the axis that mirrors real software work, today&apos;s strongest coding agents are weak and get…</description><pubDate>Tue, 16 Jun 2026 00:00:00 GMT</pubDate></item><item><title>AI Re-Priced the Whole Canon. It Didn&apos;t Repeal It.</title><link>https://rachid-chabane.com/en/blog/ai-repriced-the-canon/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/ai-repriced-the-canon/</guid><description>Every convention I inherited bet that one cost bound software: editing code a human had to read, and a machine reader plus a stochastic author now split that cost in two. This…</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Chunk on the Syntax Tree, or Don&apos;t Bother? What the Benchmarks Say</title><link>https://rachid-chabane.com/en/blog/chunk-on-the-syntax-tree-or-not/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/chunk-on-the-syntax-tree-or-not/</guid><description>Structure-aware AST chunking helps code RAG, but the cAST headline overstates how much the chunking rule itself earns. cAST reports a 4.3-point Recall@5 gain on RepoEval and a…</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Your Context Window Is a Ceiling, Not a Budget</title><link>https://rachid-chabane.com/en/blog/your-context-window-is-a-ceiling/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/your-context-window-is-a-ceiling/</guid><description>The context length printed on a model&apos;s spec sheet is a marketing ceiling, not an operating budget, so I size a prompt for what the model can actually use, not for what caching…</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Orchestrating coding agents with deterministic workflows</title><link>https://rachid-chabane.com/en/blog/deterministic-coding-agent-workflows/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/deterministic-coding-agent-workflows/</guid><description>Break an engineering task into verifiable steps, and let the agent fail early rather than late.</description><pubDate>Sat, 30 May 2026 00:00:00 GMT</pubDate></item><item><title>Hybrid RAG: reciprocal rank fusion in practice</title><link>https://rachid-chabane.com/en/blog/hybrid-rag-reciprocal-rank-fusion/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/hybrid-rag-reciprocal-rank-fusion/</guid><description>Combine BM25 and vectors without tuning ten weights: rank is enough.</description><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate></item><item><title>Publication guardrails: an automated fact-checking pipeline</title><link>https://rachid-chabane.com/en/blog/publication-guardrails-fact-checking/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/publication-guardrails-fact-checking/</guid><description>Before publishing, the agent must prove its claims, sources attached.</description><pubDate>Sat, 23 May 2026 00:00:00 GMT</pubDate></item><item><title>Quantizing an open model without breaking it</title><link>https://rachid-chabane.com/en/blog/quantizing-open-model/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/quantizing-open-model/</guid><description>GPTQ, AWQ, GGUF: what quantization actually costs, measured.</description><pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate></item><item><title>Evaluating a tool-using agent: beyond success rate</title><link>https://rachid-chabane.com/en/blog/evaluating-tool-using-agent/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/evaluating-tool-using-agent/</guid><description>An agent that completes the task but wrecks the state hasn’t succeeded.</description><pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate></item><item><title>Serving an open-source LLM in production: the real cost</title><link>https://rachid-chabane.com/en/blog/serving-oss-llm-production/</link><guid isPermaLink="true">https://rachid-chabane.com/en/blog/serving-oss-llm-production/</guid><description>vLLM, continuous batching, KV-cache: where the VRAM really goes.</description><pubDate>Mon, 04 May 2026 00:00:00 GMT</pubDate></item></channel></rss>