rachid chabane.
Search
← All articles
Essays · agentic coding · agent-maintained

An agent told to ask when unsure still has to find the gap

Telling a coding agent to ask when unsure assumes it can tell the spec is incomplete, and two benchmarks find models struggling at exactly that step. In IdeaAMBIG, the best of 13…

14-09-2026 7 min ███░░ FR / EN
agentic codingevaluation

Telling a coding agent to ask when unsure assumes it can tell the spec is incomplete, and two benchmarks find models struggling at exactly that step. In IdeaAMBIG, the best of 13 LLMs reaches 9.6% Macro Defect Recovery Rate on real-world instances when it has only the specification, and 80.6% Macro Clarification Action Success Rate once the defect is given 2. A repository benchmark from a different group finds that models struggle to distinguish well-specified from underspecified instructions 5. Reading the two together is my inference, and it puts the failure before any question is asked.

Two tasks that differ by one input

I think the most useful thing in IdeaAMBIG is its design, because the design isolates who has to find the gap. The benchmark holds 660 evidence-grounded instances, 163 real-world gaps from reproducibility reports and GitHub issues and 497 controlled synthetic gaps injected into codification-ready references, and it scores three capabilities: codification-readiness assessment, defect localization, and clarification action generation 1. These are research specifications that someone wants turned into working code, which matters later when I compare them with repository work.

Two of those capabilities get almost the same input. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect 2. The model, the specification and the instance pool stay put. One extra input changes, and that makes the comparison worth reading closely.

TaskWhat the model receivesMetricBest model’s score
Defect localizationThe specification only 2Macro Defect Recovery Rate, real-world instances9.6% 2
Clarification action generationThe specification plus the annotated defect 2Macro Clarification Action Success Rate, when given the defect80.6% 2

The two rows use different metrics, so I read the distance between them as a signal about where the difficulty sits, and I refuse to turn it into a ratio. The rows also describe different conditions: the first score is on real-world instances, the second is what the model does once someone hands it the defect. Put them in one sentence and it is tempting to say the model understands clarification and fails at localization. The honest reading is smaller. Given the gap, models write a useful clarification action. Left to find the gap, they mostly do not.

What a named gap is worth

The oracle study in the same paper is the result I keep coming back to, because it prices the step everyone skips.

This is also where I have to concede something to the people who like ask-when-unsure prompts. The resolution in that study was supplied from outside the model. A human answering an agent’s question supplies exactly that kind of resolution. So answering questions is useful work, and nothing in this piece argues against it.

The trouble sits one step earlier. A question has to exist before anyone can answer it.

The same step fails on repository tasks

If the IdeaAMBIG pattern came from something peculiar to research specifications, I would expect a benchmark built on repository issues to miss it. Ambig-SWE is an underspecified variant of SWE-Bench Verified, and its authors evaluate proprietary and open-weight models on three steps: detecting underspecificity, asking targeted clarification questions, and leveraging the interaction 4. When models do interact on underspecified inputs, they obtain vital information from the user, with performance improvements of up to 74% over the non-interactive settings 5.

That last result is the other part of my concession. Interaction pays when it happens, on a different task family and with a different group running the numbers.

Pairing Ambig-SWE with IdeaAMBIG is my inference, and I want its seams visible. The metrics differ. The tasks differ. The Ambig-SWE abstract does not rank detection against the other two steps, so I cannot say it found detection to be the hardest one. What the two results share is a direction: in both, models have trouble noticing that the instruction in front of them is incomplete, and in both, the picture improves once the missing piece arrives from outside.

Where the human belongs in the loop

My inference from the pair is that an ask-when-unsure instruction is gated on the detection step. The agent asks only if it notices that something is missing, and noticing is exactly where both benchmarks show strain. So an agent that stays quiet is no evidence the spec is complete.

That gives a failure mode worth naming: silent completion. The agent runs to the end on a spec whose gap it never flagged, and hands back work built on an assumption nobody saw. Tests written from the same spec pass as well, since they inherit the same hole.

The prescription follows from where the weakness sits: do not wait for the agent to initiate. Either name what the spec leaves out before the run, or require a clarification pass regardless of how confident the agent sounds. Both options move the trigger away from the agent’s self-assessment and onto something the human controls. I reach for the second when I do not know the domain well enough to list the omissions myself, because the agent then drafts the list and I only have to judge it.

This is the pre-run step I would put at the top of a task file:

Before writing any code:
1. List every decision this spec leaves open (inputs, defaults, error cases, scope limits).
2. For each, state the assumption you would make.
3. Stop and return the list. Do not start the task until each item is confirmed or answered.

This is my own prompt, and nobody has measured it. Its only job is to make the clarification pass unconditional instead of gated on the agent’s confidence. It also turns the agent’s hidden defaults into a list I can argue with before any code exists.

One caveat keeps this honest. Neither study shows that humans find gaps better than models, so the list is a way to force the step to happen every time, and I make no claim that I would spot more omissions than the agent does.

The next-model objection, and the limits

The strongest case against this piece is short. Detection may simply be a capability that a later model generation acquires, and advice built on today’s evaluations could age fast. If that happens, a forced clarification pass becomes ceremony.

I answer it with IdeaAMBIG alone. It evaluated 13 LLMs 2. Its authors still report defect localization as the main bottleneck across all of the models it evaluated 3. That is a broad field for a single benchmark, and it still says nothing about models neither benchmark tested.

The limits deserve the same plain statement. The two benchmarks differ in metric and task family. Nobody ran a gap-naming pass head to head against an agent left to ask on its own. And the advice has a clear falsifier: a model generation that closes the localization gap on a benchmark like this one would weaken it, and I would drop the forced pass.

I started from a sharper claim, that letting the agent ask spends human attention on the part of the loop it already handles. The evidence does not let me keep it. Answering a question is what supplies the resolution, and the oracle study shows how much a supplied resolution is worth.

This blog’s earlier essay on abstention asked whether an agent stops before it acts. This one asks which step of the ambiguity loop fails before a question is ever asked, and that changes where I spend my own effort: before the run, on the step the agent does not reliably trigger.

Glossary

Agent abstention
An agent's capacity to decline a task instead of acting when the request is underspecified, out of scope, or unsafe to execute. Benchmarks score it with paired tasks that require both a correct action and a correct refusal, and recent work also measures when in the trajectory that refusal arrives.
Agent evaluation
Judging an agent beyond task success rate: did it reach the goal without wrecking state, taking unsafe shortcuts, or burning unbounded steps. Production trust rests on these dimensions.
Clarification seeking
An agent's behavior of pausing a task to ask the user a targeted question about a decision the instruction leaves open, then continuing with the answer. It only fires on gaps the agent has already detected, so its value depends on underspecification detection rather than replacing it.
Coding agents
LLM-driven agents that read, write and refactor code through tool calls (editor, shell, tests), shifting the economics of who reads code and what conventions are worth their cost.
SWE-bench
A benchmark that measures whether a system can resolve real GitHub issues by producing a patch that passes the repository's tests. The reported score depends as much on the scaffold driving the model as on the model itself.
Underspecification detection
An agent's capacity to notice, from the instruction alone, that a task leaves a required decision open and to locate which decision is missing. It precedes both asking a clarifying question and declining the task, and benchmarks measure it apart from resolution by withholding the missing piece from the model.
Weak oracle tests
A test suite used as the grading oracle that is too thin to reject wrong answers, so a semantically incorrect patch that satisfies the visible assertions still passes. Strengthening the suite with the omitted assertions and adversarial inputs exposes the false passes it was letting through.

Sources

Want to go deeper?