A Small Phrase That AI Access Seems to Delete

Admitting uncertainty is one of the cheapest intellectual tools available: "I don't know" costs nothing and prevents a great many errors. A new body of research suggests that tool is being quietly disabled by the mere availability of a language model. In five experiments covering 3,132 participants, researchers found that people who could consult AI advice became almost completely unwilling to withhold judgment — even in a setting where the AI was deliberately and reliably unreliable.

The work is credited to Marcoccia, Quattrociocchi and Capraro in 2026, and the pattern it describes is stark. In one experiment, the share of questions participants declined to answer fell from 44 percent without AI access to just 3 percent with it. They answered far more. They were right far less.

How the Experiments Were Built

The design matters, because it rules out the most flattering explanation for the behavior. The researchers did not test whether people sensibly lean on a competent assistant. They tested what happens when the assistant is confidently bad.

Questions Chosen to Break the Model

The team selected questions where their chosen model, Step 3.5 Flash, was almost always wrong. These concerned fine visual details from films — for instance, the color of a team uniform in Bend It Like Beckham. The authors note that such minutiae rarely appear in online text, which makes them prime targets for hallucination: the model has little to retrieve and plenty of room to invent.

To confirm the setup was not simply testing a broken tool, the researchers also observed that GPT-5.5, Claude 4.6 Sonnet and Gemini 3.5 Flash handled most of the other questions correctly, though they still stumbled on the harder ones. The critical point is that because the AI advice was mostly wrong, any deference participants showed cannot be framed as reasonable delegation to a dependable instrument.

More Answers, Fewer Correct Ones

Studies 1a and 1b let participants decide for themselves whether to ask an AI before answering. In the control condition, with no AI available, they withheld judgment on 36 percent and 44 percent of the questions respectively. Once an AI was available, those abstention rates collapsed to 6 percent and 3 percent.

Study 2 added a second measurement — confidence — and produced what may be the most uncomfortable result in the set:

In Study 2, AI access drove judgment suspension to near zero (a) while confidence rose sharply (b). Correctness dropped at the same time (c). Financial incentives reduced how often participants sought AI advice (d). | Image: Marcoccia, Quattrociocchi, Capraro (2026)
In Study 2, AI access drove judgment suspension to near zero (a) while confidence rose sharply (b). Correctness dropped at the same time (c). Financial incentives reduced how often participants sought AI advice (d). | Image: Marcoccia, Quattrociocchi, Capraro (2026)
  • Confidence with AI access reached 75.9 points on a 100-point scale, compared with 29.6 points without it — roughly two and a half times higher.
  • Correct answers fell from 27.6 percent to 10.0 percent in the same comparison.
  • Across all studies, participants without financial incentives who had AI access answered correctly 9.2 percent of the time, versus 27.5 percent without AI.

In other words, people answered more questions with AI at hand and were correct roughly a third as often. The researchers characterize the shift in Study 2 as three simultaneous movements: judgment suspension dropped to nearly zero, confidence rose sharply, and correctness declined. Certainty and accuracy moved in opposite directions.

Confidence Without Competence

The combination is familiar from human psychology — people who know least often feel most sure — but here it appears to be induced on demand by a tool rather than emerging from long-held expertise. Participants were not reporting confidence in the AI; they were reporting confidence in their own answers, delivered after exposure to advice that the researchers had engineered to be wrong most of the time.

That dynamic helps explain why the effect is worth tracking beyond the lab. Real-world AI assistants are most likely to hallucinate precisely where information is thin, obscure or rarely documented — the same conditions the study used to bait the model into error. The gap between fluency and factuality is not evenly distributed; it clusters in exactly the territory where a person's own uncertainty would normally be highest.

Incentives Help, but Do Not Fix the Problem

Studies 2 through 4 introduced money. Participants earned 10 cents for each correct answer, lost 10 cents for each wrong one, and received nothing at all for saying "I don't know." That scoring scheme makes abstention genuinely costly in relative terms, yet the researchers still wanted to know whether the penalty structure for errors would push people back toward admitting ignorance.

They pre-registered the hypothesis that financial incentives would increase willingness to abstain, but that the availability of AI would weaken that effect. The available account of the findings points in a mixed direction: incentives did reduce how often participants sought AI advice, and the researchers describe them as helpful, but not a cure. The pull of an immediate answer appears to survive a direct financial argument against guessing.

What Remains Uncertain

The researchers still see a general tendency for people to defer to AI output, but they are candid that the boundaries of their result are unproven. Whether the same collapse in abstention holds as strongly outside movie trivia — in medicine, law, engineering or everyday factual questions — is an open question. It is possible that some domains activate enough wariness to preserve the instinct to say "I don't know." It is also possible that the effect is stronger where stakes are higher and answers feel more consequential.

What the study does establish is a mechanism worth watching: the presence of an answer changes the psychology of the person who might otherwise have admitted not having one. The cost is measurable. With AI in the room, participants traded a large share of their accuracy for a large share of their confidence — and gave up the one answer that is always honest.

This article is based on reporting by The Decoder. Read the original article.

Originally published on the-decoder.com