AI uncertainty study shows why warning accuracy matters
Puzzle study separates useful warnings from weak model signals

A September 30 preprint tests how AI warnings affect judgments about puzzle moves. The task concerned which piece an instruction meant when descriptions or available alternatives could be confused.
Its human study asked 210 people to assess recorded GPT 4.1 interactions rather than participate in live collaboration.
With the AIโs original messages, participants intended to accept 78% of wrong moves. Detailed descriptions plus uncertainty cues reduced that to 36%. Crucially, researchers knew which moves were wrong and placed cues on a subset of those errors.
The targeting problem
The primary comparison included 168 people. Another 42 assessed warnings triggered by the modelโs own uncertainty estimates. This exploratory condition produced weaker discrimination than ideal targeting. Differences in discrimination from descriptions alone were not statistically significant.
Separate evaluations of GPT 4.1, GPT 5 and GPT 5.5 found that uncertainty tracked vague instructions more reliably than confusing alternatives.
The findings do not establish a deployable fix or benefits in live collaboration.
Illustrative tangram photograph by Gorkaazk under Creative Commons CC0. Converted to WebP. The wooden pieces are unrelated to the study materials.



