SycoLens study shows why sycophancy rankings depend on the test
An arXiv preprint separates answer corrections from agreement under pressure.

A preprint posted on October 5 argues that measuring how often AI changes an answer can blur useful corrections with agreement under pressure.
The study introduces SycoLens, a protocol that replays an earlier answer with one scripted followup in a stateless interaction. Each trial is compared with a matching replay that removes the pressure line.
In its ten model main panel on moral scenarios, rankings under opposing opinions and challenges offering no alternative answer were almost unrelated under its open restatement format.
The paper recommends reporting prompt wording, answer format and whether changes improve accuracy.
These controlled replays use narrow task collections. They do not establish a general ranking for real conversations. Headline tests were selected after initial results, and code and data are promised upon publication.
Archival August 2007 photograph Almost done by alq666 via Wikimedia Commons under Creative Commons Attribution ShareAlike 2.0. Reduced resolution and WebP format. The photograph and resized versions retain that license. It illustrates computing infrastructure and does not show SycoLens or its evaluation setup.



