Anthropic studies how Claude’s values vary by model and language

Ask Claude for feedback on a business plan in Hindi and again in Russian, and you may walk away with different impressions of the same idea. Anthropic’s new Claude values research shows that is not a vibes claim. It is a measurable pattern across hundreds of thousands of chats.
In a July 13 research post, Anthropic reports analyzing 309,815 anonymized Claude.ai conversations in which users gave Claude a subjective task. The sample drew equally from Sonnet 4.6, Opus 4.6, and Opus 4.7, and from the 20 most common languages on Claude.ai — roughly 5,000 conversations per model-language pair. The finding: Claude’s expressed values shift by model tier and by language, even after controlling for task, topic, and user-expressed values.
What four value axes Anthropic measured
Anthropic’s earlier “Values in the Wild” work had surfaced more than 3,000 distinct values in Claude’s responses. That list is too large to reason about. Here, researchers clustered related values, labeled 339 high-level values per conversation with a privacy-preserving tool, then compressed the patterns into four axes that capture about 15% of the remaining variation:
- Deference vs. Caution — accommodating what someone wants versus guarding against risk and harm
- Warmth vs. Rigor — positivity and care versus accuracy and precision
- Depth vs. Brevity — explaining in depth versus doing only what was asked
- Candor vs. Execution — foregrounding uncertainty versus producing a polished, confident answer
Each axis is a number line between two clusters of values that tend to trade off in real chats. Claude can still show warmth and rigor in the same reply. In practice, leaning hard one way usually means less of the other.
How Sonnet and Opus diverge
The model differences line up with how people already talk about these tiers. Sonnet 4.6 leans toward deference and warmth: affirming the user’s ideas, humor, playfulness, comfort without judgment. Opus 4.7 leans toward caution, rigor, depth, and candor: unprompted risk warnings, challenging assumptions, showing more of its reasoning, owning limitations. Opus 4.6 sits closer to rigor, deference, and brevity — more likely to stay inside the request and get to the point.
Anthropic is careful: the differences are small relative to variation across conversations, but they are structured and detectable. Staff and Claude.ai users had already described Opus 4.7 as hedge-prone and Sonnet 4.6 as warm. The axes recover those impressions from labeled behavior, not from launch blogging.
Language shifts builders should eval
Language is the sharper product risk. Warmth vs. Rigor and Candor vs. Execution move the most. Deference vs. Caution and Depth vs. Brevity stay more stable.
Anthropic’s highlights:
- Most warmth in Hindi and Arabic (polite language, humor, affirming the user’s work)
- Most rigor in English and Russian (challenging assumptions, correcting details, asking for evidence)
- Most deference and brevity in Arabic; most caution and depth in English
- Most candor in Dutch; most execution in Indonesian
Same class of request, different value mix. Anthropic does not yet know how much of this is desirable cultural adaptation versus uneven training data. Some languages may be overweighted toward professional writing. Some may simply have less data for consistent character training. The post treats that as an open research question, not a solved alignment story.
What this changes for multilingual product QA
If you ship a Claude-backed product in more than one language, monolingual English evals are incomplete. A critique that feels sharp and useful in English may read softer — or harsher — when the same workflow runs in Hindi, Arabic, or Russian. That hits feedback tools, coaching agents, content moderation helpers, and any flow where “tone” is part of the product.
Analysis: Anthropic is measuring character drift in production traffic and publishing the axes instead of burying them in a system card footnote. The constitution still sets the high-level target. These four lines show where deployment already diverges without anyone deliberately steering it. Builders should add language-stratified value checks to QA now. Anthropic still has to decide which shifts to fix, which to keep, and how to steer them on purpose.



