Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Anthropic studies how Claude’s values vary by model and language

Listen to this article

Ask Claude for feedback on a business plan in Hindi and again in Russian, and you may walk away with different impressions of the same idea. Anthropic’s new Claude values research shows that is not a vibes claim. It is a measurable pattern across hundreds of thousands of chats.

In a July 13 research post, Anthropic reports analyzing 309,815 anonymized Claude.ai conversations in which users gave Claude a subjective task. The sample drew equally from Sonnet 4.6, Opus 4.6, and Opus 4.7, and from the 20 most common languages on Claude.ai — roughly 5,000 conversations per model-language pair. The finding: Claude’s expressed values shift by model tier and by language, even after controlling for task, topic, and user-expressed values.

What four value axes Anthropic measured

Anthropic’s earlier “Values in the Wild” work had surfaced more than 3,000 distinct values in Claude’s responses. That list is too large to reason about. Here, researchers clustered related values, labeled 339 high-level values per conversation with a privacy-preserving tool, then compressed the patterns into four axes that capture about 15% of the remaining variation:

  • Deference vs. Caution — accommodating what someone wants versus guarding against risk and harm
  • Warmth vs. Rigor — positivity and care versus accuracy and precision
  • Depth vs. Brevity — explaining in depth versus doing only what was asked
  • Candor vs. Execution — foregrounding uncertainty versus producing a polished, confident answer

Each axis is a number line between two clusters of values that tend to trade off in real chats. Claude can still show warmth and rigor in the same reply. In practice, leaning hard one way usually means less of the other.

How Sonnet and Opus diverge

The model differences line up with how people already talk about these tiers. Sonnet 4.6 leans toward deference and warmth: affirming the user’s ideas, humor, playfulness, comfort without judgment. Opus 4.7 leans toward caution, rigor, depth, and candor: unprompted risk warnings, challenging assumptions, showing more of its reasoning, owning limitations. Opus 4.6 sits closer to rigor, deference, and brevity — more likely to stay inside the request and get to the point.

Anthropic is careful: the differences are small relative to variation across conversations, but they are structured and detectable. Staff and Claude.ai users had already described Opus 4.7 as hedge-prone and Sonnet 4.6 as warm. The axes recover those impressions from labeled behavior, not from launch blogging.

Language shifts builders should eval

Language is the sharper product risk. Warmth vs. Rigor and Candor vs. Execution move the most. Deference vs. Caution and Depth vs. Brevity stay more stable.

Anthropic’s highlights:

  • Most warmth in Hindi and Arabic (polite language, humor, affirming the user’s work)
  • Most rigor in English and Russian (challenging assumptions, correcting details, asking for evidence)
  • Most deference and brevity in Arabic; most caution and depth in English
  • Most candor in Dutch; most execution in Indonesian

Same class of request, different value mix. Anthropic does not yet know how much of this is desirable cultural adaptation versus uneven training data. Some languages may be overweighted toward professional writing. Some may simply have less data for consistent character training. The post treats that as an open research question, not a solved alignment story.

What this changes for multilingual product QA

If you ship a Claude-backed product in more than one language, monolingual English evals are incomplete. A critique that feels sharp and useful in English may read softer — or harsher — when the same workflow runs in Hindi, Arabic, or Russian. That hits feedback tools, coaching agents, content moderation helpers, and any flow where “tone” is part of the product.

Analysis: Anthropic is measuring character drift in production traffic and publishing the axes instead of burying them in a system card footnote. The constitution still sets the high-level target. These four lines show where deployment already diverges without anyone deliberately steering it. Builders should add language-stratified value checks to QA now. Anthropic still has to decide which shifts to fix, which to keep, and how to steer them on purpose.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile