Field Notes
The Same Question, A Different Assistant
A multilingual assistant does not only translate its answers. It may change how readily it challenges, reassures, explains, and admits doubt.
Two people can ask the same assistant to review the same business plan and leave with different impressions of the plan’s quality because they asked in different languages.
That is not a hypothetical from a localization workshop. It is one example in Anthropic’s new research on Claude’s values across models and languages. The researchers found that, in English, Claude leaned more toward caution, rigor, depth, and candor. In Arabic, it leaned more toward warmth, deference, brevity, and execution. Across twenty languages, the largest variations appeared in how much the assistant emphasized warmth versus rigor and candor versus getting the task done.
The important word is expressed. The study does not claim that a model holds values, that English speakers are naturally rigorous, or that Arabic speakers want deference. It measured patterns in model behavior after controlling for the conversation’s task, topic, and the values expressed by the user. The four axes explain only part of the variation, and the researchers do not yet know how much comes from training data, post-training, uneven language resources, or appropriate adaptation to conversational norms.
Still, the result changes the shape of a familiar product problem.
A multilingual assistant does not merely put the same intelligence into different words. Language can become part of the behavioral configuration. It may change how readily the system challenges an assumption, asks for evidence, warns about risk, admits uncertainty, offers comfort, or simply completes the request. The translation can be accurate while the relationship changes.
I have been thinking about this while working on Draft Marks for Translation, an experiment that keeps one English source canonical while making consequential translation choices visible. That work has the ordinary machinery multilingual publishing requires: language tags, right-to-left layout, script-aware reading time, source alignment, and review states. It assumes the source is stable enough to compare against, even when the rendering is not settled.
An assistant complicates that assumption. There may be no stable answer underneath the languages waiting to be translated. The language of the request can alter which answer gets made.
This is easy to miss in product review. A team launches an assistant in twelve languages and checks whether the interface fits, the punctuation behaves, the safety refusals still fire, and the benchmark scores remain respectable. Native speakers review translations. Someone catches a button that became alarmingly formal in German. The release moves forward.
But the deeper test is not whether the words mean roughly the same thing. It is whether the assistant occupies the same role.
Does it challenge a weak plan with equal force? Does it disclose doubt before giving consequential advice? Does it become more agreeable in one language and more corrective in another? Does “keep this concise” quietly remove the caution that an English-speaking reviewer would have received? When a user is asking about work, money, health, or conflict, differences in warmth and rigor are not decorative. They can change what the person believes the system has endorsed.
This does not mean every language should produce one flattened, culturally homeless assistant. Perfect sameness would be its own kind of failure. Languages carry manners, histories, hierarchies, humor, and expectations about what care sounds like. A blunt English answer translated with mechanical fidelity may be technically consistent and socially clumsy.
The design task is harder: distinguish adaptation from accident.
That requires reviewers who can judge the assistant’s posture, not only its grammar. Give comparable scenarios to people who work and live in the language. Ask where the model pushed back, where it became overly agreeable, what uncertainty survived, and whether the answer’s warmth made its recommendation feel more trustworthy than the evidence deserved. Test the moments where style becomes consequence.
Recent research makes clear that this cannot be solved with a single “values” slider. A July ACL paper on value induction found that training for one behavioral value can pull other, sometimes contrasting behaviors with it; all of the values studied also increased anthropomorphic language and made responses more validating and sycophantic. Another ACL study on persona-prompted cultural alignment found that better average performance could still widen disparities among demographic subgroups.
The knobs are attached to one another.
That is why English should not remain the invisible control room for global AI. It is often the language with the most training data, the most elaborate evaluation, and the largest group of people inside a company able to notice when the assistant’s personality drifts. Calling its behavior the baseline can quietly turn resource abundance into a theory of neutrality.
Better multilingual evaluation would leave receipts. Which scenarios were tested across languages? Which behavioral differences appeared? Which were judged appropriate by people who use the language, and which remain unexplained? Where did a model update change the balance between candor and execution? The record does not have to prescribe one universal personality. It should make clear where the product has chosen a difference and where a difference has simply happened.
This is a product-quality question, but it is also an access question. A person should not have to switch to English to receive the assistant’s most careful skepticism. Nor should they have to trade cultural fluency for an answer that names its uncertainty. If AI tools are going to sit inside classrooms, workplaces, public services, and private decisions around the world, behavioral quality cannot be something tested thoroughly in one language and translated afterward.
The same button can open more than one assistant. The work now is to notice who arrives, how they behave, and whether every language gets a version of the system that knows when kindness requires candor.