GYROval: Measuring Cultural Values in Language Models
This presentation introduces GYROval, a benchmark designed to measure how large language models express cultural value orientations across different contexts. Unlike traditional benchmarks that look for objectively correct answers, GYROval examines whether models maintain consistent relative rankings when evaluated across different decision roles, topics, languages, and sampling conditions. The research reveals that while models preserve their broad ordering, their absolute value scores can shift substantially depending on how scenarios are framed—a finding with significant implications for deploying language models in culturally sensitive applications.Script
When a language model recommends a hiring decision or advises on family matters, it applies a cultural value orientation. But here's the puzzle: how do you measure values when there's no objectively correct answer? The researchers behind GYROval built a benchmark that tests whether models maintain their relative cultural rankings across 224 different scenarios, spanning seven life domains and four decision roles.
Every item presents two legitimate options with no answer key. A model acting as an advisor might score 0.67 on the emancipative pole, but that same model judging another person's behavior might score 0.80. The benchmark measures whether models preserve their rank order even when these absolute scores shift by over 20 percentage points.
The design crosses seven themes, four decision modalities, and eight value facets, producing 224 unique cells. Twenty language models were evaluated in English, and 11 were additionally tested in Russian at two sampling temperatures. The result is a stability test: do models maintain their relative positions when the context changes?
The answer is yes, but with important nuance. Rank concordance across temperatures reached 0.96, meaning models kept nearly identical orderings. Language shifts between English and Russian produced high agreement at 0.91, while decision roles showed the lowest stability at 0.69. All four factors exceeded random chance by wide margins.
A mechanistic sub-study revealed that these value distinctions live inside the models. Linear probes decoded value axes with accuracies between 68% and 97%. When researchers injected value-direction vectors into model activations, they changed choices in 50% to 100% of cases, demonstrating that internal representations causally influence outputs.
Here's what matters for deployment: relative rankings hold, but absolute scores do not. A model's cultural value score can shift by over 20% of the measurement scale depending on whether it's advising, acting, or judging. Report both the ranking and the conditions, and learn more about cutting-edge AI research like this at EmergentMind.com, where you can explore papers and create your own explanatory videos.