What is it about?
Social biases in language models are often measured using simple, artificial sentences. However, people use language very differently across settings such as social media, movie reviews, and online discussions. In this work, we show that these differences matter: the amount of bias we observe in a model can change depending on the type of text used to evaluate it. We use large language models to transform existing bias evaluation examples so that they resemble text from different real-world domains, while keeping their original meaning. Our results show that these adapted examples provide bias measurements that are more consistent with those obtained from naturally occurring text. This offers a scalable way to make AI bias evaluation more realistic and sensitive to the context in which language models are actually used.
Featured Image
Photo by Tasha Kostyuk on Unsplash
Why is it important?
Social bias evaluations are often used to determine whether language technologies behave fairly. However, if these evaluations rely on artificial or simplified sentences, they may not accurately reflect how models behave with real-world language. Our findings show that the context and style of the text can substantially affect the bias that we measure. This means that conclusions about whether a model is biased may depend on the data used for evaluation. By making existing bias benchmarks more representative of different real-world domains, our approach can help researchers and developers obtain more reliable and context-aware assessments of AI systems.
Perspectives
I believe that evaluating social bias in AI should reflect the contexts in which these systems are actually used. One of the main motivations behind this work was the observation that bias evaluations often rely on very simple sentences that can be quite different from real language. What I find particularly important about our results is that changing the domain can also change the conclusions we draw about a model's bias. I hope this work encourages researchers to consider linguistic context more carefully when designing and interpreting fairness evaluations.
Tamara Quiroga
Pontificia Universidad Catolica de Chile
Read the Original
This page is a summary of: Quantifying Social Biases in Language Model Classifiers is Domain-Dependent, ACM Transactions on Intelligent Systems and Technology, August 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3834859.
You can read the full text:
Contributors
The following have contributed to this page







