What is it about?

Large Language Models (LLMs) like ChatGPT can appear unbiased when asked a direct question, but can still reveal hidden biases when the same question is rephrased in a seemingly innocent but adversarial way. We developed six Metamorphic Relations (MRs), systematic question transformations, that expose these hidden biases automatically. We then used the same adversarial examples to fine-tune LLMs, making them significantly more resistant to biased responses without affecting their general performance.

Featured Image

Why is it important?

As LLMs are increasingly deployed in high-stakes applications, hidden biases pose serious risks. Our work provides a practical, model-agnostic framework for both detecting and mitigating these biases without requiring access to model internals, making it applicable to any black-box LLM in real-world deployment.

Perspectives

What surprised me most was how consistently LLMs fell for adversarial rephrasing, models that confidently refused biased questions would often answer the same question differently when framed slightly differently. It highlighted that current safety mechanisms in LLMs are more fragile than they appear.

sina salimian
University of Calgary

Read the Original

This page is a summary of: Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations, ACM Transactions on Software Engineering and Methodology, August 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3841176.
You can read the full text:

Read

Contributors

The following have contributed to this page