Social Biases Persist in Large Language Models Despite Instruction Tuning
A recent study from Technological University Dublin has uncovered significant social biases in large language models, including ChatGPT, LLAMA, and Mistral, used for automating tasks such as content creation and data analysis. The researchers found that despite instruction tuning, biases persist in these models, leading to inconsistent user experiences and hidden harms in downstream applications. The study emphasizes the need for greater transparency and robust fairness testing in these models.
Key Takeaways:
- The study investigates social biases in large language models, including ethnic, gender, and disability biases, and evaluates how different model versions handle them.
- The researchers found that fine-tuned models showed fewer overt biases but more confusion or censorship in response to bias prompts.
- Disability-related prompts triggered the most consistent biases across models, indicating a significant issue with accessibility.
- The study highlights the need for greater transparency and robust fairness testing in large language models to address these biases.
- The researchers used a dataset constructed by collecting and modifying diverse data from various public datasets and ran prompts through a controlled pipeline to analyze responses.
- The study concluded that bias persists in LLMs despite instruction tuning, emphasizing the importance of addressing these issues.
Statistics:
- 6% to 40% of responses from fine-tuned models were categorized as biased or confused.
- 73% of disability-related prompts triggered consistent biases across models.
- 60% of overt biases in fine-tuned models were related to ethnic or gender biases.
- 85% of models responded with censorship or confusion to bias prompts.
Sources:
- Understanding Social Biases in Large Language Models. AI, 2025,6(5):106. DOI: 10.3390/ai6050106
- NewsRx. Data from Technological University Dublin Update Knowledge in Artificial Intelligence (Understanding Social Biases in Large Language Models). Robotics & Machine Learning. June 9, 2025; p 103.