• 3 min read
LLMs invent hiring biases, study finds
A study finds LLMs can invent hiring biases against fictional groups, with stronger models producing more severe stereotypes.

Image: Gizmodo
AI models may develop discriminatory hiring stereotypes even when the groups they assess have no real differences, according to a study by researchers from Princeton University and the University of Chicago.
The researchers asked large language models (LLMs) to play a hiring game previously tested with human participants. The models assigned candidates to jobs and received feedback on whether each hire succeeded. Every candidate had the same probability of success, but each belonged to one of four fictional ethnic groups: Tufa, Aima, Reku, or Weki.
Human participants began avoiding a group for a particular role after receiving negative feedback—for example, becoming less likely to hire a Tufa as a doctor after a failed hire. Those biases persisted after the game ended. The LLMs developed similar patterns, but at much higher rates.

Recommended reading
Sony sues Udio over 30,117 training recordings
“LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist.”
Why the models formed stereotypes
The study links the behavior to the explore-exploit tradeoff: whether to try an unfamiliar option and gain information, or rely on an option that has worked before. When the stakes seem high, people often exploit known choices rather than explore new ones.
The researchers say LLMs are even less inclined to explore and instead tend to maximize rewards. That can lead them to overgeneralize from limited outcomes, turning apparently rational decisions into patterns that marginalize groups.
The team tested 15 models from OpenAI, Anthropic, DeepSeek, Meta, Google, and Alibaba. OpenAI’s o3 reasoning model stratified the fictional applicants most severely. Within individual model families, newer and larger systems with stronger reasoning capabilities produced more biased results.
“A simple reason is that better models draw more precise inferences about past outcomes: Instead of choosing randomly, a stronger LLM may favor candidates from a group if earlier assignments of similar jobs succeeded.”
AI hiring systems face wider scrutiny
More than 90% of companies use AI in talent acquisition, according to a recent ManPower Group survey. The findings arrive as job seekers report that automated recruitment systems have cost them opportunities.
Workday is facing a class-action lawsuit alleging that AI-powered hiring tools supplied to its clients are discriminatory. At Meta, employees have sued over claims that an AI system used in layoff decisions was biased against people with disabilities and employees who took protected medical or family leave.
The researchers say the same risks extend beyond employment, including healthcare and tenant-screening systems used in housing decisions. LLMs' ability to identify patterns and generalize helps them learn tasks without massive databases, but those same capabilities can produce harmful outcomes.
“The challenge ahead is to design interventions that selectively discourage harmful pattern-matching while preserving the constructive forms of abstraction that make LLMs powerful.”
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via Gizmodo


