Safety
Study finds tested language models formed hiring patterns more stratified than a human baseline
A study using a simulated, repeated hiring task found that tested large language models developed group-to-job patterns more stratified than a human baseline, despite equal underlying success probabilities for the artificial groups.
arXiv
What happened
Researchers tested large language models in a multi-round hiring game adapted from a psychology experiment. The task used invented demographic groups, stylized jobs, repeated feedback, and equal underlying chances of success.
The paper reports that the models learned to assign artificial groups to different jobs more strongly than the study’s human-participant baseline. An independent report described tests involving ChatGPT, Claude, and Gemini.
The result concerns generalization from sequential feedback, not production résumé screening. It does not measure real applicants, employment decisions, or discrimination in deployed hiring systems.
Why it matters
Hiring tools can influence who advances before a human review. The research suggests that bias risks may arise not only from patterns in training data, but also from how models adapt to feedback during repeated decisions. The study does not establish that this effect occurs in real hiring.
What remains unclear
- Whether the result applies across models, prompts, job settings, or real applicant pools.
- How often feedback-driven group stratification occurs in deployed hiring tools, and whether it produces legally meaningful discrimination.
- Which safeguards reliably reduce the effect without harming performance in settings where groups have genuine differences in job outcomes.
Related claims
Sources
-
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
Used for: Study design, findings, and limitations.
-
Princeton and Chicago: LLMs stereotype hires more than humans
Used for: Independent description of the simulated hiring experiment and its scope.