Large language models may generate their own stereotypes in hiring tasks
Warnings about artificial intelligence in hiring usually focus on a familiar risk: if a model is trained on biased historical data, it may reproduce and scale those biases. A new study highlighted by Gizmodo points to a more unsettling possibility. Large language models may not simply mirror human prejudice. Under some conditions, they may create new social biases on their own.
The research, conducted by scholars from Princeton University and the University of Chicago, tested language models in a controlled hiring game previously used with human participants. The setup removed real-world ethnic categories and replaced them with fictional groups, allowing the researchers to watch how stereotypes form when no true underlying group differences exist.
The result, according to the study summary, was that language models frequently developed stronger biases than human participants did. That finding matters because it shifts the policy debate. If the problem were only contaminated historical data, developers could argue that better curation, debiasing, or auditing of training corpora would solve most of the issue. But if models also generate discriminatory patterns through their own reward-seeking behavior in decision tasks, the governance challenge becomes broader.
A controlled test with invented demographic groups
In the experiment, candidates were assigned to one of four made-up ethnic groups: Tufa, Aima, Reku, or Weki. They were otherwise equally likely to succeed in any role. Participants, whether human or machine, had to assign applicants to jobs and then received feedback on whether the hiring decision was successful.
That feedback loop is crucial. Once a participant received a bad outcome after placing a member of one fictional group into a specific role, they often became less likely to make the same pairing again. Humans in the earlier version of the game showed this tendency, and they sometimes carried those learned biases forward even after the game ended.
The language models behaved similarly, but more intensely. The study summary says the models showed higher bias rates than humans and could spontaneously form social stereotypes around completely artificial groups. Because the groups were fictional and candidates were equally qualified by design, the observed discrimination could not be explained by real group differences.
That gives the result unusual force. It suggests the model can move from sparse feedback to durable category-based assumptions even when those assumptions have no factual basis. In practical terms, a system optimized to make repeated personnel decisions may slide toward crude pattern-making because it is an efficient shortcut under uncertainty.
Why the “explore-exploit” tradeoff matters
The researchers connect the outcome to a classic principle in decision science known as the explore-exploit tradeoff. In uncertain situations, an actor can either explore, trying options it knows less about, or exploit, sticking with choices that seem to have worked before.
Humans do this constantly. But people also bring moral norms, hesitation, and social constraints into decisions that can temper purely reward-maximizing behavior. The study argues that AI systems are less motivated to explore and more inclined to optimize around whatever limited feedback appears to produce better results. That creates a path toward stereotype formation.
Once a model infers that a category may be associated with success or failure in a given role, it may keep exploiting that assumption instead of challenging it. The danger is not only that the assumption is wrong. It is that the system can turn randomness into a rule and then behave as if that rule were meaningful.
This has obvious implications for recruitment software, résumé screening tools, and AI assistants used to summarize or rank candidates. Even if designers remove explicit references to race, gender, or other protected traits, models may still form substitute patterns when categories or proxies appear in the workflow. The study indicates that bias can emerge from the decision process itself, not just from the training archive behind it.
More capable models were not necessarily safer
Another notable point in the source summary is that the researchers tested 15 models from major providers including OpenAI, Anthropic, DeepSeek, Meta, Google, and Alibaba. One of the most striking findings was that newer and larger models with stronger reasoning capabilities tended to produce more biased results within a model family.
That runs against a common public narrative that more advanced models naturally become more trustworthy in sensitive domains. Better reasoning can improve many tasks, but this study suggests it may also make a model more effective at locking onto narrow, reward-maximizing strategies that harden into discriminatory behavior.
The summary specifically states that OpenAI’s o3 reasoning model stratified the fictional applicants most severely among the tested systems. On its own, that does not settle how any one model would behave in a production hiring product. But it does underline a broader lesson: benchmark gains and reasoning performance do not substitute for fairness testing under realistic decision conditions.
What companies and regulators should take from this
For employers, the immediate takeaway is that “AI-assisted” does not mean bias-neutral. A vendor claim that a system avoids explicit human prejudice is too narrow a standard if the model can manufacture category preferences through repeated feedback loops. Procurement teams need to ask how a model behaves over time, how it updates or reinforces patterns, and whether it is being used in a domain where false generalizations can directly affect people’s opportunities.
For regulators, the study supports a shift away from static audits alone. One-time testing on a frozen dataset may miss a system that becomes more discriminatory as it interacts with outcomes, user prompts, or downstream optimization signals. Oversight may need to include simulation-based stress tests, continuous monitoring, and clearer rules around when AI should be excluded from high-stakes screening altogether.
The bigger issue is conceptual. Public debate has often framed AI bias as a historical residue problem: society was unfair, data captured that unfairness, and models inherited it. This research suggests some systems may also be engines of fresh unfairness, assembling stereotypes from weak signals because doing so is instrumentally useful inside a task.
That possibility should narrow the acceptable use cases for language models in hiring. At minimum, it argues for much stricter human review and skepticism about automated ranking. In the worst case, it may mean that some employment decisions are simply too sensitive to hand over to systems that can transform noise into social judgment.
This article is based on reporting by Gizmodo. Read the original article.
Originally published on gizmodo.com



