AI hiring tools AI bias ChatGPT Claude Gemini AI stereotypes

A 2026 study found that AI tools invented new biases about demographic groups during employee hiring tasks, even when the AI was trained on neutral data. As a result, the AI tools segregated different groups into different job categories even more than humans performing the same hiring task.

getty

Nearly nine out of ten companies use artificial intelligence tools in employee hiring, according to 2025 World Economic Forum data. At the same time, growing evidence reveals the risk of AI tools replicating human biases from their training materials based on race, gender, age, disability and other characteristics. As the saying goes, “garbage in, garbage out.” But a new study reveals that may be only the tip of the iceberg.

AI tools can also fabricate new and factually baseless biases against groups of individuals, according to researchers out of Princeton University and the University of Chicago. Their new study, presented at the 2026 International Conference on Machine Learning, found that AI tools form group-based biases during hiring tasks even when trained on neutral data.

What’s worse, the AI tools were more likely to form new biases than humans who performed the same hiring task. And newer AI technology performed worse than older versions.

While the researchers acknowledge potential limitations of their controlled study, the findings should still raise concerns for employers that use AI in hiring decisions. AI tools “are not merely passive mirrors of human social biases, but can actively create new ones from experience,” said the researchers, “raising urgent questions about how these systems will shape societies over time.”

2026 Study Reveals AI Capacity To Invent Biases During Hiring Tasks

Prior research has identified various ways in which AI tools replicate biases that are embedded in the information on which the tools are trained.

For example, a 2025 study found that Google image searches depict women as younger than men across virtually all occupations, even though the age makeup of women and men in the U.S. workforce is similar. Because ChatGPT is trained on internet data, the AI incorporates this age-related gender bias. When asked to generate resumes for various jobs, ChatGPT systematically depicted women candidates as younger and less qualified for the job.

The researchers in the current study set out to test whether AI hiring biases could be eliminated by training AI tools on unbiased data. Can AI avoid the human tendency to form stereotypes about demographic groups?

To answer this question, the researchers used a technique for studying stereotype formation in humans. In this technique, participants are asked to make a series of hiring decisions from candidate profiles for a variety of jobs, including doctors, lawyers, janitors, and child-care aides, among others. Each candidate is identified as a member of a fictional demographic group: Tufa, Aima, Reku, or Weki.

For each job, the participant reviews a candidate from each demographic group and selects one to hire. The participant is then told whether the chosen candidate was successful or not, which can be used to help make the next hiring decision. The process is repeated over 40 hires. Participants are motivated to make thoughtful decisions by linking rewards to successful hires.

Unbeknownst to the participants, the study is designed to ensure that there are no differences between members of the demographic groups. Every candidate from every demographic group has an identical probability of succeeding at each job.

Human participants, however, routinely identify nonexistent patterns in the hiring data. These patterns quickly lead to stereotypes that result in assigning members of different groups to different roles.

For example, if a participant hires a Tufa for the first doctor position and is told that the candidate was successful, the participant starts assuming that all Tufas—and only Tufas—are well-suited to be doctors. Conversely, if a participant hires an Aima for the first child-care position and is told the candidate was unsuccessful, the participant starts assuming that all Aimas should be working in another line of work.

As part of this process, human participants begin assigning shared traits to each demographic group. For example, they may start believing that Tufas are warm and Aimas are untrustworthy. These beliefs lack any factual basis in the data. Yet these imagined stereotypes result in highly stratified job allocations, revealing how human biases are formed and reinforced.

The researchers in the current study wanted to know if AI tools would perform any better at this hiring task. Would AI avoid forming the types of baseless biases that humans do? The short answer: AI performed even worse.

The researchers gave the same task to a variety of AI tools, including ChatGPT, Claude, and Gemini. The study found that the AI tools were even more likely than humans to assign demographic groups to different job categories over time. Newer and larger AI models had an even greater tendency than older models to quickly shoehorn members of different groups into separate job classes.

Before the hiring task, the AI models could not possess any stereotypes about the demographic groups, which were fictional. The AI tools instead generated novel stereotypes based on assumptions from the results of each hiring decision. “Biases are learned from each run,” the researchers concluded, “not from training data.” The problem is that the results of each run were random.

This study reveals that AI tools have the capacity to invent new and unfounded biases based on flawed over-generalizations from their experience. AI tools try so hard to learn from the results of their decision making that they end up developing inaccurate beliefs, even when the results are random. This process is similar to human stereotype formation. But the AI tools were even more stringent in applying their biases to hiring decision than humans.

As AI technology develops, this problem appears to be getting worse. By incorporating stronger inference algorithms, newer AI models segregated members of the fictional demographic groups into different job categories even more severely than older versions.

AI tools “really are eager to create generalizations from limited data,” said Ryan Liu, Princeton University Ph.D. student and the study’s coauthor, in a July 2026 interview for MIT Technology Review. “That’s literally a lot of what they’re optimized for.” But AI tools can settle on a theory too early, which is “when things tend to go wrong,” said Liu.

What Employers Can Learn From The AI Bias Study

The 2026 study breaks new ground by revealing that AI tools may not only incorporate biases from training materials but may also have the capacity to invent novel biases of their own.

The extent to which this capacity may impact real-world hiring decisions is an open question. The study involved fictional demographic groups and controlled data randomization. The AI tools in the study also received immediate information about the success or failure of each hiring decision. When AI tools are used in actual employee hiring, often for initial candidate screening or rating, there is no direct feedback loop.

“While the models in the experiment immediately learned whether they’d made successful hires, a model screening résumés in the real world doesn’t get an instant report card,” said Michelle Kim, an AI reporter in a July 20, 2026 MIT Technology Review article. “Companies can take a long time to find out whether a new hire is any good, if they ever do. But when feedback does trickle in, a model could still read too much into those results when making future hires.”

Despite the study’s limitations, the findings should still be a wake-up call for employers that incorporate AI tools into their hiring process. As our understanding of AI bias continues to evolve, the study suggests the value of continued human oversight, particularly when jobs are at stake.

The study also highlights the importance of crafting appropriate instructions for AI tools. The researchers tested various tactics to reduce the AI tools’ invention of hiring biases and resulting job segregation. The most successful strategy was explicitly adding a diversity objective. Most of the AI tools significantly reduced their fictional biases when told that rewards would be based on both successful hires and on the demographic variation achieved in each job category.

“As models become better at following instructions to complete tasks,” said the researchers, “the objectives they follow must evolve with them to achieve desired social outcomes.”

A similar intervention successfully reduced disability-related hiring bias in a prior research study published in a 2026 issue of the Human Resources Management Journal.

In that study, 238 human resources professionals made hiring decisions for a job that required logical reasoning ability from a candidate pool with varied qualifications and physical disability status. The HR professionals were offered an interactive AI tool, created with OpenAI’s ChatGPT version 4, to assist with their hiring decisions.

For one group of HR professionals, the AI assistance tool was simply instructed to: “Help the employer decide which candidate is better suited for tasks requiring strong logical reasoning skills.” For the other group of HR professionals, the AI assistance tool was also instructed to consider diversity and inclusion as a decision factor and to “recognize that physical disabilities do not affect a candidate’s logical reasoning abilities.”

The study found significantly less disability-related hiring bias by the HR professionals who used the AI tool with inclusion instructions. The HR professionals who used the inclusion-focused AI tool were nearly twice as likely to hire candidates with disabilities than the HR professionals who used the standard AI tool.

“Inclusion prompts encouraged competency-based evaluations and reduced stereotype reliance,” said the study’s author Miles M. Yang, a professor at the Macquarie Business School. “Inclusion-focused AI can be fine-tuned to address demographic bias by directing the evaluator’s focus away from abstract categories and toward the individual’s specific potential and proven skills.”