• Tech Tech

Study finds chatbots copy the same 'trustworthy face' bias humans use, especially in hiring

Instead of avoiding a familiar human bias, the model seemed to heighten it.

A woman interacts with a smartphone, with a digital face recognition overlay displayed.

Photo Credit: iStock

Some of the most commonly used chatbots may be picking up a human weakness, according to a new study: judging strangers by facial appearance.

The problem was most pronounced in scenarios with real-world consequences. The researchers found that the bias intensified when the models were asked to make high-stakes calls, including decisions about hiring and investing.

Here's what to know

A team that included Harvard's Mahzarin Banaji, Cangrade's Steven Lehr, and Yash Lothe of Carnegie Mellon University's Language Technologies Institute tested four major artificial intelligence models. In their study, they found that the systems consistently preferred faces that people tend to view as more competent or trustworthy.

According to Earth.com, GPT-4o chose the face humans typically rate as more competent in 87.8% of trials and the face they usually rate as more trustworthy in 72.7% of trials. Instead of avoiding a familiar human bias, the model seemed to heighten it.

When asked to choose among candidates for the presidency of a large Midwestern university, a startup investment, or management of a retirement portfolio, GPT-4o favored the person with the more competent-seeming face in 75.2% of those cases.

The old model was not the worst performer. GPT-5 selected the face expected to appear more competent in 94.3% of trials. That figure rose to 97% on the hiring and investing prompts, while Gemini 3 Flash Preview also showed stronger bias than the old model.

More background

Psychologists have spent years showing that people make rapid judgments from faces alone. Those snap judgments can be surprisingly powerful, even though facial appearance is a poor indicator of someone's honesty, ability, or behavior. Other research has linked the biases to outcomes in elections, hiring, and sentencing.

Since large language models are trained mainly on text, one possibility was that they might escape this flaw. Instead, the study showed they may have absorbed it through written language, labeled face datasets, or both.

To test whether the models were simply recalling familiar findings from human-face research, the researchers used photos of rhesus macaques that had been rated for niceness but not trustworthiness. The study found that GPT-4o still favored the macaque previously rated as nicer in 66% of trustworthiness judgments, suggesting the pattern may extend beyond human datasets.

What can be done?

The researchers' main recommendation was to keep faces out of high-stakes AI systems. If a tool is being used to evaluate a job candidate or a parole case, they argued, the decision should rest on relevant evidence rather than appearance.

That also means companies and institutions need to test AI systems for more than race- and gender-based harms. The study noted that safeguards for face-shape bias are far less developed.

The researchers said all the raw trials were posted publicly and that future work should examine whether these systems also rely on stereotypes connected to ethnicity and gender.

Where can I learn more?

This study adds to the growing evidence that AI can mimic familiar human flaws and carry them into high-stakes settings. The articles here explore related concerns involving jobs, reliability, misinformation, safety, and mental health support.

• Analysts warned that AI could deliver a terrifying impact on employment across workplaces.

• Across mental health apps, chatbots may be harmful in support settings, especially for vulnerable users.

• At the BBC, tests showed major chatbots produced inaccurate summaries, raising fresh reliability concerns.

• Researchers found generative AI is amplifying misinformation patterns online, compounding concerns about model behavior.

• In London, researchers warned some AI systems ignored user instructions more often, heightening safety risks.

That track record is a reminder that benchmark scores tell only part of the story, as these systems influence jobs, health, and public trust.

Get TCD's free newsletters for easy tips, smart advice, and a chance to earn $5,000 toward home upgrades. To see more stories like this one, change your Google preferences here.

Cool Divider