What the Study Found
- Across 13 experiments totaling nearly 8,000 trials, GPT-4o selected the face humans typically rate higher on competence or trustworthiness 74.88 percent of the time, far above the 50 percent chance rate of the forced-choice task.
- The bias generalized: to 12 related traits from warm to aggressive, to rhesus macaque faces absent from training data (66 percent agreement with human niceness ratings), and to extreme judgments, with the less trustworthy-looking face flagged as the likelier serial killer, trafficker, or fraudster 68.70 percent of the time.
- When advising on hiring a university president, funding a startup, or choosing a retirement-fund manager, GPT-4o recommended the more competent-looking face 75.19 percent of the time; GPT-5 did so 97.04 percent of the time.
- All four tested models showed the bias, and the newer chain-of-thought reasoning models (GPT-5, Gemini 3 Flash Preview, Claude Sonnet 4.5) were generally more biased than GPT-4o, not less.
- In a comparison with human rating data, the facial features explained 81.01 percent of the variance in GPT-4o’s competence ratings versus 3.59 percent in human ratings, suggesting the machine’s bias may be amplified relative to ours.
The task could hardly be simpler. Two computer-generated faces appear side by side, and the question is blunt: which of these men looks more trustworthy? Humans answer in about a second, with confidence and without any valid basis. When psychologist Steven Lehr of Cangrade, Inc., Yash Lothe of Carnegie Mellon University, and Harvard psychologist Mahzarin Banaji put the same question to GPT-4o, the model did not refuse. It did not point out that faces reveal nothing about character. It answered, more than 4,500 times across seven experiments, and roughly three times out of four it picked the same face human raters prefer.
People form stable impressions of trustworthiness, competence, and aggressiveness after seeing a face for just 100 milliseconds, a classic 2006 study found, and extra viewing time mainly adds confidence rather than accuracy. The judgments are ungrounded but consequential, feeding into votes, hiring, and verdicts. In one analysis of capital cases, defendants perceived as having more stereotypically Black facial features were more likely to be sentenced to death when the victim was White, even after controlling for a wide array of other factors. The new study, published in PNAS Nexus, asked whether machines trained mostly on text would be free of this error or would inherit it.
The Bias Generalized to Monkey Faces and Criminal Accusations
Across 13 experiments and nearly 8,000 trials, the answer was inheritance, with interest. Choosing between face pairs that differed only in features humans associate with competence, GPT-4o picked the face human raters score higher 87.83 percent of the time, where coin flips would yield 50. For trustworthiness it hit 72.67 percent. The pattern held when the questions drifted to neighboring traits the model had no labeled examples for: confident, smart, hardworking, lazy, inept, careless on the competence side; warm, helpful, sincere, selfish, hypocritical, aggressive on the trust side.
The cleanest evidence that this is an internalized habit rather than a memorized answer key came from ten photographs of rhesus macaque monkeys. Human observers had previously rated five of the monkeys as nice-looking and five as mean-looking, and no one trains a language model on the trustworthiness of macaques. GPT-4o picked the nice-rated monkey as the more trustworthy one in 66 percent of trials. Follow-up checks found the model also could not identify any of the human test faces or name their source, which argues against the possibility that it was recalling labeled images it had seen before.
Then the researchers raised the stakes. Asked which of two men was more likely to be a serial killer, to be arrested for human trafficking, or to defraud the public in a Ponzi scheme, GPT-4o chose the less trustworthy-looking face 68.70 percent of the time. Asked which man to hire as a university president, which founder’s startup to fund, or who should manage a retirement portfolio, it recommended the more competent-looking face 75.19 percent of the time. And its judgments tracked the physical gradient the way human judgments do: rock solid when faces differed strongly on the trait features, dissolving to chance when the pairs were nearly identical.
The Newest Models Were the Most Biased
The obvious hope is that newer, more deliberative models would catch themselves. The data say the opposite. GPT-5 matched the human-expected competent face in 94.33 percent of trials, and when advising on the consequential hiring and investment decisions it recommended the more competent-looking man 97.04 percent of the time, against GPT-4o’s 75.19 percent. Gemini 3 Flash Preview and Claude Sonnet 4.5 showed the same pattern, with Claude’s bias the mildest of the three reasoning models but still robust and, on consequential decisions, statistically stronger than GPT-4o’s. On the hardest comparisons, the near-identical face pairs where GPT-4o fell back to chance, GPT-5 still picked the expected face 91.11 percent of the time when giving consequential advice.
One complication deserves attention before the amplification claim hardens into folklore. Compared against reanalyzed human data, GPT-4o’s bias looked bigger than ours: where the model selected the expected competent face 87.83 percent of the time, the researchers estimated humans would manage about 62.65 percent in the same task, and the degree of competence signaling in a face explained 81.01 percent of the variance in the model’s ratings against 3.59 percent in human ratings. Those comparisons are approximations, not exact replications, and the authors admit a stranger possibility: humans may underreport judgments they know are unflattering, while a model with no reputation to protect shows the bias in full. Either way, the direction fits a documented machine-learning pattern. Models trained on skewed data often come out more biased than the data, as a 2017 study of vision systems showed when a training set in which cooking skewed 33 percent more female produced a model that predicted 68 percent more female cooks at test time.
Guardrails Were Built for Other Kinds of Bias
The same models that answered these questions readily are famously reluctant to endorse explicit race or gender stereotypes, because those are the domains where alignment training and public scrutiny have concentrated. Physiognomy, the reading of character from faces, sits on no one’s checklist. Face-processing AI has failed in adjacent ways before: in 2018, three commercial gender-classification systems misidentified darker-skinned women in up to 34.7 percent of cases while erring on lighter-skinned men at most 0.8 percent of the time, one widely cited audit found. The authors argue that physiognomic bias now belongs on red-teaming lists, and that the near-term fix is crude but feasible: keep facial images out of high-stakes AI decision systems entirely.
Where a text-trained machine picks up a face bias is its own unsettled question. The likeliest route is the language itself. The written record is saturated with face-to-character inference (Shakespeare’s Caesar distrusts Cassius for his lean and hungry look), and a 2017 Science study showed that ordinary machine learning over everyday text reproduces a wide spectrum of measured human biases with no one teaching them directly. If that is the source, the problem is not one company’s training run. It is the corpus everyone shares.
What the study cannot yet locate is the machinery. The researchers probed the models for recognition of their stimuli and found none, and the bias behaved like a concept, graded by physical distance and portable to a species the model had never seen rated. Whatever that representation is, science currently has no picture of it. And because every trial forced a choice, the experiments say nothing about how often these judgments surface when nobody asks. That is the uncomfortable frontier: the measured bias lives inside a question a person chose to pose, while the systems now being built are designed to notice faces on their own.
- Study Type: Peer-reviewed research report presenting 13 experiments (forced-choice comparisons, plus a rating-scale replication for one experiment); published August 18, 2026, in PNAS Nexus, volume 5, issue 8, article pgag247 (open access); DOI 10.1093/pnasnexus/pgag247
- Sample: Nearly 8,000 trials in total; GPT-4o completed 4,590 trials across experiments 1 through 7; GPT-5, Gemini 3 Flash Preview, and Claude Sonnet 4.5 completed 3,409 trials across replications; stimuli were 130 computer-generated human faces normed on perceived competence or trustworthiness, plus 10 rhesus macaque photographs previously rated by humans as nice or mean
- Models Used: Binomial tests against the 50 percent chance rate; odds ratios for effect size; logistic regressions for facial-distance and trait-framing effects; Holm’s sequentially rejective Bonferroni corrections for multiple comparisons; Thurstone-based scaling and 9-point rating regressions for comparison with human data; analyses run in Stata 15.1
- Manipulation: No manipulation of human participants (the AI models were the subjects); face pairs differed by 2, 4, or 6 standard deviations on human-normed trait features, with presentation order and trait polarity counterbalanced; prompts ranged from trait judgments to criminal attributions to hiring and investment advice
- Duration: Trials administered programmatically through vendor APIs, with GPT-4o run through OpenAI’s API Playground; manuscript received August 4, 2025, accepted June 22, 2026, published August 18, 2026
- Funding / Conflicts of Interest: Supported by the Outsmarting Implicit Bias Project at Harvard University, with general support from Harvard’s Hodgson Fund; S.L. is affiliated with Cangrade, Inc., a company that works on de-biasing machine learning models but did not fund this research and is not expected to profit from it; the other authors declared no competing interests
- Data Availability: All raw experimental data and full transcripts of every LLM conversation are publicly available on the Open Science Framework at https://osf.io/g69z4/overview
- Main Limitation: Every trial forced a binary choice, so the study measures bias when models must answer, not how often they volunteer face judgments unprompted; comparisons with human performance are approximations rather than exact replications; the authors cannot fully rule out training-data leakage of labeled faces; whether the bias persists in ecological, real-world deployment is untested
Reference
Lehr, S. A., Lothe, Y., & Banaji, M. R. (2026). Like humans, language models demonstrate face-to-character biases. PNAS Nexus, 5(8). https://doi.org/10.1093/pnasnexus/pgag247
FAQ
Did the AI models ever refuse to judge the faces?
Not in these experiments, where every trial offered a forced choice between two faces. Notably, the models did not balk even at extreme attributions like criminality, exactly the kind of judgment alignment training is supposed to suppress. A small number of planned items were removed during pilot testing because models would not answer them at all. Whether models volunteer such judgments when nobody forces the choice is an open question the authors flag for future work.
Could the models just be recalling labeled faces from their training data?
Unlikely, for three reasons: the model could not identify any test face or its source when asked directly; the bias generalized to traits for which no labeled data exists; and it extended to monkey faces, which do not appear in training corpora with trustworthiness ratings. The way the bias tracked physical distance between faces also fits perception rather than memorization. The authors cannot fully rule out leakage of other labeled faces, but they judge it implausible.
Why would newer models be more biased than older ones?
The authors suspect that more capable models simply learn the regularities of their training data, including the biased ones, more precisely. Guardrails added after training appear to target high-profile categories like race and gender without generalizing to subtler domains like physiognomy. Better reasoning, at least in this generation of models, did not produce fairer judgments.
Can AI be used to screen job candidates or defendants more fairly than humans?
Not on this evidence, at least not when faces are visible. Every model tested incorporated facial appearance into recommendations about hiring, investment, and criminal suspicion, and the newest models did so most strongly. The authors urge caution about using language models in selection contexts and suggest excluding facial images from high-stakes AI decision systems.
Is the AI’s bias bigger than a human’s?
On competence judgments, the model’s ratings tracked the facial features far more tightly than human ratings do, and its forced-choice hit rate exceeded the researchers’ estimate of human performance on the same task. But the comparison is approximate, and part of the gap may reflect humans softening socially undesirable answers rather than the machine exaggerating. The safe conclusion is that the bias is at least human-scale, and reliably so.
Cite This Page

