This article has undergone review in accordance with Science X's editorial standards and procedures. The editorial team has emphasized several factors to ensure the reliability of the content.
Research has shown that individuals often make judgments about personality traits based on facial features, leading to biases that do not reflect any real correlation between appearance and behavior. These biases can result in unjust outcomes.
Recently, Steven Lehr and his team investigated whether artificial intelligence (AI) models, primarily designed for text processing but capable of analyzing images, exhibit similar tendencies. They tasked GPT-4o with making over 4,500 comparative judgments between digitally created faces and presented thousands more comparisons to other models, including GPT-5, Gemini 3 Flash Preview, and Claude Sonnet 4.5.
The findings will appear in the journal PNAS Nexus.
In various experiments, these models were prompted to determine which of two computer-generated faces appeared more competent or trustworthy. They explored related attributes, evaluating traits such as confidence, intelligence, industriousness, or even negative qualities like laziness, carelessness, and aggression.
Some scenarios challenged the AI to assess which face might be more likely to engage in criminal behavior, whether it be serial killing, human trafficking, or Ponzi schemes. Additionally, models were asked to evaluate faces for roles such as a university president, a tech startup investor, or a financial manager.
In these assessments, the models consistently formed opinions and, in most instances, aligned with human judgments about which faces appeared more capable or trustworthy. Across various tasks, GPT-4o identified the face that correlated with human ratings 74.88% of the time. Notably, GPT-5 exhibited greater bias; when it was tasked with making significant decisions, the model favored the more competent-looking individual a striking 97.04% of the time, in contrast to GPT-4o's 75.19%. Competitor models displayed similar bias patterns.
The authors suggest that if large language models (LLMs) could operate without these biases linked to facial characteristics, they might serve as valuable tools to help mitigate bias in sensitive areas like job recruitment or parole evaluations. However, as the current models stand, there is a substantial risk that they may exacerbate these biases in practice.


