AI models such as GPT-4o, Gemini, and Claude exhibit the same prejudices as humans when evaluating faces, often judging some as more competent or trustworthy without any valid reason. A recent study shows that these systems rely on unfounded visual impressions to make critical decisions, such as hiring or criminal assessments. This highlights a significant flaw in AI security, as developers struggle to eliminate these deeply rooted biases. AI models designed to process language and images often reflect the biases present in their training data, which can include human prejudices and the opinions of their creators.
The study, led by Steven A. Lehr, Yash Lothe, and Mahzarin R. Banaji, and published in the journal PNAS Nexus, examined this phenomenon through thirteen experiments involving nearly 8,000 trials on several large language models. The researchers asked a simple question: would AI systems reproduce a well-known human error or overcome it? Their findings suggest that AI models do reproduce these biases, often more intensely than humans. The AI models evaluated faces for competence and trustworthiness using subtle variations in facial structure, even though these features provide no real information about a person's abilities or moral value.
During the experiments, the AI models consistently judged certain faces as more competent based solely on facial structure. The researchers used computer-generated faces designed to show specific variations that influence human judgments. For instance, GPT-4o selected the face humans typically judge as the most competent in about 88% of cases and the most trustworthy in about 73% of cases. These percentages are significantly higher than chance. When the faces showed more pronounced physical differences, the AI's choices became even more consistent, selecting the "expected" face as the most competent in 98% of cases.
The study also found that AI models extended these biases beyond humans. In one experiment, GPT-4o was shown photographs of rhesus macaques and identified the faces labeled as "sympathetic" by humans as more trustworthy in about 66% of cases. This suggests that the AI had learned a general pattern linking certain facial features with reliability, even in different species. The researchers further tested the AI's ability to judge criminality and suitability for roles based solely on facial features, with the AI consistently favoring faces it deemed more trustworthy or competent. This raises concerns about the fairness of AI-based hiring systems that rely on such superficial criteria.
AI Models Exhibit Facial Bias Similar to Humans in Assessing Competence and Trustworthiness
AI-rewritten from original reportingHow it works
ai-biasfacial-recognitionmachine-learninghuman-biasai-ethicsalgorithmic-bias



