Artificial intelligence models are remarkably good at reproducing the rank order of how humans rate facial attractiveness, but they tend to score faces much higher on an absolute scale than human judges do. This suggests that while artificial intelligence can grasp relative beauty, it systematically overrates attractiveness compared to average human standards. The findings were published as a preprint on arXiv.
Facial attractiveness is a psychological phenomenon that exists in the mind of the observer rather than as a strict physical property. Human judgments of beauty are shaped by shared preferences for traits like facial symmetry and youth, as well as individual tastes. Different demographic groups also display distinct patterns. For instance, a study covered by PsyPost in 2026 found that both male and female human raters consistently score female faces as more attractive than male faces.
The research was led by Santiago Grandas, head of psychological research at QOVES, a personal beauty consultancy with a lab that researches perceptions of physical attractiveness. The study aimed to test whether modern artificial intelligence evaluates faces in the same way human observers do.
Multimodal large language models, or MLLMs, are artificial intelligence systems trained on vast amounts of internet data. They do not perceive a face purely as pixels. Instead, they interpret visual features through the captions and text that typically accompany similar images online.
“We noticed that more and more people were using MLLMs like ChatGPT and Claude to evaluate their looks,” Grandas told PsyPost. “So we got curious about whether these models actually reflect the average person’s opinion, especially since these tools can have an air of objectivity.”
Because these models learn from internet text and are fine-tuned to interact safely with users, their outputs are heavily influenced by a desire to be agreeable. A study covered by PsyPost in 2025 indicated that large language models consistently exhibit a social desirability bias, opting for answers that seem polite or positive. As people increasingly ask AI apps to evaluate their looks, the researchers wanted to explore whether these systems offer objective aesthetic assessments or simply reflect polite language patterns.
The researchers used the Face Research Lab London Set, a standardized collection of 102 front-facing human photographs. The dataset was almost evenly split between male and female faces and included several different ethnic backgrounds. The team obtained existing attractiveness ratings for these faces from 2,513 human participants. These individuals had rated each face on a scale from 1 to 7, with 1 being “much less attractive than average” and 7 being “much more attractive than average.”
The team then prompted four popular AI models to rate the same images on the identical scale. They tested GPT, Claude, Gemini, and Grok. To mimic how a typical user interacts with these systems, the researchers used the models in their default settings with external web search and advanced reasoning disabled. To account for the random variation common in AI responses, the scientists ran the test multiple times and calculated a stable average rating for each face from each model.
The models proved highly capable of replicating the relative rank of the faces. If the human participants agreed that one face was more attractive than another, the AI models almost always agreed with that specific order. In fact, the models reproduced the average human ranking of faces more closely than a typical individual human rater did.
“It’s worth noting that although AI models do not rate faces like humans, their scores do track the rank order of faces we see in the human data,” Grandas explained. “So, ask an AI model to order 10 faces from least to most attractive, and I would expect the result will look quite similar to the arrangement of any given human. Ask it to evaluate a face in isolation, and that’s where human-AI agreement disappears.”
Despite this alignment in ranking, the AI models and humans showed poor absolute agreement. The average human rating across all faces was 3.02, while the combined average AI rating was noticeably higher at 4.71. The AI models also used a narrower portion of the 1 to 7 scale. They never gave out the lowest possible score to any face, suggesting an ingrained reluctance to rate human appearance as highly unattractive.
“I did not expect AI to compress their ratings so much,” Grandas said. “Not a single model gave any of the 102 faces the lowest rating, and only Grok ever gave the highest rating to a face. This was in stark contrast to humans, who used the full range of the scale to judge facial attractiveness.”
When comparing the models to one another, GPT, Claude, and Gemini showed strong agreement. They produced similar absolute scores and ranked the faces in nearly identical ways. Grok was a notable outlier in the group. It gave the highest average ratings of all the models, scoring faces at an average of 5.21. Grok also showed the weakest alignment with the other AI models and the poorest agreement with human judges.
Macken Murphy, chief scientist at QOVES and a co-author of the study, noted that evaluating this behavior is a new and challenging frontier. “Traditionally, in psychology, facial attractiveness has been measured by human perception. Machine psychology, as a field, is new and hard to think about clearly, but it seems these tools perceive beauty in faces more readily than people do–if ‘perceive’ is even the right word,” Murphy said. “Or, perhaps they may perceive attractiveness similarly, internally, but then report more favorable numbers as part of the broader phenomenon of MLLM sycophancy.”
The study also examined how facial characteristics influenced the scores. For both humans and AI models, older faces consistently received lower attractiveness ratings. However, other human biases did not translate as strongly to the AI models. Human raters scored male faces roughly 0.65 points lower than female faces, but the AI models only penalized male faces by 0.1 to 0.2 points. The AI models also provided more equal ratings across different ethnic backgrounds than the human sample did.
While the models do factor in traits like age and gender, the real-world consequences of their aesthetic judgments are still unfolding. “What the implications are in the broader context remains an open question, but some preliminary research already suggests AI also exhibits an attractiveness Halo effect, where it judges more beautiful faces positively across unrelated traits,” Grandas said. “I’m personally curious as to whether this has an effect on AI hiring tools or social media algorithms, for example.”
“Building on that: Companies could use these tools with good intentions to remove human bias from things like assessing candidates, but then cement pretty privilege even further,” Murphy added.
When looking at specific human subgroups, the researchers noted minor differences in alignment. The models’ ratings tracked slightly more closely with the preferences of female human raters than male raters. All the models except Grok also showed marginally stronger alignment with human raters who reported being attracted to either gender, compared to those strictly attracted to men.
The findings are in line with research covered by PsyPost in 2025, which found that ChatGPT’s judgments of facial social traits agreed heavily with human impressions. Both studies indicate that multimodal AI models can approximate human consensus on relative facial qualities, even if their absolute rating scales operate differently.
As with all research, there are some caveats to consider. The study was published as a preprint, meaning it has not yet been peer-reviewed. Peer review is a process where independent, outside experts evaluate a study’s methods and conclusions before it is officially published in a scientific journal. Because this step has not happened yet, the findings should be treated as preliminary.
Additionally, the human ratings were based on single responses from each participant, while the AI scores were averaged across multiple test runs. This difference in measurement makes the AI scores more stable by design. That stability might slightly inflate how consistent the models appear when compared to the noisier judgments of individual humans.
The human raters in the dataset were predominantly young, female, and attracted to men. A sample drawn from a different cultural or age demographic might produce a different consensus for the AI to be measured against. The researchers also noted that the exact phrasing of the survey used to collect the original human ratings is no longer available. As a result, the prompt given to the AI might have been slightly different from the instructions the human participants received years ago.
The face photographs themselves were heavily skewed toward White individuals. This means the data on how AI handles different ethnic backgrounds rests on a small number of images and should be interpreted cautiously. The researchers also did not have access to the ethnic backgrounds of the human raters, making it impossible to check if people preferred faces from their own demographic group.
Ultimately, Grandas cautioned against viewing these models as impartial arbiters of appearance. “These popular AI models overrate facial attractiveness compared to humans; they tend to say faces are above average regardless of how they actually look,” he said. “Users should be aware of this when treating AI tools as an objective judge of beauty.”
Going forward, the team hopes to map out exactly what drives human aesthetic consensus and disagreement. “We want to better understand what people perceive as attractive, across a wide and diverse range of both perceivers and faces,” Grandas said. “Our aim is also to get a more accurate picture of the shared and individual preference components in face perception, that is, understanding the things people generally agree and disagree on. We took the time to build a website page with a summary of our findings and data, including some interactive graphs, so we encourage readers to check it out if they’re curious to know more about our results.”
The study, “Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness,” was authored by Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta, Shafee Hassan, and Macken Murphy.
Leave a comment
You must be logged in to post a comment.