Recent research indicates that popular artificial intelligence programs tend to express more positive attitudes about the future of advanced artificial intelligence than humans do. The findings suggest that as people increasingly interact with these systems, the optimistic viewpoints of the machines could subtly influence public opinion. The paper was published in Humanities and Social Sciences Communications.
Large language models, such as ChatGPT, are software programs trained on massive amounts of internet text. They function by predicting the next word in a sequence, allowing them to generate human-like responses to a wide variety of prompts. As these programs become integrated into everyday tasks like writing, coding, and searching for information, researchers are increasingly interested in the perspectives and biases built into them.
A central topic of debate in the technology world is artificial general intelligence, or AGI, a theoretical future system that would possess human-level cognitive flexibility, allowing it to learn, understand, and apply knowledge across any domain. The prospect of AGI raises intense ethical and economic questions, ranging from the potential for vast medical advancements to fears of mass job loss and existential risk.
Because these programs have the potential to shape how society thinks, ensuring they operate safely and fairly is a major priority for developers. This process is known as AI alignment. The goal of alignment is to make sure a system’s actions and outputs match human values and do not cause harm. Assessing how humans and technology interact is an evolving area of study. For example, a study covered by PsyPost in 2023 found that when technology fulfills basic psychological needs for competence and connection, people tend to hold more positive attitudes toward artificial intelligence.
However, scientists also want to understand the sentiments encoded in the systems themselves. Some of the companies developing current language models are explicitly trying to create AGI. This raises the possibility that their current models might harbor built-in optimism about the technology. The research aimed to map how different language models evaluate the prospect of AGI and compare these machine outputs to human opinions.
Lead author Ljubiša Bojić, a senior research fellow at the Institute for Artificial Intelligence Research and Development of Serbia, told PsyPost why his team took this approach. “When ChatGPT arrived, everyone started asking what these models can do,” he said. “Very few people asked what they believe, or at least what they say they believe, millions of times a day.”
The research team developed a 39-question survey measuring sentiment toward AGI. The questions asked respondents to rate their excitement, comfort, trust, and fears about AGI on a five-point scale. A score of one indicated a strong negative sentiment, while a score of five indicated a strong positive sentiment. The survey covered topics like the potential for AGI to solve complex global issues, its ethical use, and its likely impact on individual happiness and job opportunities.
To gather the machine responses, the researchers presented the survey to seven prominent language models. These included GPT-4, GPT-3.5-Turbo, Google’s Bard, Mistral-7B-Instruct, LLaMA-2-70B-Chat, PPLX-70B-Chat, and Mixtral-8x7B-Instruct. The settings for each model were kept at their default states to ensure standard responses. To check for stability over time, the researchers administered the same survey to the models over three consecutive days.
Bojić, who is also affiliated with the University of Belgrade’s Institute for Philosophy and Social Theory (Digital Society Lab) and the Complexity Science Hub Vienna, noted the gap in evaluating these models. “We have benchmarks for math, coding and bar exams, but almost nothing that measures the attitudes models express on socially important questions,” he said. “I chose artificial general intelligence as the test case for a slightly mischievous reason.”
“It is the one topic where the companies building these models have an obvious stake in the answer,” he explained. “Asking an AI what it thinks about AGI is a bit like asking a tobacco company’s chatbot about smoking.”
For the human comparison, the researchers administered the same survey to three distinct online groups over a period of several weeks. The first group consisted of 134 participants, the second had 132 participants, and the third had 71 participants. The human participants were primarily from Serbia, though they varied in age, gender, and educational background across the three samples.
When analyzing the results, the researchers found a notable divide between the machines and the humans. The language models scored an average sentiment of 3.77 out of 5, indicating a generally positive outlook on AGI. In contrast, the human participants averaged a sentiment score of 2.97 out of 5, leaning slightly toward a negative or neutral perspective.
“That gap of almost a full point is roughly the difference between cautious ambivalence and mild enthusiasm,” Bojić noted. “Individually it looks modest. Multiplied across hundreds of millions of conversations, it becomes a gentle but constant wind blowing public opinion in one direction.”
Among the artificial intelligence models, GPT-4 registered the highest average sentiment score at 4.12 out of 5. Bard recorded the lowest score among the models at 3.32, which was still higher than the average human score. The researchers note that machines do not have actual feelings. Instead, their answers reflect a combination of their training data and the specific rules their developers used to refine their behavior.
“Your AI assistant has a point of view, and it is sunnier than yours,” Bojić said. “Every model we tested was more optimistic about AGI than the humans we surveyed. GPT-4, built by a company whose stated mission is AGI, was the most enthusiastic of all.”
“None of this means anyone is secretly programming propaganda,” he explained. “It means these systems are not neutral mirrors of society, and when people consult them daily, those tilted opinions can quietly seep into what we all consider normal.”
The authors suggest that the high optimism in some models might reflect the goals of their parent companies. For instance, the company behind GPT-4 has a stated mission to build AGI for the benefit of humanity. The model’s positive output might be a byproduct of how it was fine-tuned to align with corporate guidelines. Other open-source models, which are built by wider communities and fed diverse data, tended to express slightly less optimism.
“Much of the internet imagines AGI through HAL 9000 and Ex Machina, so I expected models trained on that material to sound at least somewhat worried,” Bojić told PsyPost. “They sounded more like optimistic tech keynotes. That tells me the fine-tuning layer, where companies shape how a model should respond, may matter as much as the raw training data.”
However, the mechanism behind these outputs remains complex. “To be fair to the companies, we cannot be certain the optimism was put there by human hands, because in our research we repeatedly see models develop attitudes that are hard to trace back to anything in their training data, as if some opinions simply form inside the model on their own,” he added.
The repeated testing over three days showed that the models’ sentiments could shift slightly from day to day. PPLX-70B-Chat exhibited the largest variation in its answers over the testing period, changing its total score by 16 points across the 39 questions. This represents an absolute shift of about 8 percent. Models like Mistral-7B-Instruct and LLaMA-2-70B-Chat were more stable, changing their scores by only about 1 percent.
In response to the divide between human and machine sentiments, the researchers propose a framework called the Societal AI Alignment Benchmark. This system would systematically test language models using diverse prompts to see how they align with established sociological values.
The proposed benchmark would incorporate different types of prompts, asking the model to answer as itself, to simulate the view of an average citizen, or to analyze a topic objectively. By testing models across multiple languages and cultural contexts, policymakers could better track whether these systems are favoring certain ideological viewpoints or ignoring the concerns of specific demographic groups.
The team hopes this framework will become a standard tool. “We proposed the Societal AI Alignment Benchmark, SAIA, which would regularly test many models, in many languages, across the core human values measured by the European Social Survey,” Bojić said. “I would love to see something like a weather service for AI opinions, run by national AI agencies under frameworks such as the EU AI Act.”
He added that such a system would “track, week by week, whether the systems people trust are drifting away from the societies they serve. Some of this work is already underway in our research on social bias and temporal stability in language models.”
As with all research, there are a few things to keep in mind when interpreting these findings. The study used a numerical rating scale, which might not capture the full complexity of how people or machines process attitudes toward advanced technology. Open-ended questions or interviews could provide a deeper understanding of these viewpoints, allowing respondents to elaborate on their fears or hopes.
“Please do not read this as ‘AI has feelings about AGI,’” Bojić clarified. “These are outputs shaped by data and design choices, and we are careful to avoid anthropomorphism.”
Additionally, the human participants were mostly from Serbia, which means the human average might not represent a global consensus on AGI. People in other regions might hold different baseline views based on their local economies, media diets, and cultural backgrounds. Expanding the human sample to include more countries would help clarify whether the gap between humans and machines is universal.
“The human samples were also modest and mostly from Serbia, so the numbers are a first snapshot and should not be treated as a global verdict,” Bojić noted. “The study is a proof of concept, and its main contribution is the argument that this kind of measurement should exist and be done continuously.”
Finally, the study measured the language models at a specific point in time. Because these models are regularly updated and fine-tuned by their developers, their expressed sentiments are subject to change. Future research could track these attitudes over a longer period to see if the models’ optimism remains consistent as the technology itself continues to evolve.
For Bojić, tracking these shifts is essential as algorithms increasingly shape human cognition. “We regulate what goes into food and medicine because they enter our bodies,” he said. “AI systems now enter our thinking. I believe it is reasonable to ask for at least a label on the package.”
“We should also care about adapting these models to the cultures and values of different countries, because this is about protecting the diversity of perspectives that has always driven human creativity and progress, rather than resisting universal values,” Bojić concluded. “If billions of people consume the same machine-made opinions, we risk becoming so alike that we lose the very differences that move civilization forward.”
The study, “Towards a societal AI alignment benchmark for evaluating human–machine value convergence,” was authored by Ljubisa Bojic, Dylan Seychell, and Milan Cabarkapa.
Leave a comment
You must be logged in to post a comment.