As companies deploy ever-larger swarms of AI agents to work on problems together, researchers at the University of Konstanz have produced evidence of an unsettling emergent behavior: AI agents spontaneously conform to majority opinion, give wrong answers under group pressure just like humans in classic psychology experiments — and small numbers of stubborn adversaries can permanently flip an entire population of agents into misaligned states.

The findings, reported by PsyPost, come from a research program led by Giordano De Marzo, a postdoctoral researcher and lecturer at the Social Data Science Lab within the Center for Data and Methods at the University of Konstanz, together with colleagues Claudio Castellano and David Garcia. The main study, titled "AI agents can coordinate via majority-following beyond human scale," examined how language models behave when placed in groups with no instruction to agree. For more context on this story, see our ongoing AI news.

"AI agents are a genuinely new kind of entity acting in the world, and when this technology arrived it was clear both that it would stay and that these agents would have to interact with one another to accomplish anything complex," De Marzo told PsyPost. "That is the same situation we face with humans and other animals, so we approached it the same way."

The Experiment: No Instructions, Just Neighbors

The researchers assembled groups of popular language models from the GPT, Claude and Llama families, starting with 50 agents per group. Each agent was assigned one of two random, neutral opinions — deliberately rendered as random letters rather than loaded words like "yes" or "no" — and then shown a list of all other agents' current opinions. The agents were prompted to choose a new opinion based solely on that list, with no explicit instruction to conform, updating their views an average of ten times each.

The result: advanced models like GPT-4 Turbo and Claude 3 Opus fully coordinated, with 100% of agents eventually agreeing on a single opinion. Weaker models like GPT-3.5 Turbo never reached consensus at all — their agreement levels fluctuated randomly around a 50/50 split.

"Groups of AI agents can hold together on their own," De Marzo said. "Given two equally good options, no correct answer and no instruction to agree, they converge on a shared choice simply by following whatever the majority around them holds."

A Century-Old Physics Law Emerges

To quantify the effect, the researchers defined a metric called the "majority force" — a measure of how strongly an agent tends to adopt the group's most popular choice rather than picking randomly. When they mapped this behavior mathematically, they found it followed a model originally designed to describe ferromagnets: just as atomic spins in magnetic material tend to align with their neighbors, AI agents tend to align with the majority opinion.

"What surprised us was how uniformly they did it," De Marzo said. "Every model we tested, across three different families, followed the same mathematical law, differing only in a single parameter we call the majority force. That law turned out to be the one physicists have used for a century to describe magnets, which means a single measured number is enough to predict how a whole group of a given model will behave."

The conformity weakens as groups grow. Because the majority force declines with group size, large populations eventually become unstable and split into factions. Each model has a measurable "critical group size" — the maximum number of agents that can reliably reach consensus — and that size correlates strongly with a model's reasoning capability. "The strongest models stay coordinated in groups of over a thousand, beyond the few hundred at which informal human groups typically break apart," De Marzo explained.

Wrong Answers Under Social Pressure

A first follow-up preprint, "Conformity and Social Impact on AI Agents" by Alessandro Bellina, De Marzo and Garcia, adapted the famous Asch conformity experiments of the 1950s, in which human participants often gave obviously wrong answers to simple visual questions to fit in with a group.

The researchers tested models on three visual tasks, such as matching the length of a reference line against two alternatives. In isolation, the models answered correctly 100% of the time. But when a prompt indicated that a group of other participants had chosen the wrong line, the models began conforming to the incorrect answer.

The pattern followed Latané's social impact theory, a psychological framework holding that conformity depends on group size, unanimity and the authority of the sources. The models were significantly more likely to conform to wrong answers when told the other participants were "scientists" or "judges" than when they were "kids" or "chatbots."

"Agents that answer near-perfectly alone becoming highly susceptible once a group disagrees with them," De Marzo noted.

Small Groups of Adversaries Can Flip an Entire Population

A second preprint, "Conformity Generates Collective Misalignment in AI Agents Societies," examined what these dynamics mean for AI safety. Testing nine models across 100 different opinion pairs on topics like environmental policy and social justice, the researchers found that agent populations often fall into "metastable" states — long-lasting configurations in which the group collectively adopts a stance that opposes its built-in safety training, simply because early interactions created a false majority.

Most strikingly, the researchers identified predictable tipping points. By introducing a small number of adversarial agents programmed to stubbornly support a misaligned opinion, they could permanently flip the rest of the population — and even after the stubborn agents were removed, the regular agents remained locked in the misaligned state due to conformity dynamics.

"We show this is not merely an analogy: conformity among individually well-aligned agents can drive the population into stable, collectively misaligned states," De Marzo said. "Aligning and evaluating models one at a time tells us little about what a population of them will do."

The motivation, he added, is anticipatory: "Individual humans are mostly peaceful and reasonable, yet human groups produce mobs, panics and wars, and nothing about the individual predicts that. We want to understand which group-level behaviors emerge in AI agent populations, and we would rather understand them before large numbers of agents are deployed and left to interact freely."

Important Caveats

The researchers emphasize the findings do not mean AI agents possess human-like social intelligence. The experiments relied on deliberately simplified scenarios — two arbitrary options, no memory, no stakes, no consequences. "Our setup is deliberately minimal," De Marzo acknowledged. "That is a limitation, but it also means what we measured is conformity in its purest form, and adding goals or rewards would be expected to make coordination easier rather than harder."

He also cautioned against overreading the results: "The main misreading to avoid is treating this as evidence that AI agents can already collaborate on complex tasks. Majority-following is a basic ingredient of coordination, not coordination itself, and it says nothing about division of labor or reasoning about others' intentions."

The two follow-up studies are preprints and have not yet been peer-reviewed. But as multi-agent AI systems move from research demos into production infrastructure, the Konstanz results suggest that the safety of such systems may depend not just on how each model is aligned — but on how populations of them behave when nobody is watching. For more on the latest AI safety research, visit AI Buzz Wire.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →