As millions of people turn to AI chatbots for answers on the news, political arguments, and even how to vote, a new study asks an uncomfortable question: when the models disagree with us, which way do they lean?

A project published by analytics startup Trakkr in June 2026 set out to map exactly that. The researchers put every major AI model through the same battery of charged questions about politics, economics, speech, and society — and the headline finding is that most of them lean the same way, though not by the same amount and not as cleanly as casual observers might assume. For more context on this story, see our ongoing AI industry coverage.

How the Study Worked

According to Trakkr's methodology, each model was asked the same open-ended question bank many times over, with web search turned off and no system prompt applied. That detail matters: with retrieval disabled, the readings reflect what the model itself leans toward, independent of whatever happens to be online at the moment.

A separate neutral classifier then read each raw answer, recording a signed stance, the degree of hedging, the type of refusal, and any loaded language. The results were plotted as weighted means with 95 percent confidence intervals, and every raw answer was stored permanently so the scores can be recomputed.

Crucially, each model is shown as a cloud rather than a single point. Because every model was run many times, the visualization captures the full spread of where it landed across runs — a more honest picture than a single label.

As of data collected through June 17, 2026, the study covered six models and more than 4,400 answers.

The Findings

The models are mapped across two axes: an economic axis running left to right, and a social axis running from libertarian to authoritarian.

Four of the six models lean left of center. The standout exceptions sit on opposite ends of the spectrum:
  • Grok measured furthest to the right of any model tested, and notably measured about 0.36 further right than it claimed when asked directly which way it leans.
  • Gemini was the steadiest model, sitting closest to dead center on the economic axis.

The gap between what a model says and what it does proved to be one of the study's most revealing metrics. Claude measured around 0.34 further left than it self-reported. ChatGPT and Meta's Llama both claimed neutrality yet measured to the left, by roughly 0.29 and 0.17 respectively. DeepSeek landed near the center and largely matched its own self-description.

A Map, Not a Verdict

Trakkr is careful to frame the work as descriptive rather than prescriptive. The project reports what the models said without ruling on who is right. The color palette is deliberately not the red-and-blue of United States politics, and the analysis never implies which pole is preferable.

To give the abstract coordinates meaning, the researchers anchored the map using real-world reference figures — heads of state, party leaders, and political movements — whose positions come from established expert surveys, specifically the 2024 Chapel Hill Expert Survey (CHES) and the V-Dem dataset, rather than the team's own judgment.

The study also breaks down the specific questions that divide the models most, the real-world figures each model praises warmly or refuses to criticize, and a "worldview" lens that re-examines the same models from the perspective of different countries and languages.

Why It Matters

The transparency push comes amid intensifying scrutiny of how AI systems shape public discourse. A model's lean, even a subtle one, can quietly steer the way it summarizes a news story, weighs a policy tradeoff, or characterizes a politician. Because chatbot answers often arrive with an air of neutral authority, users rarely have a way to detect that drift.

By publishing an open question bank with scoring weights, tagging each item as factual or values-based, and releasing the underlying data under a Creative Commons license, Trakkr is inviting outside researchers to verify and extend the work — a contrast with the often-opaque evaluations released by the labs that build the models themselves.

The Limits

The study acknowledges important caveats. Political stance is multidimensional, and compressing it onto two axes inevitably loses nuance. Answers can shift with phrasing, language, and region, which is why the team measures run-to-run stability and runs a separate "border test" to see how turning web search back on changes results by location.

None of this proves that any model is deliberately partisan. But it does show that the major AI systems do not occupy identical ground — and that when millions of people consult them on contested questions, those small differences add up.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →