A new research report argues that the sources grounding Perplexity's software recommendations are dominated by pages built to be read by models rather than people — including three sites that together published 215,128 machine-generated "best software" pages. The report, published Wednesday by Trellner Research, found that nearly 60 percent of the citations behind Perplexity's answers in 380 software categories pointed to domains ranked outside the web's top 100,000 sites. The findings land amid intensifying AI industry coverage of how search engines and answer engines can be quietly steered by engineered content.

What the Study Measured

Trellner put 380 buyer-intent software categories — from "CRM software" to "museum collection management software" — to two of Perplexity's web-grounded models, perplexity/sonar and perplexity/sonar-pro, running through OpenRouter. The test used one prompt per category per model: 760 calls in all, each asking for a ranked top five in JSON with each product's official homepage domain.

Every call returned a parseable answer. That produced 3,800 recommendation slots naming 1,807 distinct products, backed by 7,534 citations across 2,055 distinct domains. Trellner then checked every cited domain against the Tranco list of the world's most-visited websites and the Wayback Machine, and fetched each of the 1,502 vendor homepages the models supplied.

Where the Citations Land

The headline numbers are stark. Of the 7,534 citations, 59.8 percent pointed to domains ranked worse than #100,000 in the Tranco top-1 million list, and 23.4 percent pointed to domains that do not appear in the top million at all. The median Tranco rank among citations that did point at a ranked domain was 71,611.

The most-cited sources were a mix of the familiar and the obscure. G2 led with 291 citations, followed by Reddit with 261. But the third-largest source — cited 194 times, ahead of Gartner's 158 — was guideflow.com, a vendor's marketing blog for interactive product demos that competes in none of the categories it was cited for. Its blog supplied the grounding for answers about 3D rendering software, IVR software, RFID software and architecture practice software alike. Wikipedia, by comparison, was cited three times in 7,534 citations.

Newer, Obscure, and Growing Fast

The report adds a temporal dimension: the unranked domains are also much newer. The median first Wayback Machine capture for unranked cited domains was 2020, against 2011 for ranked ones, and 16.6 percent of archived unranked domains first appeared in 2025 or later — versus 1.6 percent of ranked domains. In other words, a meaningful slice of what grounds Perplexity's answers did not exist until very recently.

215,128 Pages Built to Be Read by Models

The report's most striking claim concerns three sites Trellner says operate under apparently common control. Two of them gave their homepage the HTML title "Facts & Grounding Page" — grounding being the very retrieval step these models perform — and together with a third site they have published 215,128 machine-generated "best [category]" pages. None of the three domains existed before December 2023.

Trellner is careful with the framing. The report notes that nothing about the vendor blog's content marketing is deceptive in itself — companies publish large blogs all the time — and that the measurement is about what the retrieval layer does with it: "a vendor's own listicles about markets it does not operate in became the third-largest evidence base" for purchase recommendations.

Why It Matters for AI Search

The study is a quantitative snapshot of a shift SEO practitioners have been warning about: optimization has moved from ranking in search results to being retrieved by answer engines. Pages written for models — dense, formulaic, and engineered to look authoritative — can become the evidence base for recommendations that consumers treat as neutral. On HN the report drew hundreds of upvotes, with commenters debating whether this reflects a failure of AI search or simply the internet adapting to it.

Important Caveats

Trellner is explicit about scope. Only Perplexity was measured, "and nothing here should be read as a claim about any other engine." Google was left out entirely because grounding a Gemini model through OpenRouter routes it through OpenRouter's own web-search plugin, which would contaminate the measurement. The categories were written before any results were seen and never revised, and the full dataset — prompts, citations and domain lookups — accompanies the report.

Those caveats cut both ways. They make the findings narrower than a blanket indictment of AI search, but they also make them harder to dismiss: within its stated scope, the measurement is straightforward, repeatable, and uncomfortable for anyone who assumed grounded answers rest on the open web's most trusted shelves.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →