For two decades, search marketing trained companies to think about visibility as a ranking problem: identify the query, improve the page and try to move toward position one. Conversational AI is breaking that mental model. A new company benchmark says major AI engines selected the exact same top vendor in fewer than 1.5% of 1,500 commercial prompts, suggesting that the idea of one universal “number-one brand” may make progressively less sense as product discovery moves into answer engines.
The figure comes from SEOPulse and was examined by NetContentSEO on September 3. SEOPulse published the benchmark alongside the launch of its enterprise AI Visibility & Intelligence Platform. The company says its system monitors consumer-facing experiences across services including ChatGPT, Gemini, Perplexity, DeepSeek and ByteDance's Doubao rather than relying solely on model APIs.
The headline number needs an important qualification. This is company-reported research released as part of a product launch, not an independent academic benchmark. The public announcement says 1,500 commercial prompts were analyzed but does not disclose enough information about the category mix, model versions, markets, repetition methodology or prompt distribution to treat 1.5% as a universal law of AI search. The more defensible conclusion is directional: different answer engines can recommend very different brands for the same commercial intent.
Independent research points to the same fragmentation
The exact degree of disagreement changes dramatically depending on how it is measured. That is itself revealing. BrightEdge's AI Catalyst research, for example, found substantial aggregate overlap among the brands surfaced by ChatGPT, Perplexity, Gemini, Google AI Mode and Google AI Overviews, but major differences by industry.
BrightEdge reported pairwise overlap in top brands ranging from 36% to 55% across its overall dataset. When divided by category, however, retail, travel and technology showed much stronger convergence, while healthcare and finance were considerably more fragmented. In BrightEdge's interpretation, AI systems tend to converge more when users are transacting and diverge more when they are researching complex subjects.
A separate 2026 study from FogTrail found another pattern in B2B software. Across repeated waves of 20 questions submitted to five AI engines, agreement on the top recommendation hovered around half of the queries, while unanimous agreement was substantially rarer. Project-management questions were especially fragmented, whereas analytics recommendations showed more consolidation.
These studies are not contradictory because they are measuring different things. Comparing whether engines mention many of the same top-30 brands is very different from asking whether every engine selects the exact same company in first place. The narrower the definition of consensus becomes, the easier it is for agreement to collapse.
AI recommendations are not one ranking system
The underlying reason is architectural. ChatGPT, Gemini, Perplexity and other assistants do not consult one shared index and apply one shared ranking formula. They use different models, retrieval systems, indexes, source-selection rules, freshness mechanisms, product data and interface instructions. Some answers rely heavily on live web retrieval; others can mix retrieval with knowledge already represented in the model.
The prompt also matters more than marketers accustomed to short keywords may expect. “Best CRM” is a broad category question. “Best CRM for a 200-person European SaaS company that needs Salesforce integration, EU data residency and a predictable annual price” describes a substantially different decision. Conversational interfaces encourage users to supply precisely that kind of context.
As prompts become more specific, recommendation markets can fragment into smaller intent clusters. A brand does not merely compete to be “best CRM.” It competes to be a plausible answer for particular company sizes, geographies, integrations, budgets, industries and use cases.
This creates something closer to a portfolio of recommendation markets than a traditional results page. A company can be strong on one engine, weak on another and dominant only for a subset of commercial prompts. Averaging all of that into one AI visibility number can conceal the strategic information marketers actually need.
A mention, a citation and a recommendation are different outcomes
SEOPulse's announcement adds another useful distinction. The company says more than half of the answers in its benchmark mentioned brands, while only about 10% both mentioned and cited a brand. That gap matters because AI visibility has several layers.
A model can mention a company from its existing knowledge without providing a source. It can recommend a company while citing an independent review site. It can cite the company's website as factual evidence without recommending its products. Or it can name a competitor first while using the company's own content to explain the category.
Those outcomes have different commercial value. Citation can generate referral traffic and demonstrate that a source is entering the retrieval layer. Recommendation can influence consideration even without a click. A simple mention may improve awareness but say little about whether the assistant would choose the brand when a buyer asks for a shortlist.
This is why conventional web analytics alone cannot describe AI discovery. Some of the influence occurs inside the generated answer and may never produce a session on the brand's website. Measurement therefore needs to distinguish at least presence, recommendation, citation and the context in which the brand appears.
The source behind the answer may be the real optimization target
When an AI engine repeatedly favors a competitor, the useful question is not merely how to rewrite a product page. It is where the engine is getting the evidence that supports that competitor.
Recommendation systems can draw on product documentation, reviews, specialist publishers, community discussions, comparison sites, news coverage and other third-party material. A company's own website is only one component of that information environment. If trusted external sources consistently describe a rival as the category leader, first-party optimization alone may not overturn that pattern.
This pushes AI-search strategy closer to digital PR, reputation management and entity building. Accurate product information still matters, but so does independent corroboration. Companies need consistent facts across the web, credible expert coverage and sources that clearly associate the brand with the problems it actually solves.
The fragmentation between engines makes source analysis even more important. One assistant may repeatedly retrieve a specialist publication that another barely uses. Another may favor structured product information or fresher sources. Rather than searching for a universal GEO ranking factor, teams can inspect the evidence environment of the specific engine where they are weak.
Real sessions and APIs can produce different answers
SEOPulse is also emphasizing a methodological issue that will matter increasingly for the AI visibility industry: whether monitoring tools test consumer interfaces or model APIs. The two are not necessarily equivalent.
A public AI product can have system instructions, web search, personalization, location signals, shopping integrations and retrieval layers that are not reproduced by sending a prompt to a raw developer API. A benchmark conducted entirely through APIs may therefore measure model behavior without accurately reproducing the recommendation experience a consumer sees.
Real-session testing has disadvantages of its own. Interfaces change frequently, personalization can introduce noise and generative responses vary between runs. That means credible monitoring requires repeated observations rather than a single screenshot. A brand that appears first once and disappears on the next four runs does not have a stable number-one position.
This variability is one reason marketers should be cautious with dashboards that report precise-looking AI visibility percentages without explaining their sampling methodology. Prompt selection, geography, language, repetition frequency and what counts as a recommendation can all materially change the result.
International brands face an even more fragmented map
The problem becomes larger outside the English-language Western market. SEOPulse says its platform includes DeepSeek and Doubao alongside Western AI products and supports multilingual, multi-region monitoring. That reflects an important reality: global AI discovery is unlikely to consolidate around one universal information ecosystem.
Changing language can change the sources retrieved, the brands recognized and the local evidence available to the model. A company that is prominent in English-language publications may have a weak footprint in Italian, Japanese or Chinese sources. The same product category can also contain completely different incumbent brands by geography.
For multinational companies, AI visibility therefore has at least three dimensions: engine, market and language. Add different commercial intents and the measurement problem expands quickly. The answer cannot realistically be to track every possible prompt. Brands need a controlled panel representing the questions that matter most to their customers.
Prompt tracking is becoming the successor to rank tracking
Traditional rank trackers monitor a stable set of keywords over time. The emerging AI equivalent is a stable set of representative prompts. The concept sounds similar, but the design challenge is harder because conversational queries can encode much richer intent.
A useful prompt set should separate discovery questions from product comparisons and high-intent vendor selection. It should represent important customer segments, markets and use cases without becoming so large that the resulting visibility score loses meaning. Teams also need to preserve enough prompts unchanged over time to distinguish real movement from changes in the measurement panel.
Business context matters as much as frequency. Appearing in 80% of low-intent educational answers may be less valuable than being recommended in 30% of prompts used immediately before a purchase. A mature AI visibility metric will eventually need weighting based on commercial importance rather than simply counting appearances.
The new search landscape has many first places
SEOPulse's under-1.5% result should remain attached to SEOPulse's own test set until fuller methodology is available. Independent studies show that consensus can be much higher in certain categories and under different definitions. But together, the evidence points toward a durable change in search strategy.
There may no longer be one meaningful “AI ranking” for a brand. There are positions inside multiple engines, multiple languages, multiple markets and multiple intent clusters. Some of those positions will converge because the market has obvious leaders. Others will remain contested because the systems draw on different information environments.
That fragmentation sounds like a measurement headache, but it also creates opportunity. A company that cannot displace an entrenched leader across conventional Google results may discover that one AI engine already treats it as a credible alternative for a specific use case. The absence of universal consensus means recommendation markets are still open.
The strategic goal is therefore not to chase every generated answer. It is to identify the engines and prompts that matter commercially, measure them repeatedly, understand which sources shape the recommendations and improve the underlying evidence. Search once offered brands a single results page to fight over. AI is replacing that with many simultaneous versions of first place.