AI Visibility Snapshot
Get my snapshot — $19

Do ChatGPT, Claude, Gemini, Perplexity and Grok recommend the same brands?

Mostly at the top, and much less underneath. Across 22 software categories and 768 recorded answers, all five engines named the same leading brand in 22 of 22 categories. But only about 37% of a typical top-20 was named by all five, and the average pair of engines overlapped on just 64.4% of the brands they named.

Original measurement · 154 prompts × 5 engines · 768 answers · Recorded · Free to reuse with attribution (CC BY 4.0)

The headline numbers

What this actually means

The convenient story for a company that sells multi-engine reports would be that the engines disagree wildly. That is not what the data says, and we are not going to pretend otherwise. If your question is "who leads my category in AI answers", one engine will usually tell you correctly.

The picture changes as soon as the question is about a specific brand rather than the leader. Roughly 63% of every leaderboard is not unanimous, and that unstable region is where almost every brand that is not the category leader actually sits. A brand can look healthy on one engine and be absent from another, which is why our sample report found a single brand ranging from 43% to 86% mention rate across engines on the same day. Both findings are true at once: stable at the top, noisy everywhere else.

By engine

Leaderboard slots is how many of the ranked brand positions across all 22 categories each engine contributed to. Solo picks are brands that only that engine named.

EngineLeaderboard slotsSolo picksNamed the category leader
ChatGPT329122 of 22
Claude308722 of 22
Gemini340422 of 22
Perplexity322422 of 22
Grok344022 of 22

No engine is a dramatic outlier. The spread between the most and least prolific engine is small, and every engine identified every category leader. Claims that one assistant is systematically "better" at brand recall are not supported by this dataset.

Which engines agree with each other

PairOverlap in brands named
ChatGPT vs Grok72.1%
Gemini vs Grok68.5%
ChatGPT vs Perplexity65.2%
Claude vs Grok65.1%
Perplexity vs Grok64.4%
ChatGPT vs Gemini64%
ChatGPT vs Claude62.9%
Gemini vs Perplexity62.7%
Claude vs Gemini60.4%
Claude vs Perplexity58.3%

The most and least settled categories

Categories where few brands were named by all five engines are the ones where AI visibility is still up for grabs. Categories with many unanimous brands have a consensus that is harder to break into.

Most contested: appointment scheduling software (4), time tracking software (4), accounting software for small business (5).

Most settled: online course platforms (13), SEO tools (12), VPN services (10).

The number in brackets is how many brands out of the top 20 were named by all five engines.

Method and limitations

Seven buyer-intent prompts per category, each sent to ChatGPT, Claude, Gemini, Perplexity and Grok through their current fast production models, giving 768 usable answers out of 770 attempted. Brands were extracted from the raw answer text and ranked by how many answers named them. The complete prompt set and scoring rules are on the methodology page.

Limitations worth stating. Language models are stochastic, so a rerun would not reproduce these figures exactly; the effect is largest in the tail, which is precisely where we report the least agreement, so treat the tail numbers as directional. The categories are English-language business software and the findings may not transfer to consumer goods, local services or regulated markets. Leaderboards are capped at the top 20 brands per category, so "brands named by all five engines" is measured within that cap. Every underlying category page is published with its own data and date, so the inputs are inspectable rather than asserted.

This data is free to reuse with attribution. If you cite it, a link to this page is enough.

Frequently asked questions

Do different AI engines recommend the same brands?

Partly. Across 22 software categories measured on August 20, 2026, all five engines named the same top brand in 22 of 22 categories. But only about 37 percent of each category top-20 was named by all five, and the average pair of engines agreed on just 64.4 percent of the brands they named. Engines agree on the leader and diverge underneath it.

Which AI engine recommends the most brands?

Grok filled the most leaderboard slots across the 22 categories, and Claude the fewest. The gap is modest, so no engine is dramatically more or less generous than the others.

Does this mean checking one AI engine is enough?

Only if you are asking who leads a category. If you are asking whether your own brand is recommended, one engine is not enough: the leader is stable but the rest of the list is not, and most brands sit in that unstable part. Our sample report found one brand ranging from 43 percent to 86 percent mention rate depending on the engine, on the same day.

How was this measured?

Seven buyer-intent prompts per category were sent to each of the five engines, giving 768 recorded answers. Every brand named in the raw answer text was extracted and ranked. The prompts, the scoring rules and the known limitations are published in full on the methodology page, and every category page shows its own underlying data.