Do ChatGPT, Claude, Gemini, Perplexity and Grok recommend the same brands?
Mostly at the top, and much less underneath. Across 22 software categories and 768 recorded answers, all five engines named the same leading brand in 22 of 22 categories. But only about 37% of a typical top-20 was named by all five, and the average pair of engines overlapped on just 64.4% of the brands they named.
The headline numbers
- 22 of 22 categories had a leader that every engine named. Category leadership in AI answers is remarkably settled.
- Only 7.4 brands out of a typical 20-brand leaderboard were named by all five engines — about 37%.
- 64.4% mean overlap between any two engines on which brands they named at all.
- 1,312 distinct brands were named across the 22 categories, an average of 60 per category.
- Only 3.6% of leaderboard places were held by a brand that a single engine named. Total divergence is rare; partial divergence is the norm.
What this actually means
The convenient story for a company that sells multi-engine reports would be that the engines disagree wildly. That is not what the data says, and we are not going to pretend otherwise. If your question is "who leads my category in AI answers", one engine will usually tell you correctly.
The picture changes as soon as the question is about a specific brand rather than the leader. Roughly 63% of every leaderboard is not unanimous, and that unstable region is where almost every brand that is not the category leader actually sits. A brand can look healthy on one engine and be absent from another, which is why our sample report found a single brand ranging from 43% to 86% mention rate across engines on the same day. Both findings are true at once: stable at the top, noisy everywhere else.
By engine
Leaderboard slots is how many of the ranked brand positions across all 22 categories each engine contributed to. Solo picks are brands that only that engine named.
| Engine | Leaderboard slots | Solo picks | Named the category leader |
|---|---|---|---|
| ChatGPT | 329 | 1 | 22 of 22 |
| Claude | 308 | 7 | 22 of 22 |
| Gemini | 340 | 4 | 22 of 22 |
| Perplexity | 322 | 4 | 22 of 22 |
| Grok | 344 | 0 | 22 of 22 |
No engine is a dramatic outlier. The spread between the most and least prolific engine is small, and every engine identified every category leader. Claims that one assistant is systematically "better" at brand recall are not supported by this dataset.
Which engines agree with each other
| Pair | Overlap in brands named |
|---|---|
| ChatGPT vs Grok | 72.1% |
| Gemini vs Grok | 68.5% |
| ChatGPT vs Perplexity | 65.2% |
| Claude vs Grok | 65.1% |
| Perplexity vs Grok | 64.4% |
| ChatGPT vs Gemini | 64% |
| ChatGPT vs Claude | 62.9% |
| Gemini vs Perplexity | 62.7% |
| Claude vs Gemini | 60.4% |
| Claude vs Perplexity | 58.3% |
The most and least settled categories
Categories where few brands were named by all five engines are the ones where AI visibility is still up for grabs. Categories with many unanimous brands have a consensus that is harder to break into.
Most contested: appointment scheduling software (4), time tracking software (4), accounting software for small business (5).
Most settled: online course platforms (13), SEO tools (12), VPN services (10).
The number in brackets is how many brands out of the top 20 were named by all five engines.
Method and limitations
Seven buyer-intent prompts per category, each sent to ChatGPT, Claude, Gemini, Perplexity and Grok through their current fast production models, giving 768 usable answers out of 770 attempted. Brands were extracted from the raw answer text and ranked by how many answers named them. The complete prompt set and scoring rules are on the methodology page.
Limitations worth stating. Language models are stochastic, so a rerun would not reproduce these figures exactly; the effect is largest in the tail, which is precisely where we report the least agreement, so treat the tail numbers as directional. The categories are English-language business software and the findings may not transfer to consumer goods, local services or regulated markets. Leaderboards are capped at the top 20 brands per category, so "brands named by all five engines" is measured within that cap. Every underlying category page is published with its own data and date, so the inputs are inspectable rather than asserted.
This data is free to reuse with attribution. If you cite it, a link to this page is enough.
Frequently asked questions
Do different AI engines recommend the same brands?
Partly. Across 22 software categories measured on August 20, 2026, all five engines named the same top brand in 22 of 22 categories. But only about 37 percent of each category top-20 was named by all five, and the average pair of engines agreed on just 64.4 percent of the brands they named. Engines agree on the leader and diverge underneath it.
Which AI engine recommends the most brands?
Grok filled the most leaderboard slots across the 22 categories, and Claude the fewest. The gap is modest, so no engine is dramatically more or less generous than the others.
Does this mean checking one AI engine is enough?
Only if you are asking who leads a category. If you are asking whether your own brand is recommended, one engine is not enough: the leader is stable but the rest of the list is not, and most brands sit in that unstable part. Our sample report found one brand ranging from 43 percent to 86 percent mention rate depending on the engine, on the same day.
How was this measured?
Seven buyer-intent prompts per category were sent to each of the five engines, giving 768 recorded answers. Every brand named in the raw answer text was extracted and ranked. The prompts, the scoring rules and the known limitations are published in full on the methodology page, and every category page shows its own underlying data.