Bob Michaels/ai
The Head-to-Head series · 03 · Bob MichaelsJuly 2026

AI already picked your head-to-head competitors. Do you know them?

  • When AI answers a buyer, it names the comparison set inside the answer. That set is your real competitive field, and it is discovered, never assumed.
  • On AI assistants the discovered field diverges from your search rankings. In one SaaS snapshot study, 81% of the brands ChatGPT recommended did not rank in Google's top ten for the same keyword.
  • One screenshot is not the field. The same question almost never returns the same list twice; what stays stable is how often each name appears.
  • The discovery protocol: discovery, recommendation, and comparison questions, across five platforms, recording who gets named and how often.
  • Check for entity confusion while you are there. AI sometimes blends similarly named companies, and you inherit the other company's reputation.

Every company keeps a competitor list. It lives in the pitch deck, shapes the battle cards, and decides whose pricing page gets watched. Here is the uncomfortable question: who wrote the other list, the one AI systems actually use when a buyer asks who to choose? Because that list exists, it is in production right now, and most companies have never seen it.

The two lists diverge more than people expect, especially on the assistants where buyers ask for picks. In an April 2026 snapshot study of 150 SaaS companies by a marketing agency, 81% of the brands ChatGPT recommended did not rank in Google's top ten for the same keyword. One platform and one month, so hold the digits loosely; the direction is the point. Google's own AI Overviews stay much closer to classic rankings, but the assistants assemble their field their own way, and that field rarely matches the slide from the last board meeting.

The lens

Everything runs through the head-to-head: who does AI recommend, you or your competitor, and why. First measure who is winning. Then change it with the Content is Code methodology: query, decode, engineer, verify.

This is the third piece in the series built on that statement, and the methodology in it is plain: query the platforms with real buying questions, decode why the winning answers win, engineer that evidence on your own domain, verify by retesting. The opener argued the lens: AI answers deliver recommendations, so the metric that matters is who wins the comparison. The second piece showed why the standard content playbook never sees that comparison. This one is about the first thing the query step produces: the field. You cannot win a head-to-head against competitors you have never identified.

The field is discovered, never assumed

When an AI platform answers a buying question, it does the competitor analysis inline: it names a set of options, frames each one, and usually picks. Whoever is in that set is competing for your customer at the moment of decision, whether or not they are in your deck.

Your deck's list is not worthless. It encodes real market knowledge, and it usually overlaps the discovered field. Treat discovery as an audit of the list, never a replacement for market sense; the finding is the difference between the two.

And the discovered set has a structure your intuition will not supply. A June 2026 preprint tracking 100,000+ prompt responses across 100+ brands found first-run brand appearance stacked in tiers by brand gravity, meaning how large and widely known the brand already is: global household names appeared in about 73% of first runs, mid-market brands 44%, and niche brands 11%. A vendor preprint, held loosely, but the shape matches what I see in client work: the field AI assembles mixes obvious rivals with adjacent players, national brands you consider out of category, and sometimes companies you have never heard of. The surprise is the point. If the discovered field matched the deck, discovery would be a formality. It rarely does.

One screenshot is not the field

The biggest mistake people make after hearing this is running one query and treating the answer as the verdict. The evidence says the single answer is nearly worthless. In January 2026, SparkToro and Gumshoe published a study in which 600 volunteers ran 12 buying questions nearly 3,000 times across ChatGPT, Claude, and Google's AI. The questions leaned consumer (headphones, chef's knives, novels), so carry the digits into B2B with care. The odds that two runs returned the same brand list were under 1 in 100. The same list in the same order: roughly 1 in 1,000.

The same study found the signal underneath the noise. However the questions were phrased, the leading brands in a category kept showing up, appearing in 55 to 77% of responses. In one niche, a single firm appeared in 85 of 95 runs. The authors' conclusion matches the discipline I run: appearance frequency across dozens of repeated runs is a meaningful measure; any single-run “AI ranking” is noise. Your field is the set of names that recur.

The discovery protocol

Here is the protocol I use to extract the field. It costs a few hours and no tooling.

Three question types, in real buyer phrasing. Discovery questions: the best options in your category, with the qualifiers a buyer would add (budget, industry, city). Recommendation questions: who should I hire for this specific job. Comparison questions: you against a named alternative, and the alternative against you. For a regional accounting firm the trio might be “best accounting firms for construction companies in central Texas,” “who should a mid-size contractor hire for job costing and audits,” and “compare us against the firm the last answer recommended.”

Five platforms. ChatGPT, Claude, Perplexity, Gemini, Grok. They retrieve differently and they disagree; the disagreement is data.

A three-column record, plus a tally. For every answer: who was recommended, who else was named, what reasons were given. Then count the recurrences. A name that appears in a quarter or more of your runs is in your field. A name that appears once is noise until it recurs.

Run the pass more than once before you conclude anything. The consistency study above ran each question 60 to 100 times per platform; you do not need that scale to start, but treat a dozen runs per question as the floor before you call any name recurring. When you are ready to make it a standing instrument, lock the questions and repeat them on a cadence. The locked series is what turns a discovery exercise into a metric that can move.

The entity confusion trap

While you are recording names, watch for a failure mode that can cost you deals silently: entity confusion. AI systems sometimes blend similarly named companies, and the blend hands you the other company's attributes: their industry, their city, their reviews, sometimes their lawsuit. Picture a regional services firm that shares its name with an out-of-state company in a rougher industry: a buyer asking about the local firm can get the other company's service area, ratings, and headlines woven into the answer as fact. That is the shape of the failure, and nothing on your website ever shows it happening. I take this failure mode seriously enough that the production research engine I architected runs explicit wrong-candidate detection. Similar-name blending shows up in live research runs, and a system that cannot catch it reports the wrong company's record as yours.

The check is cheap. Ask each platform to describe your company, cold. Flag every detail that belongs to someone else, and note which sources the answer leaned on when it went wrong. Blended with a stronger brand, you hold borrowed shine that one correction takes away. Blended with a weaker or troubled one, you bleed trust in conversations you never see. Either way you want to know today.

What to do with the field

Put the discovered field next to the deck and log three things: names AI includes that you do not, names you watch that AI never mentions, and any confusion cases. Give the confusion cases an owner immediately: the fix starts at the sources the wrong answers leaned on, and it is the most time-sensitive finding on the page. That one page replaces an assumption with an observation, and decisions change when it lands.

Then use it. The discovered field is the opponent list for your head-to-head series, and the losing comparisons in that series tell you exactly what evidence to go build. How you decode a winning answer and engineer that evidence on your own domain is the next piece in this series. First measure who is winning. Then change it: query, decode, engineer, verify.

Questions worth asking next

How do I find out which competitors AI names against my business?

Ask the five major platforms three kinds of questions in real buyer phrasing. Discovery questions ask for the best options in your category. Recommendation questions ask who to hire for a specific job. Comparison questions weigh you against a named alternative. Run them on ChatGPT, Claude, Perplexity, Gemini, and Grok, then record every company named and how often it recurs across repeated runs. The names that keep appearing are your field. The names that appear once are noise.

Why does the AI competitor list change every time I ask?

Because the platforms assemble answers fresh each run. A January 2026 study of nearly 3,000 runs across ChatGPT, Claude, and Google's AI, mostly on consumer buying questions, found the odds of two runs returning the same brand list were under 1 in 100. The stable signal sat underneath: leading brands in a category still appeared in 55 to 77% of responses however the question was phrased. Treat any single answer as an anecdote. Measure frequency of naming across repeated, locked questions instead.

What is entity confusion in AI answers?

Entity confusion is when an AI system blends your company with a similarly named one and hands you their attributes: their industry, their location, their reviews, sometimes their failures. It shows up in live research runs often enough that the production research engine I architected carries explicit wrong-candidate detection. Check for it by asking each platform to describe your company and flagging any detail that belongs to someone else.

Sources

  1. Rand Fishkin and Patrick O'Donnell (SparkToro/Gumshoe), "New Research: AIs are highly inconsistent when recommending brands or products," January 28, 2026. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/
  2. Matt Shirley (EMGI Group), "The SaaS AI Citation Gap Report 2026," April 2026. https://emgigroup.com/blog/saas-ai-citation-gap-report/
  3. Pratyush Kumar (Ranqo), "Generative Engine Optimization at Scale," arXiv preprint 2606.20065, June 2026. https://arxiv.org/abs/2606.20065

About the practice behind this guide

I am Bob Michaels, a Web and AI Systems Architect in Austin, Texas. I have built the web since 1994, and today I run AI visibility measurement, complete web presence transformations, and custom AI system builds for organizations that want one accountable owner across all three. The first engagement starts exactly here: a head-to-head assessment that discovers the competitor field the five major AI platforms actually use, and measures who wins it.

Evaluating me for an AI leadership role instead? The work record is here.

← All writingJuly 7, 2026 · 8 min read