Bob Michaels/ai
The Head-to-Head series · 04 · Bob MichaelsJuly 2026

How you win the AI head-to-head

  • A losing head-to-head is a finding, and the response is a loop: query, decode, engineer, verify.
  • Decode means reading the winning answers for the themes, evidence, and sources they lean on. That is the brief for what you build next.
  • Engineer on your own domain first. In one large vendor study, corporate websites took about 78% of the citations in AI answers.
  • Consolidate one message across your domain, your social presence, and your machine layer, so the comparison engine reads one sharp story.
  • Verify on the locked question series and hold the results honestly. One program I can name moved from a 6.2% to a 36.2% win rate in eleven weeks. Another program stayed flat. Both belong in the record.

By this point in the series you may have run the measurement and found the answer that stings: when AI compares you against the competitors it names, you lose. Good. A measured loss beats an invisible one, because a losing comparison is a finding, and findings have a method. The needle is movable: academic benchmark work on generative engine optimization, published at KDD 2024, found content optimization alone improved visibility in AI-generated answers by up to 40%, with the effect varying by domain. Qualified, benchmarked, and enough to retire the idea that AI answers are weather.

I have also watched the needle move in the field. One client program went from a 6.2% to a 36.2% AI recommendation win rate in eleven weeks. Directional evidence, and the full record with its limits is below. This closing piece of the arc is the method behind it: the four moves I run, in order, and the honesty that has to travel with them.

The lens

Everything runs through the head-to-head: who does AI recommend, you or your competitor, and why. First measure who is winning. Then change it with the Content is Code methodology: query, decode, engineer, verify.

The lens says the recommendation is the metric. The scoreboard says coverage alone never moves it. The discovery protocol hands you the competitor field. This piece is what you do with a losing comparison, and the four moves are plain English: ask the platforms real buying questions, study why the winning answers win, build the missing evidence on your own site, and ask the same questions again to see what changed.

Move one: query on a locked list

The loop starts where the measurement series already lives: a locked set of real buying questions, the same questions repeated unchanged on a cadence, across ChatGPT, Claude, Perplexity, Gemini, and Grok. The losing comparisons in that series are your work queue. Nothing else in this method makes sense without the lock, because you cannot measure movement against questions that keep changing.

Move two: decode why the winner wins

Read every winning answer like a brief, because that is what it is. Three questions per answer. What themes does the platform repeat when it praises the winner: speed, compliance, specialization, price, proof? What evidence does it cite: case pages, reviews, documentation, third-party coverage? And where does that evidence live: the winner's own domain, a best-of list, a review platform?

The reasons are the movable part. A June 2026 preprint tracking 100,000+ prompt responses found the framing of a brand flips about 6.7 times more often than whether the brand appears at all. Held loosely as a vendor preprint, the implication matches my client work: you are rarely fighting to exist in the answer. You are fighting over what the answer says, and the sources it says it from. Decode those sources and you know exactly what your domain is missing.

Move three: engineer the evidence on your domain

Your own website is where the leverage concentrates. The same preprint found corporate websites take about 78% of the citations in AI answers. Third-party surfaces matter, best-of lists alone drew roughly a fifth of citations, but the surface AI cites most is the one you already own and can change this week.

Engineering means building the evidence the winning answers lean on, on your own pages. Liftable language is the standard: what you do, for whom, with what proof, stated so plainly a machine can quote it whole. Google's guidance asks for unique, non-commodity content, and that is what a decoded brief produces: pages that answer the specific comparison, carrying the proof only you can publish. This is Content is Code in the literal sense. Content is the material AI reads when it decides what to recommend, so you engineer it with the same discipline as software: specified by the decode, built, shipped, tested by the retest.

Consolidate: one message everywhere a machine reads

The engineering move has a companion rule: one story. Your domain, your social presence, and your machine layer (the metadata, structured data, and machine-facing files AI systems parse) should carry the same claims in the same language. A scattered story averages out in a comparison while your competitor's stays sharp. The correlation evidence points the same way. In a May 2025 study of 75,000 brands, branded web mentions were the strongest measured factor in AI Overview visibility, at a moderate 0.664, three times backlinks. The authors stress correlation rather than causation. Broad, consistent presence of one message is associated with showing up. The machine layer makes that message discoverable and callable. It earns nothing on its own: Google needs no special AI file, and evidence that does not exist cannot be marked up.

Move four: verify on the locked series

Rerun the locked questions. Record the same three columns: who was recommended, who was named, what reasons were given. The delta against your baseline is the result, and it is directional by nature, because the market keeps publishing while you do. Platforms change retrieval, competitors ship pages, and a single retest is one more point estimate. The trend across retests is the signal.

What the needle looks like when it moves

The record I can publish. Soapbox Bulletin, a client I can name, moved from a 6.2% to a 36.2% AI recommendation win rate over eleven weeks of exactly this loop. The mechanism was nothing exotic: measure the losing comparisons, publish the evidence the winning answers were citing from elsewhere, retest the same questions, repeat. Win rate here means the share of locked head-to-head questions the platforms answered in the client's favor. Directional evidence, not a controlled study, and I report it that way every time. A financial services firm I keep anonymous took four out of four head-to-head wins in a strict URL-grounded series against named competitors.

And the honesty that keeps the record trustworthy: on another program I published into the gaps and the win rate stayed flat. Nothing holds the rest of the market still while the benchmark runs. That flat program is the reason the method retests instead of declaring victory, and it is the reason no one should promise you guaranteed recommendation gains. Evidence moves the odds. The full write-ups, limits included, live on the case studies page.

The loop is the work

Set expectations by brand gravity. A niche firm will not out-mention a household name across the open web, and it does not need to. It needs to win the specific comparisons real buyers ask, and those are won with specific evidence. Run the loop on a cadence, feed the 30-day work plan with what the decode finds, and hold every result to the directional standard this series has used throughout.

First measure who is winning. Then change it: query, decode, engineer, verify. That is the whole method, and the whole series.

Questions worth asking next

How long does it take to change which company AI recommends?

Plan in quarters, measure in weeks, and treat any promise of guaranteed gains as a red flag, because nothing holds the rest of the market still while you work. The measured record I can report: one named program moved from a 6.2% to a 36.2% AI recommendation win rate over eleven weeks of the loop, as directional evidence rather than a controlled study. On another program I published into the gaps and the win rate stayed flat. That flat result is exactly why the method retests instead of declaring victory.

What content actually moves AI recommendations?

Evidence that answers the comparison, placed where AI reads it. Decode the winning answers in your category and you get the brief: the themes, proof points, and sources the platforms lean on. Then build concrete pages on your own domain that state what you do, for whom, with what proof, in plain liftable language. Academic benchmark work found content optimization alone moved visibility in generative engine responses by up to 40%, varying by domain. Generic coverage is fuel; comparison-answering evidence is the payload.

Do llms.txt or schema markup make AI recommend you?

No. The machine layer, meaning metadata, structured data, machine-facing files, and a callable endpoint, makes your evidence easier for machines to read and act on, and that is worth doing well. Google states plainly that no special AI file is required to appear in its AI features, which run on the same core ranking systems as search. Treat the machine layer as access infrastructure for evidence that has to exist first, never as a shortcut around building it.

Sources

  1. Pranjal Aggarwal et al., "GEO: Generative Engine Optimization," KDD 2024, arXiv 2311.09735. https://arxiv.org/abs/2311.09735
  2. Pratyush Kumar (Ranqo), "Generative Engine Optimization at Scale," arXiv preprint 2606.20065, June 2026. https://arxiv.org/abs/2606.20065
  3. Louise Linehan and Xibeijia Guan (Ahrefs), "An Analysis of AI Overview Brand Visibility Factors (75K Brands Studied)," May 26, 2025. https://ahrefs.com/blog/ai-overview-brand-correlation/
  4. Google Search Central, "AI features and your website". https://developers.google.com/search/docs/fundamentals/ai-optimization-guide

About the practice behind this guide

I am Bob Michaels, a Web and AI Systems Architect in Austin, Texas. I have built the web since 1994, and today I run AI visibility measurement, complete web presence transformations, and custom AI system builds for organizations that want one accountable owner across all three. The loop in this article is the engagement: a head-to-head assessment first, then the decode, the engineering, and the retest, run with you until the trend is yours.

Evaluating me for an AI leadership role instead? The work record is here.

← All writingJuly 9, 2026 · 8 min read