Bob Michaels/ai
Case studiesMeasured client outcomes

I built the systems that got these outcomes.

Three engagements, three numbers, and the system behind each one. Trinzik delivered the AI engagements on the platform I architected and built. Two clients are anonymized by sector, and the one named client is already public.

How I hold a number

Every number here is client-verified, and I will walk any of them line by line. Where the evidence is directional rather than causal, the page says so on the same screen as the number.

100% of measured engagements renewed. Zero declined.

The renewal record is context, not a headline. It comes from measured engagements only, and it is stated here rather than at the top of the page for that reason.

6x

AI recommendation win rate

6.2% to 36.2% in 11 weeks.

6.2%36.2%11 weeks

Soapbox Bulletin

The system that produced it

I built the five-platform Research Engine that measured it and the Content is Code editorial loop that acted on it.

Soapbox Bulletin was losing recommendations it never saw. Buyers were asking ChatGPT, Claude, Perplexity, Gemini, and Grok who to use, the models were naming somebody else, and nothing in the existing search reporting could say why.

I built the measurement first. A locked prompt library ran the category's real buying questions across all five platforms on a fixed cadence, with official-domain verification so a same-named company could not be counted as a win, forced reasoning so the model had to say why it chose what it chose, and normalized records so week 11 could be compared to week 1 without arguing about the method.

Then I ran the editorial loop against what the measurement exposed. The models were rewarding evidence the brand had never published. We published it, on the brand's own domain, under human approval, and retested against the locked benchmark rather than against a feeling.

The win rate moved from 6.2% to 36.2% over 11 weeks. Trinzik delivered the engagement on the platform I built.

You get

A defensible read of why the models prefer a competitor, and a publishing loop that goes after it and retests.

What it does not prove

This is directional evidence, not a controlled study. Nothing holds the rest of the market still while the benchmark runs. On another program I published into the gaps and the win rate stayed flat, which is the reason I retest instead of declaring victory.

4

Head-to-head wins, strict URL-grounded series

Four out of four, against named competitors.

04strict URL-grounded

Financial services firm, anonymized

The system that produced it

The Research Engine ran the series on locked prompts with official-domain verification, so every win traces to a cited source.

A financial services firm wanted to know how it stood against a named competitor field when a buyer asked an AI system to compare them. Not a share-of-voice score. The actual head-to-head question a real buyer asks.

The series ran on locked prompts, one matchup at a time, under the strictest read the engine has. A finding only counted when the URL the model cited appeared in the set of URLs the provider actually retrieved. Ungrounded answers were dropped and kept in the record as dropped evidence rather than quietly promoted into a win.

The firm took four out of four. Every one of them traces back to a cited source that can be opened and read.

You get

A matchup-by-matchup record of how you land against the competitors you actually lose to, with the citation behind every result.

What it does not prove

Four matchups is a small series, and a head-to-head read is a point-in-time measurement. It says what the models did on those questions, on those dates, under that gate. Rerunning it is part of the work, not a favor.

27x

Peak ad unit against the display benchmark

1.47% average CTR across 2.14M impressions, versus a 0.05% to 0.10% display benchmark.

0.10% benchmark1.47%2.14M impressions

Texas public university, anonymized

The system that produced it

Display units I designed and built, scored against the standard display benchmark instead of against themselves.

A Texas public university needed enrollment campaigns to perform in a channel where almost nothing performs. Display advertising averages a tenth of a percent click-through, and most reporting hides that by scoring a campaign against its own prior campaign.

I designed and built the units, and I scored them against the published display benchmark instead of against themselves. The peak unit ran 1.47% average click-through across 2.14 million impressions, roughly 27 times the top of the standard benchmark range.

The number holds up because the denominator is honest. 2.14 million impressions is enough volume that the rate is not an artifact of a small sample, and the comparison is to the industry floor everyone else is standing on.

You get

Work that is measured against the outside benchmark rather than against its own last run.

What it does not prove

This is a creative and placement result on one campaign family at one institution. Click-through is not enrollment, and I do not claim the pipeline behind it.

Every number here survives an audit.

Claim posture

Named only when already public

Soapbox Bulletin is named because the work is public. Every other client is described by sector and nothing more, including in conversation.

Attribution stays precise

I built or architected the systems. Trinzik delivered the client engagements on that platform. I do not take credit for the delivery, and I do not hide that I built the thing underneath it.

The limit runs with the number

Directional evidence is labeled directional. A small series is labeled small. A number with an honest denominator survives an interview; one without it does not.

Want this read run on your category?

The first engagement is a head-to-head assessment: how ChatGPT, Claude, Perplexity, Gemini, and Grok describe, cite, compare, and recommend you against a real competitor field, on locked prompts with the citation gate on.

Named engagements, architecture detail, and references are available under NDA. Tell me what you are evaluating and I will share what is relevant.