Est.
MeasurementLong read

Reporting the Measurement Gap Honestly to Clients

Agencies must tell clients that conversational AI advertising measurement simply doesn't work yet.

Senior Contributor · · 9 min read
Cover illustration for “Reporting the Measurement Gap Honestly to Clients”
Measurement · September 26, 2026 · 9 min read · 1,947 words

Money is moving into conversational AI advertising faster than anyone has built the tools to measure what it's doing. eMarketer projects AI chatbot ad spending will jump 1,641% in 2026, reaching $0.96 billion for pure chatbot inventory, inside a total US AI ad spending figure on track to hit $68.25 billion by 2030. OpenAI's ad revenue went from zero to a nine-figure annualized run rate in six weeks. That speed is the whole story: budgets are flowing into chat interfaces faster than the infrastructure meant to account for them, and agencies still pitching this channel with search-era measurement logic are already behind.

An ad product does not go from nothing to nine figures in an annualized run rate inside six weeks on novelty spend or test budgets. That kind of ramp only happens when serious advertisers decide, immediately, that the channel is worth real money. Demand arrived before the measurement stack was ready for it, and that mismatch, not the growth number itself, is what agencies need to be building their client conversations around.

Calibrate the headline figure before getting excited about it, though. More than 80% of total AI ad spending in 2026 still shows up adjacent to AI-generated content, things like AI Overviews sitting above traditional search results, rather than inside pure chatbot conversations. So when a client says "AI advertising," find out which part they mean first. The inventory, the pricing, and the measurement all behave differently depending on the answer, and the pure-chat slice, the one eMarketer flags for that 1,641% jump, is smaller and newer than the headline number suggests. It's also the slice with the least measurement precedent behind it.

The structural differences between conversational AI advertising and search and social

Search advertising targets a query. Social advertising targets a profile. Conversational AI advertising targets a moment inside a live exchange, and that's a genuinely different unit of account, not a variation on the other two. There's no persistent audience segment sitting the way there is in a social platform's ad set. Advertisers describe the kinds of conversations their product belongs in, and the system matches placement against the topic of an active dialogue rather than a stored interest tag or a keyword bid.

Placement timing changes what a click even means. In a chat interface, the sponsored card typically appears after the model has already answered the question, so the user already has a recommendation before any brand shows up. On a search results page, the sponsored links sit above the organic ones, competing for attention before the user has an answer. A click on a chat ad happens in a different psychological moment than a click on a search ad. Treating the two as comparable units is a category error, and it's the error most agencies are currently making by default.

Ad load only widens that gap. Google Search can surface up to seven sponsored results on a single page. LLM interfaces typically show one. Each placement in a chat surface carries far more weight per impression. The click-through benchmarks built over two decades of search advertising simply don't map across. ChatGPT ad CTR runs around 0.9%, against roughly 6.4% for Google Search, a gap that reflects a fundamentally different ad unit rather than underperformance. That's a chatbot doing a different job under a different attention structure, and any agency handing a client the 0.9% figure without that context is setting them up to read it as failure. It's a different unit doing a different job under a different attention structure, and any agency handing a client the 0.9% figure without that context is setting them up to read it as failure.

The specific things that genuinely cannot be measured reliably yet

Answer engines, Google's AI Mode, ChatGPT, Perplexity, are shaping purchase decisions without sending a trackable click back to a brand's website. Industry analysts call this the AI attribution gap, and as of mid-2026, there is no direct AI citation attribution data available in Google Search Console or GA4. Google has taken some steps toward surface-level reporting, but cross-platform citation tracking still doesn't exist. Google has taken early steps, labeling certain "preferred sources," but nothing in that chain is measurable end-to-end yet. An agency that tells a client otherwise is either misinformed or betting the client won't ask a follow-up question.

Platform-reported metrics move under everyone's feet, too. A widely cited guide claimed the ChatGPT Ads dashboard exposed only two metrics, while OpenAI's own documentation at the same moment listed seven, since the Help Center had already updated by mid-July. The two-metric claim was stale before it reached print. That's the operating condition of this whole category right now: secondary coverage goes out of date faster than it gets read, so checking against primary documentation is a weekly habit. It's a weekly habit, and agencies that skip it end up repeating claims that were already wrong when they were written.

There's no baseline to lean on, either. OpenAI's own position is that ChatGPT Ads has no cross-advertiser benchmarks, not by industry, not by campaign objective, not by format. So when a client asks whether a given CTR is good for their category, the honest answer in 2026 is that there's no industry number to check it against. Agencies that invent one to sound authoritative are manufacturing comfort that collapses the moment performance dips.

The asymmetry between large and small clients

The attribution gap doesn't land evenly, and pretending it does is its own kind of dishonesty. A large enterprise client usually has a data science team positioned to catch what GA4 misses, cross-referencing platform data against internal warehouses and building proxy models of its own. An SMB owner pulling up GA4 every Monday morning has none of that. That owner is making real budget calls off a dashboard that's now systematically incomplete, in the same direction, week after week, and has no internal team to flag it.

The confidence numbers show how precarious this is even before you split by company size. Only 30% of CMOs report confidence in their ability to measure marketing ROI, yet 64% base future budget decisions on past ROI performance anyway. That's a direct contradiction sitting inside most marketing organizations: distrust the measurement, keep using it to decide where the next dollar goes. The trend is moving the wrong way, too. Only 49% of marketers could demonstrate ROI on their AI investments in the prior year, and confidence has continued to lag as spend rises. Confidence is falling as spend rises, which is close to the opposite of what should happen if the measurement tools were catching up to the money.

That asymmetry should drive how reporting gets built. An enterprise client with in-house analytics can absorb a framework with real methodological nuance built in. An SMB client needs something smaller: fewer metrics, clearly labeled, with the caveats sitting directly next to the number rather than buried in an appendix nobody opens.

Measuring and tracking in 2026: an honest framework alongside the gaps

None of this means throwing the report out. Platform-reported clicks, impressions, CPM, and CPC are real and worth tracking, but they aren't a closed attribution loop, so present them as leading indicators, not proof of outcome. That distinction matters more in a client meeting than almost any other sentence in this piece.

Cost benchmarks are usable even without a full conversion chain behind them. Starting CPC bids run roughly $3 to a few dollars more for ecommerce, and considerably higher for software and finance verticals, which gives clients a real budget anchor before a single conversion number comes in. Landing page conversion rates for traffic arriving from ChatGPT ads are measurable at the destination even when the upstream source signal is incomplete, running roughly 4% to 7% for commercial verticals, with top performers clearing 8%.

The most complete public reference point as of mid-2026 comes from a single $60,000 account, Opascope, which returned $89,000, a 1.49x blended ROAS. Cite that to a client as the best available data point. One account is a sample size, not a category norm, and presenting it as anything more sets up a bad conversation in the next quarterly review.

Structuring the client conversation while acknowledging gaps without losing confidence

Credibility breaks in both directions here. Papering over the gap with invented benchmarks feels good in the pitch meeting and turns into a liability the moment performance doesn't match the story. Over-qualifying every metric with hedges and disclaimers reads as an agency that doesn't actually understand its own product. Neither posture survives a second quarter of client scrutiny.

The workable middle ground has three parts, said out loud, every time: name the gap, name the proxy standing in for it, and name the timeline over which real signal should show up. Something like this can't be attributed to a closed conversion loop yet, so here's the leading indicator being tracked in its place, and here's the window over which that indicator should start moving. That sentence does more for a client relationship than a dashboard full of numbers no one can fully stand behind.

Framing the gap as an industry condition, rather than an agency shortfall, helps too. Measurement teams across the industry spend more time stitching together siloed data than generating insight from it. The fragmentation a client feels is an industry condition, not a failure specific to whoever runs their account. It's the state of the market, and naming that shifts the tone of the conversation from defensive to informative.

Channel conditions can shift fast, and that belongs in the conversation as well. Perplexity pulled back from advertising in February 2026, citing user-trust concerns. Contract terms and campaign pacing should account for the possibility that a platform's ad posture changes on short notice, because in this category, it already has.

Where the measurement infrastructure is headed

AI is entering marketing attribution from two directions at once, and both are still under construction. Machine learning models are refining how credit gets assigned across touchpoints inside existing attribution platforms. Separately, conversational activity is a new source of buyer behavior that those same attribution systems were never built to see. Both are expanding through 2026, on parallel but distinct tracks, and conflating them is a mistake agencies keep making when they pitch "AI measurement" as a single feature rather than two separate problems.

AI chat is landing on a measurement ecosystem that already has enormous coverage gaps. It's landing on one that already has enormous coverage gaps. Seventy-seven percent of industry respondents say gaming is underrepresented in existing attribution models, 50% say the same about commerce media, and conversational AI is simply the newest blind spot in a system that already had several before chat interfaces existed.

Awareness of where this is heading is close to universal, even where confidence hasn't caught up. Ninety-six percent of buyers say they're aware of agentic AI being used for ad buying and campaign execution, so the knowledge gap has effectively closed. What's left is a confidence gap and an infrastructure gap, and those close on a slower timeline than awareness does.

There are early signs the infrastructure layer is getting built rather than just discussed. Verve Group became the first open-market advertising platform to operationalize high-fidelity intent data pulled from AI chat interfaces for programmatic activation, a concrete step toward turning conversational signal into something a media buyer can actually transact against. That's a sign the plumbing is under construction, not a sign it's finished. Agencies that build client reporting around what's provably true today, while staying ready to adopt better infrastructure as it lands, will be the ones still trusted when measurement finally catches up to the spend.

Sources

  1. The Rise of LLM-Powered Ads: How AI is Redefining Personalized Marketing | by Keevan Store | StartupInsider | Medium
  2. Verve Group launches industry-first targeting capability activating conversational intent signals from major LLM environments
  3. emarketer.com
  4. emarketer.com
Filed underMeasurement

More in Measurement