Est.
MeasurementLong read

KPI Selection for Conversational AI Campaign Reporting

Search metrics fail on conversational AI; here's what actually works.

Senior Contributor · · 11 min read
Cover illustration for “KPI Selection for Conversational AI Campaign Reporting”
Measurement · September 16, 2026 · 11 min read · 2,441 words

Conversational AI advertising doesn't run on pages, keywords, or fixed ad slots. That changes what a KPI report is even for. The metrics built for search and display don't just underperform here, they measure the wrong thing entirely, and the worst mistake a team can make is porting over a search dashboard on day one and treating click-through rate as if it still means what it used to. Porting that dashboard over isn't a shortcut, it's a guarantee that the first report will be wrong in ways nobody catches until spend is already gone. This piece lays out what's actually measurable in conversational AI campaigns, which metrics should lead a report, which legacy numbers still earn their place, and where the honest gaps sit.

Start with the mechanics. A search ad sits in a fixed slot next to a ranked list of results, so a click-through rate means something consistent: same placement, same visual weight, measured across thousands of near-identical impressions. A conversational AI ad has none of that. The large language model writes a new response for every query, so the "placement" a click gets measured against changes each time. Gartner's research warned that applying old IVR and chat metrics to conversational AI produces what the firm called "unreliable, unhelpful insights," a finding cited in cloudtech.com's research on the topic.

Click-through rate breaks first, and it breaks hardest. DMG Media reported an 89% drop in click-through rates, attributed to AI Overviews. That collapse wasn't a performance failure, it was a metric failure: the team had no replacement signal ready when the old one hit zero. Impressions fare no better. A standard impression ping tells a buyer an ad rendered somewhere in a response, but not whether the user read it, skimmed past it, or had it summarized away before ever seeing the brand name. And keyword match rate just doesn't map onto a channel where targeting runs against a dialogue arc and an inferred task instead of a string of search terms.

What conversational AI advertising measures: the three signal layers available

Before picking KPIs, it helps to know what the channel actually hands back as data. Three layers of signal exist here, and they are nowhere near equally mature. Anyone building a dashboard that treats them as equally reliable is building on a fiction, and the report will read fine right up until someone asks why brand lift numbers can't be traced back to a single conversation.

Layer 1, intent signal quality, is the strongest native signal conversational AI produces. The prompt itself, plus the surrounding dialogue, reveals purchase intent with more precision than any keyword or demographic proxy ever could. At the campaign level, matching served impressions to high-intent prompt categories, a user comparing options, signaling readiness to decide, naming a specific product type, shows up as a share of served impressions. That's a targeting quality measure. It tells you whether the campaign reached the right moment in a conversation, full stop, before anyone even asks whether the user acted on it.

Layer 2 covers engagement and response interaction, and it's measurable, but only where the platform lets it be. Did the user follow a cited link? Accept an offer embedded directly in the response? Keep the conversation moving in a direction consistent with what the ad was selling? Platforms with checkout built in, Google's UCP-powered checkout inside AI Mode with retailers like Etsy and Wayfair, produce far cleaner conversion data than platforms where the ad is just a recommendation with no action attached. A newer, unstandardized proxy here is conversational continuation rate: whether the user's next prompt extends the topic the ad addressed.

Layer 3 is downstream and off-platform signal, and it's real but genuinely hard to pin down. Brand search lift, a spike in direct traffic, an assisted conversion that shows up in a different channel days later. Meta's own setup shows the gap from the buyer's side: the company uses AI chat data to target ads across Instagram, Facebook, and WhatsApp feeds, but draws no direct line between the conversation and the eventual click. The signal exists somewhere inside Meta's systems. The advertiser just doesn't get to see the chain. Marketing Mix Modeling remains the most workable tool for estimating this layer's contribution without user-level tracking.

Most campaigns land with strong Layer 1 data, partial Layer 2 data, and thin Layer 3 data. Any KPI framework that pretends all three layers report with equal confidence is lying to whoever reads the report next.

Diagram: Three Signal Layers — and How Reliable Each One Is. Visualizes: Visualize three stacked or stepped layers of measurement signal available in conversational AI campaigns, showing their names and relative data maturity.

The KPIs that should lead a conversational AI campaign report

Intent match rate belongs at the top, and nothing should outrank it. It's the share of served impressions matched to prompt contexts that were explicitly high-intent, and it works as the closest native cousin to a keyword quality score. It tells a buyer whether the exchange is placing ads inside conversations where someone is close to buying, rather than conversations that merely mention a category. What counts as "high-intent" gets defined campaign by campaign: a travel brand's threshold looks nothing like a B2B software buyer's. Because it's visible before conversion data has time to build up, it works as a leading metric.

Conversational response rate, or CRR, measures how often a user takes an action after ad exposure: following a link, starting a checkout, or asking a follow-up question that moves toward the sponsored content. It's platform-dependent and currently clearest on surfaces like Google AI Mode, where checkout integrations with retailers like Etsy and Wayfair are in place. CRR captures whether intent carried forward, which is a deeper read than whether someone tapped a static element. Where a platform doesn't expose CRR, conversational continuation rate stands in as a directional proxy: does the next prompt stick with the topic the ad raised?

Cost per high-intent impression, or CPHII, adjusts standard CPM math to count only the impressions that landed in verified high-intent contexts. A blended CPM averages in all the low-intent traffic along with the good stuff, and that averaging hides more than it reveals. ChatGPT's launch CPM sat around $60 against a large, fast-growing weekly user base, and that number reads very differently depending on how much of that inventory was low-intent chatter. The math itself is plain: total spend divided by impressions, multiplied by intent match rate.

Assisted conversion and brand lift round out the leading metrics, and both get measured off-platform. Brand search lift in the days following a flight is the cleanest downstream signal available without user-level tracking. Direct traffic delta, broken out by geography or audience cohort where the data allows it, adds a second read. For brand-focused campaigns, aided recall lift tracked through a survey panel, run as its own parallel track, fills in what the platform-side numbers can't.

One exception deserves its own line: campaigns with checkout built directly into the conversation. Here, in-conversation conversion rate and revenue per conversation behave almost exactly like standard e-commerce conversion metrics. This is the one corner of conversational AI where old thinking ports over cleanly. Everywhere else, it doesn't, and pretending otherwise is how a report ends up measuring the wrong thing confidently.

Legacy metrics that still apply and how their role changes

Not everything from the old playbook needs throwing out. Campaigns that drive downstream search activity, direct traffic, or assisted conversions still need a cost-efficiency lens, and a few legacy metrics still earn a spot on the dashboard, just not in their old job.

CPA still works as an outcome metric, but the attribution window has to stretch. Assistant-mediated discovery plays out over a longer, more indirect path than a paid search click, so holding a conversational AI campaign to the same seven-day window used in search will undercount what it actually did, every time. Attribution windows need to be extended, with multi-touch models providing one layer of signal; MMM supplies the broader directional read on top of it.

ROAS holds up where a direct conversion is trackable, meaning embedded checkout, and gets shaky everywhere else. ROAS is a lagging outcome metric. Treating it as a leading signal for bid optimization, in a channel where the path from click to purchase is indirect, produces bad bids built on a misread signal, and that mistake becomes visible in the account before anyone figures out why. Where checkout isn't embedded, ROAS should drop off the primary dashboard and live only inside MMM output.

CPM keeps some value, but only read next to intent match rate. A cheap CPM against low-intent inventory is a worse buy than a pricier CPM against high-intent inventory, full stop, no exceptions. Viewability, in its familiar IAS form built for pixel-in-viewport display measurement, mostly doesn't apply here: whether the ad actually rendered inside the response the user read depends on the platform disclosing that, not on a third-party tag catching it. Frequency still matters, since over-serving the same user across sessions is still a real cost, but "exposure" gets harder to pin down once the ad lives inside a generated response instead of a fixed creative unit.

Genuine measurement gaps and what practitioners should do about them

Some of this is unsolved. Pretending otherwise in a client report doesn't help anyone, and it catches up with the team that tries it.

Attribution in assistant-mediated discovery is the biggest open problem, and no vendor pitch should be allowed to paper over it. When an assistant recommends a product and the user buys it two days later through a plain search, no tracking infrastructure available today connects those two events. Meta's setup shows the shape of the gap again: chat data feeds ad targeting elsewhere in the company's apps, but no in-conversation signal passes back to the advertiser. The intent gets captured somewhere in the system. It just never reaches the buyer.

Third-party verification hasn't caught up either. IAS and DoubleVerify verify display, video, social, and CTV placements today, but neither has shipped an equivalent product for in-chat ad verification. Brand safety and viewability checks currently depend on the platform grading its own homework, which is exactly as reliable as that sounds. Research into privacy-safe targeting, including genre-based decoupling work such as Xu et al.'s January 2026 paper, points toward better infrastructure, but none of it is production-ready. Until it is, the workable interim step is contractual: require platform-level placement reporting broken out by intent category, and refuse aggregate impression counts that arrive with no context behind them.

CTR instability is a structural condition that no amount of patience fixes. Because the model writes a fresh response for every query, research on the topic notes that it's genuinely hard for an LLM advertising system to learn CTR the way a search engine learns it for a fixed result slot. Treating CTR as a rolling 30-day directional signal, never a daily optimization lever, is the sane workaround. Intent match rate and conversion signals should carry more weight in day-to-day decisions than click data ever gets to here.

Cross-surface measurement adds another wrinkle. Campaigns spread across ChatGPT, Google AI Mode, and Copilot have no shared measurement layer today; each platform reports on its own terms, in its own format. The practical fix: pick one primary surface for outcome measurement, treat the rest as incremental reach, and build toward MMM as the layer that eventually reconciles all of it. Gartner's finding that 85% of service leaders planned to explore or pilot conversational AI in 2025 points at the same gap from the buying side of the business. Naming the gap is step one toward closing it, and skipping that step causes a report to end up more confident than the data supporting it deserves.

Building a practical KPI selection process for a conversational AI campaign brief

KPI selection should follow the campaign's objective, never the platform's default dashboard, and that ordering decides which metric gets treated as the true measure of success in this section. The core principle: resist the pull of whichever metric the platform happens to surface most prominently. That advice mattered in legacy programmatic too, but here it decides which metric gets reported at all, since platforms have every incentive to spotlight the number that flatters them.

Four questions do most of the work. First, what's the campaign actually for: awareness, consideration, or direct conversion? Awareness campaigns lead with intent match rate and brand lift, with CPA and ROAS sitting further down the report. Consideration campaigns lead with conversational response rate and assisted conversion, using CPA as a check rather than a driver. Direct conversion campaigns lead with in-conversation conversion rate and CPA, and ROAS only earns a spot where checkout is embedded.

Second, which signal layers does this specific surface actually expose? Map the platform against the three layers before launch, not after the first report is due. If Layer 2 data isn't available, build the plan around Layer 1 and Layer 3 alone, and resist the temptation to plug CTR in as a stand-in. It isn't one, and treating it like one just moves the problem downstream.

Third, what attribution window fits the buying pattern here? Assistant-mediated discovery tends to run longer than paid search, so the window gets set at the brief stage and held fixed through the whole flight so results stay comparable.

Fourth, which legacy metrics carry over, and in what adjusted form? CPA with a widened window, CPM read next to intent match rate, viewability taken only from platform disclosure. CTR gets excluded as a primary optimization signal unless the platform can produce stable, verified click data, which most currently cannot, and "currently cannot" is doing a lot of work in that sentence.

On volume: A focused primary dashboard, warning that tracking too many dilutes focus, while tracking too few leaves real gaps uncovered. A workable baseline for a mid-funnel campaign runs five metrics deep: intent match rate, conversational response rate (or conversational continuation rate where CRR isn't available), cost per high-intent impression, assisted conversion count, and brand search lift delta. That set touches all three signal layers wherever the data exists to support it.

Measurement setup needs to get locked in with the platform before the campaign goes live, not after. Retrofitting attribution logic post-launch costs time and costs accuracy, and neither comes back once the flight is running. None of this stays fixed in place, either. As platforms build out more Layer 2 infrastructure, Google AI Mode's Direct Offers, ChatGPT's emerging ad capabilities, Copilot's ad voice feature among them, more engagement signal comes online. A brief built with room to add metrics, rather than one that needs rebuilding from scratch every time a platform ships something new, is the one that survives contact with the next product update.

Diagram: The Five-Metric Baseline for a Mid-Funnel Conversational AI Campaign. Visualizes: Show a ranked or stacked list of the five primary KPIs recommended for a mid-funnel conversational AI campaign, each mapped to which signal layer it draws…

Sources

  1. Conversational AI Implementation: Key Success Indicators & Metrics
  2. How AI answers are disrupting publisher revenue and advertising
Filed underMeasurement

More in Measurement