Intent-Relevance Scoring as a Campaign Quality Signal
Conversational AI demands a new quality metric that reads intent from prompts instead of keywords.

Legacy quality scoring was built for a specific shape of query, one with a few words, a discrete keyword, and a discrete ad unit sitting next to it on a page. Conversational AI breaks that shape entirely, and the industry's answer is intent-relevance scoring, a signal that measures how closely an ad matches the live intent inside a user's prompt rather than how closely it matches a keyword. This piece explains what that scoring actually does, how it works inside the ad stack, and why it's replacing click-through rate as the metric that matters most.
Google Quality Score is the clearest example of a system built for that older shape. Accounts with high Quality Scores pay notably lower CPCs than the industry median; accounts with low scores pay significantly more. That gap exists because the input is narrow and legible, consisting of a keyword group, a few related terms, and copy written specifically for the intent packed into those words. Responsive search ads reward exactly this kind of tight match, but the system only functions because a search query has already compressed intent down to three or four words.
A ChatGPT prompt doesn't compress anything. It runs long, sometimes a full paragraph, and it can carry a problem, a budget, a use case, and a mood all at once. A keyword match score has nothing to parse there. It's not a matter of tuning the model harder or adding more keyword variants; the mismatch is structural. Click-through rate breaks down too, for a related but separate reason: once an assistant can research, compare, and recommend inside the conversation itself, a lot of the real influence on a purchase decision happens before any click, so last-click attribution simply doesn't see it.
What conversational intent looks like at the prompt level
The scale of this shift is hard to overstate. ChatGPT became the fastest mobile app ever to reach 1 billion monthly active users, and Google Gemini has climbed to 650 million monthly users. That's an unprecedented volume of intent signal now flowing through chat interfaces rather than search boxes. Roughly 20% of ChatGPT queries already carry direct commercial intent, according to reported OpenAI figures.
What makes that 20% different from search intent isn't just volume, it's density. A user who tells an assistant what they want to buy, why they want it, what's stopping them, and what they're willing to spend, all inside one message, is handing over a different class of signal than a three-word query ever could. The markers that matter most inside that message are fairly specific: a named brand and model, price or deal language, action verbs like "buy" or "order," and constraints such as a budget or a stated use case. Those are the features any scoring system has to catch.
Multi-turn conversation adds a distinct layer of complexity. What used to take several searches and a retargeting cookie to reconstruct, concern, openness, comparison, readiness to buy, now unfolds in a single thread, turn by turn. A platform reading only a keyword sees none of that progression. Verve Group's analysis of more than 1 billion daily signals found that over 20% of users now begin their pre-purchase journey inside an AI chat rather than a search bar; for travel queries specifically, that figure climbs to 37%.
What intent-relevance scoring measures
Intent-relevance scoring is a real-time signal that measures how closely an ad's content, category, and offer line up with the live intent expressed in a prompt. It operates at the level of the conversation. Think of it as doing for chat what Quality Score does for search: compressing a messy, human-shaped input into something a bid system can act on. The comparison helps, but it shouldn't be pushed too far, because the input here is richer and the math behind it is considerably harder.
It also helps to state what the score is not. It isn't a relevance check applied to creative after the auction closes. It isn't a brand safety filter. And it isn't a click-probability model, even though it feeds into bid price. Research published in 2026 (arXiv 2607.27686), which tried to build a synthetic click-intent signal for LLM-native ads using a shared evaluator trained with an ordinal score-distance loss, produced a smooth expected-intent output rather than a binary pass or fail, modeling intent as a gradient. That's the direction the field is heading.
Peer review of that same paper flagged a real problem, though: the gap between what the model measures and true conversational fit creates a signal-to-noise issue that the research itself flags as an open problem. The scoring methodology is promising. It is not settled science.
There's a privacy dimension too. Verve's integration of LLM intent signals, announced in March 2026, runs on aggregated, high-level metrics and pseudonymized behavioral patterns; no raw message content and no individual conversation data gets disclosed or stored. Any credible scoring system operating in this space has to work inside that same constraint.
How scoring works mechanically inside the LLM ad stack
The LLM ad stack has four layers: demand and auction, context and targeting, creative generation, and measurement and attribution. Intent-relevance scoring sits right at the seam between the first two. At the context-and-targeting layer, the platform (or an intermediary) reads the conversation itself, its topic, the need being expressed, the constraints named out loud, and how deep the thread already runs, then maps that against available ad candidates before the auction even clears.
The inputs feeding that score break down into a few types: a blunt topic category at the surface level, sharper intent markers that this surface category produces (brand names, price language, action verbs, use-case constraints), and turn structure, meaning how far into a multi-turn exchange the user already is. Deeper threads tend to carry more refined, more committed intent than a single opening message.
This changes how the auction itself behaves. The winning ad isn't necessarily the highest bid; a lower bid paired with a strong intent-relevance score can beat a higher bid with weak conversational fit, the same structural effect Quality Score has always had on search auction clearing prices. An emerging model called conversation depth optimization takes this further, targeting multi-turn dialogues that lead to conversions and treating turn depth itself as a scoring dimension, on the logic that sustained engagement signals something stronger than a single prompt ever could.
Control here matters for buyers. The demand-and-auction layer and the context-targeting engine both sit with the platform, meaning OpenAI, Google, Microsoft, not the advertiser. Advertisers operate closer to creative and measurement. That split limits how much scoring transparency any buyer can realistically expect to get. Agentic optimization systems now running on top of these stacks analyze over 200 signals and reallocate budget every 15 to 30 minutes, and intent-relevance scoring feeds into that process as one input among many, its weight compounding across a campaign rather than applying to a single impression in isolation.
Low ad load makes relevance the primary lever, not reach
Ad load in conversational AI is thin compared to search or social, and that thinness changes the economics of every placement. Perplexity tested sponsored placements before abandoning advertising entirely in February 2026. Google's AI Overviews, by contrast, now show ads alongside roughly a quarter of all responses, up from 5.17% in early 2025. Fewer placements means each one carries more weight than an equivalent slot in a search results page ever would.
Users trust AI-generated answers more than they trust traditional search results, and a paid placement borrows against that trust. An irrelevant ad doesn't just get ignored the way a bad banner ad does; it damages the platform's credibility directly. That's part of why Perplexity walked away from advertising in February 2026, after pausing the effort following the August 2025 departure of its head of advertising and shopping. A Perplexity executive put the underlying problem this way: "A user needs to believe this is the best possible answer." Ads appearing in that answer lead users to doubt the response's integrity.
Search absorbs a bad match without much damage. A low Quality Score ad still shows up, it just costs more and ranks lower down the page. A conversational interface with very few placements per session has no such cushion; a low-relevance ad that wins an auction in that thread can end up wasting the limited slots available. In a high-volume environment, relevance acts as a multiplier on performance. In a low-load conversational one, it behaves more like a threshold: fall below it and the placement causes harm, clear it and the ad earns the trust it's borrowing. That's the real argument for retiring CTR as the primary signal. CTR measures what happened after an ad showed up. Intent-relevance score predicts whether it should have shown up.
Platforms' efforts to operationalize ad quality in AI environments
OpenAI began testing ads on February 9, 2026, for logged-in adult users in the US on the Free and Go tiers. Its Ads Manager opened to US advertisers on May 5 and has since expanded to Australia, the UK, Canada, New Zealand, Brazil, South Korea, Japan, Mexico, and other markets, putting OpenAI in the position of buying across more than 40 countries. The company reached a $1 billion annualized ad revenue run rate in under 200 days.
OpenAI's stated ad principles hold that ads don't influence ChatGPT's answers, are clearly labeled, and appear separately from the response itself; the Plus, Pro, Business, and Enterprise tiers stay ad-free. Criteo signed on as OpenAI's first ad tech partner, and Dentsu, Omnicom, Publicis, and WPP joined as launch holding company partners.
Google's ads inside AI Overviews launched on mobile in the US in October 2024, expanded to desktop in May 2025, and reached 11 more countries by December 2025. Ads inside Google's AI Mode are also live, running against a large daily user base. Google's long-running Quality Score infrastructure gives it a head start on relevance filtering generally, though how that machinery translates into a conversational context isn't something the company has documented publicly.
Microsoft Copilot has carried ads since 2023, inherited from Bing Chat, and has since expanded into "Compare & Decide" ad formats, shopping campaigns, and an in-conversation checkout feature called Copilot Checkout that launched in January 2026. Microsoft's advertising network overall pulls in more than $20 billion a year.
Perplexity and Anthropic sit apart from this. Perplexity confirmed a full exit from advertising in February 2026, aiming instead for a substantial sum in annualized subscription revenue as the deliberately ad-free alternative in the category. Anthropic has committed to keeping Claude ad-free. Neither of these should read as a failure of conversational ads as a channel; they're a rejection of specific implementations that prioritized speed to monetization over relevance architecture.
None of the major platforms solve the challenge of scoring relevance consistently across surfaces, though. Each one scores relevance and targets on its own terms, and none shares its methodology. An advertiser running campaigns across ChatGPT, Google AI Overviews, and Copilot has no shared relevance signal to lean on, no way to optimize across surfaces, and no independent measurement layer sitting above all three. Verve Group's March 2026 announcement is the clearest public attempt so far at building a third-party layer for cross-surface conversational intent, and it points at a real gap: a demand-side platform built specifically for conversational AI, one that can read context and operate across multiple surfaces rather than inside just one, is the buying-side equivalent of the scoring logic each platform already runs internally.
Optimizing a campaign against intent-relevance score rather than CTR
Search advertising already proved the underlying logic here. Fewer keywords per ad group, with copy written specifically for the intent behind those keywords, produces higher Ad Relevance scores than broad, generic groupings ever do. The same principle carries over to conversational ad units: narrower, more precisely scoped placements, each built around a specific intent cluster, will outperform broad creative running across an entire account.
Building those clusters starts with the strongest purchase-intent markers, a named brand or model, price language, action verbs, budget constraints, use-case specificity, rather than demographic buckets or keyword categories. The creative brief should start from "what is this person actually trying to decide right now?" not "who is this person, broadly speaking?"
Creative quality itself becomes a quality signal in a way it never quite was in display. With very few placements per session, an ad can't just be a banner dragged over from a different channel; it has to read in the conversational register, native to the exchange and answering something the user is already asking about.
Conversation depth belongs in the bidding logic too. Deeper multi-turn threads tend to carry sharper, more committed intent, so campaigns chasing high-intent moments should weight bids toward later turns in a thread rather than treating every first message the same as a fifth one. That's the emerging conversation depth optimization model in practice.
Measurement needs rebuilding from the ground up. Standard last-click attribution misses influence that happens entirely inside the conversation, so advertisers should plan on incrementality testing and first-party attribution from day one, since platforms hand back only aggregated performance data and no individual conversation-level detail. In search, high Quality Score accounts generate two to three times more clicks on the same budget; the structural logic in conversational AI runs parallel, where a consistently strong intent-relevance score should lower the effective cost per placement and raise the odds of winning one of the few slots a session actually offers.
Buyers evaluating a DSP for this channel should ask one blunt question: can it read conversational context at auction time, or does it only apply audience segments built from past behavior? A DSP that brings reach but can't parse what the conversation in front of it is actually about has no way to act on intent-relevance scoring, no matter how large its inventory looks on paper.
What remains genuinely unresolved in intent-relevance scoring
The training-label circularity deserves to stay front and center. The 2026 arXiv research (2607.27686) reported that a trained evaluator hit 79% relevance sensitivity against 60 to 67% for zero-shot frontier judges, a meaningful improvement on paper. But peer review of that same work found the evaluator had largely learned the relevance gate already built into its own training labels, which raises a real question of whether the score measures conversational fit or simply reflects back its own starting assumptions.
Privacy sets a hard boundary too. Meaningful scoring depends on reading what's actually inside a conversation, and doing that at the individual level runs against both regulation and the aggregated, pseudonymized architecture that credible platforms have publicly committed to. That tension limits how fine-grained scoring can get, and nobody has resolved it publicly yet.
Cross-surface comparability is its own open wound. Every platform running ads inside AI environments defines relevance on its own terms and keeps its methodology closed. A buyer running the same campaign across ChatGPT, Google AI Overviews, and Copilot is really working with three incompatible quality signals side by side, because no industry standard exists yet to translate one into another.
Attribution for assistant-mediated discovery is the last open question, and maybe the hardest one. When an assistant researches, compares, and recommends without the user ever clicking anything, there's no pixel to fire and nothing for a standard tag to catch. The influence on the eventual purchase affects conversion, but it stays invisible to the measurement tools built for a clickable web. Incrementality testing looks like the right direction, but the methodology for doing it well, consistently, across platforms, hasn't settled yet.


