Est.
MeasurementLong read

Conversational AI Ad Performance Benchmarks by Vertical

Early data shows conversational AI ads perform vastly better in e-commerce than health.

Staff Writer · · 12 min read
Cover illustration for “Conversational AI Ad Performance Benchmarks by Vertical”
Measurement · September 21, 2026 · 12 min read · 2,780 words

Conversational ad performance depends on what a person actually types into a chat window, not on a page they land on or a keyword they've searched before. That distinction is the entire premise of this piece: intent lives inside the prompt itself, and prompts vary enormously in specificity, urgency, and how close they sit to an actual purchase. A prompt like "best flights to Lisbon in October" carries different weight than "symptoms of low iron," and no single benchmark, no matter how large the sample, captures both. This article works through what early data says about e-commerce, finance, travel, health, and a handful of smaller categories, and it's upfront about what remains unmeasured. Chatbot-native ad spending is projected by eMarketer to hit roughly $0.96 billion in 2026, up 1,641% from where it started, so treat everything below as an early read, not a settled baseline.

How conversational AI ads work before any vertical enters the picture

Ads inside a chat assistant don't sit in a sidebar or a banner slot. They get built into the response itself, so the ad and the answer are the same object rather than two separate things stacked on a page.

Targeting works without cookies, device fingerprints, or the kind of third-party audience files that programmatic buyers have used for two decades. OpenAI's model for ChatGPT infers relevance from a user's own conversation history plus "context hints" that advertisers supply, and a machine learning layer matches those hints to prompts that are semantically similar. Every question a user asks, and every answer the model gives back, becomes a fresh signal evaluated in real time. There's no static profile sitting in a database waiting to be pulled.

Three, really four, distinct strategies have emerged across the major platforms, and they don't look alike. OpenAI turned ads on for ChatGPT on February 9, 2026: sponsored product cards now show up below responses for Free tier users in the US, at a CPM around $60, roughly triple Meta's going rate, with a minimum buy-in of $200,000. Plus, Pro, Business, Enterprise, and Education subscribers see none of it. Adobe, Ford, and Target signed on as launch partners, and WPP, Omnicom, Dentsu, and Publicis are already buying inventory on behalf of clients. OpenAI reports more than 900 million weekly active users, which is the scale that made agencies show up on day one.

Microsoft Copilot has run ads since 2023, inherited from Bing Chat, and its "ad voice" feature builds a conversational bridge between the AI's answer and the sponsored message rather than bolting one onto the other. Perplexity went the other direction entirely: it tested sponsored follow-up questions starting in November 2024 with brands including Indeed and Whole Foods, stopped taking new advertisers in October 2025, lost its head of advertising and shopping that same August, and abandoned the ad business in February 2026 to chase a substantial annualized subscription revenue target as an ad-free product. For now Perplexity is targeting substantial annualized subscription revenue as an ad-free product. Google's Gemini is reported to be exploring advertising, Gemini's global monthly active users are up roughly 30% since August, and its US footprint is projected to reach 85.7 million users by 2029. Meta is instead using what people say to its AI as a signal to aim ads elsewhere. It's using what people say to its AI as a signal to aim ads elsewhere, in the Instagram feed, which is a structurally different bet than the other four.

The auction mechanics are still being invented. A framework called LLM-Auction, introduced at ICLR 2026, treats ad placement as a preference-alignment problem rather than a straightforward bid-and-win auction, aiming for allocation efficiency without adding extra inference cost. Separately, genre-based decoupling frameworks partition bidding around coarse semantic clusters instead of individual user prompts, which is as much a privacy fix as a targeting one; it's built specifically to satisfy FTC and GDPR disclosure requirements without exposing what any single person actually typed. Ads across these platforms are labeled and visually set apart from organic answers, which matters more in some verticals than others, and that point comes back later in the health section.

None of this is neutral toward category. A targeting mechanism built on live prompt language amplifies the categories where people type specific, decision-stage questions, and it does much less for categories where the questions stay vague. That's the structural reason some verticals are already outperforming others, not a situational one.

The cross-platform baseline: what Copilot's data establishes for conversational ads overall

Microsoft Advertising published research in 2025 that stands, at the moment, as the most public cross-vertical benchmark available for conversational ads. That research found Copilot users produce click-through rates 73% higher than traditional search and conversion rates 16% stronger, and their path from question to purchase runs 33% shorter. Fewer steps, higher completion, at every stage Microsoft measured.

The explanation Microsoft gives is contextual matching, not a better crowd of users. Copilot's ad formats adapt to what a person actually typed, which raises relevance in a way a static search results page can't. That said, media planners should note who's in the room: Copilot's user base skews toward higher-income households, a detail that matters most in categories like finance and travel where average order value tracks with income.

A separate Adobe Analytics survey, cited inside the Microsoft research, found that 53% of US consumers plan to use generative AI for online shopping, and 39% already have. The audience is already moving into the channel on its own. It's already moving there on its own, and advertisers are the ones catching up.

This is Microsoft's own data about its own platform. Independently audited, cross-vertical benchmarks for conversational AI advertising barely exist yet, so treat the 73% lift as directionally real and platform-reported, not as a law that holds everywhere ads run. It's the ceiling of what aggregate numbers can tell a media planner. Everything past this point is about what happens once you stop averaging and start looking at a single vertical at a time.

Diagram: Conversational Ads vs. Traditional Search: Copilot's Cross-Vertical Benchmark. Visualizes: Show three performance lifts that Microsoft Advertising's 2025 research found for Copilot users versus traditional search: click-through rates 73%…

E-commerce: high-volume, low-AOV prompts are where conversational AI ads perform best today

E-commerce prompts tend to name a product and a constraint at the same time: "best wireless headphones under $150," "running shoes for flat feet." That phrasing sits close to the pre-decision stage of a purchase, which is the highest-intent cluster a conversational system can catch.

Lapis's study of ChatGPT ads puts typical conversion rates at 4 to 7% across verticals broadly, and finds that clicks from ChatGPT convert somewhere between several times higher and roughly quadruple the rate of Google Search in matching categories. That's a wide range, and it should be, because e-commerce itself splits sharply by price point.

Digital Applied's 2026 benchmark work on AI-generated creative found ROAS parity for products under $100 in average order value: AI-matched ads perform at or above human-produced creative for impulse and low-consideration purchases. Past $100 in average order value, conversion rates start to drop, and the gap widens further for anything over $500. Conversational AI ads, in other words, are already strong for fast-moving consumer goods and considerably weaker for high-consideration retail, at least with the creative and targeting tools available right now.

The traffic numbers back up the direction of travel. Adobe data cited by Triple Whale shows AI-referred traffic to US retail sites up 4,700% year over year, and NVIDIA reported that more than 80% of retailers are using or piloting generative AI in some form. Agentic AI, meaning software that takes purchasing actions on a user's behalf rather than just answering questions, could influence up to $1 trillion in US retail revenue by 2030 over a longer horizon. That shifts the whole unit of analysis from a "prompt" to an "agent action," and attribution for that hasn't been solved by anyone yet.

For now, the practical read for e-commerce advertisers is straightforward: put low-to-mid AOV product lines into conversational AI campaigns first, use the multi-turn conversation for comparison-stage targeting, and hold premium, high-AOV creative back for channels where brand storytelling still carries more weight than a matched product card.

Diagram: E-Commerce Conversion Cliff: How AOV Shapes Conversational Ad Performance. Visualizes: Illustrate how conversational AI ad performance in e-commerce drops sharply as average order value rises, using three AOV bands from Digital Applied's…

Finance: high purchase intent, but regulatory caution shapes what can run and where

Financial prompts, things like "best high-yield savings accounts" or "should I refinance now," are about as commercially explicit as conversational queries get. A person asking that question has already told the system they have a financial need and are looking to act on it.

The limiting factor isn't intent quality, it's compliance. Financial advertising carries disclosure requirements and suitability rules that predate any of this technology by decades, and those rules interact with a native-ad format, one built into the answer itself, in ways nobody has fully worked out yet. Copilot's income skew, that roughly 20% of surveyed users at $100,000 household income or above, makes it a relevant inventory source for products where income eligibility actually matters, and planners should weigh that against product fit rather than assuming every conversational platform serves the same audience.

ChatGPT's privacy-first targeting model cuts the other way for finance specifically. Because there's no third-party data and no income-band checkbox to target against, the language inside the prompt has to carry more of the targeting weight than it would in a traditional programmatic buy. That's precisely where the genre-based decoupling method, described in recent academic research, becomes relevant: it lets a financial advertiser reach a coarse cluster of high-intent conversations without ever touching the personally sensitive content of an individual prompt.

Finance, then, sits in an odd spot: the intent signal is arguably the cleanest of any vertical covered here, but the guardrails around it are still being built. Treat it as high-potential and early-stage at once, not as ready to scale without a compliance review first.

Travel: the vertical where multi-turn conversation structure most closely mirrors the purchase journey

Travel planning was conversational before conversational AI ever existed. Someone refines destination, dates, budget, lodging type, and activities across several back-and-forth turns, and each turn narrows the field further.

That makes travel prompts unusually dense as targeting signals. By the third or fourth turn of a trip-planning conversation, a user has effectively told the system their destination, their travel window, and a rough sense of their budget, all without filling out a single form field. No keyword search delivers that kind of layered signal in one shot.

Copilot's 33% shorter customer journey (per the Microsoft Advertising research cited earlier) means something particular in travel, where the traditional research-to-booking funnel can span days of tab-switching between airline sites, hotel aggregators, and review pages. Compressing that into a single conversation thread is a bigger structural change here than in almost any other category. And the same income data point from Copilot, roughly 20% of users at $100,000 household income or higher, maps cleanly onto premium travel products: business class fares, luxury hotel bookings, high-end guided tours.

Once AI agents start booking flights and hotels on a traveler's behalf without a human clicking "confirm," the agentic trend shifts the whole question of ad placement. Once AI agents start booking flights and hotels on a traveler's behalf without a human clicking "confirm," the whole question of ad placement shifts. A brand now gets chosen by an agent acting for the user rather than by where an ad sits inside a conversation, a question nobody in the industry has answered yet. Travel is likely one of the strongest early performers in this channel precisely because intent builds across a conversation instead of arriving in one search term, and advertisers testing now are testing before CPMs catch up to that reality.

Health: intent is high but user trust is load-bearing, and that changes the performance calculus

Health prompts, "symptoms of low iron," "best sleep aids without dependency," carry real urgency. Somebody typing that question usually has an actual, present need, not idle curiosity. Purchase intent runs high as a result. But the person asking is also often anxious, uncertain, or in a vulnerable moment, and that changes how any advertising lands next to the answer.

Health prompts rank among the higher-intent categories that conversational AI ads can reach. Research has indicated that user perception of ad authenticity can weigh on purchase intent. In most categories that's a modest headwind. In health, where the entire relationship depends on the user trusting the source of information, that perception risk gets amplified rather than absorbed.

Healthcare is already one of the largest adopters of conversational AI generally. Mordor Intelligence found customer support made up the largest segment of the chatbot market by 2024, and Accenture estimates the US healthcare economy could save roughly $150 billion annually through conversational AI adoption. That's an enormous base of existing AI interaction, and it represents enormous potential ad-adjacent volume, if users can come to trust the ads rather than being alienated by them.

This is where the labeling and disclosure rules from earlier stop being a compliance checkbox and start being the actual performance variable. OpenAI has stated that ads don't influence ChatGPT's underlying responses, and for health advertisers that separation, editorial answer versus sponsored message, is not a footnote. Whether users trust the format depends on it. Health advertisers should lean toward clearly labeled formats tied to genuinely useful content, and stay well away from anything that reads as exploiting a distress signal. The conversion ceiling in this vertical means advertisers can expect real limits on returns, but getting the approach wrong carries a reputational cost higher here than almost anywhere else on this list.

Where other verticals stand and what their prompt patterns signal

B2B software prompts, "best CRM for a 50-person sales team," "compare project management tools," look structurally like high-intent e-commerce queries: specific, comparison-stage, commercially loaded. Omneky's 2025 platform data shows AI-personalized B2B creative on LinkedIn getting notably higher engagement than generic creative, which is directional support for the idea even though that data comes from social advertising, not a conversational assistant. B2B sales cycles run long and involve multiple stakeholders, so a single chat exchange is unlikely to close anything. Conversational ads in this category are best understood as top-of-funnel and lead-generation tools, not direct-response engines.

Automotive sits closer to travel in prompt behavior but closer to high-end retail in outcome. "Best SUVs for a family of five under $50K" is a common, valuable research query, but the actual purchase almost always closes offline, at a dealership, weeks later. Digital Applied's 2026 finding on conversion drop-off for purchases over $500 applies directly here. Ford's presence as a ChatGPT launch advertiser shows the category is already testing the channel, but the honest framing is research-assist, not conversion-driver.

Retail and CPG outside of pure e-commerce look like the strongest structural fit in the whole set. Target signed on alongside Adobe and Ford as a ChatGPT launch partner, and Fortune Business Insights data cited through Nextiva research puts retail and commerce ahead of every other industry in conversational AI adoption by market share. Combine that with the ROAS parity finding for low-AOV, impulse-adjacent products, and retail looks like it's simply further along than most categories, both in audience readiness and in the kind of purchase conversational ads are already good at closing.

What early benchmarks leave unresolved: measurement, attribution, and the data gap

Every number in this piece comes from a platform reporting on itself, an early-stage academic paper, or a survey with a specific, narrow scope. That's not a flaw unique to conversational AI advertising, most new ad formats get measured this way before independent auditors catch up, but it should be named rather than glossed over.

Attribution is the harder problem created by the data gap. When an ad is embedded inside an AI-generated answer rather than clicked from a results page, standard last-click and multi-touch models don't map cleanly onto what happened. Did the user convert because of the sponsored card, or because the AI's own answer was persuasive and the ad just happened to sit beside it? Nobody publishing benchmarks right now has fully separated those two effects, and the agentic commerce shift described earlier, AI systems taking purchase actions directly, will make that separation harder, not easier, as it develops.

None of that undercuts the vertical patterns laid out above. E-commerce's split by AOV, travel's multi-turn signal density, finance's compliance bottleneck, health's trust sensitivity: those patterns appear consistently across the data available. What's missing is independent, audited measurement at the scale that search and social advertising took years to build. Budget decisions made against these benchmarks should be made with that gap in mind, treating today's figures as a strong early signal about where to test, not a promise about what the return will be once the channel matures.

Sources

  1. AI Advertising Statistics 2026: 30+ Key Facts & Fig… — Omneky
  2. AI in Ecommerce Statistics: 32 Stats Every Online Retailer Should Know in 2026 | Triple Whale
  3. 50+ Conversational AI Statistics for 2026
  4. Conversational AI redefines audience engagement
  5. digitalapplied.com
  6. ChatGPT Ads Conversion Rate Benchmarks by Industry (2026 Data Study) | Lapis
  7. buildmvpfast.com
  8. beet.tv
Filed underMeasurement

More in Measurement