Est.
FeaturesLong read

Bid Adjustment Logic in Conversational AI Auction Environments

Conversational ads need bid logic built for dialogue, not search queries.

Contributing Editor · · 11 min read · Updated
Cover illustration for “Bid Adjustment Logic in Conversational AI Auction Environments”
Features · September 1, 2026 · 11 min read · 2,546 words

AI advertising stopped being a pilot program the day ChatGPT launched advertising on February 9, 2026, moving from a $200,000 enterprise-only pilot to open self-serve access in 86 days. That is a market deciding, in under three months, that the inventory is ready to sell at scale, whether the pricing infrastructure underneath it is ready or not. The evidence suggests otherwise, and the industry is bidding on conversational ad inventory using a pricing model built for a different kind of query entirely.

The audience backing that decision is not small. Reuters reporting puts AI search ad revenue on a path from just over $1 billion in 2025 to nearly $26 billion by 2029. At that growth rate, an inefficient auction stops being a rounding error and becomes a structural cost, one that compounds every quarter it goes unaddressed.

Microsoft Copilot, Google's AI Mode, and ChatGPT are all running live auctions right now, and each has chosen a different shape for the ad itself. ChatGPT's unit is the chat_card, a sponsored card sitting below the AI's answer, while Copilot introduced an "ad voice" feature that builds a conversational bridge between the model's response and the sponsored message that follows it, and Google places ads alongside AI Overview results. Each format implies a different insertion point, a different window of relevance, and a different problem for anyone trying to price what a bid on that placement is actually worth. The pricing infrastructure underneath these formats has not caught up to units that did not exist two years ago, and most of the industry is still bidding as though it had.

How intent accumulates across a conversation and why a single prompt is the wrong unit of analysis

Diagram: The Conversational Funnel: Intent Value Across a Single Session. Visualizes: Visualize how purchase intent and bid value escalate across five turns in one conversation session.

Search auctions fire once, on a self-contained string. Type "best running shoes for flat feet" and the query carries everything the auction needs to price it. A conversation works differently: each prompt is conditional on what came before it, and reading one turn in isolation throws away most of what makes that turn valuable. Treating a conversational prompt like a search query is among the more consequential mistakes being made in this market right now.

Walk through a real session. Turn one, "What are the best running shoes for flat feet?" is category-level, informational, cheap to serve an ad against and worth little, while turn three, "How do the Brooks Adrenaline and the ASICS Gel-Kayano compare on cushioning?", introduces a comparison, a consideration set, mid-funnel behavior. Turn five, "Which one ships fastest if I order today?", is transactional; the user is close to a purchase. Yet nothing about the text of turn five, read on its own, tells an auction that. Its value comes entirely from the four turns that built up to it, and an auction that scores turn five alone has already lost the information that makes it worth bidding on.

Profound's analysis of more than 50 million ChatGPT prompts in 2025 found that only 9.5% classified as commercial and 6.1% as transactional. The bulk of volume, informational prompts at 32.7% and task-completion prompts making up the largest share, is not purchase-adjacent at all. Bid on topic matching alone and an advertiser pays the same rate to reach a browser as to reach a buyer. That is a failure mode a conversational auction has to be built to avoid, and it is one most current bidding systems are quietly committing on every session. The signal worth paying for is the direction the conversation is heading, more than the content of any single line in it.

That direction is not a straight line, either. A user can show strong purchase intent in turn two and drift back into research mode by turn four. Bid logic that only tracks position in the sequence, rather than the trajectory of intent across it, will misprice both ends of that swing. The conversation often functions as an entire purchase funnel compressed into one sitting, unbranded and exploratory at the start, narrow and specific by the end. Treating every turn as equivalent guarantees mispricing somewhere in that session.

The signals that replace keywords: what prompt-level data actually contains

A keyword is a match; a prompt is a statement, and statements carry more than a topic: intent, named entities, sentiment, how specific the language is, and the conversational role of the line itself, whether it is a question, an instruction, or a comparison request.

Several distinct signal types live inside that sequence. There is semantic intent, what the person is actually trying to get done in their own words rather than an advertiser's taxonomy, and there is entity specificity: has the user moved from "running shoes" to "Brooks" to "Adrenaline GTS 23, size 10"? There is funnel language, the gap between "what is" and "versus" and "buy," and there is sentiment toward brands the model has already surfaced earlier in the same session, along with momentum, how quickly intent is escalating turn over turn.

The model itself does the heavy lifting of turning that into a usable score. In the LERA framework (LLM-Enhanced RAG for Ad Auction), the LLM is queried directly to produce logits over candidate ads, and those logits become the refined relevance score feeding the auction, effectively turning the model's own confidence into an auction input. The question a conversational auction has to answer is whether that ad matches where this specific user is, in this specific conversation, right now, a materially harder problem than anything keyword matching ever attempted. A user naming a competitor's product in turn three is worth more to an advertiser than a hundred impressions of generic category browsing, because naming a name reveals the actual consideration set, not just an area of interest.

What is missing from this data matters just as much as what is in it. There is no stable identifier carried across sessions, no historical behavioral profile in the traditional ad-tech sense, and no page context to fall back on. Bid logic here works off signals that exist for the length of one session and then vanish for good.

How current auction frameworks are being redesigned for conversational environments

Three distinct approaches show up in the research literature, and each solves a different piece of the puzzle. Each addresses a failure mode the other two leave open, and no single one of the three covers all three failure modes alone.

The first is the segment auction with an LLM acting as ranker, the LERA approach described above. The model produces relevance scores over candidate ads, those scores combine with advertiser bids, and the framework is structured to preserve truthful bidding as the rational strategy for advertisers trying to maximize their own utility. It is built to extend naturally to multiple ad insertions inside one long, dynamic response.

The second is the dynamic optimal stopping auction, described in the LLM-OSDA framework. Instead of committing to a fixed insertion point, the auction integrates Bellman optimal stopping with the allocation and pricing logic, effectively waiting for the right moment in a conversation before firing. The mechanism separates the timing decision from the bidding logic, committing to an insertion moment before resolving which ad wins. The framework is designed so that allocation properties support truthful bidding as the rational strategy for advertisers. On a simulated conversational advertising corpus, this approach improved net revenue by 11% over the strongest fixed-timing baseline. This framework tackles a question most bidding logic today has no mechanism for answering: when in the conversation to bid at all, not just what to bid.

The third, and the most structurally different from anything in search or display, is the neuron auction. This moves the auction away from surface text entirely and into the model's internal representation space. The auctionable unit in this approach becomes the intervention budget: the degree to which an advertiser can influence the model's internal representations. The design aims to guarantee incentive compatibility while pricing out interventions aggressive enough to degrade the response for the user. Relevance and placement, in this model, happen at the weight level rather than the token level, a genuinely different object to bid on than anything a search engine has ever sold.

The thread running through all three: the language model is an active component of the pricing mechanism itself, a real departure from a keyword auction, where the search engine's ranking algorithm and the ad auction remain, for the most part, separate systems.

What bid adjustment logic must look like when timing and context are both variables

Search and display bid adjustments are multipliers set on known dimensions ahead of time: device, geography, hour of day, membership in an audience list. Everything the multiplier needs is available before the auction fires.

A conversation works on different terms. The dimensions that matter, intent state, entity specificity, funnel position, only exist once the conversation is already underway. Static multipliers on device and daypart keep some residual value, but they now explain a much smaller share of what an impression is worth, and any bidding system still leaning on them as the primary lever is pricing off thinner data than the moment calls for. The dominant lever has to become a dynamic, in-session adjustment tied to the intent state the conversation has revealed so far. A bid on turn five should not equal a bid on turn one in the same session if the user has moved from browsing to comparing to ready-to-buy. A flat bid across both turns treats a browser and a buyer as identical, and that is precisely the mispricing search-based logic was never built to catch.

Timing itself becomes a bid variable, not just a placement variable, which is the core insight behind the LLM-OSDA framework. Fire too early, into high-volume but low-intent turns, and the budget burns on people who were never close to buying; fire too late, after the user has already resolved their own question, and the window closes before the ad ever gets seen. The right moment to insert is a function of where intent is heading, not which numbered turn the conversation happens to be on.

Autobidding is the operating model this points toward, and there is not much of an alternative that holds up. Research in the ACM SIGecom community argues that handing the bidding process to the platform is the practical path for LLM advertising, since scoring intent at the token level in real time is not something a human bid manager can do by hand. What an advertiser sets under this model is a value signal: which intent states are worth paying for, what level of entity specificity should trigger a higher bid, which funnel stage maps to which target cost per acquisition. What the platform's autobidder resolves is the actual bid in each individual auction, informed by the live state of that one conversation. Still, none of it works if the creative on the other end does not match; a high bid fired on a transactional-intent turn is wasted money if the ad copy behind it is still written for awareness.

Why the generative externality problem makes standard quality scoring inadequate

Standard ad auctions place a winning ad next to content that exists whether the ad shows up or not, and the editorial page does not change shape because an ad won the slot next to it.

Conversational AI works under a different arrangement, and this is the point most quality-score frameworks miss entirely. The ad is inserted into, or alongside, a response the model is generating in that moment, which means winning the auction can change what the user actually reads. The LERA research names this directly: an inserted ad can alter the flow, tone, specificity, and length of the entire response, an effect the paper calls the generative externality. A CTR-weighted quality score, the standard tool from search and display, measures fit between an ad and its surrounding context, but says nothing about what the ad does to the context once it is inserted, and that blind spot is disqualifying, not a minor gap to patch later with a lower score after the fact.

An adequate quality score in this environment has to account for relevance to the live intent state rather than a static topic category, for how much the insertion degrades or preserves the usefulness of the AI's answer, and for the cost to user experience that a heavy-handed insertion carries. That cost needs to be priced into the auction directly, not flagged after the fact once the damage is already done. The neuron auction handles this structurally, by making the size of the intervention itself the thing being bought, so a heavier intervention simply costs more.

For advertisers, quality score is becoming less a function of historical click-through rate and more a function of fit with the conversation's current state, turn by turn. Legacy quality signals, built for pages that do not change shape around the ad, do not transfer cleanly to a page being written in real time by the same system selling the ad slot on it. OpenAI's current safeguards, keeping ads away from health, mental health, and political topics, and excluding users under 18, are policy guardrails that draw a line around where ads cannot go. Even so, they do not answer the harder technical question of what an ad does to a response once it is allowed to appear inside one.

How conversational intent data changes which impressions are worth bidding up

Diagram: The Conversational Intent Ladder: From Browse to Buy. Visualizes: Visualize a ranked, ascending ladder (or funnel/slope) showing how intent — and bid value — escalates across a conversation.

The question that determines bid value has changed shape. It used to be who is this user and what page are they standing on; now it is what this specific conversation has revealed about where this person sits in their own decision process. Advertisers still bidding on the old question are overpaying for the wrong impressions and underpaying for the right ones.

That produces something close to a valuation ladder. At the low end sits category exploration: no named brands, informational phrasing, the "what is" and "how does" register that Profound's data shows makes up the largest share of conversational volume by a wide margin. These turns are cheap and should stay cheap; bidding them up just burns money against users who have not yet formed a preference. Further up the ladder, once a session introduces a named competitor product or moves from category language into comparison language, "versus," "better than," "compare," the intent signal strengthens, and the fair price of that impression rises with it. Near the top sits transactional language: shipping timelines, price, checkout, size, quantity, phrasing that shows the user has already narrowed the field and is deciding how to complete a purchase, not whether to consider one.

The trajectory across a session matters as much as the label on any single turn. An impression late in a conversation that has been climbing steadily toward transactional language deserves a materially higher bid than the same phrasing appearing in a session that has been drifting back toward research mode. Bid logic built on keyword matching cannot tell the difference between those two sessions, because it never looks at the path in the first place. Bid logic built for conversation has to, because the path is where the value actually lives, and any auction design that ignores it is pricing on faith rather than evidence.

Sources

  1. arxiv.org
  2. arxiv.org
  3. sigecom.org
  4. trylapis.com

More in Features