Est.

Creative Versioning by Conversation Stage and Intent Depth

Two axes—conversation stage and intent depth—unlock better ad timing than funnel models.

Staff Writer · · 12 min read
Cover illustration for “Creative Versioning by Conversation Stage and Intent Depth”
Conversational Creative · September 15, 2026 · 12 min read · 2,605 words

Creative versioning in conversational AI advertising fails for one recurring reason: it treats the funnel as a single line from awareness to purchase. Most teams still track only how far a conversation has gone. That's the wrong single variable. The system needs two axes tracked at once: how far the conversation has gone, and how deep the user's stated intent actually runs. A user three turns into a chat can still be window shopping, and a user on their first prompt can arrive with the decision already made. Miss that distinction and the same ad copy that closes a sale in one context will drive a user away in another.

What prompt-level signals actually reveal about intent depth, and why they are richer than any keyword

Search advertising built an entire industry on inferring intent from a keyword. That industry undersold what a conversation actually contains, and the underselling was the whole problem. A conversation gives up far more, because the user has narrated the situation in full sentences rather than typing three words into a box. Specificity of constraint matters most here, since a prompt naming a brand, a SKU-level attribute, a budget range, or a delivery window carries more signal than a general category query ever will.

Comparison framing tells a second story. "Which is better" reads differently than "tell me about," and both read differently than "I need to decide between X and Y by Friday." That last phrasing has a deadline in it: the user has already run whatever research they intend to run. Negative space matters too. When a user rules something out explicitly, prior evaluation already happened somewhere else, maybe in a browser tab that's still open.

Action verbs separate browsers from buyers just as cleanly. "Find me," "book," "compare prices," "help me choose" signal a different posture than "explain," "what is," "how does." Socio-economic cues embedded in language, price sensitivity phrasing, premium-brand name-dropping, shift how a model behaves toward a user, too. Research out of Princeton (Wu, Liu et al.) found that large language models already shift sponsored recommendations based on conflict-of-interest conditions, surfacing certain options and downplaying unfavorable comparisons in ways that favor company incentives over user welfare.

Stage adds a layer that intent depth alone can't capture, and this is the part keyword-era targeting never had to solve. A vague prompt on turn eight, after the assistant has already surfaced options and narrowed a category, carries more downstream purchase probability than a specific-sounding prompt on turn one from someone still finding their bearings. Deep intent paired with late stage is the single highest-value moment in the conversation. Shallow intent paired with early stage is the worst possible place to drop a conversion-oriented message, full stop. Keyword lists targeted collapsed strings. Conversational intent maps reward brands that show up with a precise value proposition instead of a broad behavioral profile built for a different medium.

How the two-axis grid organizes creative decisions before any copy is written

Diagram: The Two-Axis Grid: Stage vs. Intent Depth. Visualizes: Visualize a 2×2 quadrant grid with 'Conversation Stage' on the horizontal axis (Early → Late) and 'Intent Depth' on the vertical axis (Shallow → Deep).

Cross the two variables and four quadrants fall out. Conversation stage runs early to late on one axis, intent depth runs shallow to deep on the other. Early-stage-shallow is a user orienting themselves with no purchase signal yet. Early-stage-deep is a user who arrived with a mission on message one. Late-stage-shallow is a user who's been talking for a while but is still comparing loosely. Late-stage-deep sits closest to a decision, and it carries the highest creative stakes of any cell in the grid.

Each quadrant needs its own objective, its own format, its own offer logic. A single copy variant slotted into an otherwise identical unit doesn't cut it, and treating it as sufficient is the most common mistake in this entire discipline. The grid isn't a ladder a user climbs in order, either: someone can walk into a conversation at late-stage-deep on the first message, then drift back toward shallow intent two turns later when a new question broadens the scope again. So the system has to re-evaluate both axes on every turn, not once at the start of the session.

Treat the framework as a decision tool, not a funnel diagram. Anyone who still builds creative versioning around a linear funnel is optimizing for a shape conversation doesn't actually take, and that mismatch is where most wasted spend in this channel originates. Done right, versioning means having the correct asset ready for wherever the user happens to be standing, not herding them down a path someone designed in a slide deck.

Creative posture for early stage, shallow intent: earning the right to exist in the conversation

This is the browsing state: the user learning a category rather than shopping a product. A promotional offer here, a CTA, a price point, any of it signals that the ad is working against what the user is actually trying to do in that moment. Most brands get this quadrant wrong by treating "early" as merely "less far along" rather than as a different job entirely, and that misreading is the single most avoidable error in the whole framework.

The right posture is informational and low-friction: creative that adds to the assistant's answer instead of redirecting it. An inline informational card works. A branded definition or explainer works. A "learn more" prompt that extends the topic the user is already inside of works. Copy should stay in third-person category language rather than first-person brand claims, and offer logic should sit this one out. Discounts and urgency cues do active harm here, because they tell the user the brand is trying to sell to someone who never asked to be sold to.

The job in this quadrant is retention: salience, a name and a frame planted early enough to get activated later, once intent actually deepens. Research out of Princeton (arXiv:2604.08525) found that ads in AI chatbots which violate cooperative conversational norms, recommending irrelevant products, making false claims, hiding sponsorship, erode user trust. That research doesn't frame the violation around purchase-intent stage specifically, but the mechanism applies here with extra weight: the user's cooperative expectation of the assistant sits at its highest point before any purchase intent has surfaced at all.

Creative posture for early stage, deep intent: the user arrived ready but the conversation hasn't validated it yet

Some first prompts arrive fully loaded. "Best noise-canceling headphones under $200." "Which travel insurance covers pre-existing conditions." "Compare X and Y for a small business." The conversation is new, but the need behind it isn't, and treating this user like an early-funnel browser wastes the one advantage the brand has: a user who already did the thinking.

Creative here should move toward differentiation faster than the stage alone would suggest. Copy needs to answer the stated criteria directly, attribute by attribute: "covers pre-existing conditions from day one" beats "comprehensive travel protection" every time, because the user asked a specific question and deserves a specific answer, not a marketing paragraph dressed up as one. A comparison-friendly format, a carousel or comparison card, fits naturally, since the user already framed the decision as a comparison before the assistant said a word. Offer logic can carry a soft CTA: "see full details," "check eligibility," something that extends the conversation instead of trying to close it before it's earned a close.

The failure mode in this quadrant is under-serving. An awareness-register unit dropped into a deep-intent first prompt reads as a brand that didn't bother to read what the user actually wrote, and that's a worse look than saying nothing at all. This is also where conversational ad systems hold a real structural edge over search: the full prompt sits right there, so the creative can answer the actual sentence in front of it instead of guessing at intent from three keywords.

Creative posture for late stage, shallow intent: the user who browses but hasn't landed

Several turns deep, the assistant has surfaced options, and the user keeps circling: price ranges, alternatives, use cases, without ever converging on one. The diagnostic tells here are questions that broaden instead of narrow, "what else is out there," "are there cheaper options," "what do people use this for," combined with an absence of any personal constraint language.

A hard close in this quadrant reads as pressure, plain and simple, and it's the wrong call almost every time it's made. Price, urgency, a checkout link, any of it breaks the relationship the assistant has spent several turns building with the user, and it rarely converts anyway. The better move is to re-anchor the conversation around a single differentiating frame, something that gives the user a decision criterion they didn't have a moment ago. "Most travelers choose X when Y matters most" does more work here than a countdown timer ever could. A branded follow-up prompt is a strong format choice, because it extends the conversation on the brand's terms without trying to end it prematurely. Offer logic should hold the discount back and lean on social proof or a use-case match instead.

The goal is to move the user from shallow to deep, not to force a conversion out of a shallow state. Forcing it rarely works, and it tends to leave a residue of brand irritation that outlasts the session.

Creative posture for late stage, deep intent: the decision moment and what not to waste it on

Named brand, stated constraint (budget, date, location, use case), an action verb like "I want to book" or "I'm ready to," and several prior turns that have already narrowed the field: this is the only quadrant where urgency, pricing, and a direct CTA belong, and even then only when the offer genuinely answers what the user has already said.

Copy shifts entirely here. First-person brand voice becomes appropriate. The offer should tie to the user's stated criteria specifically, "a three-night minimum that fits the window you described" rather than a generic percentage off. The CTA should complete an action rather than initiate one: "book now," not "learn more." An interactive poll or a decision-support card fits the format need well, since it gives the user a frictionless next step without asking them to leave the conversation to finish the task.

Criteo data from US retailers puts LLM-referred user conversion at roughly 1.5 times the rate of other referral channels. This is the quadrant where that premium gets earned, or spent for nothing. And the Princeton research bears repeating: a majority of LLMs tested already recommend sponsored products in ways that favor company incentive over user welfare. That's a mistake, not a strategy. At the exact moment a user is genuinely ready to buy, a transparent, well-matched creative is both the ethical call and the better business decision, because users who feel served convert now and come back later. Anyone treating this quadrant as just another spot for aggressive discounting has misread what got the user here in the first place.

How copy, format, and offer logic each version differently, and why all three must move together

Diagram: Three Levers, One Quadrant: Copy, Format, and Offer Logic. Visualizes: Visualize three parallel tracks — Copy Register, Format, and Offer Logic — each stepping through four stages that map to the four quadrants (Early/Shallow → Early/Deep…

Three levers move independently, and versioning only one of them produces broken creative, not stronger creative. Copy register tracks intent depth more than stage: informational, categorical, differentiating, transactional, in that order. Format tracks stage more than intent depth: inline card, branded follow-up prompt, comparison card, interactive action-completing unit. Offer logic tracks the intersection of both: no offer, decision criterion, soft CTA, hard CTA with a matched incentive.

Change copy without touching format and offer defaults, and the result comes out internally inconsistent: a transactional headline sitting inside an informational card, or an action-completing CTA paired with exploratory copy that never earned the right to ask for a close. In an LLM environment, format isn't cosmetic, and treating it as a skin to swap out is where a lot of otherwise-decent copy gets wasted. An inline card dropped into a long exploratory answer reads completely differently from the same card following a narrow, criteria-specific response. The academic literature has a term for this: a generative externality, where an inserted ad can shift the tone, length, and specificity of the surrounding assistant response. Format choice affects the answer itself, not just the ad unit riding alongside it.

For production teams, this means a creative brief has to name the target quadrant explicitly, not just the campaign objective. A brief that says "drive conversion" without specifying expected stage and intent depth produces one asset that fits its intended cell and misfires everywhere else it gets served.

How the ad system reads both axes in real time and routes to the right creative version

Routing on two axes requires the system to read two distinct categories of signal at auction time. Stage signals: turn count, topic continuity across turns, what the assistant has already surfaced, whether the user has repeated or refined a question. Intent-depth signals: entity extraction at the prompt level, named brands, stated constraints, action verbs, and negative constraints, the things a user has already ruled out.

This is a structurally different targeting input than a keyword, full stop. The system reads the full semantic content of a conversation, not a collapsed search string. Generalist demand-side platforms haven't historically had direct access to that conversational context, though that's starting to shift: emerging integrations are beginning to let DSPs buy LLM ad inventory while the LLM platform itself handles conversational placement and the DSP supplies user-profile signals. That gets a DSP onto the surface, but it doesn't hand the DSP native ability to read stage or intent-depth signals on its own, and that gap matters more than the headline partnership suggests.

Single-surface AI ad networks read context well on their own turf, but can't offer the cross-surface reach that would let one versioning logic apply consistently across ChatGPT, Copilot, and whatever comes next. Some emerging cross-platform networks claim to serve across multiple AI surfaces at once, though the claims are early and unproven at scale, and ought to be treated that way until the data says otherwise. Perplexity walked away from advertising entirely in February 2026, which removes it as an active surface where the same user might otherwise be reachable.

On the auction side, research out of Peking University, Alibaba Group, and Shandong University, published as the LERA framework, shows how LLM-native auction infrastructure can use the model's own logits over candidate ads as a relevance score, combining conversational context and bid signal in a way legacy auction design has no equivalent for. For ad ops teams, the practical consequence is this: creative libraries need to be tagged by quadrant, not campaign objective, so routing logic can select the right version at auction time without a human touching each impression.

What changes in intent depth when the AI is acting as an agent, not just an answer engine

Agentic flows break the two-axis model in a specific way. The user states intent once, "book me a flight to Berlin under $600 for next Thursday," then steps back. The agent executes. The user isn't weighing options across a dialogue anymore. They've delegated the whole task in a single instruction.

Intent depth in that instruction sits at its ceiling from the first message. And "conversation stage" as a concept starts to lose its meaning here, because there's no gradual deepening left to track. The delegation statement already contains everything the ad system would otherwise have spent several turns inferring, including the budget, the date, the destination, and the decision to act, all delivered at once. The creative system now serves someone who has already stopped comparing choices. It's serving someone who has already decided, and who is now watching, or not even watching, to see whether the task gets done well.

Sources

  1. Ads in AI Chatbots? An Analysis of How Large Language Models NavigateConflicts of Interest
  2. LERA: LLM-Enhanced RAG for Ad Auction in Generative Chatbots
  3. ChatGPT Ads in 2026: Early Results, Best Practices, and How to Get Started

More in Conversational Creative