CTA Structure in Conversational Ad Contexts
CTA logic built for search fails in conversational AI because intent unfolds differently.

A question about logistics is different from a question about price, which is different from a question about how two options compare, and each phase creates a distinct moment with a distinct appropriate next action. The query is an opening, not a decision gate, and that single difference is why CTA logic built for search breaks down the moment it enters a live AI conversation.
Why CTA logic built for search fails in a live conversation
The channel rewards a CTA that opens a deeper next step rather than one that grabs a fast click, since engagement depth metrics carry more weight than traditional click signals. The user has turned a need into a short string of words, and that act of compression is itself a signal of readiness to act. A conversational AI exchange skips that step entirely. Search CTA logic treats every user as standing at a decision gate, but a conversational user is somewhere in the middle of a journey, often exploratory, and the answer they just received may have raised new questions rather than closed any.
That difference is visible in the structure of the exchange itself, not just in how the user feels about it. A sponsored link competing with nine others on a results page is fighting for a click that is already forming in the user's head. A sponsored card sitting below a finished AI answer has nothing left to fight for in that sense, because the question has already been answered. The CTA's job changes accordingly: it has to earn a next step rather than catch a click that was going to happen anyway. ChatGPT's 2026 ad architecture makes this literal. The sponsored card sits below the answer and never inside it, the model's response is not for sale, and the commercial unit lives next to something the user has already received for free. Microsoft Copilot takes a related but distinct approach, reading the whole session rather than a single query to decide when a commercial moment has actually arrived, and its "ad voice" feature adds a short transitional line explaining why a sponsored section is showing up before the ad block appears. Both designs assume the same thing: the user has already been served. Everything a CTA does from that point forward has to work as a request.
What the conversation itself already contains
What makes this channel different isn't just that the moment of intent compression hasn't happened yet. It's that the conversation, by the time an ad might appear, usually already contains more useful information than a search query ever could. A live exchange routinely holds the user's budget, their constraints, the shortlist they've assembled, and the options they've already ruled out, none of which a keyword or a demographic bucket can capture with the same precision. That richness should directly shape what a CTA asks the user to do next.
Targeting systems built for this channel read the full trajectory of the conversation rather than a single line, what was asked several exchanges back, how the model responded, which follow-up questions surfaced, and where the thread seems to be heading. That trajectory is the context any CTA has to sit inside. Because intent here is stated outright rather than guessed at, a user will often describe the problem directly, name the options already under consideration, and even describe what a good answer would look like. A CTA that ignores all of that and defaults to "Shop Now" or "Get a Quote" throws away the clearest signal the channel offers.
The trajectory also carries information about phase. A question about logistics is a different moment than a question about price, and a question about price is different again from a question comparing two options directly, with each phase calling for its own next action. Consider a user who never types "project management software" but spends several turns describing problems with team coordination, missed deadlines, and the friction of managing a remote team. A keyword system would miss that user entirely, yet the intent signal is about as strong as it gets. The CTA that works in that exchange has to speak to the problem being described, not to a product category the user never named.
How conversational targeting is decided, and the structural constraints that follow for CTA design
None of this is a matter of tone alone. The way these platforms decide which ad to show in the first place puts real mechanical weight behind getting the CTA right. Relevance-weighted, context-driven matching treats an ad's copy, its targeting description, and its CTA as one combined signal of fit, so a CTA that clashes with the conversation, in tone or in substance, can cost the advertiser the placement itself and not just the click.
The main lever advertisers control is a natural-language context hint, capped at 280 characters, describing the conversations an ad belongs in, and the system matches that description against the live exchange inside a relevance-weighted second-price auction. A more relevant ad can beat a higher bid outright, which turns copy precision, CTA wording included, into a real competitive lever rather than a matter of creative taste. Research from Turner-Smith et al. backs this up directly: their NaiAD evaluation reads title, copy, CTA, and landing page together as one unit, and the relevance evaluator they tested agreed with human judgment in 86% of pairwise comparisons across five annotators. A CTA promising a transactional outcome the conversation hasn't earned yet will register as low-relevance under that kind of evaluation.
Microsoft Copilot's "ad voice" system reinforces the same point from a different angle, triggering ad units based on the direction of the whole session rather than the most recent query alone. A CTA needs to hold together across an arc rather than being optimized for one line of dialogue. Duke Fuqua's threshold-based timing research names a further constraint: if the best available ad would generate value below a set threshold, the model keeps the conversation going without serving anything; once an ad clears that threshold, the model stops learning and serves it. That means an ad only appears once the conversation has reached a point of real relevance. A CTA that fails to match that moment has been handed an unusually warm opening and let it go to waste.
CTA wording as answer continuation, not marketing interruption
The clearest working principle to come out of all this is that a CTA inside a conversation should read like a plausible next line in the exchange rather than a shift into sales register. The model has just delivered a complete answer, and the user is still thinking in dialogue, still in an exploratory frame of mind. A CTA that suddenly switches to imperative sales language, "Buy Now," "Claim Your Offer," breaks that register and tells the user that something other than the conversation has just elbowed its way in.
The contrast is easy to see side by side. "Want a side-by-side comparison of these options?" or "See pricing for the team size you described" both read as continuations of what the model was already doing, because they're phrased as questions or offers tied to specifics the user already supplied. "Get Started Today" does none of that; it could be dropped into any conversation about anything, and that's exactly the problem. Utility-first copy that functions like a fragment of the answer itself, short, specific, tied to the actual topic on the table, consistently performs better than copy written to sound like an advertisement.
This has a direct implication for how CTA copy gets built at scale. Templated variants should be organized around the intent clusters that come out of the context hint rather than around the product alone, since a CTA written for a user still framing their budget looks different from one written for a user in a shortlisting conversation, even when the underlying product is identical. Treating every stage of a conversation as if it called for the same line of copy wastes the signal that makes this channel different from search.
Where the CTA should sit relative to the conversation's stage
Wording is only half the job. The CTA must also match the phase of the conversation it lands in, because even a well-written CTA served at the wrong moment reads as pressure the user didn't ask for. A conversation that starts as a restaurant recommendation can evolve into trip planning, then into a budget discussion, then into questions about transportation, and each of those stages is its own distinct advertising opportunity. The CTA that fits a budget conversation has no business showing up next to the original restaurant question.
Three phases are worth distinguishing, because each one calls for a different kind of ask. In the informational phase, the user is still building a mental model of the space, so the right CTA extends that learning, something like "See how this works for your use case," rather than pushing toward a decision the user isn't ready to make. In the comparative phase, the user has already narrowed things down and is actively weighing a short list, so the CTA should reduce friction on the specific dimension under consideration, as in "Compare plans side by side". In the decisional phase, the user has signaled real readiness: a budget has been named, constraints have been laid out, and alternatives have already been set aside. Only at this stage does a fully transactional CTA like "Start your trial" actually fit the moment.
This isn't just a stylistic suggestion. It follows directly from the Duke Fuqua threshold mechanism described earlier: the platform itself withholds the ad until the conversation has built up enough value to clear the bar. An advertiser whose CTA is matched to phase is working with that mechanism rather than fighting it, since the platform has already done the work of waiting for the right moment; the CTA just has to avoid wasting it.
Pacing and CTA density versus search and display norms
Conversational placements also reward a much lighter touch than banner ads or search results ever called for. A single, well-timed ask consistently outperforms a stack of competing options or urgency cues borrowed from formats built for a different kind of attention. Display and search ads often pile a primary CTA on top of secondary options and countdown language, "Limited time," "Only 3 left," because the user's attention is scarce and the window to catch it is short. A user deep in a multi-turn conversation has already given sustained attention across several exchanges, so that scarcity logic simply doesn't apply.
The numbers back this up directly: conversational traffic shows form start rates 20 to 40 percent higher than equivalent display traffic, which suggests the channel rewards a CTA that opens up a deeper next step rather than one engineered to grab a fast click. A single CTA matched to the right phase and the right register consistently earns more engagement than two or three competing for the same moment, and urgency language in particular tends to backfire here, since a user who went to the trouble of asking an AI a detailed question is already in a deliberative frame of mind, and urgency cues read as an attempt to short-circuit exactly that deliberation.
None of this argues for passive or vague copy. Restraint in this channel means precision: one clear, specific ask, placed at the moment the conversation has earned it, carries more weight than three competing lines ever could. Strength, in this context, is a CTA confident enough to say one thing well instead of hedging across several.
The trust constraint that CTA design cannot ignore
Writing a CTA that blends naturally into the conversation solves an engagement problem, but it opens a trust problem at the same time, and in a channel where the user showed up because they trust the AI's answers, that trust is the entire value of the placement. Work by Tang et al. from 2025 found that a meaningful share of readers fail to recognize a sponsored recommendation even when it carries a disclosure label. Copy that folds too smoothly into the model's own voice can rack up impressions while quietly taking away the user's ability to tell they're looking at a paid placement at all.
The Michigan researchers frame this as the central ethical tension of the format: the line between genuinely useful assistance, surfacing a relevant product, cutting down on information overload, and influence that operates on the user without their awareness. The design answer to that tension is to keep the tone dialogic, utility-first, and matched to phase, while making the structure around it, the label, the visual separation, unambiguous. Tone and disclosure aren't competing goals if the separation is built into the format itself rather than left for the copy to handle alone.
ChatGPT's architecture already enforces this split: the sponsored card is labeled and placed below the answer rather than inside it, so the advertiser's task is to write a CTA that earns trust inside that disclosed frame rather than trying to blur the frame itself. Perplexity's decision to wind down its ad program offers the clearest evidence of what happens when that constraint gets underweighted. The company told the Financial Times that sponsored placement risked making users suspicious of the entire answer, not just the ad sitting inside it, which is as direct a warning as the industry has produced about what's actually at stake.
What attribution and measurement look like for CTAs in this channel
CTA performance in conversational AI advertising cannot yet be read cleanly through the same attribution models as search or display, and designing CTAs without acknowledging this leads to mispriced bids and mistaken creative conclusions. A user can see a CTA, ask several more follow-up questions, leave the chat entirely, and convert somewhere else later on, and last-click attribution has no way to credit the CTA that started that chain, a dynamic that might be called the conversation gap.
Research from Turner-Smith et al., produced jointly at CMU and Tsinghua, goes further and identifies click-through intent as something current methods simply cannot measure reliably in LLM advertising. Behavioral logs aren't available, human annotators can't calibrate consistently against each other, and even frontier LLM judges tend to confuse genuine intent with polished writing. In practical terms, the standard metrics advertisers use to compare CTA variants in search and display don't yet have a working equivalent in this channel.
Until that gets solved, engagement depth signals, form starts, how long a session continues, whether users come back, offer a more honest read on whether a CTA worked than raw click-through rate does. It is a partial substitute for attribution, one that won't settle every argument about which CTA variant earned its placement. It's the most reliable signal the channel currently offers, and any framework for CTA design in conversational advertising has to be built with that limitation in view rather than around it.
Sources
- LLM Ads Explained: How AI Advertising Works in 2026
- Large Language Model Advertising in 2026: Who’s Winning the AI Attention War?
- Understanding Contextual Targeting in ChatGPT Ads: A 2026 Deep Dive
- Generative AI Advertising as a Problem of Trustworthy Commercial Intervention
- Evaluating and Pricing Advertisements in AI-Generated Responses
- TeamCMU at Touch\'e: Adversarial Co-Evolution for Advertisement Integration and Detection in Conversational Search
- Advertising in the Age of AI Conversations


