Est.

QA Checklist Before an AI Placement Goes Live

Test AI ad placements across unpredictable contexts, not fixed pages.

Contributing Editor · · 12 min read
Cover illustration for “QA Checklist Before an AI Placement Goes Live”
Campaign Setup · September 4, 2026 · 12 min read · 2,698 words

A QA checklist built for banner ads or search campaigns does not transfer to AI placements, because the thing being checked has fundamentally changed. AI placements insert an ad into a generated response rather than a fixed page, so the "content" around the ad is different every time it runs, and that instability is what breaks the old approach. What follows is a framework for the checks a team actually needs to run before a single impression goes out.

Traditional pre-launch QA assumes a surface that holds still: a page, a search results layout, a feed slot. The ad sits next to content that already exists, and the reviewer confirms it renders correctly inside that slot. That assumption breaks completely once the ad lives inside a language model's output. Screenshot-based creative review cannot capture a response that changes shape depending on the prompt, and keyword blocklists have nothing to block against, since there's no page topic, only a conversation moving in real time. Brand safety tools built around URL-level filtering have no URL to anchor to. Bid verification built for a known ad unit with known placement rules runs into an auction that behaves differently when it sits inside a language model.

The stakes are higher than a missed click-through rate would suggest. Users treat AI conversations as private and personally responsive, closer to a conversation with a person than to browsing a results page. An ad that reads as tone-deaf inside that conversation doesn't just underperform; it reads as a violation of the exchange the user thought they were having. That's a brand perception problem, not a targeting problem, and it demands its own discipline rather than a patched-up version of the old checklist.

How AI placements are assembled at runtime, and what that means for what can go wrong

The mechanic at the center of all this: an ad isn't placed on a page, it's inserted into a response the system generates on the spot. The language model, or an auction layer sitting above it, matches available ad inventory to the intent it detects in the user's prompt. That single sentence hides three separate places where the placement can go wrong.

First, context detection. The system reads the conversation and infers what the user actually wants; if that inference is off, the ad that gets selected is off too, no matter how well the ad itself was built. Second, auction and bid logic. The ad that wins isn't necessarily the most contextually fitting one, it's whichever ad clears the auction, and while well-designed systems try to optimize relevance and revenue together, nothing guarantees the two stay aligned. Third, response integration. The ad sits inside or beside text the model is writing in real time, so if the tone or subject of that surrounding text shifts mid-response, the ad can end up in a context nobody planned for.

Research published by Zhao and colleagues in late 2025 tackles this by folding auction logic directly into the generation process, using reinforcement learning so the model learns to weigh response quality against ad revenue as one combined objective. That's the direction production systems are heading, but plenty of live deployments haven't gotten there yet, which means the three failure points above are still live risks in most current setups, not solved problems.

Retrieval-augmented generation systems add a further wrinkle. Because the response pulls from documents retrieved at the moment of the query, the material surrounding an ad placement can shift even within a single campaign that hasn't changed at all on the advertiser's end. The practical takeaway: a "placement" here is a probabilistic event, a range of possible contexts an ad might land in, not a single fixed slot a reviewer can screenshot once and sign off on. QA has to test across that range, and the range itself depends on format: contextual recommendations, in-chat sponsored placements, and native or display units each integrate differently and fail in different ways, so the scope of testing has to match the format, not a generic template.

Verifying context-matching: does the ad actually belong in the conversations it's entering

The question at the heart of this section is simple to state and hard to answer well: across the whole range of prompts that could trigger this ad, does the ad feel like it belongs in that conversation, or does it feel dropped in from somewhere else?

Answering that starts with prompt sampling. Build a test set of representative prompts spanning upper-funnel informational questions, mid-funnel comparison questions, and lower-funnel transactional ones, then confirm the ad fires on the right ones and stays suppressed on the wrong ones. Push further with intent boundary testing: construct prompts that sit right next to the target intent but shouldn't trigger the ad at all, a competitor query, a complaint, a purely informational question with zero purchase signal, and confirm none of them match.

Some categories carry informational and transactional signals inside the same prompt, which means the ad needs to be configured for the right funnel stage within that mixed signal, not just matched to the general topic. Worth noting too: users arrive at AI chat with a narrower consideration set than they bring to a search engine. They're often already comparing two specific options rather than browsing broadly, so the ad's messaging should assume that narrower frame instead of pitching broad awareness to someone who's already three-quarters of the way to a decision.

Publisher category targeting deserves its own look. Confirm the categories selected in the campaign actually line up with where the target audience is asking relevant questions; travel queries, for instance, start inside AI chat at notably higher rates than many other categories, so a travel campaign's publisher mix should reflect that. And before any of this testing starts, write down what "a correct match" means for this specific campaign. Skip that step and reviewers will disagree endlessly over borderline cases, because everyone is quietly applying a different definition.

Reviewing the ad's integration into the AI response: tone, placement, and label clarity

A banner ad's effectiveness doesn't depend on the article next to it. An ad inside an AI response is different: its performance is tied to the text surrounding it, so a reviewer has to judge the ad in that context, never in isolation.

Start with tone. Does the ad's voice match the register of the response around it, or does it read like a hard sell that landed in the middle of a measured, helpful answer? Tone mismatch in these environments is a measurable failure rather than a matter of taste, because the ad's reception is tied directly to the text surrounding it. Test this across different response styles too, since the same ad copy that lands fine in a direct factual answer can land badly inside a cautionary or empathetic one.

Labeling matters just as much. The placement needs to read clearly as advertising, not buried in the response and not phrased ambiguously. This is a compliance requirement, but it's also a trust requirement, and the trust side matters more here than in most ad formats: users experience these conversations as personal, so an undisclosed ad isn't a policy gap, it's a breach of what the user thought was happening. Confirm the label survives when the response structure shifts. If it's injected at a fixed position but the surrounding text can restructure itself, check that the label doesn't disappear or turn unreadable in some generated versions.

Look at the call to action next. Does it ask for something that fits where the user actually is? A "buy now" CTA dropped into a purely informational exchange is both a wasted conversion opportunity and, arguably, a brand suitability issue in its own right, since it signals the advertiser wasn't paying attention to the conversation at all. Finally, follow the ad past the click. The landing page should continue the same conversational thread the user was in, not dump them onto a generic product page that ignores everything that came before, so QA the full handoff, not just the ad unit sitting inside the chat.

Brand suitability in AI environments: what blocklists and URL filters can't cover

Brand safety infrastructure, as it exists today, was built to filter pages and domains. AI responses have neither of those things, so the filtering logic has to get rebuilt starting from the conversation itself rather than borrowed wholesale from the old system.

That rebuild starts with defining conversation-level suitability criteria before launch: not which websites to avoid, but which topics, emotional registers, or user situations should keep this ad from appearing at all. Worth naming a few concrete exclusion triggers as a starting point: language that signals distress, competitor complaint threads, medical or legal urgency, and politically charged subjects that brush up against the product category even loosely. Sensitivity also varies by format, and in-chat sponsored placements sit closer to the actual response text than contextual recommendations do, so the suitability bar for in-chat placements needs to sit higher too.

Publisher-level suitability needs the same scrutiny. A general-purpose assistant, a vertical tool built for travel planning or health questions, and a coding assistant carry different conversational norms and different user expectations, so confirm the publisher mix in the campaign actually matches what the brand considers acceptable. The IAB Tech Lab noted in mid-2025 that the fundamentals of the open web are under ever-increasing pressure as AI disrupts how publisher content gets surfaced and classified, which is a useful reminder that existing taxonomies weren't built with conversational AI in mind and probably don't cover it cleanly.

Whatever gets decided here needs a name attached to it. Document who determined which conversation contexts are acceptable, and lay out an escalation path in advance for whatever edge case inevitably surfaces after launch.

Auditing the bidding setup: what changes when the auction lives inside a language model

In standard programmatic buying, the auction finishes before anyone generates content. In LLM environments, auction logic and content generation are increasingly happening together, and that changes what "bidding correctly" actually means for a campaign.

A few things to check before spend goes live. Confirm the pricing model actually fits the campaign's goal: CPC and CPA suit performance objectives, CPM suits awareness, and running CPM against a transactional prompt context (or CPC against a broadly informational one) sets the campaign up to waste money from day one. Check pacing logic too, since conversational query volume swings less predictably by hour and by day than search volume does; confirm the pacing settings won't burn through budget during low-intent windows before the higher-intent traffic even shows up. Look at bid floors and confirm they're set relative to the actual publisher category and conversation type in play, rather than imported wholesale from a search or display campaign's defaults.

If the buy runs programmatically, check the OpenRTB 2.6 integration specifically: confirm the bid request fields tied to conversational context are being passed through and read correctly on the other end. A missing or malformed context signal means the demand-side platform is bidding blind, no matter how carefully the campaign was configured upstream.

There's a deeper tension worth naming here too. Research into LLM auction design, including work from Hajiaghayi and colleagues on RAG-based systems, shows that relevance from the retriever is a factor alongside bid price when allocating ads within generated responses — meaning a pure price auction may not serve the campaign's actual goals. Confirm the platform's auction model actually accounts for relevance rather than running as a pure price auction dressed up in new language. Last check: since an AI conversation can run many turns rather than closing out in one search session, frequency caps need to be defined at the conversation level. A cap set only at the session or day level can still let the same ad surface repeatedly inside one long dialogue with a single user.

Technical validation: rendering, tracking, and signal integrity across AI surfaces

Rendering isn't a given here the way it is in a browser. AI interfaces differ widely in how they display anything that isn't plain text, so confirm the ad unit actually renders correctly on the specific surface it's targeting, not just inside a generic preview tool that may not reflect the real environment at all.

Creative checks need to go deeper than a desktop mockup. Character limits and truncation behave differently inside conversational interfaces than a desktop preview suggests, so long copy that looks fine in preview may get cut off in the live chat window. Link behavior needs its own check too: confirm URLs render as clickable and formatted correctly, since some AI surfaces handle links in ways that diverge from a standard browser. And since a large share of AI assistant use happens on mobile, QA needs to run on mobile renders directly rather than assuming a desktop pass covers it.

Tracking brings its own set of problems. AI placements are cookieless by design, so conversion tracking has to run through a method the platform actually supports: first-party signals, SDK events, or an attribution window both sides agree on ahead of time. Attribution in a conversational setting doesn't behave like search attribution either, since a user might act on something they learned in an AI conversation hours or days later, through an entirely different channel. Define what counts as a conversion and settle on the attribution window before launch, not after the reporting comes back looking strange. This attribution gap is a genuinely unresolved problem at the industry level right now; QA can't fix it, but it can make sure the team made explicit, documented decisions about how it's measuring, instead of discovering the gap for the first time in a post-campaign report.

Signal integrity deserves a direct check too: confirm conversational context signals are passing accurately from the supply-side platform or exchange through to the demand-side platform. A broken or incomplete bid request means the ad can serve into the wrong context no matter how carefully the targeting itself was built. And run a standard tag and pixel audit, but check the firing conditions specifically for a conversational interface, since tags that fire on page load may not fire correctly at all inside single-page or app-embedded AI surfaces.

Stakeholder sign-off and the QA documentation habit that prevents post-launch disputes

A checklist only does its job if someone owns each line of it. Assign every QA category, context-matching, response integration, brand suitability, bidding, technical validation, to a named role with explicit authority to sign off before the campaign moves into production. Without that, the checklist is just a document nobody's actually accountable to.

Sign-off should produce a paper trail, not just a verbal go-ahead. Keep the prompt test set used for context-matching review and note who approved it as representative, and keep the suitability criteria defined for the campaign, including anything excluded beyond whatever the platform defaults to. Keep the pricing model, bid logic, and frequency cap decisions along with the reasoning behind each one. Keep the attribution methodology the team agreed to, including an honest account of what it will and won't be able to measure, and keep a running list of known edge cases or open questions that need watching after launch.

For agencies specifically, this documentation is the artifact that protects everyone, agency and client both, the moment a placement does something unexpected. "The team checked for this" only holds up if it's written down somewhere with a name attached, not recalled from memory during a heated call three weeks after launch.

Set the post-launch monitoring cadence at this same stage, not later. AI placements drift as the underlying language model's response patterns evolve over time, so a campaign that passed every check at launch still needs a re-review at defined intervals; stability should never be assumed just because the first week looked clean. As more ad budget shifts toward these placements, the teams that build this kind of repeatable, documented process now are the ones that move faster on every campaign after the first one. The first run through this checklist is always the hardest, and the documentation from that first run becomes the institutional memory that makes the second one easier.

Sources

  1. arxiv.org
  2. arxiv.org
  3. wearebrain.com
  4. gendiscover.com
Filed underCampaign Setup

More in Campaign Setup