Creative Submission Requirements for AI Ad Platforms
Advertisers must write conversational copy that blends into AI responses, not traditional ad assets.

Creative submission requirements for AI ad platforms don't resemble anything advertisers learned buying display, video, or search. There's no aspect ratio, no file weight cap, no fifteen-second duration limit to hit. The ad is text that has to survive inside a conversation, and that single fact rewrites the entire spec sheet. Standalone chatbot ad spending is projected to hit $0.96 billion in 2026, up more than 1,600% year over year, which means a lot of advertisers are about to buy into a channel with no shared understanding of what "creative" even means here.
What an AI ad actually is, and why that changes everything about how it must be built
A conversational AI ad is a paid, targeted message inserted into an AI assistant's response, matched to the intent of the conversation happening right now, not to a webpage, a keyword list, or a stored user profile. Worth separating this from three things it keeps getting confused with. It isn't a chatbot built to sell products, the conversational-marketing idea that's been around since roughly 2016. It isn't AI-generated ad creative either; that's a production tool for making images and copy faster, not a placement channel. And it isn't a search ad sitting next to an AI-written summary. That's still search, just with a different coat of paint.
What makes this format genuinely different is that the ad has to work as language inside someone else's sentence. It's read, not looked at. It has to flow with the response around it instead of interrupting it the way a banner interrupts a page. Research out of Peking University and Alibaba Group (the LERA studies) describes this as a "generative externality": drop an ad into an LLM's response, and it can shift the tone, the specificity, even the length of everything the model says around it. That's not true of a banner ad sitting in a sidebar. A banner doesn't change the article next to it.
So every word choice in the ad has consequences beyond the ad itself. Copy that reads as an obvious sales pitch doesn't just fail to convert, it can visibly warp the surrounding answer, and users notice. What gets submitted to these platforms isn't a video file or a JPEG. It's a structured data payload: a title, a body, a call to action, a URL, and a disclosure string. That's the whole canvas.
The structured ad object: what fields platforms actually ask you to submit
Across the platforms running these formats, the object handed to publishers follows the same rough shape, according to public integration documentation: title, body copy, CTA text, destination URL, and a disclosure label. Advertisers aren't submitting a finished creative. They're filling in fields.
The title has to be short and declarative, naming the brand or product without reaching for superlatives. Body copy needs to read like informative prose, something closer to what a knowledgeable friend would say than what a slogan writer would pitch. The CTA is a short phrase, "See options," "Compare plans," "Learn more," built to work as plain text rather than as a button designed to grab the eye. The destination URL is submitted as part of the structured object and subject to platform approval. And the disclosure string, usually just "Sponsored," is typically locked by the platform. Advertisers don't get to touch it.
Some platforms will also accept an image for card-style renders alongside the required text fields. The text fields carry the core of the ad; the image, when it exists, supports it. Character limits shift from platform to platform and format to format, but the discipline stays constant: say less, say it precisely, cut the filler.
OpenAI's approach places ads below ChatGPT's responses as a visually distinct unit, so the ad never touches or alters the answer itself. Criteo runs a different model with its Prompt Smart Ads on ChatGPT, assembling ad content dynamically rather than relying on advertiser-written copy. In that setup, the advertiser doesn't write ad copy at all. It supplies a clean product feed, and the platform assembles the words. That's a real shift in what "submission" means: from writing sentences to structuring data.
How targeting context shapes the copy requirement, and why one version doesn't work
Legacy display runs the same creative everywhere a placement fires. Conversational AI doesn't work that way, because the conversation itself is the targeting signal. There's no cookie tracking behavior across sessions; what the person is asking, right now, decides which ad has a shot at appearing.
Someone asking "what's the best protein powder for endurance running" sits in a different mental spot than someone asking "is creatine safe." The product behind both queries might be identical. The copy that fits each one won't be. That means advertisers need a matrix of copy, not a single approved asset: map out the conversation contexts likely to trigger the campaign, write body copy for each one that reads as responsive rather than inserted, and match the CTA to whatever stage of the decision the user seems to be in, whether that's early research or ready to buy.
Microsoft's reporting on Copilot gives a sense of what's at stake when context-matching actually works: 73% higher click-through rates and 16% stronger conversion rates than traditional search. That gap doesn't come from a better logo or a punchier headline. It comes from the copy meeting the moment. Criteo's dynamic-generation model is really just the logical endpoint of this idea: hand the platform structured data, let it write copy suited to each prompt. Advertisers who show up with one static block of copy and expect it to hold across every context are fighting the grain of how the channel actually works.
Tonal requirements: what "non-intrusive" means when the ad is made of words
Every platform's policy language says the ad has to be informative, not intrusive. That's a principle, not a spec, and turning it into actual copy decisions takes some work.
Intrusive, in a text-only environment, tends to look like a specific set of habits: superlatives that snap the reader out of the conversational tone ("the #1 rated," "revolutionary"), copy clearly written for a button that has no visual weight in this format, urgency language ("limited time," "act now") that clashes against what's otherwise an informational exchange, and a first-person brand voice fighting against the calm, neutral register the AI has already established. Non-intrusive copy does the opposite. It adds something the user would plausibly want given what they just asked, matches the flat, specific tone the platform has set, and names the next step instead of barking an order.
The cost of getting this wrong isn't hypothetical. Research from Princeton and the University of Washington (Wu, Liu, and colleagues) found that LLMs given ad incentives sometimes recommended sponsored products that ran nearly twice the price of alternatives in test conditions, pushed sponsored options into the flow at moments that disrupted the user's own purchasing path, and in some cases concealed pricing in comparisons that would've favored a competitor. That's what happens when promotion overrides helpfulness at the system level. Separate research from the University of Michigan (Tang and colleagues, 2025) looked at the user side and found that labeling ads in chatbot responses made the chatbot feel less trustworthy and more intrusive to users, who described the whole concept as manipulative. Leave the ad unlabeled, and people mostly can't spot it. Label it, and many users still miss it anyway. Either way, a labeled ad has to earn its spot in the conversation on the merits of what it says, or it does damage rather than good.
A useful test for any copy brief: write it as though answering "what would a genuinely helpful expert say here," then strip out anything that expert couldn't say with a straight face.
Disclosure requirements and how they constrain creative decisions
Every major platform running ads inside AI responses requires a clear label. OpenAI's approach calls for clear labeling and separation from the answer itself. The word "Sponsored" (or whatever the platform's chosen phrasing is) usually can't be removed, restyled, or tucked out of sight by the advertiser. It's fixed.
That constraint reshapes the creative brief in a specific way. The ad has to be worth reading even after the user has already clocked it as paid placement, because that's the order in which they'll encounter it: label first, message second. The headline needs to deliver something real, not just a brand name, since a user who spots "Sponsored" is deciding in that instant whether to keep reading. The body copy can't lean on the AI's ambient credibility either. It has to hold up entirely on its own.
Buyers running campaigns across several AI surfaces at once run into a further wrinkle: disclosure standards aren't consistent surface to surface. One platform's required phrasing or placement might not match another's, which means cross-surface campaigns need copy built to satisfy each surface's disclosure rule individually, not one version stretched across all of them. Brand-safety and suitability filtering happens before the ad ever renders, at the platform level. But passing that filter and earning a user's trust are two different bars. Copy can clear the first and still fail the second if the tone lands wrong.
Image and rich-media assets: when they are required, optional, or irrelevant
Text is the baseline everywhere. The structured object, title, body, CTA, URL, disclosure, is required on every platform running these ads. Image assets aren't universally accepted, and where they are, they're secondary.
When platforms do take images, they're for rendered card formats: inline cards, carousels, anything with a visual container sitting inside the chat window. Several native formats appear in conversational environments, including inline cards, branded follow-up prompts, carousels, and interactive polls, and each carries its own asset rules. Inline cards and carousels generally want clean, product-forward images that hold up at small sizes without any text baked into the image itself, since the text fields are already doing that job. Branded follow-up prompts and interactive polls, by contrast, typically carry no image at all. The whole unit is words.
A few mistakes show up often enough to flag directly. Advertisers submit banner-style images with text overlaid on them, which just duplicates what's already in the text fields and looks broken once the image shrinks down. Others submit assets sized for social platforms, a 9:16 vertical or a 1:1 square, when the container the platform actually renders expects something else entirely. And some treat the image as the main event, when in this format it's supporting material at best. For advertisers feeding a dynamic-creative system like Criteo's Prompt Smart Ads, the real submission work shifts away from designed assets and toward ensuring the underlying product data is accurate and complete, because that data is what the system builds the ad from.
Technical requirements at the platform and integration level
Speed is non-negotiable. A sub-250ms p95 response time is the standard service-level target for ad calls inside LLM interfaces, and an ad object that can't clear that window doesn't get delayed, it gets dropped. That has a direct effect on asset choices: images need to be light and fast-loading, because ad objects that can't be delivered within the response window get dropped rather than delayed.
The auction mechanics differ from display too. In LERA-style two-stage systems, of the kind described in the Peking University and Alibaba Group research, an embedding-based filter narrows down candidate ads first, and then the language model itself scores the remaining candidates for relevance, combining that score with the bid to pick a winner. Relevance to the actual conversation is part of the auction math, not a separate quality layer bolted on afterward. An advertiser with a high bid and copy that doesn't fit the context still loses.
Pricing models run the familiar range, CPM, CPC, and others depending on the campaign's goal, and quality signals like contextual relevance can feed into delivery efficiency and cost. Sloppy creative doesn't just underperform, it gets more expensive to run. Brand-safety filters sit ahead of delivery on every platform, and each one defines its own filter categories, which effectively decide which conversations a given ad is even eligible to appear in. For campaigns bought through a demand-side platform spanning multiple AI surfaces, matching creative to conversational context remains essential. A platform that can't account for that context will still serve the ad, just without the match, and both the relevance score and the user's experience take the hit.
Building a submission workflow that accounts for all of this
Start by mapping the conversations the campaign is actually buying into. What prompts or question types trigger the targeting? What stage of decision-making is the user likely in in each case? What would a genuinely useful expert say to that person at that moment? This mapping work happens before a single word of copy gets written.
From there, write to the fields, not to a layout. Draft the title, body, and CTA as three separate pieces of text, each one strong enough to stand without leaning on the others. Run every draft through the tonal test: would a neutral expert actually say this, or is it just asserting value without offering anything? Produce a version for each conversation context identified in the mapping step, since one version rarely covers the range.
Then prepare whatever assets the chosen formats call for. Check which formats on the target platforms take images at all, and what container size each expects. Strip any overlaid text from images entirely, since the text fields are already carrying the message. Confirm the file sizes are light enough to clear that 250ms delivery window without a hitch.
Before anything goes out, run a compliance pass. Confirm the disclosure string is genuinely platform-locked and not something the creative accidentally overrides. Check each platform's content policy for restricted claims or categories. For anything running through a product-feed model, audit the feed itself: pricing accuracy, image quality, complete metadata, because that feed is the actual creative in that setup.
Finally, plan the iteration cycle differently than legacy media trained advertisers to expect. In display and social, creative fatigue sets in fast, with industry estimates putting the performance drop somewhere between 20 and 40% over a few weeks. Conversational AI runs on a different mechanism. The same piece of copy can perform completely differently depending on which conversation it lands in, so the real lever isn't refreshing tired creative, it's tightening the match between copy and context. Reporting needs to break performance out by conversation context, not just by ad unit, or the whole optimization signal gets lost in an average.
Campaigns spanning multiple AI surfaces need a buying platform built for this specifically, one with direct publisher relationships and the ability to read prompt-level signals and apply the right creative variant accordingly. A generalist buying tool, however capable elsewhere, wasn't built to do that, and it shows in the results.


