How to test ChatGPT ads: a budget and test plan
- Daily campaign budget minimum for EUR accounts is EUR 15.00 (15,000,000 micros); other currencies have their own minimum, stated only in the rejection error.
- A campaign budget is a daily OR a lifetime limit, never both, and Ads Manager treats the daily figure as an average: actual daily spend can run up to 2x it and seven-day spend up to 7x it.
- There is no A/B testing feature, campaign duplication or draft state in the API; a test is two or more ad groups or campaigns you build and compare yourself.
- An ad group carries one bid, one billing event and its own context hints, so it is the natural unit for an angle, bid or hint test.
- Conversions need at least a day to settle in reporting; impressions and clicks are fresh within minutes.
- Campaigns support both included and excluded locations, which is what makes a geo split settable on this platform.
- Custom audiences are not available for campaigns targeting the EEA or Switzerland.
Testing ChatGPT ads does not need a benchmark you do not have. It needs a budget sized from a target you set, a test structure that isolates one variable at a time, and decision criteria written down before you see a single number. This guide covers all three, plus how to run a holdout if you want to measure real lift rather than just clicks, using only what the platform documents and no invented figures.
- Compute budget needed as target conversions multiplied by the cost per conversion you are willing to accept, then check it against the daily minimum and the platform's 2x/7x pacing swing.
- Give every angle its own ad group: it is the level that holds its own hints, bid and billing event.
- Decide test length and the amount of data you need before launch, not after you see the numbers.
- Use location include and exclude to build a geo holdout if you want to measure lift, not just clicks.
- Write the scale, cut or iterate rule down before you see a single result.
How much budget do I need to test ChatGPT ads?
Budget needed equals the number of conversions you want to see during the test, multiplied by the cost per conversion you are willing to accept while you are still learning: budget needed = target conversions x expected CPA you set yourself. Neither number comes from a published benchmark; you choose both based on what you can afford to spend to get an answer.
As an illustration only, with made-up round numbers: if you want to see 20 conversions during the test and you are willing to accept up to EUR 40 per conversion while you learn, the test needs roughly EUR 800. Once you have that figure, check it against two platform facts before you commit to it. First, the daily campaign budget minimum for EUR accounts is confirmed at EUR 15.00 a day; other currencies have their own minimum, shown only in the error if you set the budget too low, so create the campaign with a placeholder budget first if you need to see the real number for your currency. Second, Ads Manager treats the daily budget as an average: actual daily spend can run up to twice it and seven-day spend up to seven times it, so your test budget needs enough headroom to absorb a genuinely busy day without running out before the test period ends.
How long should a test run before I judge it?
Long enough to cover at least one full weekly cycle, so weekday and weekend behaviour both get counted, and long enough that the conversions you are measuring have had time to settle: the API needs at least a day for conversions to appear in reporting, on top of whatever time your own funnel takes between a click and a result.
Beyond that, there is no fixed number to reach for. A common approach in any test, on any channel, is to decide in advance how much data is enough, for example a minimum number of conversions per ad group, and to hold that line rather than calling a winner from the first day's spend. Set that threshold for yourself as part of the decision criteria below, before the test starts.
How do I structure ad groups so the test stays clean?
One angle per ad group: a single set of context hints describing one situation, one bid, one billing event and the creative for that angle, because the ad group is the level that carries its own hints, bid and billing event, so a result there can be tied to that one angle.
A test that mixes two angles, two bids or two creative styles inside one ad group cannot tell you which change did the work. Keep everything else constant while you vary one thing: hints for an angle test, chat card title and body for a creative test, bid or budget for a pacing test. The guide on context hints covers how to write hints that describe one situation clearly, which is what makes an angle test attributable in the first place.
| Test type | Vary | Hold constant | Read the result from |
|---|---|---|---|
| Angle / hint test | Context hints in the ad group | Bid, budget, creative | Ad group cost per conversion |
| Creative test | Chat card title and body | Hints, bid, budget | Ad-level clicks and click-through rate |
| Budget or bid test | Budget or fixed bid | Everything else | Spend pace and cost per outcome |
| Geo holdout | Locations included and excluded | Creative, hints, bid | Difference between the conversion change in the test region and in the held-out region |
How do I size a ChatGPT test against my Google and Meta budgets?
Apply the same formula to every channel, with the CPA you set for each: split a fixed testing budget across channels in proportion to how much you need to learn from each one, not in proportion to what you already spend elsewhere.
As an illustration only, with made-up round numbers: a team with EUR 3,000 to test three channels might put EUR 800 into ChatGPT ads because it is new and unproven, EUR 1,200 into Google because it needs a larger sample to beat an existing baseline, and EUR 1,000 into Meta for a creative test. The split is a decision about where the uncertainty is greatest, not a fixed ratio to copy. Whatever you choose, run all three for a comparable period so the comparison is fair, and keep the platform-specific minimums and pacing swings from the section above in mind for the ChatGPT side.
What is a holdout or geo split, and how do I set one up here?
A holdout compares a region or segment where ads run against a comparable one where they do not, so the difference you measure is the ads' effect rather than whatever else was already happening; on this platform it is built with a campaign's location targeting, which supports both included and excluded locations.
In practice: pick two markets you consider broadly comparable in size and behaviour, run the campaign with the test market as an included location and the control market as an excluded location, and make sure no other campaign reaches the control market. Spend nothing in the control market for the length of the test, then compare each market's change in conversions against its own pre-test baseline over the same window, using your own analytics, since OpenAI has not documented a built-in lift or incrementality report. Custom audiences are not available for campaigns targeting the EEA or Switzerland, so an EU-focused holdout has to be built on location targeting rather than an audience split. The reconciliation between what the platform reports and what your own analytics shows, plus a fuller treatment of incrementality testing, is covered in measuring ChatGPT ads with GA4, UTMs and incrementality tests.
What decision criteria should I set before I launch?
Three things, written down before the test starts: the target cost per conversion or return that counts as a pass, the minimum amount of data you need before you trust the number, and what you will do if the result is a clear win, a clear loss, or genuinely unclear.
Deciding these in advance matters because a test that is judged as it runs tends to get stopped the moment it looks good or bad, which is exactly when the sample is smallest and least reliable. Put the target, the minimum sample, and the three possible actions (scale, cut, extend) in the same place you will look when the test ends, and hold to them even if the first few days' numbers are tempting to act on early.
What do I do with the results: scale, cut or iterate?
Scale a clear winner in the size of budget change your own guardrails allow, pause a clear loser rather than letting it keep spending, and give a genuinely unclear result either more time or a wider set of hints before you call it either way.
Scaling in large jumps undoes the discipline of the test: a winner that looked good on a EUR 800 budget is not guaranteed to look the same at ten times the spend, so increase gradually and re-check. Adsonomy's own default guardrails, a maximum budget increase of 25 percent and a maximum decrease of 50 percent per change, are one example of pacing that growth without a big jump undoing what the test found; the guide on budgets, pacing and bidding covers the full set. A losing ad group can simply be paused, since everything on the platform is created paused and can be paused again at any time. An unclear result is not a failure of the test; it usually means the sample was too small or the angle needs a genuinely different set of hints, not a bigger budget on the same one.
What does a launch checklist look like?
Six checks, in the order you would actually do them, so nothing about the budget, the structure or the decision rule is left to be worked out mid-test.
- Set your target CPA or ROAS and compute budget needed = target conversions x expected CPA.
- Check that figure against the platform's daily minimum (EUR 15.00 a day confirmed for EUR accounts) and its pacing swing (up to 2x a day, 7x a week), so a busy day cannot exhaust the test budget early.
- Build one ad group per angle, each with its own hints, bid and budget, so the results can be attributed cleanly.
- Decide the test length and the minimum amount of data you need in advance, since conversions take at least a day to settle.
- If you want to measure lift rather than just clicks, set up a geo holdout with location include and exclude before you launch, not after.
- Write down what a win, a loss and an unclear result each mean, and what you will do about each, before you see any numbers.
Frequently asked questions
Is there a minimum budget for a ChatGPT ads test?
For EUR accounts the confirmed minimum daily campaign budget is 15,000,000 micros, EUR 15.00. Other currencies have their own minimum; the platform states it only in the rejection error if you set the budget too low, so check it when you create the campaign rather than assuming a figure.
Can I run a native A/B test on the platform?
No. There is no A/B testing feature, campaign duplication or draft state in the Ads API. A test is two or more ad groups or campaigns you build separately and compare yourself, ideally changing only one thing between them.
How much should I budget for a first ChatGPT ads test versus Google or Meta?
Use the same formula for every channel: budget needed equals the number of conversions you want to see during the test, multiplied by the cost per conversion you are willing to accept while learning. Set both numbers yourself; there is no published benchmark to borrow instead.
Can I exclude a region as a holdout?
Yes. Campaigns support both included and excluded locations, so you can run ads in a test market and hold a comparable market at zero spend, then compare each market's change in conversions against its own pre-test baseline over the test window.
Why do my first few days of results look noisy?
Conversions need at least a day to settle in reporting, and daily spend can run up to twice the campaign's average budget on a given day. Set your test length and the amount of data you need before you judge it, rather than reading the first days on their own.
Adsonomy runs ads inside ChatGPT for shops and service businesses. Control edition: you run it with rules. Autopilot: the AI manager runs it inside your guardrails. Launching 1 October 2026.
Join the waitlist