Home / Blog / How to industrialize creative testing with GenAI pipelines

How to industrialize creative testing with GenAI pipelines

, 8 min read

GenAI pipelines can generate and test hundreds of creative variants each week. Building one that delivers reliable winners takes a specific, repeatable system.

How to industrialize creative testing with GenAI pipelines
George Milton / Pexels

Three things a GenAI pipeline needs underneath it

A GenAI pipeline doesn't fix a broken ad account. Three things need to be in place before it's worth building: a stable offer with proven unit economics, a tracking system that records conversions at the creative level - Keitaro, Binom, or a homegrown S2S setup - and a weekly budget of at least $500-$1,000 per geo, below which reaching statistical significance is unlikely.

A 'test' needs a firm definition too: a creative that's earned at least 50 conversions or $200 in spend before any judgment gets made. Anything short of that is noise. And if the tracking can't separate spend and conversions by creative ID, that's what gets fixed first.

Metrics come last, and not every creative needs to win on ROAS. Some win on CTR, some on CPM, some on conversion rate, so primary and secondary KPIs should be set per funnel stage. Top-of-funnel video might run on CTR above 2% as the primary metric and CPM under $10 as the secondary; bottom-of-funnel images flip to ROAS above 2x primary, conversion rate above 3% secondary.

Build a creative brief database

It starts with collecting every winning creative from the past 6 months, sorted by vertical, geo, and funnel stage. For each one: the hook, the format - UGC, explainer, demo, static - the call to action, and a working theory of why it worked. That collection is the seed data everything else builds on.

That database feeds into structured prompt templates - something like 'Generate 10 video hooks for a weight-loss offer in DE, targeting women 35-50. Hooks should use pain-point opening, then transition to solution. Reference these winning CTAs: [list].' Output quality tracks directly with how specific the brief is.

A Google Sheet holds 20+ of these templates across different formats and stages, each with placeholders for geo, audience, offer angle, and brand guidelines. Creative briefs turn into a repeatable process this way, rather than staying a fresh creative leap every single time.

Set up your GenAI generation pipeline

Two generation streams run in parallel: video and image creative, and ad copy. Video comes from Runway, Pika, or a custom fine-tuned model, producing 30-60 second clips; images come from Stable Diffusion or Midjourney against consistent style references; copy comes from GPT-4 or Claude, run through the prompt templates.

A short script or a no-code tool like Make or Zapier automates the generation itself. The usual setup is a Python script that reads the prompt sheet, sends the requests to each API, and saves outputs into folders named by campaign and creative ID, producing 50-100 variants a batch at $0.10-$0.50 per video and $0.05-$0.10 per image.

Nothing goes straight from output to ad account without a filter first. Three pass-fail criteria decide what survives: visual clarity, no artifacts or broken text; message fit against the brief; and regulatory compliance, no false claims or restricted content. A junior VA can filter 200 outputs an hour, and the typical reject rate runs 30-50%.

Tools and costs

Text-to-video runs through current-generation models - Runway's Gen-4 line or Pika Labs - at roughly $0.10-$0.50 per clip depending on length and model. Images come from Stable Diffusion, free if self-hosted, or Midjourney. Copy runs through a frontier LLM API at around a cent per 1,000 words.

A hundred video creatives, rejects included, cost $10-$50 total. Images run $5-$10 for the same volume, and copy barely registers - against a $200-$500 in-house designer day.

Structure your ad account for mass testing

A flat account structure drowns in its own data. A three-level hierarchy holds up better: campaign by geo and funnel, ad set by audience and interest, ads by creative ID - with each ad set carrying 3-5 creatives minimum and never more than 10, since beyond that, reaching statistical significance takes longer.

Creative names should carry a code: generation batch, format, variant number. 'VA_DE_01_UGC_04' reads as Video Ads, Germany, Batch 01, UGC format, variant 04 - a naming scheme that makes it possible to pull reports by batch and compare how each generation round actually performed.

Ad set budgets need to guarantee a minimum spend per creative. A $1,000-a-week campaign split across 5 ad sets puts $200 on each; 3 creatives per ad set brings that to roughly $67 a week per creative, which is enough for significance in most Tier-1 geos within 7-10 days.

Budget allocation example

A $2,000 weekly budget splits three ways across the funnel: 40% to top ($800), 35% to middle ($700), 25% to bottom ($500).

At the top of the funnel, that's 8 ad sets at $100/week each, 4 creatives per set, working out to $25 per creative per week. In Tier-1 geos, that buys 200-500 impressions per creative weekly - thin for CTR testing, but enough for a CPM comparison.

Run a disciplined testing cadence

The cadence runs on a fixed weekly clock. Monday morning, 20-30 new creatives launch across 5-10 ad sets. Wednesday, early signals get checked and anything under 50% of ad set average CTR or over 2x average CPM gets paused. Friday, underperformers that haven't hit 3 conversions get killed outright. The following Monday, survivors get evaluated and winners get scaled.

A creative earning 10+ conversions at ROAS above 1.5x in its first week is a scaling candidate - it moves to a dedicated scaling campaign with a bigger budget. One that flatlines, no conversions after $50 spent, is dead. There's no reviving it; the move is on to the next batch.

Each batch's velocity gets logged too - generation date, launch date, first conversion date, total spend to that first conversion. A batch that takes longer than 3 days to produce a first conversion is telling you something: the prompts or the audience targeting need adjusting.

Weekly testing cadence example
DayActionDecision rule
MondayLaunch 20-30 new creativesAll pass quality check
WednesdayEarly signal checkPause if CTR <50% avg or CPM >2x avg
FridayKill underperformers<3 conversions after $50 spend
MondayEvaluate and scale>10 conversions, ROAS >1.5x

Analyze results and feed back into prompts

A creative performance report runs every two weeks - filtered by format, hook type, CTA, and visual style, with average CTR, CVR, and ROAS calculated for each attribute. That's what shows what's actually working in the vertical right now.

Those findings feed straight back into the prompt templates. If 'before/after' hooks are beating 'problem/solution' hooks in the dating campaigns, the template shifts to prioritize before/after. A 'prompt version history' tracks which iteration produced which result.

Generation model version matters too - a new Runway update can quietly change output quality. An underperforming batch is worth checking against whether the model version changed underneath it, and rolling back if that's the cause.

The mistakes that waste a testing budget

Over-generating without a quality filter wastes the whole exercise - push 500 unvetted creatives into an account and the budget burns on garbage before any winner has a chance to surface. Manual filtering or a CLIP-based scoring model both work; skipping that step doesn't.

Changing too many variables at once breaks attribution. A test should vary one element per creative, hook or format or CTA, never all three together, or there's no way to know which change actually produced the win.

Audience-creative fit gets ignored too. A creative that kills with 25-year-old men in Brazil can flop entirely with 45-year-old women in Japan, so mixing audiences in the same ad set only makes sense when the creative itself is segmented to match.

Stopping too early throws away real signal. A 20% lift in CVR at 95% confidence needs roughly 300 conversions per variant to call - anything under 50 conversions is close enough to random noise that it shouldn't change a decision.

Winner rate and cost per test in a normal week

In a typical week with 100 new creatives across 3 geos, 5-15 show early promise - CTR or CPM in the top 20%. Of those, 2-5 turn into real winners, ROAS above 2x after 2 weeks; the rest end up average or fail outright.

A winner costs roughly $100-$300 in ad spend on top of $5-$30 in generation costs - 5-15x cheaper than producing the same volume through in-house designers or an agency.

The first winner typically shows up 1-3 weeks after the pipeline is set up. After 2-3 rounds of prompt refinement, the win rate climbs from around 2% to roughly 5-8% of launched creatives.

FAQ

How many creatives should get tested per week?

20-30 a week is the starting point on a $1,000-$3,000 budget, scaling to 100-150 once the prompts and filtering are reliable. The limiting factor is the capacity to analyze results and act on them.

Does building a GenAI pipeline require a developer?

No. No-code tools like Make or Zapier can connect the API calls directly. For heavier automation, batch generation against custom models, basic Python or a freelance coder for 2-3 days covers it.

How long before the pipeline pays for itself?

Setup runs $200 and generation another $50 a week, so the pipeline has cost $400 by the end of the first month. One winning creative clearing $400 in profit pays for all of it, and at 20-30 creatives a week that usually lands in the third or fourth week.

What about platform policy compliance?

GenAI creatives trigger moderation flags more often than hand-made ones. Every output needs to run through the platform's own compliance checker - Meta's ad policy tool, Google's ad approval - with a human reviewing before anything launches.

I can do this on your product

I consult on acquisition, funnels and retention - including hard verticals.

Ioann Putevoy
Ioann Putevoy
Head of Traffic & growth lead. I build products and take them to market - see the portfolio.

Bring me a product that needs to find its market