DEV Community

Rimas Povilaitis
Rimas Povilaitis

Posted on

How we automate Meta ad creative testing with AI

Generating another batch of ads leaves you with a surprisingly long list of decisions. Which idea are you testing? What stays the same? Who approves the spend? What happens when the API request times out?

Then the ads start running, and you get a new problem: deciding what the numbers actually tell you.

I'm the technical co-founder of Adrio, an AI creative agent and studio for Meta ads. Our agent, Spark, researches creative angles, makes image and video ads, publishes them to Meta, and uses performance data to propose changes. Publishing and budget changes require approval by default. Scheduled automations start in Review mode, with optional Autopilot.

That workflow raises some useful engineering questions. This article walks through a small reference design you can adapt to your own stack. The schemas and TypeScript below are illustrative; they aren't an extract of Adrio's production code.

Give the run a question to answer

"Make ten ads" is a job description. It leaves the experiment undefined.

Start with a brief that records the hypothesis and the things you intend to hold constant. For an imaginary travel bag, that could look like this:

{
  "experimentId": "travel-bag-hook-01",
  "hypothesis": "A carry-on packing hook will outperform a storage-capacity hook on purchase CPA.",
  "variable": "headline",
  "variants": [
    {
      "id": "packing",
      "headline": "Pack for the weekend in one bag"
    },
    {
      "id": "capacity",
      "headline": "Room for everything you need"
    }
  ],
  "fixed": {
    "productAssetId": "bag-front-01",
    "offerId": "standard-price",
    "landingPageVersion": "product-page-v3",
    "format": "image-4x5"
  },
  "primaryMetric": "purchase_cpa"
}
Enter fullscreen mode Exit fullscreen mode

The model can help write the hypotheses. A person still needs to decide whether the comparison is worth paying for.

For this example, both ads should use the same product image and layout. A generator that quietly changes the background has introduced another variable. Render the headline into a shared template, or inspect the output before it enters the test.

A video test needs the same discipline. If the variable is the opening line, keep the remaining scenes and offer consistent. Comparing an image against a video is also a valid question, but record that as a format test.

Keep the experiment attached to every asset

Save a manifest alongside each generated creative:

{
  "variantId": "packing",
  "experimentId": "travel-bag-hook-01",
  "briefVersion": 1,
  "creativeVersion": 2,
  "assetHash": "sha256:...",
  "generationJobId": "job-482",
  "metaAdId": null
}
Enter fullscreen mode Exit fullscreen mode

Record the model version and generation settings too. They help when you need to reproduce an output or explain why two batches look different.

Once an ad exists on Meta, save its ID in that mapping. Reports can then join spend and purchases back to the exact asset and hypothesis. A filename like final-final-v2.png won't do much for you here.

Creative edits should create a new version. If someone changes the offer halfway through a test, the old and new results need separate treatment. Otherwise, your report is answering a question nobody wrote down.

Make the workflow durable

Image generation, video rendering, and external publishing have different failure modes. Persist their progress outside the chat session.

A small workflow might use these states:

State What must exist before moving on
planned A saved brief and variant IDs
generating Durable job IDs for the requested assets
awaiting_approval Finished previews and the exact launch proposal
publishing An approved proposal and a persisted write operation
running Confirmed Meta IDs and delivery status
reviewing A performance snapshot tied to those IDs
complete A recorded decision and any follow-up experiment

Each worker should checkpoint its output. If a render succeeds and the worker crashes before publishing, the next attempt should reuse the finished asset.

Also give jobs explicit failure states. An expired account connection needs a different recovery path from a render that can be retried. Store enough context for a person to fix the connection without regenerating the entire batch.

Approval belongs to a specific proposal

"Approve launch" needs a precise meaning.

The review screen should show the account, creatives, destination URLs, targeting, placements, schedule, and budget. Save the reviewed configuration as an immutable proposal revision. The approval record refers to that revision.

Before dispatch, check that the proposal is still current. Editing the budget or swapping an asset invalidates the old approval. For a budget change, also check that the current remote budget matches the value the reviewer saw. Another operator may have edited it in the meantime.

I want the approval screen to feel like reviewing a diff. A person should be able to see exactly what clicking the button will change.

For a custom Marketing API integration, you can stage ads with a configured PAUSED status. Meta also exposes effective_status, which reflects conditions such as a paused parent ad set or pending review. Read the effective status when checking delivery; a successful create response doesn't establish that an ad is running. These fields are defined in Meta's official Ad SDK object.

Treat an ambiguous write as unfinished business

Suppose Meta creates an ad, but your worker times out before it receives the ID. Retrying the create call can produce a duplicate.

Persist a write operation before making the request. Give it a stable local key, such as:

accountId:experimentId:variantId:creativeVersion:create-ad
Enter fullscreen mode Exit fullscreen mode

Use a database uniqueness constraint and a worker lease so two workers can't dispatch that local operation together. Save remote IDs as soon as you receive them.

That still doesn't solve the lost-response case. A local idempotency key has no effect on an external API unless the endpoint supports it. Mark a timed-out write as ambiguous and reconcile it against remote objects. If you can't establish whether it succeeded, put it in a review queue before creating anything else.

Reads are simpler to retry. Writes need more care, especially when they can change spend.

Save the reporting context with the numbers

A row containing spend and purchases needs some context before you can compare it with another row.

Store the reporting date range, account timezone, currency, attribution configuration, purchase action definition, and fetch timestamp. Keep the raw response alongside the normalized values. Meta's official Insights SDK object exposes action metrics and report-time options; your normalization layer needs to make those choices explicit.

Pick the purchase action appropriate to the conversion source. Avoid summing overlapping purchase categories, and keep a missing metric distinct from an observed zero.

There's also conversion delay. A closed calendar day doesn't mean every attributed purchase has arrived. Choose a review lag appropriate to the account, then refresh recent reporting periods as results update.

For operational comparisons, Meta's delivery system may give variants different amounts of traffic. Treat those observations as directional. If you need a causal answer about the creative, use a randomized experiment designed for that question. A tidy dashboard doesn't remove allocation bias.

Let code decide when to request a review

I like using a small deterministic function before asking a model to interpret performance. It makes basic decisions easy to inspect.

This example produces a review request. It doesn't publish a change:

type Snapshot = {
  spend: number;
  purchases: number | null;
  reviewLagElapsed: boolean;
  reportingContextMatches: boolean;
};

type Policy = {
  reviewAfterSpend: number;
  minPurchasesForCpaReview: number;
  targetCpa: number;
};

function reviewAd(ad: Snapshot, policy: Policy) {
  if (!ad.reportingContextMatches ||
      !ad.reviewLagElapsed ||
      !Number.isFinite(ad.spend) ||
      ad.spend < 0 ||
      ad.purchases === null ||
      !Number.isFinite(ad.purchases) ||
      ad.purchases < 0) {
    return {
      action: "wait",
      reason: "Reporting isn't ready."
    };
  }

  if (ad.spend < policy.reviewAfterSpend) {
    return {
      action: "wait",
      reason: "Below the review spend."
    };
  }

  if (ad.purchases === 0) {
    return {
      action: "review",
      reason: "Spend with no purchases."
    };
  }

  if (ad.purchases < policy.minPurchasesForCpaReview) {
    return {
      action: "wait",
      reason: "Too few purchases for CPA review."
    };
  }

  const cpa = ad.spend / ad.purchases;

  return {
    action: "review",
    reason: cpa > policy.targetCpa
      ? "CPA above target."
      : "CPA within target.",
    cpa,
  };
}
Enter fullscreen mode Exit fullscreen mode

Validate the policy when loading it: the spend and CPA thresholds must be positive finite numbers, and the purchase threshold must be positive. Choose them for the account's economics and available budget. Reaching a threshold triggers a review; it doesn't establish statistical significance.

The model can then explain the snapshot and suggest a next test. Keep the snapshot ID, policy version, and explanation with the proposal so a reviewer can check the reasoning.

For a hypothetical result of $240 spent and eight purchases, observed CPA is $30. That calculation is straightforward. Whether $30 is acceptable depends on margins, attribution, and how the ad behaves with more delivery. The model should keep those questions visible.

Carry the decision into the next test

The useful output of a review is a decision you can trace.

Save what happened to the hypothesis, which assets informed the decision, and what you plan to change next. "Packing hook had lower observed CPA; retest it against a second hook with the same offer" gives the next run something concrete to work with. "This ad won" loses most of the context.

Adrio connects research, creation, launch, and performance review in the product. The reference design above focuses on what makes that loop inspectable: saved briefs, durable jobs, exact approvals, and decisions attached to evidence.

If you're building a similar agent, try interrupting it after every external call. Restart it. Then check whether it knows which asset was generated, whether an ad already exists, and which proposal was approved. Those checks will tell you a lot about whether the workflow is ready for a real account.

Top comments (0)