Orcool

Strategy · Creative infrastructure

The real fight in AI advertising is not the model.

The short version

  • Most AI-ad pitches lead with the generation model: which video engine, which avatar, which lip-sync. That layer is the one that becomes a commodity fastest, because any vendor can call the same model the week it ships.
  • In our own 1,061-ad durability panel, the format teams ship the most, dynamic and automated, held for a median of just 33 days at 57% of total volume, the shortest lifespan of any format measured.
  • In a separate blind ABCD scoring test, an ad built from real user language scored 0.76 against 0.49 for a polished single-tool competitor. The gap was attention and connection, not render quality.
  • Both results point the same direction: the layer that decided the outcome was what fed the model, and what happened after it shipped, not the model itself.

Everyone is racing the middle layer

Ask most AI ad tools what makes them different and you get a model name. A newer video engine, a more realistic avatar, a faster render queue. It is an easy pitch, because it is demoable in thirty seconds and it changes every few months, which makes for a good headline.

It is also the layer with the least defensibility. A model is an API call. Whatever advantage a faster or cleaner generator gives one vendor this month, a competitor gets by switching providers the next. Nobody keeps a proprietary hold on how good AI video generation is, because nobody built it in house.

Three layers, one race

Strip an AI ad down and there are really three layers stacked on top of each other, and only one of them is the model.

signal what to build from  →  generation the model, the render  →  durability does the winner hold

Signal is the brief before the brief: what real users already say, what competitors are already running, which angle a category has not tried yet. Generation is the part everyone photographs, the actual video coming out. Durability is what happens after launch, whether the ad keeps earning its spend for two weeks or two months.

The middle layer gets the marketing budget. Our own numbers say the first and third layers are where the outcome actually gets decided.

What our own numbers already showed

We were not looking for this pattern when we found it. Two separate, unrelated tests both landed on the same conclusion from different angles.

The first was a durability read across 1,061 creatives in a single Meta panel, broken out by format (see the 6% problem). Static images held for a median of 101 days. Video held for 66. Dynamic, automated creative, the format built specifically to be generated fast and cheap at volume, held for a median of just 33 days, despite making up 57% of total volume shipped. The format most optimized for generation speed had the shortest life of the three.

The second was a blind scoring test on Google's ABCD framework (see your user reviews are the best ad brief you already have). Two AI-built ads for the same app, one built from a generic script through a polished single-tool generator, one built from real user reviews through an agent pipeline. The reviews-based ad scored 0.76. The polished one scored 0.49. Both were rendered competently. The gap was in the first three seconds and in whether the words sounded like something a person would actually say, not in how smooth the footage looked.

Neither result was about which model rendered the video. Both were about what fed the model going in, and in the durability panel, what happened to the ad once it was live.

The ad that wins is rarely the best rendered one. It is the one built from the right signal, and it stays a winner only as long as someone is actually watching whether it holds.

Why the model layer commoditizes first

This is not a complaint about generation models. They keep getting better, and that is good for everyone using them. It is the reason the middle layer cannot be a durable edge: improvements there spread to every buyer of the same API at the same time. Nobody owns a private version of a foundation model's video quality.

What is harder to copy overnight is a live, current picture of what a specific brand's users actually say, what a specific category's competitors are actually running this week, and a running measurement of whether last month's winner is still winning. Those three things change per brand, need constant updating, and do not ship as a product update from a model provider. That is the part of the stack worth competing on.

What to actually watch as a buyer

If a vendor's pitch stops at the model name, ask two different questions instead. First: where does the script come from before generation starts, and is it built from something a real user said or a real competitor is running, or from a generic prompt. Second: what happens to the ad after it goes live, does anything track whether it is still earning its spend in week three, or does the tool consider its job done at render.

Those two questions describe the signal layer and the durability layer. Our numbers say that is where the actual difference between a working ad and a wasted one gets made.

Frequently asked questions

Isn't a better generation model still an advantage?

A small one, and a short-lived one. Generation quality is set by whichever model a vendor calls, and any competitor can call the same model the week it ships. It is a real edge, just not a durable one, because it is not owned by the team using it.

So the generation model does not matter at all?

It matters, it just is not where the competition actually happens. In both of our internal tests, the ad with the better underlying script and signal beat the ad with the smoother render. The model was a delivery step, not the decision.

What should I actually ask an AI ad vendor if not which model they use?

Ask how the brief gets built before generation starts, and ask what happens after an ad goes live. Those two answers describe the signal layer and the durability layer, the two places our data shows the real gap sits.

How do you measure whether an ad's advantage is durable?

By tracking how long a creative keeps running before it gets pulled or fatigues, not just how it scores on day one. Our 1,061-ad panel measured exactly that, by format, and the results did not line up with which format looked the most produced.

Methodology note: the durability figures (101 / 66 / 33-day medians, 57% volume share) come from our internal read of a 1,061-ad Meta panel, first published in "the 6% problem." The ABCD scores (0.76 vs 0.49) come from a separate blind head-to-head test, first published in "your user reviews are the best ad brief you already have." Both are internal, single-panel results, not third-party audited studies. We are citing our own prior work rather than restating new figures here.

See your own signal and durability layer, in your chat

Add Orcool to Claude, ChatGPT, or Cursor. Paste one URL, sign in with a magic link, and your competitor hooks, category patterns and your own reviews come back as a tool call. Free for two weeks, no card.

Connect the MCP