Blog

The attention test: choosing a partner for ad, creative and message testing

· Fintan Smith

Whether the content is a social media ad, a piece of brand creative or a political message, it faces the same two gates: earn attention, then change minds. This article sets out what good testing looks like, compares nine providers, and offers five questions every buyer should ask.

Everything competes in the feed

A brand film, a static creative and a political message used to live in different worlds. They now compete in the same place: a social feed, against everything else in it.

That shared environment imposes a shared test. Every piece of content must first win a fraction of a second of attention, and only then gets the chance to persuade. The format differs. The problem does not.

The practical consequence for buyers is simple. A provider that measures persuasion but not attention, or attention but not persuasion, is testing half the problem. The right question is not "who tests ads?" or "who tests messages?" but "who can tell me whether this content will both get watched and change minds?"

Attention is the gating metric

The scarcest resource in modern communications is not budget or reach. It is attention.

The economics are stark. Attention research consistently finds that a large share of paid digital impressions receive little or no active attention at all, and that the ads which are looked at typically hold the eye for a second or two. Platforms have engineered feeds around this reality, ranking content on watch time, completion and interaction. An ad the algorithm judges unengaging is shown less and costs more per person reached.

The implication for testing is uncomfortable but unavoidable. A test that forces respondents to watch content measures a situation that never occurs in the market. Forced-exposure scores systematically flatter weak creative, because the hardest part of the job, earning the view, has been done for the respondent by the researcher.

The best-performing organisations therefore apply a two-gate standard to every piece of content:

  1. 1.Gate one: attention. Does the content earn engagement in a realistic environment, measured the way platforms measure it, with respondents free to scroll past?
  2. 2.Gate two: effect. Among those exposed, does the content cause a measurable shift in the outcome that matters, whether purchase intent, campaign support or voting intention, relative to a randomised control group?

Content that passes one gate but not the other fails in market. A persuasive message nobody watches has no reach. A watched ad that persuades nobody has no return. Only a test that measures both, on the same content, with the same sample, tells a leader what will happen when the budget is spent.

What good looks like: five criteria

In our experience advising political parties, government departments, agencies and brands, five criteria separate testing that changes decisions from testing that decorates them.

  1. 1.Causal evidence, not correlation. The methodological gold standard is the randomised controlled trial: respondents are randomly allocated to see the content or a control, and the difference between groups is the causal effect. Scores, norms and pre-post comparisons cannot distinguish what the content did from what the audience already believed.
  2. 2.Implied preferences over expressed ones. What people do is a better guide than what they say. Respondents overstate their attention, their interest and their intent to share, because an expressed preference is an answer given to a researcher, while an implied preference is a choice made when nobody appears to be asking. Good testing observes behaviour: whether people stop scrolling, how long they watch, what they skip, and what they choose when options carry real trade-offs. A survey question about whether someone would share the content is an opinion, not a measurement.
  3. 3.Speed matched to the decision cycle. Campaign iteration cycles are now measured in days. A test that reports in three weeks evaluates content the organisation has already replaced. Leading providers return results in hours, which changes testing from a compliance step into an iteration tool.
  4. 4.Comparability across markets. Multinational campaigns fail when each market tests differently. The same creative should run through the same design, with the same measures, in every market, with translation handled inside the platform rather than through a chain of local agencies. Anything else doubles cost and destroys comparability.
  5. 5.A feedback loop, not a verdict. Roughly half of tested content underperforms. The value of a test lies in what happens next. Best practice pairs the quantitative result with structured analysis of open-text feedback, so the audience itself explains why content worked or failed and informs the next version. The audience becomes a collaborator in the creative process rather than a scorecard at the end of it.

The provider landscape

Nine providers dominate the conversation for UK and international buyers. They cluster into three groups.

Predictive scoring platforms. System1 scores emotional response against a large historical database and is the established choice for brand advertisers benchmarking against category norms. DAIVID applies AI-predicted attention and emotion at scale and speed. Both produce a score rather than a causal estimate: they predict performance from patterns in past ads, they do not measure what a specific piece of content causes.

Attention and measurement specialists. Lumen Research leads on eye tracking and attention panels, showing precisely where people look and for how long. This is the strongest pure read on gate one, but it is measurement rather than experiment, and says nothing about persuasion.

Full-service and panel houses. Kantar's Link remains the most widely used copy test globally, with deep norms and pricing to match, suited to large FMCG advertisers on longer timelines. Ipsos offers Creative|Spark within a full-service agency, a sound choice for organisations wanting a single supplier. Zappi provides fast, automated concept screening at volume. YouGov combines a large panel with strong political data and rapid fieldwork, better suited to monadic testing and tracking than to multi-arm experiments.

Experimental providers. Swayable runs true RCTs on video and message content and reports causal lift, with a strong reputation in US political work and premium pricing. Convergent Opinion, our firm, is the only provider on this list built around both gates simultaneously: every piece of content, whether a video ad, a static creative, a message or a policy framing, receives a randomised controlled trial and a behavioural attention measure in a simulated feed, combined into a single Cut-Through Score, the causal shift multiplied by the engagement earned.

How Convergent approaches the two-gate problem

Four design choices define our model, and each maps to a criterion above.

First, the platform is self-serve. Clients upload content, define an audience and launch from their desk, with results in around four hours from £600 per piece of content. Researcher support is available but not required, which removes the agency bottleneck from the iteration cycle.

Second, attention is measured natively. We have reverse-engineered how platforms measure engagement, so content is scored on the same behavioural signals in TikTok-style, Facebook-style and X-style feed environments.

Third, the audience is built into the creative loop. Hundreds of viewers explain in their own words why content landed or fell flat. Our models read every response and return a plain recommendation with supporting verbatims, so the next version is written with the audience rather than at it.

Fourth, markets are comparable by design. Translation with human review is built into the platform at roughly eight hours per study, allowing the identical creative to be tested across 12 or more markets in one design with one set of measures. Bayesian modelling underpins the analysis throughout, so subgroup findings carry honest uncertainty rather than false precision.

Five questions for any provider

Leaders evaluating providers should ask:

  1. 1.Are you measuring implied or expressed preferences? Observed behaviour and real trade-offs, or ratings and claimed intentions?
  2. 2.How is attention measured? On the same behavioural signals platforms use, or through a proxy question?
  3. 3.What is the turnaround, and the marginal cost of another variant? Testing that cannot keep pace with iteration will be skipped under deadline pressure.
  4. 4.How does multi-market work? One design and one platform, or a re-brief in every country?
  5. 5.What do we get when content fails? A number, or an explanation strong enough to shape version two?

The organisations that answer these questions well test everything they publish, whatever its format, against the same two gates, attention and effect, and choose partners built to measure both.