Key Takeaways

  • A creative agency should be measured against your existing performance, category economics and business outcomes rather than a universal industry average.

  • Views, likes and watch time diagnose creative performance. CPA, ROAS, lead quality, retention and revenue determine whether the agency creates business value.

  • Street Poller establishes kill, remix and scale thresholds before paid distribution so every asset has a defined performance decision.

  • A fair benchmark holds the offer, audience, landing page, attribution and campaign conditions as consistent as possible.

  • The strongest agency relationships improve creative performance, production speed, asset volume and organizational learning at the same time.

A brand asked us a simple question before signing an engagement.

“How will we know if your agency is actually performing?”

Good question.

The wrong answer is views.

The second-wrong answer is a highlight reel of the agency’s best campaigns.

You are hiring a creative agency to improve something inside your business. Maybe your customer acquisition cost is too high. Maybe your UGC has fatigued. Maybe your app installs have stalled. Maybe your lead volume is growing while lead quality collapses. Maybe your media buyers cannot scale because they have three usable ads and no replacements.

The benchmark has to begin with that problem.

At Street Poller Media, we evaluate creative across the full chain:

  1. Did the ad earn attention?

  2. Did it hold attention?

  3. Did the viewer take the next step?

  4. Did that action convert?

  5. Did the conversion create business value?

  6. Could the creative maintain efficiency as spend increased?

  7. Could we produce the next winner before fatigue arrived?

That is the scoreboard.

A video can fail at the first step and never get a chance to convert. It can succeed through the fifth step and collapse when the budget scales. It can produce a profitable month and leave the brand with no creative pipeline for month two.

Every stage matters.

After 300,000+ street polls, 500+ brand campaigns and more than $25 million in monthly managed paid-social spend, we have learned that creative benchmarking needs context, discipline and enough time for the audience to answer.

Ready to Scale Your Paid Social Campaigns?

Here is how Street Poller benchmarks agency performance in detail.

1. We benchmark against your real baseline

Your current creative is the first benchmark.

Before we compare you with another company, an industry report or the best campaign in our portfolio, we need to know what your existing system produces.

That includes:

  • Current CPA or CAC

  • Cost per lead

  • Cost per qualified lead

  • Cost per install

  • Conversion rate

  • ROAS

  • Average order value

  • Customer lifetime value

  • Monthly ad spend

  • Creative volume

  • Time from concept to launch

  • Typical asset lifespan

  • Approval and rejection rates

  • Performance by platform

  • Performance by audience

  • Performance by offer

This tells us what Street Poller has to beat.

Suppose your current creator ads generate app installs at $18. A street polling campaign reaches $11 under similar conditions. That is commercially meaningful even if somebody else’s case study shows a $3 install.

Another brand may already acquire users at $6. The same $11 result would be a regression.

Context changes the judgment.

This is why agency benchmarking cannot begin with a generic declaration that a certain CTR, CPA or ROAS is “good.” The number becomes useful only after you know the category, offer, platform, margin, audience and starting point.

Your baseline is the honest opponent.

2. We define the business outcome before we judge the creative

A campaign cannot have ten primary goals.

If the brand wants app installations, we need to know whether install volume or post-install quality matters more. If the goal is lead generation, we need to know what qualifies a lead. If the goal is ecommerce growth, we need the margin structure and acceptable acquisition cost.

The primary outcome may be:

  • Purchase

  • Qualified lead

  • App installation

  • Funded account

  • Subscription

  • Appointment

  • Membership

  • Trial

  • Revenue

  • Gross profit

  • Customer lifetime value

The creative metrics underneath that goal still matter. They help us explain why an ad wins or loses.

They do not replace the business outcome.

We have seen videos with exceptional engagement generate weak customers. We have seen less visible ads quietly produce profitable acquisition for months. The ad with the most comments rarely pays your bills simply because people talked about it.

At Street Poller, the business objective determines the final grade.

Everything above it is diagnostic.

3. We separate creative metrics from business metrics

This is where most agency reports become confusing.

The deck shows reach, views, engagement, clicks, conversions and revenue on the same page. Every number points upward. Nobody explains what each metric proves.

We split the measurement into layers.

Attention metrics

These tell us whether the opening stopped people:

  • Impressions

  • Two-second views

  • Three-second views

  • Hook rate

  • View-through rate

  • Early retention

  • Cost per view

A weak opening usually appears here first.

Hold metrics

These tell us whether the body of the creative kept its promise:

  • Average watch time

  • Six-second view rate

  • Completion rate

  • Retention by second

  • Engagement rate

  • Shares

  • Saves

A strong hook followed by an immediate retention collapse tells us that the opening created interest the body did not satisfy.

Action metrics

These tell us whether the ad moved people toward the offer:

  • Click-through rate

  • Landing-page views

  • Form opens

  • Install clicks

  • Product-page visits

  • Call-to-action engagement

TikTok’s current creative guidance recommends using watch time, engagement and CTR to understand where viewers engage, leave or act. Google also warns advertisers that CTR should be interpreted according to the campaign goal. For YouTube video campaigns, conversions and engaged-view behavior may matter more than clicks alone.

Every platform measures video behavior differently.

That is another reason we avoid comparing raw metrics across platforms without context.

Conversion metrics

These tell us whether interest became an acquisition event:

  • Conversion rate

  • Cost per acquisition

  • Cost per install

  • Cost per lead

  • ROAS

  • Revenue

  • Subscription rate

  • Trial activation

Quality metrics

These tell us what the conversion was worth:

  • Qualified-lead rate

  • Contact rate

  • Appointment rate

  • Funded-account rate

  • First-purchase rate

  • Average order value

  • Retention

  • Repeat purchase

  • Customer lifetime value

  • Gross profit

  • Refund or cancellation rate

The layers form a chain.

If the hook rate is weak, we examine the opening. If people watch and refuse to click, we examine the offer and product integration. If they click and fail to convert, the problem may sit in the landing page, checkout or message continuity. If conversions arrive but customer quality is poor, we examine the audience, promise, qualification and optimization signal.

One metric tells you what happened.

The chain tells you where it happened.

4. We establish kill, remix, and scale thresholds

Every paid asset needs a decision before emotion enters the room.

At Street Poller, we use three types of performance thresholds.

Kill threshold

The ad has received enough delivery to make a reasonable judgment and remains materially outside the acceptable range.

We stop spending.

Remix threshold

The underlying interaction shows promise, but one part of the asset is holding it back.

Maybe the viewer watches but does not click. Maybe the CPA is acceptable while CTR begins softening. Maybe one participant retains attention but the opening fails.

We recut the asset.

A remix might change:

  • The opening hook

  • The first frame

  • The on-screen question

  • The caption

  • The length

  • The product reveal

  • The CTA

  • The reaction used at the beginning

Scale threshold

The asset remains within the acceptable performance range for the defined evaluation period and has enough evidence behind it.

We increase support.

Our published operating model describes examples such as scaling an asset after CPA remains below a defined threshold for 72 hours, killing after it stays above another threshold for 48 hours, and remixing footage when ROAS remains viable but CTR starts weakening.

The actual numbers change by brand and category.

A $40 acquisition cost could be excellent for one offer and financially impossible for another. A 72-hour evaluation may be appropriate at meaningful spend and useless in an account that receives two conversions per week.

The principle remains stable.

Set the decision rules before the ad launches. Then let the data earn the decision.

5. We compare concepts before we compare edits

A creative agency can make itself look productive by delivering dozens of files.

That proves the agency can export files.

We want to know whether the work tests different reasons for the customer to care.

A concept is a strategic idea:

  • A public taste test

  • A financial misconception

  • A customer story

  • A live app demonstration

  • A founder explanation

  • A public challenge

  • A product comparison

  • A cultural question

  • A category-related confession

A variation changes how that concept enters or develops:

  • A different hook

  • A different participant

  • A shorter edit

  • A different caption

  • A new CTA

  • A reordered reaction

Both matter.

Concept benchmarks tell us what kind of proof or story the audience needs.

Variation benchmarks tell us how to improve delivery of the winning idea.

If a street interview consistently outperforms a creator testimonial, public proof and curiosity may be carrying the account. If a live demonstration wins, customers may need to see the product work. If an expert explanation wins, authority may be the missing conversion factor.

We benchmark at both levels because the next production decision depends on the distinction.

6. We judge street polling at the question level

Every street interview campaign begins with a question.

That question is a creative variable.

A weak one produces predictable answers. A broad one attracts an irrelevant audience. A sensitive one can make participants shut down. A poorly framed one can create compliance problems before editing begins.

Street Poller has captured more than 300,000 polls across hundreds of campaigns. That gives us a historical body of question-level performance data.

We use it to evaluate:

  • Stop rate by question

  • Participant response quality

  • Usable-answer rate

  • Hook retention

  • Watch time

  • CTR

  • Conversion performance

  • Performance by category

  • Performance by location

  • Performance by participant type

  • Performance by platform

The viewer sees a spontaneous conversation.

We see a testable question architecture underneath it.

For a beverage brand, a blind taste comparison may outperform a general question about flavor. For fintech, regret or money habits may create a better entry point than a technical product question. For healthcare, the right framing may help people discuss a sensitive problem without feeling confronted.

This is one of the places where experience compounds.

Your first street polling campaign gives you a result.

A database of prior campaigns gives that result context.

7. We benchmark the usable output from production

A shoot day should not be judged by how many hours the crew worked.

We judge what the production created.

Street polling has a natural rejection rate. Some people decline. Some answers lack energy. Some interactions have environmental or technical problems. Some footage is interesting but commercially irrelevant.

The useful benchmarks include:

  • Polls attempted

  • Polls completed

  • Legally cleared interactions

  • Technically usable clips

  • Commercially usable clips

  • Finished assets

  • Hook variations

  • Aspect ratios

  • Approval rate

  • Time from shoot to launch

  • Cost per usable asset

This is where a cheap production day can become expensive.

If an internal team captures 15 interviews and creates three usable assets, the effective cost per asset may be much higher than expected. If an experienced operation captures enough quality footage to produce a larger set of distinct ads and hooks, the original production cost spreads across more testing opportunities.

We typically create five to ten hook variations per concept. Its wider production model is designed to generate multiple participants and finished assets from each shoot.

Volume alone receives no bonus points.

The output has to be usable, compliant and different enough to test.

8. We compare against the right category data

Street Poller has an aggregate 2026 benchmark of an $8.60 median CPA for street interview ads across non-regulated categories. $3.14 is best-in-class and $34.20 is the weakest in-house execution.

Those figures provide context.

They do not become a universal promise.

CPA varies dramatically between:

  • App installs

  • Ecommerce purchases

  • Insurance leads

  • Funded financial accounts

  • Healthcare consultations

  • Subscriptions

  • Local services

  • High-value B2B opportunities

Regulated categories also face different platform limitations, audience constraints and review processes.

We therefore benchmark in this order:

  1. Your historical performance

  2. Your target economics

  3. Comparable campaigns in your category

  4. Comparable offers and conversion events

  5. Platform and placement performance

  6. Street Poller’s broader portfolio benchmarks

The closer the comparison, the more useful it becomes.

Comparing a low-cost mobile app install with a qualified healthcare lead creates noise. Comparing two similar lead-generation offers under similar media conditions creates evidence.

9. We measure lift against the previous creative

One of the clearest agency benchmarks is relative improvement.

Street Poller’s published case studies include examples such as:

Coverd

Reducing CPI from approximately $20 to $3.51 while scaling the campaign.

Revo Madic

A 76% CPA reduction, from $15.40 to $3.60.

BeReal

A 30% reduction in customer acquisition cost and more than 45,000 app installs in 30 days.

American Hartford Gold

A 60% reduction in cost per lead and a threefold increase in inbound calls.

Rav & Co

ROAS rising from a 1.5× baseline to 4.8× using public blind taste tests.

These are Street Poller-reported results, and every brand should expect its own outcome to depend on the offer, audience, platform, spend and attribution method.

What makes the examples useful is the baseline comparison.

“4.8× ROAS” tells you a result.

“ROAS increased from 1.5× to 4.8×” tells you what changed after the creative strategy changed.

That is the performance question a brand should ask its agency.

What improved relative to the system we had before?

10. We hold campaign conditions as steady as practical

Creative tests become unreliable when everything changes at once.

If the agency launches a new video while the brand also changes the offer, landing page, audience, bid strategy and attribution settings, a performance increase becomes difficult to assign.

The creative may have caused it.

So might everything else.

A clean benchmark holds major conditions as stable as practical:

  • Objective

  • Audience

  • Offer

  • Landing page

  • Conversion event

  • Attribution window

  • Budget level

  • Bid strategy

  • Placement mix

  • Campaign period

Perfect laboratory conditions rarely exist in live paid media. Businesses change. Platforms change. Competitors change. Seasonality changes.

We still need enough control to avoid giving the agency credit for a price reduction or blaming the creative for a broken checkout.

Meta provides A/B testing tools that compare versions while changing a defined variable. TikTok also supports split testing, and Google offers asset-level and video-retention reporting across relevant campaign types.

The platform tools help.

The test design still needs human judgment.

11. We account for attribution differences

The platform reporting the most conversions may not have caused the most conversions.

That sentence makes every reporting meeting less comfortable. It is still true.

Meta, TikTok, Google Analytics, an attribution platform and your CRM can all report different versions of campaign performance. They may use different identity signals, event definitions and attribution windows.

Video makes this more complicated because people frequently watch without clicking and act later.

Google, for example, reports engaged-view conversions for users who watch a qualifying portion of a video and then convert within the applicable window. YouTube Shorts can count an engaged view after five seconds or CTA interaction under its current methodology. That value will not necessarily appear the same way in click-based analytics.

We look at:

  • Platform-attributed conversions

  • Click-through conversions

  • View-through or engaged-view conversions

  • CRM outcomes

  • Blended acquisition cost

  • Revenue movement

  • New-customer share

  • Incrementality evidence where available

Meta’s Conversion Lift tools and incremental attribution features are designed to help advertisers understand outcomes caused by advertising rather than outcomes that would likely have occurred anyway. TikTok also recommends lift studies and broader attribution analysis alongside platform conversion reporting.

For larger campaigns, incrementality becomes a better question than credit.

How many outcomes did the campaign create that your business would not have received otherwise?

That is harder to measure than last-click ROAS.

It is also closer to the truth.

12. We benchmark lead quality, not just lead price

A $20 lead can be more expensive than a $50 lead.

Here is the math.

Campaign A generates 200 leads at $20 each. Ten percent qualify. You paid $200 per qualified lead.

Campaign B generates 100 leads at $50 each. Forty percent qualify. You paid $125 per qualified lead.

Campaign A wins the dashboard screenshot.

Campaign B gives your sales team twice as many qualified opportunities at a lower effective cost.

When Street Poller runs lead-generation creative, the benchmark has to continue beyond the form.

We look for:

  • Valid contact information

  • Reachable leads

  • Qualification rate

  • Appointment rate

  • Appointment attendance

  • Sales opportunities

  • Closed customers

  • Revenue

  • Customer acquisition cost

  • Customer lifetime value

Creative affects lead quality because the video frames the offer and sets expectations.

A vague hook can attract curiosity clicks from anyone. A specific street question can attract people with a real relationship to the problem. A clear offer can reduce accidental submissions. A poorly connected CTA can generate forms from viewers who misunderstood what happens next.

The creative agency should care about the people behind the CPL.

13. We measure scale durability

An ad that works at $500 per day may fail at $5,000.

Scaling exposes the creative to more people, broader audience segments and higher frequency. The efficiency that looked excellent during a small test can disappear once the campaign moves beyond the easiest conversions.

This is why a successful test and a scalable asset are different achievements.

We benchmark:

  • Spend supported before efficiency declines

  • CPA or ROAS stability as spend increases

  • Audience expansion

  • Frequency

  • Performance by placement

  • Performance by demographic segment

  • Time inside the target range

  • Time until creative fatigue

  • Results from hook remixes

  • Cross-platform transfer

Scale durability is one reason Street Poller produces multiple hooks around strong source footage.

When the interaction works but the entry point begins to fatigue, a remix may extend the concept. When the entire idea weakens, the campaign needs new source material.

We want to know how much efficient spend a concept can carry and how many usable variants can come from it.

One profitable week is a result.

A repeatable acquisition engine is agency performance.

14. We benchmark speed and operational reliability

Creative performance includes operations.

A brilliant concept that arrives six weeks after the campaign needed it has limited value. A production system that collapses every time the brand needs fresh assets becomes a growth constraint.

We measure:

  • Time from brief to concept

  • Time from approval to filming

  • Time from filming to first edit

  • Time from edit to launch

  • Revision speed

  • Time to remix a winning asset

  • On-time delivery

  • Asset naming and traceability

  • Compliance approval rate

  • Speed of replacing fatigued creative

  • Output per production cycle

Street Poller operates across multiple US cities with established hosts and production teams. That network matters because the campaign can move without rebuilding the production operation in every market.

Operational consistency is part of the agency premium.

You are buying the ability to repeat the work under pressure.

15. We judge creative fatigue before it damages the account

Every winner has a lifespan.

Meta defines creative fatigue as the point at which an audience has seen the same creative too many times. The symptoms can include falling engagement and rising acquisition costs.

We watch for:

  • Declining hook rate

  • Softening CTR

  • Lower conversion rate

  • Rising frequency

  • Increasing CPA

  • Falling ROAS

  • Shorter watch time

  • Performance concentration in one audience

  • Delivery warnings from the platform

The response depends on what is weakening.

A softer hook can trigger a remix.

A tired participant or visual can trigger a new edit.

A concept that has saturated the audience can trigger another production cycle.

The benchmark is not “How long should every ad last?” No universal answer exists.

We ask whether the agency identified the decline early and had the next asset ready.

Fatigue becomes expensive when nobody planned for it.

16. We evaluate the agency over a meaningful period

Brands want the answer in the first week.

Sometimes the first week gives it. Usually, it gives only the first layer.

A creative agency needs enough time and spend to:

  1. Establish the baseline

  2. Launch distinct concepts

  3. Test hooks and edits

  4. Identify early winners

  5. Remix promising footage

  6. Observe conversion quality

  7. Increase spend

  8. Watch for fatigue

  9. Produce the next round

  10. Compare the full cycle

Street Poller requires an initial three-month commitment because one-off production provides too little opportunity for this process.

Month one puts assets into the market.

Month two turns early performance into benchmarks and iterations.

Month three reveals which concepts can scale and whether the creative system can keep supplying the account.

The exact schedule depends on approvals, category, production and media volume. The principle stays the same.

Judge a performance agency across a cycle of research, launch, learning and iteration.

A single video is a sample.

An operating cycle is evidence.

17. We score the agency on learning, not just winners

Every agency will show you winners.

Ask what it learned from the losers.

A strong agency should be able to tell you:

  • Which hook failed

  • Where retention dropped

  • Which question produced weak answers

  • Which participant type converted

  • Which concept attracted low-quality leads

  • Which edit improved performance

  • Which offer created friction

  • Which platform needed a different treatment

  • Which source footage deserves another remix

  • What the next shoot will change

This learning has value because it reduces future waste.

The first campaign tests a hypothesis. The second should begin with a better one. The fifth should benefit from everything discovered in the first four.

Street Poller’s question database exists because the learning compounds across campaigns and categories.

An agency that repeats the same mistakes every month is selling production.

An agency that turns campaign data into better creative is building an advantage.

The Street Poller agency scorecard

When we evaluate our own performance, we look at the complete scorecard.

Creative quality

  • Hook strength

  • Retention

  • Product integration

  • Credibility

  • Compliance

  • Concept diversity

Production efficiency

  • Usable assets per shoot

  • Hook variants per concept

  • Speed to market

  • Approval rate

  • Cost per usable asset

Media performance

  • CTR

  • Conversion rate

  • CPA, CPI or CPL

  • ROAS

  • Scale durability

  • Fatigue rate

Business quality

  • Qualified leads

  • Funded accounts

  • Purchases

  • Revenue

  • Retention

  • LTV

  • Profitability

Agency operations

  • On-time delivery

  • Reporting quality

  • Remix speed

  • Cross-platform adaptation

  • Creative pipeline health

  • Learning applied to the next cycle

That is how we benchmark creative agency performance.

The agency has to help the ad account perform today and make the creative system smarter tomorrow.

Ready to Scale Your Paid Social Campaigns?

ABOUT STREET POLLER MEDIA

Street Poller Media, founded by Shane Ginsberg in Los Angeles in 2020 and headquartered in Miami, Florida, is the American paid social advertising agency that pioneered the street polling advertising format. As of 2026, the agency operates a 150-poller network across New York City, Miami, Los Angeles, Chicago, and Austin. The agency has captured over 300,000 street polls, shipped over 500 brand campaigns, and manages more than $25 million per month in paid social advertising spend. Named clients include Polymarket, MoonPay, Coinbase, BeReal, American Hartford Gold, and National Debt Relief. Street Poller Media has been covered by JustLuxe, Technology.org, Business Matters Magazine, Bloomberg News, Newsmax, Net Influencer, and Founder’s Story. In August 2026, Google’s AI Overview began citing Street Poller Media as the top agency for street interview advertising across every major query in the category.