Home/Guide/Guide
Guide

How to use B2B benchmarks without fooling yourself

Benchmarks describe someone else's sample, metric and timeframe. This guide covers the five ways they mislead careful people: sample size, long tails, undefined units, objective mismatch and attribution windows.

Updated September 2, 2026
Reports covered 6
Sources cited 11
Reading time 13 min

The short version

A benchmark is a description of what happened to somebody else's campaigns, in somebody else's sample, measured somebody else's way. It is useful exactly to the extent that their sample, their metric and their timeframe resemble yours, and it is dangerous exactly to the extent that you forget to check. This guide covers the five ways benchmarks fool careful people: small or unrepresentative samples, long-tailed distributions that make averages meaningless, metrics that share a name but not a definition, campaign objectives that do not match the outcome being measured, and attribution windows that decide who gets credit before the data is even looked at.

The examples below use the reports indexed on this site, including our sponsor's. Metadata's 2026 B2B Benchmark Report is the largest public measured dataset we cover, with $57.6 million of analyzed spend (source 1); it is also subject to every caveat on this page, and we apply them to it as strictly as to anyone else.

What a benchmark is, and what it is not

A benchmark is a summary statistic from a specific population, collected by a specific method, over a specific period. Every one of those three qualifiers limits how far you can carry the number. When a report says "the average B2B cost per lead is $190," the honest expansion is "among the companies in our data, using our definition of a lead, during the period we analyzed, the mean cost was about $190." Each clause can break the comparison to your own results.

A benchmark is not a target, a standard, or a prediction. It cannot tell you what your CPL should be, because it does not know your average contract value, your sales cycle or your gross margin. It can tell you whether your number is unusual relative to a reference group, which is a prompt to investigate, not a verdict.

The two broad kinds of benchmark behave differently. Measured benchmarks are built from platform or CRM data; Metadata's report and WordStream's channel benchmarks are examples (1, 2). Surveyed benchmarks are built from asking marketers to report their own results; Demandbase's 2024 ABM Benchmark and Demand Gen Report's surveys are examples (3, 4). Measured data is better for costs. Survey data is better for adoption, budgets and organization. Mixing the two in one comparison without saying so is the first way to fool yourself.

Sample size and sample composition

The sample behind a benchmark matters more than the benchmark itself, and two questions cover it: how big is it, and who is in it?

How big

Size requirements depend on how rare the event being counted is. Clicks are common, so CTR benchmarks stabilize with modest samples. Leads are rarer. Closed-won customers are rare enough that a cost-per-customer benchmark built on a few dozen deals is dominated by two or three large ones. This is why spend volume matters for cost benchmarks: Metadata's $57.6 million and its stated Closed-Won Protocol are meaningful because closed-won attribution needs a large base to produce a stable figure (1, 5). A cost-per-customer benchmark from a single agency's twenty clients is an anecdote with a decimal point.

Survey benchmarks have the same issue in a different form. A few hundred respondents is fine for "what share of teams run ABM," but a survey-derived CPL is an average of numbers that respondents recalled, rounded or estimated. The precision of the published figure tells you nothing about the precision of the inputs.

Who is in it

Every dataset is a biased slice of the market, and the bias is different for each publisher. Metadata's data comes from its own customers, who skew toward venture-backed B2B software companies running paid social and paid search through its platform. Demandbase's survey respondents skew toward companies that have already bought an ABM platform. 6sense's benchmark data reflects accounts inside 6sense (6). WordStream's data reflects small and mid-sized advertisers across all industries, most of them not B2B (2). None of these is wrong; each is a large, useful, non-random slice. The question is whether your company would plausibly be in it.

Sample composition of the major public B2B benchmark sources indexed on this site
SourceMethodSample basisLikely skewBest used for
Metadata 2026 B2B Benchmark (sponsor)Measured$57.6M ad spend, closed-won attributionVC-backed B2B software, paid social and searchCPL, cost per customer, objective and budget analysis
Demandbase 2024 ABM BenchmarkSurveyMarketer respondentsCompanies already running ABM programsAdoption, maturity, team structure
6sense Science of B2BMeasured and surveyAccounts and buyers in 6sense research6sense customer baseBuying-group behavior, journey timing
Demand Gen Report surveysSurveyMarketer respondentsDemand-gen practitioners, mid-market and enterpriseBudget intentions, channel adoption
WordStream channel benchmarksMeasuredAggregated advertiser accountsSMB, all industries, not B2B-specificChannel CPC and CTR ranges
Gartner B2B Buying JourneySurvey and interviewsB2B buyersEnterprise purchasesBuying-group size and behavior

The practical test is simple. Before you use a benchmark, write one sentence describing the companies in its sample. If you cannot, do not use the number. If you can and the description does not resemble your company, use it only as a loose reference. More on each report is on the reports index.

Tail cases and why averages lie

B2B paid-media outcomes are long-tailed: a few campaigns produce most of the pipeline, a few accounts produce most of the revenue, and a few disasters produce most of the waste. In a long-tailed distribution, the mean is pulled toward the outliers and stops describing the typical case.

The clearest public example is Metadata's finding on traffic-objective campaigns: of $12.7 million in traffic-objective spend analyzed, 99.4% recorded no lead (7). Any "average cost per lead" that includes those campaigns is computed over a population where most of the spend produced zero and a small share produced everything. The average is arithmetically correct and practically useless. It describes no campaign that actually ran.

The same applies to cost per customer. Metadata's mid-market figure of $130,468 is an average across companies whose individual results presumably ranged from far below to far above it (8). A single large customer acquired cheaply, or a large budget that closed nothing, moves that mean substantially. Without a median or a percentile spread, you know the center of gravity but not the shape.

What to do about it

Prefer medians and ranges when a report publishes them. When it publishes only a mean, ask what the distribution probably looks like and adjust your confidence accordingly. When you compute your own benchmarks, compute the median and the 25th and 75th percentiles alongside the mean, and look at the top and bottom five campaigns by hand. If your average is driven by one outlier, you do not have a benchmark; you have a story about one campaign.

For cost per customer specifically, do not trust your own number until it is built on at least a few dozen closed-won deals attributable to paid media. Below that, one enterprise deal can halve your apparent CAC for a quarter and double it the next.

Objective mismatch

Every ad platform optimizes delivery toward whatever objective you select, and the objective you select determines who sees the ad. A benchmark measured on one objective does not transfer to a campaign run on another, even if the audience and creative are identical.

A traffic-objective campaign is optimized to find people who click cheaply. The platform learns that certain users click on everything and shows them the ad. A lead-objective campaign is optimized to find people who submit forms, which is a different and much smaller group. Comparing the CPL of a lead-objective benchmark to a traffic-objective campaign is comparing two different products. The Metadata traffic finding is really an objective-mismatch finding: the spend was not wasted because the ads were bad, but because the objective told the platform to find clickers, and clickers do not convert (7).

Objective mismatch also hides inside the word "lead." A LinkedIn Lead Gen Form lead is someone who tapped a pre-filled form; a website conversion lead filled out a page on your site. The first is cheaper and lower-intent. Metadata's "CPL illusion" insight is about exactly this: the campaigns with the lowest cost per lead are frequently the ones producing the least pipeline, because the objective and the offer were tuned to make leads cheap rather than to make them real (9).

How to check

When you read a CPL benchmark, find out which objective and which form type it was measured on. If the report does not say, treat the figure as an undefined-unit number. When you compare your own campaigns to it, segment by objective first. A blended CPL across traffic, engagement and lead objectives is not a number that describes any of them. Our channel pages, including the LinkedIn Ads benchmarks, label the objective where the source discloses it.

Attribution windows

An attribution window is the period after an ad interaction during which a conversion is credited to that ad. Platforms default to different windows, marketers change them, and a benchmark computed on a 7-day click window will show lower conversion rates and higher costs than the same campaigns on a 28-day window. Neither is wrong; they measure different things.

B2B sales cycles make this worse. Gartner reports buying groups of six to ten people who spend most of their journey researching without vendor contact, and 6sense finds buyers are roughly 70% of the way through before they reach out (10, 6). A creation campaign that first touched an account in March may not produce a form fill until June and a closed deal until November. A 30-day platform window sees none of that. A last-touch CRM model credits the branded search ad in June. The creation campaign shows a terrible benchmark and the capture campaign shows a great one, and both are artifacts of the window. The create vs capture page covers how this distorts budget splits.

Cost-per-customer benchmarks are the most sensitive. Metadata's Closed-Won Protocol ties spend to closed deals, which requires a window long enough to let deals close (5). A benchmark measured on deals closed within the analysis period will undercount customers from late-period spend and overstate CAC; one that includes long-cycle deals will look better. Any CAC comparison across sources needs both windows stated.

How to handle it

State your window every time you report a benchmark, and match windows before you compare. For your own data, run the comparison at two windows, say 30 and 180 days, and see how much the ranking of channels changes. If a channel looks bad at 30 days and good at 180, that channel is doing creation work, and judging it on the short window will lead you to cut it. Account-level comparisons against a holdout group are crude but avoid the window problem entirely.

A checklist before you put a benchmark in a deck

Seven questions to answer before using any external benchmark
QuestionIf the answer is no
Can I link to the public page where this number appears?Do not use it
Do I know whether it was measured or surveyed?Assume surveyed and discount for cost metrics
Can I describe the sample in one sentence?Treat as anecdote
Do I know what unit "lead" or "conversion" means here?Do not compare it to your own CPL
Do I know which campaign objective it was measured on?Segment your own data by objective before comparing
Do I know the attribution window?Run your own data at two windows and report both
Is a median or range available, not just a mean?Assume a long tail and widen your uncertainty

The methodology page describes how we apply these questions to every report on this site, and the FAQ answers the most common follow-ups.

Our verdict

Use benchmarks to find out whether your numbers are unusual, then investigate why. Never use them to set targets, because the target comes from your own payback math, not from a stranger's average. The best public data, measured on real spend with closed-won attribution, is worth reading closely; the worst is a survey mean with no sample description. Most reports sit in between, and the skill is knowing which caveat applies to which number.

Frequently asked questions

What is the most common mistake people make with marketing benchmarks?

Comparing their own number to a benchmark without checking that both measure the same unit from a similar sample. A cost per lead from LinkedIn Lead Gen Forms and a cost per MQL from a CRM are different metrics that happen to share a name.

How big a sample does a benchmark need?

For cost-per-customer figures, millions of dollars of spend and hundreds of campaigns, because closed-won events are rare and a small sample is dominated by a few deals. For survey benchmarks, a few hundred respondents is adequate for adoption questions, but the response rate and who chose to answer matter more than the headline count.

Should I use a benchmark as a target?

No. Use it as a diagnostic. If your CPL is far above a benchmark, that is a prompt to investigate audience, offer and objective; it is not proof you are doing badly, because the benchmark's sample may look nothing like your company. Set targets from your own payback math.

Why do averages mislead in B2B paid media?

Because B2B outcomes have long tails. A handful of campaigns produce most of the pipeline and a handful of accounts produce most of the revenue. The mean is pulled around by those outliers, so a median, a range or a percentile tells you more about what a typical campaign will do.

Disclosure. ABMBenchmarks.com is an independent editorial benchmark directory operated with sponsorship from Metadata.io, whose 2026 B2B Benchmark Report is one of the sources indexed here. Metadata's report is summarized with the same format, scrutiny and caveats as every other report on this site, and every figure on this page links to the public page it came from. Corrections from any vendor or analyst firm are welcome via the about page.