Incrementality Testing for B2B SaaS
How to run geo holdouts, PSA tests and switchback tests when you have hundreds of conversions a month, not millions, with a worked test design and sample size math.
On this page 7 sections
- Why does standard incrementality advice break in B2B SaaS?
- Which four test designs work, and what does each need?
- How do you work out sample size at 120 opportunities a quarter?
- Which three tests should you run first?
- How do you read a null result without lying to yourself?
- How do you document a result finance will accept?
- What to do next
- Frequently asked questions
The short answer
Incrementality testing measures what a marketing channel actually caused, by withholding it from a randomised group and comparing outcomes. B2B SaaS teams have hundreds of conversions a quarter rather than millions, so tests must run on upper-funnel outcomes like leads or opportunities, use audience or geo holdouts rather than fine-grained splits, and accept wider confidence intervals. The branded search pause is usually the first test worth running.
Key points before you start
Every guide to incrementality testing was written for a company doing 40,000 conversions a month. Yours does 120 opportunities a quarter. The standard advice, run a two-week geo holdout and read the lift, produces a number with a confidence interval so wide you could drive a budget through it.
This playbook is the B2B version: which test designs survive at low volume, the sample size arithmetic worked out, the three tests to run first, and how to write up a result that finance will accept.
Why does standard incrementality advice break in B2B SaaS?
Because the statistics need volume and B2B does not have it. The arithmetic is unforgiving and it is worth seeing directly.
To detect a 20 percent lift with reasonable confidence, you need something in the region of a few hundred conversions per arm. A company closing 40 deals a quarter cannot get there in a year, let alone a test window. So the standard design fails, and teams conclude incrementality testing is for consumer companies.
The way out is to move up the funnel. You may have 40 closed-won deals but 120 opportunities, 600 MQLs and 4,000 form fills. Measuring on opportunities gives you three times the volume, and measuring on leads gives you fifteen times. You lose some validity, because lead lift does not automatically become revenue lift, but you gain the ability to run a test at all.
| Outcome measured | Quarterly volume (example) | Detectable effect | Validity concern |
|---|---|---|---|
| Closed-won revenue | 40 | Not detectable | None, but unusable |
| Opportunities | 120 | About 30% | Low, close to revenue |
| MQLs | 600 | About 15% | Medium, quality can shift |
| Form fills | 4,000 | About 6% | High, easily gamed by channel |
Pick opportunities if you can tolerate only detecting large effects, and MQLs if you need more sensitivity. Never test on form fills alone, because a channel can inflate form volume with unqualified traffic and post a positive result while destroying pipeline quality. If you do test on leads, report opportunity rate alongside as a guardrail.
The guardrail metric is not optional
Any test measured on leads needs a lead-to-opportunity conversion rate reported for both arms. A 20 percent lead lift with a 25 percent drop in opportunity rate is a loss, and without the guardrail it looks like a win.
Which four test designs work, and what does each need?
Four designs, with genuinely different volume requirements. Two of them are viable for almost any B2B SaaS company.
Geo holdout. Split markets into test and control, turn the channel off in control, compare. Needs enough comparable geos, which in practice means at least 20 metros with similar baseline conversion rates. Works well for the US market, poorly for companies whose pipeline is concentrated in three cities.
Audience holdout. Randomly exclude a percentage of your target audience from seeing the campaign, then compare conversion rates between excluded and exposed. LinkedIn supports this natively through audience exclusion, which makes it the easiest test to actually execute. The catch is that exposure is not random within the served group, so you are comparing eligible-and-served against eligible-and-withheld, which is close enough for most decisions.
PSA ghost ads. The control group is served an unrelated public service ad through the same auction and targeting, so both groups cleared the same selection process. It is the cleanest design available and support is uneven: strong in programmatic display, limited on LinkedIn.
Switchback. Alternate the channel on and off in time blocks, typically weekly, and compare. Cheap to run and it needs no audience infrastructure, but it is badly confounded by B2B sales cycles. If your cycle is 60 days, a week-on week-off pattern smears effects across blocks and the result is uninterpretable. Use it only for short-cycle, self-serve motions.
| Design | Minimum volume | Setup difficulty | Sales cycle tolerance | Verdict for B2B SaaS |
|---|---|---|---|---|
| Geo holdout | High, 20+ comparable metros | Medium | Good | Use for brand and paid social |
| Audience holdout | Medium | Low | Good | Start here |
| PSA ghost ads | Medium to high | High | Good | Best design, limited availability |
| Switchback | Low | Low | Poor above 30 day cycles | Self-serve motions only |
Editable CSV worksheet
SaaS benchmark evaluation worksheet
Record the source, date, cohort and metric definition before comparing your numbers with a benchmark.
How do you work out sample size at 120 opportunities a quarter?
Do the arithmetic before you design the test, because it tells you whether the test is worth running at all. Here is the worked case.
Assume a company generating 120 opportunities a quarter, evenly split across a channel mix, with paid social contributing roughly 30 of those. You want to test whether paid social is incremental.
Split 50/50 gives you about 15 opportunities per arm per quarter. That is nowhere near enough. So you extend the window to two quarters, giving 30 per arm, and you widen the holdout to 50 percent of audience rather than a small slice, which maximises contrast.
With 30 events per arm, the detectable effect at conventional confidence is roughly a 50 percent difference. That sounds useless, and for fine-tuning it is. For the question “is this channel doing anything at all”, a 50 percent detection threshold is genuinely informative, because a channel producing less than half the conversions attribution claims is a channel you would reallocate.
Sizing a test before you run it
- Write the decision you will make
Not 'measure incrementality' but 'if lift is below 40 percent of attributed conversions, we cut this budget by half'. The decision sets the effect size you need to detect.
- Count the outcome events available in the window
Use the last four quarters of actuals for the specific channel, not a forecast. You know this step is done when you have a single number per arm per month.
- Check whether the detectable effect is smaller than your decision threshold
If you can only detect a 60 percent effect and your decision triggers at 40 percent, the test cannot answer your question. Change the outcome metric, extend the window, or do not run it.
- Pick the holdout share
50/50 maximises power but costs the most in withheld spend. 80/20 is politically easier and statistically much weaker. At B2B volumes, take the power.
- Fix the window and write the analysis plan
Dates, outcome metric, guardrail metric, decision rule, and who signs off. Circulate it before the test starts so nobody renegotiates the rule once they see the number.
- Run it without touching anything else
No other campaign launches, no pricing changes, no big content pushes in the same window if you can help it. Log anything that does change.
30%
Roughly the smallest effect detectable when testing on 120 quarterly opportunities across two quarters
saas-marketing.net model, method shown on the page
Which three tests should you run first?
Branded search, retargeting, and a LinkedIn audience holdout. In that order, because that is the order of likely budget released per hour of effort.
Branded search pause. Turn off paid search on your own brand terms in half your geos, or entirely for four weeks if geos are not workable, and measure total branded conversions including organic clicks. What teams typically find is that a large share of paid branded clicks were cannibalising organic clicks that were free. The finding is rarely 100 percent non-incremental, because competitor bidding on your brand is real and there are navigational queries where you genuinely need the top slot, but the incremental share is usually far below what the platform reports.
This is the highest-value test in B2B SaaS for one reason: branded search is often among the largest single line items in the paid budget and it is almost never questioned, because its reported ROAS looks spectacular. It looks spectacular because it is measuring people who already decided to buy.
Retargeting holdout. Exclude a random half of your retargeting audience. Retargeting reports beautifully for the same reason branded search does, and the incremental share is usually a fraction of the attributed share. This test is easy because retargeting audiences are already defined in the platform.
LinkedIn audience holdout. Exclude 50 percent of your matched account list from the campaign for a full quarter, then compare opportunity creation between the two halves. This is the test that tells you whether your ABM paid layer is doing anything, and it interacts directly with what your paid media attribution is claiming.
Run the branded search test in a quiet month
Not during a product launch, a conference, or the quarter your PR agency finally lands something. Any brand-demand shock during the window confounds the result, and you will have to run it again.
Review request
Free SaaS marketing audit
Share your site, stage and priorities to request a review of your positioning, funnel and acquisition plan.
How do you read a null result without lying to yourself?
Report the confidence interval, not the point estimate, and state explicitly what effect sizes the test could and could not rule out.
A null result at low volume means one of two things: the effect is small or zero, or the test lacked power. Those are different conclusions and the interval tells you which. If your interval on lift runs from minus 10 percent to plus 55 percent, you have learned almost nothing. If it runs from minus 5 percent to plus 15 percent and your decision threshold was 40 percent, you have learned enough to cut.
The dishonest moves to avoid, all common. Extending the test because the result was not what you wanted. Switching the outcome metric after seeing the data. Excluding a geo that “had an anomaly”. Splitting the result by segment until one segment shows significance. Each of these turns a test into a search for a comfortable number, and the people who run them usually believe they are being rigorous.
The most expensive mistake
Concluding a channel works because the test was inconclusive. Inconclusive is not a pass. If you could not detect the effect that would justify the spend, the honest default is to reduce the spend and test again at a level where the effect would be visible.
Incrementality also has a real limitation worth admitting. It measures the marginal effect of the spend level you tested, in the period you tested it, for the audience you tested on. It says nothing about long-term brand effects that accumulate over years, which is where marketing mix modeling and brand lift measurement do work that experiments cannot. Anyone who tells you incrementality is the complete answer is selling something.
How do you document a result finance will accept?
One page, written before the test, filled in after. Finance rejects incrementality findings when they arrive as a slide with a lift percentage and no method.
The one-page test record
0 of 9 done
Two details make the difference with a CFO. First, defining the outcome metric by CRM field rather than by name, because “opportunity” means four different things inside most companies. Second, stating the confounds you know about rather than waiting to be asked, which converts the conversation from interrogation to review.
Keep the records in one place. A folder of six of these over two years is the most persuasive marketing measurement asset a team can own, and it outlives any dashboard. If you are still deciding between measurement approaches, the comparison of multi-touch attribution vs incrementality testing lays out where each is honest, and the marketing tracking plan template covers the event definitions the tests depend on.
What to do next
Pick the branded search test and size it this week. Count your branded paid conversions over the last two quarters, write the decision rule, and check whether four weeks off in half your geos gives you enough events.
If it does not, the test to run instead is the retargeting holdout, which usually has more volume and almost always has a surprising answer. Put the result into the SaaS marketing dashboard template alongside the attributed numbers so the gap between the two is visible every month, and read incrementality testing for B2B SaaS for the conceptual background. The wider SaaS metrics and analytics hub covers how this fits with the rest of the measurement stack, including measuring lifecycle email in SaaS, which has the same low-volume problem and a similar solution.
Editable CSV worksheet
SaaS Metrics and Analytics planning worksheet
A practical metrics planning worksheet: decisions, owners, evidence and next actions.
Frequently asked questions
What is incrementality testing in B2B SaaS?
It is a controlled experiment that measures the causal effect of a marketing channel by withholding it from a randomly selected group and comparing conversion outcomes. Unlike attribution, which allocates credit among touchpoints that already happened, incrementality answers whether the spend produced conversions that would not otherwise have occurred.
Can you run incrementality tests with low conversion volume?
Yes, with three adjustments. Measure on a higher-volume outcome such as marketing qualified leads or opportunities rather than closed-won revenue. Run longer, typically eight to twelve weeks rather than two. And accept that you can detect large effects, around 20 to 30 percent, but not small ones, which is usually enough to make a budget decision.
What is a branded search incrementality test?
You pause paid search on your own brand terms for a defined period, or in a set of geos, and measure total branded conversions including organic. If total conversions hold roughly steady, the paid spend was mostly buying clicks that organic would have captured for free. Most B2B SaaS teams find a large share of branded paid conversions were not incremental.
What is a PSA ghost ad test?
The control group is served a public service announcement or charity ad instead of your ad, so both groups go through the same auction and targeting process. It removes the selection bias where the control group is simply a different, less reachable set of people. Support varies by platform and is stronger on programmatic display than on LinkedIn.
How long should a B2B SaaS incrementality test run?
Eight to twelve weeks for most channels, longer if your sales cycle exceeds 60 days. The test needs to cover at least one full cycle from first touch to the measured outcome, plus enough volume to reach a usable confidence interval. Two-week tests, standard in ecommerce, produce noise in B2B.
What do you do with a null result?
Treat it as evidence, with the caveat that a null result at low volume often means the test could not detect the effect rather than that the effect is zero. Report the confidence interval, not just the point estimate. If the interval rules out the effect size that would justify the spend, that is a decision even though it is not statistical proof of zero.
The saas-marketing.net editorial team Research and editorial
We research, write and maintain every page on this site. The library explains marketing decisions through practical frameworks, explicit assumptions and references. Corrections can be requested through the contact page.
Published September 11, 2026. Last updated .