# Incrementality Testing for B2B SaaS

> How to run geo holdouts, PSA tests and switchback tests when you have hundreds of conversions a month, not millions, with a worked test design and sample size math.

Source: https://saas-marketing.net/playbooks/incrementality-testing/
Topic: SaaS Metrics and Analytics
Type: playbook
Published: 2026-09-11
Last updated: 2026-09-11
Publisher: SaaS Marketing (saas-marketing.net)
License: CC BY 4.0. Quote or republish with attribution and a link to https://saas-marketing.net/playbooks/incrementality-testing/

## Short answer

Incrementality testing measures what a marketing channel actually caused, by withholding it from a randomised group and comparing outcomes. B2B SaaS teams have hundreds of conversions a quarter rather than millions, so tests must run on upper-funnel outcomes like leads or opportunities, use audience or geo holdouts rather than fine-grained splits, and accept wider confidence intervals. The branded search pause is usually the first test worth running.

## Key takeaways

- Test on leads or opportunities, not closed-won, because closed-won volume in B2B SaaS is too small to reach significance in a quarter.
- A branded search holdout typically shows that a large share of those conversions would have happened anyway.
- Geo holdouts need roughly 20 comparable metros and eight weeks to detect a 20 percent effect at B2B volumes.
- A null result is a finding, and treating it as a failed test is how teams keep funding channels that do nothing.
- Audience holdouts on LinkedIn are the easiest test to run because the platform supports exclusion natively.
- Write the analysis plan before the test starts, including the decision rule, or the result becomes a negotiation.

---

Every guide to incrementality testing was written for a company doing 40,000 conversions a month. Yours does 120 opportunities a quarter. The standard advice, run a two-week geo holdout and read the lift, produces a number with a confidence interval so wide you could drive a budget through it.

This playbook is the B2B version: which test designs survive at low volume, the sample size arithmetic worked out, the three tests to run first, and how to write up a result that finance will accept.

## Why does standard incrementality advice break in B2B SaaS?

Because the statistics need volume and B2B does not have it. The arithmetic is unforgiving and it is worth seeing directly.

To detect a 20 percent lift with reasonable confidence, you need something in the region of a few hundred conversions per arm. A company closing 40 deals a quarter cannot get there in a year, let alone a test window. So the standard design fails, and teams conclude incrementality testing is for consumer companies.

The way out is to move up the funnel. You may have 40 closed-won deals but 120 opportunities, 600 MQLs and 4,000 form fills. Measuring on opportunities gives you three times the volume, and measuring on leads gives you fifteen times. You lose some validity, because lead lift does not automatically become revenue lift, but you gain the ability to run a test at all.

Pick opportunities if you can tolerate only detecting large effects, and MQLs if you need more sensitivity. Never test on form fills alone, because a channel can inflate form volume with unqualified traffic and post a positive result while destroying pipeline quality. If you do test on leads, report opportunity rate alongside as a guardrail.

Any test measured on leads needs a lead-to-opportunity conversion rate reported for both arms. A 20 percent lead lift with a 25 percent drop in opportunity rate is a loss, and without the guardrail it looks like a win.

## Which four test designs work, and what does each need?

Four designs, with genuinely different volume requirements. Two of them are viable for almost any B2B SaaS company.

**Geo holdout.** Split markets into test and control, turn the channel off in control, compare. Needs enough comparable geos, which in practice means at least 20 metros with similar baseline conversion rates. Works well for the US market, poorly for companies whose pipeline is concentrated in three cities.

**Audience holdout.** Randomly exclude a percentage of your target audience from seeing the campaign, then compare conversion rates between excluded and exposed. LinkedIn supports this natively through audience exclusion, which makes it the easiest test to actually execute. The catch is that exposure is not random within the served group, so you are comparing eligible-and-served against eligible-and-withheld, which is close enough for most decisions.

**PSA ghost ads.** The control group is served an unrelated public service ad through the same auction and targeting, so both groups cleared the same selection process. It is the cleanest design available and support is uneven: strong in programmatic display, limited on LinkedIn.

**Switchback.** Alternate the channel on and off in time blocks, typically weekly, and compare. Cheap to run and it needs no audience infrastructure, but it is badly confounded by B2B sales cycles. If your cycle is 60 days, a week-on week-off pattern smears effects across blocks and the result is uninterpretable. Use it only for short-cycle, self-serve motions.

## How do you work out sample size at 120 opportunities a quarter?

Do the arithmetic before you design the test, because it tells you whether the test is worth running at all. Here is the worked case.

Assume a company generating 120 opportunities a quarter, evenly split across a channel mix, with paid social contributing roughly 30 of those. You want to test whether paid social is incremental.

Split 50/50 gives you about 15 opportunities per arm per quarter. That is nowhere near enough. So you extend the window to two quarters, giving 30 per arm, and you widen the holdout to 50 percent of audience rather than a small slice, which maximises contrast.

With 30 events per arm, the detectable effect at conventional confidence is roughly a 50 percent difference. That sounds useless, and for fine-tuning it is. For the question "is this channel doing anything at all", a 50 percent detection threshold is genuinely informative, because a channel producing less than half the conversions attribution claims is a channel you would reallocate.

**Sizing a test before you run it**

**30%** Roughly the smallest effect detectable when testing on 120 quarterly opportunities across two quarters

## Which three tests should you run first?

Branded search, retargeting, and a LinkedIn audience holdout. In that order, because that is the order of likely budget released per hour of effort.

**Branded search pause.** Turn off paid search on your own brand terms in half your geos, or entirely for four weeks if geos are not workable, and measure total branded conversions including organic clicks. What teams typically find is that a large share of paid branded clicks were cannibalising organic clicks that were free. The finding is rarely 100 percent non-incremental, because competitor bidding on your brand is real and there are navigational queries where you genuinely need the top slot, but the incremental share is usually far below what the platform reports.

This is the highest-value test in B2B SaaS for one reason: branded search is often among the largest single line items in the paid budget and it is almost never questioned, because its reported ROAS looks spectacular. It looks spectacular because it is measuring people who already decided to buy.

**Retargeting holdout.** Exclude a random half of your retargeting audience. Retargeting reports beautifully for the same reason branded search does, and the incremental share is usually a fraction of the attributed share. This test is easy because retargeting audiences are already defined in the platform.

**LinkedIn audience holdout.** Exclude 50 percent of your matched account list from the campaign for a full quarter, then compare opportunity creation between the two halves. This is the test that tells you whether your ABM paid layer is doing anything, and it interacts directly with what your [paid media attribution](/guides/paid-media-attribution-for-saas/) is claiming.

Not during a product launch, a conference, or the quarter your PR agency finally lands something. Any brand-demand shock during the window confounds the result, and you will have to run it again.

## How do you read a null result without lying to yourself?

Report the confidence interval, not the point estimate, and state explicitly what effect sizes the test could and could not rule out.

A null result at low volume means one of two things: the effect is small or zero, or the test lacked power. Those are different conclusions and the interval tells you which. If your interval on lift runs from minus 10 percent to plus 55 percent, you have learned almost nothing. If it runs from minus 5 percent to plus 15 percent and your decision threshold was 40 percent, you have learned enough to cut.

The dishonest moves to avoid, all common. Extending the test because the result was not what you wanted. Switching the outcome metric after seeing the data. Excluding a geo that "had an anomaly". Splitting the result by segment until one segment shows significance. Each of these turns a test into a search for a comfortable number, and the people who run them usually believe they are being rigorous.

Concluding a channel works because the test was inconclusive. Inconclusive is not a pass. If you could not detect the effect that would justify the spend, the honest default is to reduce the spend and test again at a level where the effect would be visible.

Incrementality also has a real limitation worth admitting. It measures the marginal effect of the spend level you tested, in the period you tested it, for the audience you tested on. It says nothing about long-term brand effects that accumulate over years, which is where [marketing mix modeling](/guides/marketing-mix-modeling-for-saas/) and [brand lift](/glossary/brand-lift/) measurement do work that experiments cannot. Anyone who tells you incrementality is the complete answer is selling something.

## How do you document a result finance will accept?

One page, written before the test, filled in after. Finance rejects incrementality findings when they arrive as a slide with a lift percentage and no method.

**The one-page test record**

Two details make the difference with a CFO. First, defining the outcome metric by CRM field rather than by name, because "opportunity" means four different things inside most companies. Second, stating the confounds you know about rather than waiting to be asked, which converts the conversation from interrogation to review.

Keep the records in one place. A folder of six of these over two years is the most persuasive marketing measurement asset a team can own, and it outlives any dashboard. If you are still deciding between measurement approaches, the comparison of [multi-touch attribution vs incrementality testing](/comparisons/multi-touch-attribution-vs-incrementality/) lays out where each is honest, and the [marketing tracking plan template](/templates/marketing-tracking-plan/) covers the event definitions the tests depend on.

## What to do next

Pick the branded search test and size it this week. Count your branded paid conversions over the last two quarters, write the decision rule, and check whether four weeks off in half your geos gives you enough events.

If it does not, the test to run instead is the retargeting holdout, which usually has more volume and almost always has a surprising answer. Put the result into the [SaaS marketing dashboard template](/templates/saas-marketing-dashboard/) alongside the attributed numbers so the gap between the two is visible every month, and read [incrementality testing for B2B SaaS](/guides/incrementality-testing-b2b-saas/) for the conceptual background. The wider [SaaS metrics and analytics](/saas-metrics/) hub covers how this fits with the rest of the measurement stack, including [measuring lifecycle email in SaaS](/guides/email-marketing-attribution-saas/), which has the same low-volume problem and a similar solution.

## Frequently asked questions

### What is incrementality testing in B2B SaaS?

It is a controlled experiment that measures the causal effect of a marketing channel by withholding it from a randomly selected group and comparing conversion outcomes. Unlike attribution, which allocates credit among touchpoints that already happened, incrementality answers whether the spend produced conversions that would not otherwise have occurred.

### Can you run incrementality tests with low conversion volume?

Yes, with three adjustments. Measure on a higher-volume outcome such as marketing qualified leads or opportunities rather than closed-won revenue. Run longer, typically eight to twelve weeks rather than two. And accept that you can detect large effects, around 20 to 30 percent, but not small ones, which is usually enough to make a budget decision.

### What is a branded search incrementality test?

You pause paid search on your own brand terms for a defined period, or in a set of geos, and measure total branded conversions including organic. If total conversions hold roughly steady, the paid spend was mostly buying clicks that organic would have captured for free. Most B2B SaaS teams find a large share of branded paid conversions were not incremental.

### What is a PSA ghost ad test?

The control group is served a public service announcement or charity ad instead of your ad, so both groups go through the same auction and targeting process. It removes the selection bias where the control group is simply a different, less reachable set of people. Support varies by platform and is stronger on programmatic display than on LinkedIn.

### How long should a B2B SaaS incrementality test run?

Eight to twelve weeks for most channels, longer if your sales cycle exceeds 60 days. The test needs to cover at least one full cycle from first touch to the measured outcome, plus enough volume to reach a usable confidence interval. Two-week tests, standard in ecommerce, produce noise in B2B.

### What do you do with a null result?

Treat it as evidence, with the caveat that a null result at low volume often means the test could not detect the effect rather than that the effect is zero. Report the confidence interval, not just the point estimate. If the interval rules out the effect size that would justify the spend, that is a decision even though it is not statistical proof of zero.
