# Incrementality testing for B2B SaaS

> Run geo holdouts, channel pauses and matched market tests at low B2B volume, how long each takes to read, and when a simple mix model beats another tracker.

Source: https://saas-marketing.net/guides/incrementality-testing-b2b-saas/
Topic: SaaS Demand Generation
Type: guide
Published: 2026-09-11
Last updated: 2026-09-11
Publisher: SaaS Marketing (saas-marketing.net)
License: CC BY 4.0. Quote or republish with attribution and a link to https://saas-marketing.net/guides/incrementality-testing-b2b-saas/

## Short answer

Incrementality testing measures what a channel actually causes rather than what it gets credited for, by withholding it somewhere and comparing. B2B SaaS volume is too low for clean conversion based experiments, so the workable designs are channel pauses read against branded search and direct traffic, geo holdouts across matched metro sets, matched account holdouts for ABM, and campaign level offer splits. Read leading indicators over four to six weeks, not closed revenue.

## Key takeaways

- A four week channel pause read against branded search volume is the cheapest real experiment most SaaS teams can run.
- Geo holdouts need roughly 20 matched metros per arm before the comparison means anything at B2B volume.
- Measure leading indicators: branded search, direct sessions, demo requests, not closed won nine months out.
- A spreadsheet mix model on 24 months of spend and pipeline beats a $60k attribution platform for budget decisions.
- If a channel cannot survive a four week pause without a visible drop, it was harvesting demand rather than creating it.
- Present every result as a range with the test's power stated, or the first skeptic in the room wins.

---

Attribution tells you which touchpoint was standing nearby when someone converted. Incrementality tells you whether the conversion needed it. Those are different questions, and only one of them should drive a budget decision.

The objection every B2B team raises is volume. You have 60 demo requests a month, not 60,000 ecommerce transactions, so the textbook experiment designs do not apply. True. The designs below are the ones that still read at that scale.

## Which test designs actually work at B2B volume?

Four, and they suit different questions. Pick by what you are trying to decide, not by what your ad platform offers.

The fourth one is the odd entry. A creative split is a good experiment and a bad incrementality test, because both arms still receive the channel. It tells you which ad works better, never whether the channel works at all. Teams conflate these constantly.

### The channel pause, in detail

Turn a channel off entirely for four weeks. Watch branded search impressions, direct sessions, demo requests and qualified pipeline created. Then turn it back on and watch the recovery curve, which is often more informative than the drop.

You need a baseline first. Pull the prior 13 weeks of the same metrics, calculate the week over week variance, and write down the band inside which you will call the result null. Doing this before the test is the difference between an experiment and a story you tell afterwards.

Somebody switches the campaign back on in week two because pipeline looked soft. Get written agreement from the CRO before you start, including the specific number that would justify aborting. If you cannot get that agreement, you are not ready to run the test.

## What does pausing branded search usually reveal?

That most of it was not incremental. This is the most commonly run pause test in SaaS and the results are boringly consistent.

When you pause branded paid search, organic results typically absorb 60 to 90 percent of the lost clicks. Total branded sessions dip a few points. Demo requests often do not move at all outside normal variance. You were paying for clicks you already had.

Three conditions flip the result, and all three are worth checking before you cut the budget:

- A competitor is actively bidding on your brand terms. Check the auction insights report. If someone else has impression share above 10 percent on your name, pausing hands them the click.
- An AI Overview or a large SERP feature pushes your organic listing below the fold on brand queries. This has become common enough that the old assumption of guaranteed organic capture no longer holds.
- Your brand name is a common word. If you are called Monday or Linear, generic searchers and brand searchers blend, and the paid listing is doing real disambiguation work.

**60-90%** Share of paused branded paid clicks typically recovered by organic listings

I would run this test first at almost any SaaS company, because it is the one most likely to free budget without hurting anything. What you do with the freed budget is a separate question the [demand generation budget calculator](/calculators/demand-gen-budget-allocator/) can help frame.

## How do you size and time a test with 60 opportunities a month?

Work backwards from detectable effect. At low volume you cannot detect small effects, so decide up front the smallest lift you would act on and check whether the test can see it.

Rough rule: with weekly counts in the 10 to 20 range and typical B2B variance, a four week test can detect roughly a 25 to 30 percent change in that metric. It cannot detect 8 percent. If your hypothesis is that LinkedIn contributes 10 percent of demand, a pause test will not settle it and you should not pretend otherwise.

**Sizing a pause test**

That last step sounds bureaucratic. It is the only thing standing between you and a result everyone reinterprets to suit their existing position.

## What do you measure when revenue is nine months away?

Leading indicators with short latency, ordered by how quickly they respond.

| Indicator | Typical latency | Sensitivity | Use it for |
| --- | --- | --- | --- |
| Branded search impressions | 3-10 days | High | Awareness channels, podcast and display |
| Direct traffic sessions | 1-2 weeks | Medium | Broad reach campaigns |
| Demo requests | 2-6 weeks | High | Any bottom funnel channel |
| Qualified pipeline created | 4-10 weeks | Medium | The decision metric for most budgets |
| Self reported source mix | 4-8 weeks | Medium | Detecting channels tracking cannot see |
| Closed won revenue | 6-12 months | Low, within a test | Annual review, never a test readout |

Branded search is the workhorse. It moves fast, it is cheap to measure through Search Console and Google Ads impression data, and it is genuinely hard to fake. When a podcast sponsorship works, branded impressions move within a week of the episode dropping, which is the clearest [brand lift](/glossary/brand-lift/) signal available to a company with no budget for survey based measurement.

Pair the quantitative read with the qualitative one. Watching your self reported attribution mix shift during a pause is a second, independent read on the same question, and the two rarely disagree without reason.

Set up a weekly export of branded impressions, direct sessions and demo requests into one sheet six weeks before your first test. Half of failed incrementality programs fail because nobody had a clean baseline when the test started.

## Can you build a mix model in a spreadsheet?

Yes, and for most SaaS companies under $30M ARR it is a better use of a week than another tracking implementation. The output is not publication grade. It is directional, and directional is what a budget meeting needs.

What you need: 24 months of monthly spend by channel, monthly pipeline created, and monthly values for two or three controls such as headcount in sales, seasonality index and any major launch flags.

The build, in outline:

1. One row per month, one column per channel's spend, one column for pipeline created.
2. Apply adstock to each spend column: this month's effective spend equals this month's spend plus roughly 0.4 times last month's effective spend. That decay factor encodes the lag. Try 0.3, 0.5 and 0.7 and see which fits better.
3. Apply a saturation transform. Taking the square root or log of effective spend is crude and much better than nothing, because it stops the model claiming that tripling spend triples pipeline.
4. Run a linear regression of pipeline on the transformed columns plus controls. LINEST in Sheets or the Data Analysis pack in Excel will do it.
5. Test stability. Remove one month at a time, refit, and see whether any coefficient flips sign. Coefficients that flip are noise you should not spend against.

Twenty four data points and eight channels. That is three observations per parameter and the model will fit the noise perfectly. Collapse channels into four or five groups before you fit anything: paid search, paid social, events, content and organic, other.

Step five is the whole value. A coefficient that survives leave-one-out resampling is a channel you can defend. One that does not is a channel you should test directly. The fuller build, including handling of multicollinearity between channels that scale together, is covered in the [marketing mix modeling guide](/guides/marketing-mix-modeling-for-saas/).

Mix modelling and experimentation are complements, not alternatives. The model narrows the list of channels worth testing, the tests calibrate the model. Where this sits relative to click based measurement is the subject of [multi touch attribution versus incrementality testing](/comparisons/multi-touch-attribution-vs-incrementality/).

## How do you present results without getting picked apart?

Ranges, always. A point estimate invites an argument about the decimal place and loses the room.

The format that survives an executive meeting:

> Pausing LinkedIn for four weeks reduced demo requests by an estimated 6 to 18 percent, centred on 11 percent. The test could reliably detect effects above 12 percent, so we cannot rule out that the true effect is small. Recommendation: keep the channel, re-test in Q3 at doubled spend to look for a clearer signal.

Three things that sentence does. It gives the range. It states the test's power honestly, which pre-empts the sharpest objection. And it names the next action, so the meeting ends in a decision rather than a debate about methodology.

Never present a test result as a single number with no uncertainty attached. The first time someone catches you doing it, every subsequent result you bring gets discounted.

## The honest costs and failure modes

Pause tests cost real pipeline. If a channel is genuinely incremental, four weeks off is four weeks of lost demand, and at a six month sales cycle you feel it two quarters later. Budget for that. A team running two pause tests a year on channels representing 20 percent of spend should expect a measurable dent in one quarter's pipeline.

Geo holdouts are harder than they look in B2B because your accounts are not evenly distributed. Half your pipeline may sit in five metros, which makes matched market design close to impossible. Check the concentration before committing.

And matched account holdouts leak. Sales will call an account in the holdout group, because sales is measured on pipeline and not on your experiment. Either accept the contamination and report it, or get the holdout list respected in writing, which in practice means the CRO enforcing it.

The tradeoff worth stating plainly: incrementality testing buys you truth and costs you speed. Two well designed tests a year is a realistic cadence for a team of three, and those two will teach you more about your [demand generation](/saas-demand-generation/) engine than any dashboard. A platform that promises the same answer continuously without withholding anything is selling modelling, not measurement, which is worth remembering while reading vendor material on [paid media attribution](/guides/paid-media-attribution-for-saas/).

## Start here

Run the branded search pause first. It is the cheapest, fastest and most likely to change a budget line. Get the CRO's written agreement on the abort threshold, set up the weekly baseline export six weeks out, and pre-register the decision rule.

If it comes back null, which it usually does, you have freed budget and proven the process works. Then pick the channel you argue about most and pause that one next. The deeper mechanics of running the programme continuously live in the [incrementality testing playbook](/playbooks/incrementality-testing/), and the annual planning side belongs in your [demand generation plan template](/templates/demand-generation-plan-template/).

One more thing worth testing early: lifecycle email, which almost never gets a holdout despite being trivially easy to hold out. The method is the same as everything above, and the specifics are in [measuring lifecycle email in SaaS](/guides/email-marketing-attribution-saas/).

## Frequently asked questions

### What is incrementality testing in B2B SaaS?

It is measuring the causal contribution of a channel by removing it from part of your audience, geography or time period and comparing outcomes against a control. Unlike attribution, which divides credit for conversions that already happened, incrementality answers whether those conversions would have happened anyway without the spend.

### Can you run incrementality tests with low B2B volume?

Yes, but you change the outcome metric. Closed won deals are far too sparse. Use branded search volume, direct traffic, demo requests and qualified pipeline created, which arrive in sufficient numbers weekly. A company generating 60 demo requests a month can read a meaningful pause test in four to six weeks.

### How long should a channel pause test run?

Four weeks of pause plus two weeks of observation afterwards, minimum. Shorter tests get swamped by the lag between impression and inbound action, which in B2B is commonly two to six weeks. Run it in a period with no product launch, no conference and no pricing change.

### What happens when you pause branded search ads?

Usually organic clicks absorb 60 to 90 percent of the lost paid clicks, and total branded sessions barely move. The exception is when a competitor is bidding on your name or your organic result is pushed below an AI Overview, in which case the drop is real and immediate.

### Is marketing mix modeling realistic for a small B2B SaaS?

A simplified version is. With 24 months of monthly spend by channel and monthly pipeline created you can fit a regression with adstock in a spreadsheet. It will not be publication grade, but it will tell you which channels have coefficients that survive the removal of any single month, which is the decision you actually need.

### Should we buy an incrementality platform?

Not before you have run two manual tests. Platforms sell measurement infrastructure, and the binding constraint for most SaaS teams is experiment discipline, not tooling. Spend the first year proving you can hold a pause for four weeks without somebody switching the campaign back on.
