Incrementality testing for B2B SaaS
Run geo holdouts, channel pauses and matched market tests at low B2B volume, how long each takes to read, and when a simple mix model beats another tracker.
On this page 8 sections
- Which test designs actually work at B2B volume?
- What does pausing branded search usually reveal?
- How do you size and time a test with 60 opportunities a month?
- What do you measure when revenue is nine months away?
- Can you build a mix model in a spreadsheet?
- How do you present results without getting picked apart?
- The honest costs and failure modes
- Start here
- Frequently asked questions
The short answer
Incrementality testing measures what a channel actually causes rather than what it gets credited for, by withholding it somewhere and comparing. B2B SaaS volume is too low for clean conversion based experiments, so the workable designs are channel pauses read against branded search and direct traffic, geo holdouts across matched metro sets, matched account holdouts for ABM, and campaign level offer splits. Read leading indicators over four to six weeks, not closed revenue.
Key points before you start
Attribution tells you which touchpoint was standing nearby when someone converted. Incrementality tells you whether the conversion needed it. Those are different questions, and only one of them should drive a budget decision.
The objection every B2B team raises is volume. You have 60 demo requests a month, not 60,000 ecommerce transactions, so the textbook experiment designs do not apply. True. The designs below are the ones that still read at that scale.
Which test designs actually work at B2B volume?
Four, and they suit different questions. Pick by what you are trying to decide, not by what your ad platform offers.
| Design | Answers | Minimum scale | Time to read | Main risk |
|---|---|---|---|---|
| Channel pause | Is this channel creating demand or harvesting it | Any spend above ~$8k/month on that channel | 4 weeks pause + 2 weeks observation | Seasonality contamination |
| Geo holdout | What does this channel contribute per market | ~20 matched metros per arm | 8-12 weeks | Market matching quality |
| Matched account holdout | Does ABM spend move target accounts | 300+ accounts per arm | 1-2 quarters | Sales team leaking into the holdout |
| Offer or creative split | Which message produces qualified pipeline | ~200 conversions per arm | 3-6 weeks | Measures relative lift, never absolute incrementality |
The fourth one is the odd entry. A creative split is a good experiment and a bad incrementality test, because both arms still receive the channel. It tells you which ad works better, never whether the channel works at all. Teams conflate these constantly.
The channel pause, in detail
Turn a channel off entirely for four weeks. Watch branded search impressions, direct sessions, demo requests and qualified pipeline created. Then turn it back on and watch the recovery curve, which is often more informative than the drop.
You need a baseline first. Pull the prior 13 weeks of the same metrics, calculate the week over week variance, and write down the band inside which you will call the result null. Doing this before the test is the difference between an experiment and a story you tell afterwards.
The single thing that ruins pause tests
Somebody switches the campaign back on in week two because pipeline looked soft. Get written agreement from the CRO before you start, including the specific number that would justify aborting. If you cannot get that agreement, you are not ready to run the test.
What does pausing branded search usually reveal?
That most of it was not incremental. This is the most commonly run pause test in SaaS and the results are boringly consistent.
When you pause branded paid search, organic results typically absorb 60 to 90 percent of the lost clicks. Total branded sessions dip a few points. Demo requests often do not move at all outside normal variance. You were paying for clicks you already had.
Three conditions flip the result, and all three are worth checking before you cut the budget:
- A competitor is actively bidding on your brand terms. Check the auction insights report. If someone else has impression share above 10 percent on your name, pausing hands them the click.
- An AI Overview or a large SERP feature pushes your organic listing below the fold on brand queries. This has become common enough that the old assumption of guaranteed organic capture no longer holds.
- Your brand name is a common word. If you are called Monday or Linear, generic searchers and brand searchers blend, and the paid listing is doing real disambiguation work.
60-90%
Share of paused branded paid clicks typically recovered by organic listings
Typical observed range in brand search pause tests
I would run this test first at almost any SaaS company, because it is the one most likely to free budget without hurting anything. What you do with the freed budget is a separate question the demand generation budget calculator can help frame.
Editable CSV worksheet
SaaS benchmark evaluation worksheet
Record the source, date, cohort and metric definition before comparing your numbers with a benchmark.
How do you size and time a test with 60 opportunities a month?
Work backwards from detectable effect. At low volume you cannot detect small effects, so decide up front the smallest lift you would act on and check whether the test can see it.
Rough rule: with weekly counts in the 10 to 20 range and typical B2B variance, a four week test can detect roughly a 25 to 30 percent change in that metric. It cannot detect 8 percent. If your hypothesis is that LinkedIn contributes 10 percent of demand, a pause test will not settle it and you should not pretend otherwise.
Sizing a pause test
- Pick the read metric
Demo requests or qualified pipeline created, weekly. Not closed won. You want something that arrives within the test window.
- Measure baseline variance
Take 13 weeks of history and compute the standard deviation of the weekly count. If weekly demo requests run 15 with an SD of 5, your noise floor is high.
- Set the minimum detectable effect
Roughly two standard deviations of the weekly mean over a four week window. Write it down as a percentage before the test.
- Check the decision is worth it
If the MDE is larger than any plausible true effect, extend to eight weeks or pick a different design. Do not run an underpowered test and report the result anyway.
- Choose a clean window
No launches, no conferences, no pricing changes, no fiscal quarter end. In practice this means February to March or September to October for most B2B calendars.
- Pre-register the call
Write the decision rule in a doc: 'if demo requests fall more than X percent over four weeks we keep the channel'. Circulate it before day one.
That last step sounds bureaucratic. It is the only thing standing between you and a result everyone reinterprets to suit their existing position.
What do you measure when revenue is nine months away?
Leading indicators with short latency, ordered by how quickly they respond.
| Indicator | Typical latency | Sensitivity | Use it for |
|---|---|---|---|
| Branded search impressions | 3-10 days | High | Awareness channels, podcast and display |
| Direct traffic sessions | 1-2 weeks | Medium | Broad reach campaigns |
| Demo requests | 2-6 weeks | High | Any bottom funnel channel |
| Qualified pipeline created | 4-10 weeks | Medium | The decision metric for most budgets |
| Self reported source mix | 4-8 weeks | Medium | Detecting channels tracking cannot see |
| Closed won revenue | 6-12 months | Low, within a test | Annual review, never a test readout |
Branded search is the workhorse. It moves fast, it is cheap to measure through Search Console and Google Ads impression data, and it is genuinely hard to fake. When a podcast sponsorship works, branded impressions move within a week of the episode dropping, which is the clearest brand lift signal available to a company with no budget for survey based measurement.
Pair the quantitative read with the qualitative one. Watching your self reported attribution mix shift during a pause is a second, independent read on the same question, and the two rarely disagree without reason.
Instrument before you test
Set up a weekly export of branded impressions, direct sessions and demo requests into one sheet six weeks before your first test. Half of failed incrementality programs fail because nobody had a clean baseline when the test started.
Can you build a mix model in a spreadsheet?
Yes, and for most SaaS companies under $30M ARR it is a better use of a week than another tracking implementation. The output is not publication grade. It is directional, and directional is what a budget meeting needs.
What you need: 24 months of monthly spend by channel, monthly pipeline created, and monthly values for two or three controls such as headcount in sales, seasonality index and any major launch flags.
The build, in outline:
- One row per month, one column per channel’s spend, one column for pipeline created.
- Apply adstock to each spend column: this month’s effective spend equals this month’s spend plus roughly 0.4 times last month’s effective spend. That decay factor encodes the lag. Try 0.3, 0.5 and 0.7 and see which fits better.
- Apply a saturation transform. Taking the square root or log of effective spend is crude and much better than nothing, because it stops the model claiming that tripling spend triples pipeline.
- Run a linear regression of pipeline on the transformed columns plus controls. LINEST in Sheets or the Data Analysis pack in Excel will do it.
- Test stability. Remove one month at a time, refit, and see whether any coefficient flips sign. Coefficients that flip are noise you should not spend against.
Where spreadsheet mix models go wrong
Twenty four data points and eight channels. That is three observations per parameter and the model will fit the noise perfectly. Collapse channels into four or five groups before you fit anything: paid search, paid social, events, content and organic, other.
Step five is the whole value. A coefficient that survives leave-one-out resampling is a channel you can defend. One that does not is a channel you should test directly. The fuller build, including handling of multicollinearity between channels that scale together, is covered in the marketing mix modeling guide.
Mix modelling and experimentation are complements, not alternatives. The model narrows the list of channels worth testing, the tests calibrate the model. Where this sits relative to click based measurement is the subject of multi touch attribution versus incrementality testing.
Review request
Free SaaS marketing audit
Share your site, stage and priorities to request a review of your positioning, funnel and acquisition plan.
How do you present results without getting picked apart?
Ranges, always. A point estimate invites an argument about the decimal place and loses the room.
The format that survives an executive meeting:
Pausing LinkedIn for four weeks reduced demo requests by an estimated 6 to 18 percent, centred on 11 percent. The test could reliably detect effects above 12 percent, so we cannot rule out that the true effect is small. Recommendation: keep the channel, re-test in Q3 at doubled spend to look for a clearer signal.
Three things that sentence does. It gives the range. It states the test’s power honestly, which pre-empts the sharpest objection. And it names the next action, so the meeting ends in a decision rather than a debate about methodology.
Never present a test result as a single number with no uncertainty attached. The first time someone catches you doing it, every subsequent result you bring gets discounted.
The honest costs and failure modes
Pause tests cost real pipeline. If a channel is genuinely incremental, four weeks off is four weeks of lost demand, and at a six month sales cycle you feel it two quarters later. Budget for that. A team running two pause tests a year on channels representing 20 percent of spend should expect a measurable dent in one quarter’s pipeline.
Geo holdouts are harder than they look in B2B because your accounts are not evenly distributed. Half your pipeline may sit in five metros, which makes matched market design close to impossible. Check the concentration before committing.
And matched account holdouts leak. Sales will call an account in the holdout group, because sales is measured on pipeline and not on your experiment. Either accept the contamination and report it, or get the holdout list respected in writing, which in practice means the CRO enforcing it.
The tradeoff worth stating plainly: incrementality testing buys you truth and costs you speed. Two well designed tests a year is a realistic cadence for a team of three, and those two will teach you more about your demand generation engine than any dashboard. A platform that promises the same answer continuously without withholding anything is selling modelling, not measurement, which is worth remembering while reading vendor material on paid media attribution.
Start here
Run the branded search pause first. It is the cheapest, fastest and most likely to change a budget line. Get the CRO’s written agreement on the abort threshold, set up the weekly baseline export six weeks out, and pre-register the decision rule.
If it comes back null, which it usually does, you have freed budget and proven the process works. Then pick the channel you argue about most and pause that one next. The deeper mechanics of running the programme continuously live in the incrementality testing playbook, and the annual planning side belongs in your demand generation plan template.
One more thing worth testing early: lifecycle email, which almost never gets a holdout despite being trivially easy to hold out. The method is the same as everything above, and the specifics are in measuring lifecycle email in SaaS.
Editable CSV worksheet
SaaS Demand Generation planning worksheet
A practical demand gen planning worksheet: decisions, owners, evidence and next actions.
Frequently asked questions
What is incrementality testing in B2B SaaS?
It is measuring the causal contribution of a channel by removing it from part of your audience, geography or time period and comparing outcomes against a control. Unlike attribution, which divides credit for conversions that already happened, incrementality answers whether those conversions would have happened anyway without the spend.
Can you run incrementality tests with low B2B volume?
Yes, but you change the outcome metric. Closed won deals are far too sparse. Use branded search volume, direct traffic, demo requests and qualified pipeline created, which arrive in sufficient numbers weekly. A company generating 60 demo requests a month can read a meaningful pause test in four to six weeks.
How long should a channel pause test run?
Four weeks of pause plus two weeks of observation afterwards, minimum. Shorter tests get swamped by the lag between impression and inbound action, which in B2B is commonly two to six weeks. Run it in a period with no product launch, no conference and no pricing change.
What happens when you pause branded search ads?
Usually organic clicks absorb 60 to 90 percent of the lost paid clicks, and total branded sessions barely move. The exception is when a competitor is bidding on your name or your organic result is pushed below an AI Overview, in which case the drop is real and immediate.
Is marketing mix modeling realistic for a small B2B SaaS?
A simplified version is. With 24 months of monthly spend by channel and monthly pipeline created you can fit a regression with adstock in a spreadsheet. It will not be publication grade, but it will tell you which channels have coefficients that survive the removal of any single month, which is the decision you actually need.
Should we buy an incrementality platform?
Not before you have run two manual tests. Platforms sell measurement infrastructure, and the binding constraint for most SaaS teams is experiment discipline, not tooling. Spend the first year proving you can hold a pause for four weeks without somebody switching the campaign back on.
The saas-marketing.net editorial team Research and editorial
We research, write and maintain every page on this site. The library explains marketing decisions through practical frameworks, explicit assumptions and references. Corrections can be requested through the contact page.
Published September 11, 2026. Last updated .