How to Test SaaS Pricing
Why price A/B tests rarely reach significance in B2B, and what to run instead: cohort launches, geo splits, quote tests, and sequenced tier changes.
On this page 7 sections
The short answer
Most B2B SaaS sites cannot run a valid price A/B test. At a 3 percent trial start rate, detecting a 15 percent relative change needs roughly 14,000 pricing page visitors per variant, which is more than most B2B sites see in a quarter. The workable alternatives are new customer cohort launches, a new tier or add on used as a price probe, geographic and currency splits, controlled quote testing in sales led deals, and sequential pre and post changes with a holdout segment.
Key points before you start
Somebody in your company wants to A/B test the pricing page. It sounds reasonable, the tooling is already installed, and the test will produce a number within three weeks. That number will be noise, and the decision made on it will be worse than a decision made on judgement alone.
Here is the arithmetic, then five things you can actually run.
Why the A/B test does not work
Run the power calculation before you design anything. With a baseline pricing page to trial start rate of 3 percent, wanting to detect a 15 percent relative change at 80 percent power and 95 percent confidence, you need somewhere around 14,000 visitors per variant. That is 28,000 pricing page sessions.
Most B2B SaaS companies below 20 million ARR do not see 28,000 pricing page visits in a quarter. And the effects worth detecting are usually smaller than 15 percent, which makes it worse fast: halving the detectable effect roughly quadruples the sample.
| Baseline trial start rate | Effect you want to detect | Visitors needed per variant | Realistic for |
|---|---|---|---|
| 3% | 15% relative | ~14,000 | High traffic PLG only |
| 3% | 10% relative | ~31,000 | Very few B2B companies |
| 3% | 5% relative | ~124,000 | Effectively nobody in B2B |
| 8% | 15% relative | ~5,000 | Self serve with strong top of funnel |
There is a second problem that no sample size fixes. Price tests contaminate. A visitor who sees $49, closes the tab, comes back on a phone and sees $79 does not conclude that they are in a test. They conclude you are untrustworthy, and enterprise buyers compare notes in Slack communities constantly. Differential pricing is not usually unlawful, but the trust exposure is real and the screenshot always travels.
The test that never had a chance
A team runs a pricing test for three weeks, gets 900 visitors per variant, sees variant B at 3.4 percent against 2.9 percent, and ships it. That difference is well inside the noise band for that sample. They were not measuring price sensitivity. They were measuring which three weeks happened to contain a good campaign.
Five things to run instead, ranked
Ranked by how much you learn against how much you risk.
1. New customer cohort launch. Change the price for all new customers from a date, grandfather everyone existing, and compare the new cohort against the equivalent cohort from the prior period with seasonality accounted for. Clean commercially, messy statistically, and the best default for most companies.
2. New tier or add on as a probe. Introduce a higher tier at the price you suspect the market will bear. You learn willingness to pay without touching an existing price. If it fails, retire it quietly. This is the lowest risk option on the list.
3. Geographic and currency split. Different prices in different markets are normal, expected and defensible. Randomisation is not perfect because markets differ, but it is the closest thing to a genuine parallel test available at B2B traffic levels.
4. Sales quote testing at controlled list prices. In sales led motions, hold list price constant and vary the discount authority or the opening quote by rep territory. You get a real elasticity read from deals that already exist, and you need enough deal volume to see it.
5. Sequential pre and post with a holdout. Change the price, keep one segment on the old price as a holdout, and compare. Confounded by time but the holdout catches most of the seasonality.
| Method | Learning quality | Risk | Minimum volume | Best for |
|---|---|---|---|---|
| New customer cohort launch | Medium | Medium | 60 to 100 new customers per period | Most self serve and hybrid companies |
| New tier or add on probe | Medium | Low | Any | Testing upward willingness to pay |
| Geo and currency split | High | Low | Two markets with similar mix | Companies already selling internationally |
| Sales quote testing | High | Medium | 150+ quotes per quarter | Sales led, ACV above 20k |
| Sequential with holdout | Low to medium | Medium | Any | When nothing else is available |
Editable CSV worksheet
Save your marketing measurement plan
Keep a worksheet for your inputs, assumptions and next actions. You can also print the calculation directly from your browser.
A worked cohort launch, with real numbers
A 6 million ARR product marketing tool sells a Growth plan at $79 per seat per month. They believe $99 is supportable. Roughly 110 new self serve customers land per month, average 6 seats.
They set the launch for 1 April, grandfathering all existing accounts. March becomes the baseline month: 112 new customers, average 6.1 seats, $79, giving about $53,900 in new monthly recurring revenue.
April lands 94 new customers at 5.8 seats and $99, giving about $53,900 as well. Conversion fell 16 percent, revenue held flat. That is close to unit elastic in this range, and on that alone the change is neutral.
The decision is not neutral, though, and this is where the reasoning gets interesting. Fewer customers at the same revenue means less support load, fewer onboarding hours and a smaller free to paid funnel. Whether that is good depends on whether your growth model needs logo count for network effects and word of mouth, or revenue. A tool with no network effect should keep the price. A collaboration product probably should not.
Read the segments, not the total
In the same launch, seats per customer dropped from 6.1 to 5.8. That is mix shift: buyers started smaller to test the higher price. Look at whether the shrinkage is temporary hedging, which expansion revenue will recover, or a genuine loss of the larger accounts. The total number hides both.
Seasonality and mix shift, the two things that break every read
Seasonality first. B2B demand moves with budget cycles, quarter ends and holidays. Comparing April against March tells you almost nothing on its own. Compare year over year for the same month, or hold a holdout segment on the old price, or ideally both.
Mix shift is subtler and more dangerous. Your traffic composition changes for reasons unrelated to price: a webinar brings in a different segment, a paid campaign pauses, a comparison page starts ranking. Revenue per visitor can rise while your best segment quietly stops converting, and you will not see it in the headline.
Guard against it by fixing the reporting cut before launch. Split by source, by company size band and by plan tier, and commit to reading all three. A single aggregate number is how teams ship price changes that hurt them.
2 sales cycles
Minimum observation before judging a price change, and one renewal cycle before judging retention
Aggregated practitioner reports, saas-marketing.net estimate
Decision rules: keep, revert or expand
Write these down before the price changes. The entire value of running a cohort launch instead of pretending to run an A/B test is that you set the decision rule in advance, when nobody is defending a result.
The launch document
- State the hypothesis with a number
Not 'we think we are underpriced' but 'we expect new customer count to fall no more than 20 percent while revenue per cohort rises at least 10 percent'.
- Fix the observation window
Two sales cycles for conversion, one renewal cycle for retention. Written down, with dates.
- Define the revert trigger
A specific threshold, checked weekly. For example: new customer count down more than 30 percent for two consecutive weeks, or any segment down more than 40 percent.
- Name the person who can pull it
One person, with authority, who does not need a meeting. Committees do not revert prices.
- Fix the reporting cut before launch
Source, company size band, plan tier. Agreed in advance so nobody goes hunting for a favourable slice afterwards.
- Grandfather existing customers explicitly
In writing, with a stated duration, so support and sales give the same answer.
- Prepare the reversal message
If you revert, say the price is going back and stop. Do not explain the experiment. Customers do not want to hear they were test subjects.
- Review against the written rule
Read the launch doc before you read the dashboard. In that order.
Expansion criteria deserve a mention too. If the cohort launch works, the next question is whether to move existing customers, and that is a rollout with a communication plan rather than a continuation of the test. The price increase rollout playbook covers the sequencing, notice periods and the grandfather decision, and the pricing change readiness checklist covers what has to be true operationally before you touch anything.
Self-paced learning
Build your SaaS marketing study plan
Choose a free course and work through its published lessons at your own pace. Save the course index for later.
What research methods can and cannot add
Survey methods like Van Westendorp and Gabor Granger get proposed whenever the traffic argument lands. They are useful for bracketing a range before you commit engineering time, and they are unreliable as a source of a specific number, because stated willingness to pay overstates actual willingness to pay by a wide margin in every field where both have been measured.
Use them to answer “is $99 plausible or absurd”, not “is $99 better than $89”. Then test the narrowed range with a cohort launch. Conjoint analysis is better and needs a sample and a budget most marketing teams do not have.
The background concepts worth reading first are price elasticity, which is what you are actually trying to estimate, and price anchoring, which explains why the tier above the one you are testing changes the result. If you are still choosing a model rather than a number, SaaS pricing frameworks compared is the earlier decision, and the broader SaaS pricing strategy hub covers packaging.
The position
Stop calling these tests. Call them monitored launches with revert criteria, because that is what they are, and the honesty changes the behaviour around them.
A monitored launch with a written hypothesis, a fixed window, a segment level reporting cut and one named person who can pull it beats an underpowered A/B test in every dimension that matters. It risks less, it teaches more, and it does not require anyone to pretend that 900 sessions per variant means something.
Start by running the power calculation for your own baseline. Then look at SaaS pricing benchmarks for where comparable companies sit, and at SaaS repricing case studies for what happened to the ones who moved. If you want to work through a change with a structure around it, the SaaS pricing sprint runs the sequence end to end.
Editable CSV worksheet
SaaS Pricing planning worksheet
A practical pricing planning worksheet: decisions, owners, evidence and next actions.
Frequently asked questions
Can you A/B test pricing on a B2B SaaS site?
Technically yes, statistically almost never. The arithmetic is unforgiving: low conversion rates and modest traffic mean a test that could detect a realistic price effect would need to run for many months, during which seasonality, product changes and campaign mix all move. Teams that claim a clean price test usually stopped when the number looked good.
How much traffic do you need to test pricing?
With a 3 percent baseline trial start rate and a target of detecting a 15 percent relative change at 80 percent power, you need roughly 14,000 visitors per variant. Halve the effect you want to detect and the requirement roughly quadruples. Run the number for your own baseline before designing anything.
Is it legal to show different prices to different customers?
Differential pricing is generally lawful in most markets when it is not based on a protected characteristic, but it carries real consumer trust and disclosure risk, and consumer protection regimes in the EU and UK increasingly scrutinise personalised pricing. The bigger practical risk is reputational: somebody screenshots two prices side by side and posts it.
What is a price elasticity test in SaaS?
Any structured attempt to measure how demand responds to a price change. In practice that means comparing conversion and revenue per visitor across a price difference you deliberately created, whether across cohorts, geographies or time. True elasticity curves need several price points, which most SaaS companies never get.
How long should you observe a price change before deciding?
Two full sales cycles for conversion effects, and at least one renewal cycle before you claim anything about retention or expansion. A price increase that looks fine at 60 days can show up as elevated churn at the twelve month renewal, which is why the observation window should be written into the launch document.
Should existing customers be included in a price test?
No. Test on new customers only, and grandfather existing accounts through any change you are still evaluating. Moving existing customers is a rollout decision with its own communication plan, not an experiment, and conflating the two is how a pricing test turns into a churn event.
What is the safest way to test a higher price?
Launch a new tier or add on at the higher price rather than raising the existing one. You learn about willingness to pay without touching the price anyone already saw, the downside is contained, and if it fails you retire the tier quietly instead of explaining a reversal.
The saas-marketing.net editorial team Research and editorial
We research, write and maintain every page on this site. The library explains marketing decisions through practical frameworks, explicit assumptions and references. Corrections can be requested through the contact page.
Published September 11, 2026. Last updated .