Measuring Lifecycle Email in SaaS
How to prove email revenue with holdout groups, incrementality tests, per sequence attribution rules, and the four board metrics that survive real scrutiny.
On this page 9 sections
- Why open rate stopped being a measurement input in September 2021
- What replaces open rate, and what each replacement actually tells you
- How to design a holdout when you only get 800 trials a month
- When a permanent holdout is worth the revenue it costs
- Incrementality and switchback tests for always on sequences
- Per sequence attribution rules that stop email double counting with sales
- The four numbers to put in front of a board
- What this costs, and the three ways it fails
- Start here in the next two weeks
- Frequently asked questions
The short answer
SaaS email attribution works through incrementality, not opens. Hold back 5 to 10 percent of the audience from every always on sequence, then compare trial conversion, activation and revenue between the held back group and the treated group. The gap is your incremental MRR. Apple Mail Privacy Protection has recorded opens for unread mail since September 2021, so report click to delivered, reply rate and downstream revenue instead. Last touch credit double counts sales activity and overstates email by a wide margin.
Key points before you start
Most lifecycle email reports open with a 44 percent open rate and close with a revenue figure lifted straight out of last touch in the CRM. Both numbers are wrong, and they are wrong in opposite directions. Opens are inflated by machines fetching tracking pixels on behalf of people who never read the message, while last touch hands email the credit for upgrades that were going to happen anyway. You do not fix that with a better attribution model. You fix it by deliberately withholding email from some people and watching what they do.
Why open rate stopped being a measurement input in September 2021
Apple shipped Mail Privacy Protection alongside iOS 15 in September 2021. When a recipient turns it on, Apple’s proxy downloads remote content in advance, including your tracking pixel, whether or not a human ever looks at the message. The open fires. Nobody read anything.
That single change broke four things at once, and most SaaS teams only rebuilt one of them. Subject line tests judged on opens became coin flips. Re-engagement segments defined as did not open in 90 days started excluding the most engaged Apple users and including people who had already forgotten you exist. Send time optimisation trained itself on proxy timestamps. And deliverability monitoring lost its cheapest proxy for inbox placement.
In the B2B lists we look at, Apple Mail typically accounts for somewhere between a third and a half of recorded opens, which means a reported 44 percent open rate might represent 25 percent real attention or 40 percent, and nothing in your platform can tell you which. Treat any open based figure in a vendor case study from 2022 onward as decorative.
| Metric | What it meant before 2021 | Status in 2026 | Use this instead |
|---|---|---|---|
| Open rate | Rough attention proxy | Contaminated by proxy fetches | Click to delivered |
| Subject line A/B win | Attention lift | Noise | Click to delivered, or downstream conversion |
| Did not open in 90 days | Disengaged | Often an Apple user with images off | No click and no product session in 90 days |
| Open rate collapse on one domain | Deliverability alarm | Still a valid alarm | Keep it, but pair with Google Postmaster Tools |
The segment that quietly kills your list
What replaces open rate, and what each replacement actually tells you
Four metrics survive. Click to delivered, reply rate, spam complaint rate, and downstream conversion measured against a holdout. Each answers a different question and none of them answers all three.
Click to delivered is your creative metric. Use delivered rather than sent as the denominator, because bounces are a list problem and mixing them in makes a good email look weak. It tells you whether the message earned an action, and it is the number to judge copy, offer and send time on. Click to open ratio is dead for the same reason opens are.
Reply rate matters more in sales assisted motions than most teams admit. A plain text activation email from a named human that pulls 3 percent replies is worth more than one pulling 9 percent clicks into a docs page, because a reply creates a conversation an account executive can work. It also reads as a strong engagement signal to mailbox providers.
Spam complaint rate is now a hard operating constraint rather than a vanity concern. Google and Yahoo’s bulk sender requirements, which took effect in February 2024, set a ceiling of 0.3 percent, and in practice you want to sit under 0.1 percent. Watch it per sequence, not per account, because one aggressive winback campaign can poison delivery for your onboarding mail.
0.3%
Spam complaint ceiling for bulk senders under the Google and Yahoo requirements introduced in February 2024
Google and Yahoo bulk sender requirements
Downstream conversion is the only one that carries revenue, and it is only trustworthy against a control group. That is the rest of this page. If you want the rates other teams are hitting on the first three, our SaaS email benchmarks split them by sequence type and ACV band rather than averaging a newsletter together with a password reset.
Editable CSV worksheet
SaaS benchmark evaluation worksheet
Record the source, date, cohort and metric definition before comparing your numbers with a benchmark.
How to design a holdout when you only get 800 trials a month
Start from the effect you need to detect, not from the split you find comfortable. A reasonable approximation for a two proportion test at 80 percent power and 95 percent confidence is sixteen times the baseline variance divided by the squared difference you want to catch. With a baseline trial to paid rate of 18 percent, the variance term is 0.18 multiplied by 0.82, which is 0.1476.
Work an example. A product takes 800 trials a month, converts 18 percent of them, and the team wants to know whether the six email onboarding sequence does anything at all.
| Lift you want to detect | Split | Users per arm | Total users | Months of enrolment at 800 trials |
|---|---|---|---|---|
| 18% to 21% (+3 points) | 50/50 | 2,624 | 5,248 | 6.6 |
| 18% to 23% (+5 points) | 50/50 | 945 | 1,890 | 2.4 |
| 18% to 26% (+8 points) | 50/50 | 370 | 740 | 1.0 |
| 18% to 21% (+3 points) | 90/10 | 1,458 control, 13,122 treated | 14,580 | 18.2 |
Read the last row twice. An unbalanced 90/10 design needs nearly three times the total population to answer the same question, because the small arm governs the standard error. That is why the split you use to test is not the split you use to monitor.
So the rule at low volume is this. Run the initial test at 50/50, accept that you can only detect large effects, and be honest that a sequence producing a two point lift is invisible to you and always will be at this volume. Once the sequence has proven itself, drop to a permanent 5 to 10 percent holdout as a regression alarm rather than a measurement instrument.
Running your first sequence holdout
- Pick one sequence and one outcome
Onboarding measured on trial to paid, not on four metrics at once. You will be tempted to measure everything. Do not.
- Randomise at the account level
Hash the account ID, not the user ID, so two people in the same workspace get the same treatment. Otherwise the treated user forwards the email to the control user.
- Freeze the sequence
No copy edits, no new steps, no send time changes until enrolment closes. A mid test change makes the result unreadable.
- Enrol for a fixed window
Calculate the weeks needed from the table above and put the end date in the calendar. Do not peek at day nine and stop.
- Wait one full conversion cycle
Add the trial length plus the tail. For a 14 day trial, wait 21 to 30 days after the last enrolment before you read anything.
- Compare rates and revenue per user
Conversion rate tells you if it worked. Revenue per enrolled user tells you whether it moved anyone up a plan or only pulled cheap accounts forward.
- Publish the result including a null
A sequence with no measurable lift is a finding worth the same respect as a win. Turn it off or rewrite it, then say so in the channel.
When a permanent holdout is worth the revenue it costs
Once a sequence is live and proven, keep 5 to 10 percent of eligible accounts out of it permanently. That is our position, and it costs money, so here is the arithmetic.
Say your dunning sequence recovers 38 percent of failed payments and failed payments represent $50,000 of MRR a year. A permanent 5 percent holdout forgoes roughly 5 percent of the recovery, about $950 a year. In exchange you get a live control that tells you the day a deliverability problem, a broken merge tag or a platform migration quietly halves your recovery rate. Most teams discover those failures six weeks late through a revenue dip nobody can explain.
The holdout is worth it when three conditions hold. The sequence runs continuously rather than as a campaign. The outcome it drives is measurable inside 30 days. And the cost of a silent failure is larger than the cost of the withheld sends. Dunning, onboarding and trial expiry clear that bar comfortably. A quarterly product newsletter does not, and forcing a permanent holdout onto it is theatre.
Holdouts leak
Incrementality and switchback tests for always on sequences
User level randomisation fails in two specific SaaS situations. When several people share a workspace and forward things to each other, and when the sequence is tied to a product surface you cannot hold back, such as an in app banner that fires off the same event.
Switchback testing solves both by randomising time instead of people. Turn the sequence on for a week, off the next, and repeat. Then compare the outcome for cohorts that entered during on weeks against cohorts that entered during off weeks. You need enough blocks for the comparison to mean anything, and sixteen weekly blocks is the practical floor. Fewer than that and a single bad week of seasonality dominates the result.
Two things ruin switchbacks. Carryover, where someone entering during an off week still gets touched because they were already in a related sequence, and seasonality, where your on weeks happen to land on the first week of each month. Randomise the block order rather than alternating strictly, and check that on and off weeks are balanced across the month.
| Design | When to use it | Minimum run | Main weakness |
|---|---|---|---|
| Balanced user split | One off test of a new sequence | 6 to 10 weeks of enrolment | Leaks inside shared workspaces |
| Account level split | Any B2B product with multi seat accounts | Same as above | Needs more accounts than users |
| Permanent 5 to 10% holdout | Proven, always on sequences | Runs forever | Detects only large regressions |
| Switchback by week | Sequences tied to in app surfaces | 16 weekly blocks | Seasonality and carryover |
| Staggered start (48h delay) | Dunning and other service adjacent mail | 8 weeks | Measures timing, not the message itself |
Per sequence attribution rules that stop email double counting with sales
Set the rules before you need them, in writing, and get the sales leader to agree in the same meeting. The argument you are avoiding is the one where marketing claims 60 percent of pipeline, sales claims 70 percent, and finance stops reading either deck.
Three rules cover most cases. First, a click must precede the conversion inside the attribution window. An open never qualifies, for the reasons above. Second, if an account executive logged a call, meeting or manual email inside that same window, the deal is sales sourced and email influenced, and those go in separate columns. Third, the window varies by sequence, because a trial expiry email that works does so within 72 hours, while a winback might land three weeks later.
| Sequence | Window | Qualifying event | Credit rule |
|---|---|---|---|
| Welcome and onboarding | 30 days | Click then activation milestone | Activation lift versus holdout, no revenue claim |
| Activation nudge | 7 days | Click then feature first use | Activation lift only |
| Trial expiry | 72 hours | Click then upgrade | Email sourced unless an AE touched the account |
| Dunning | 14 days | Click then successful payment | Involuntary churn recovered, always email sourced |
| Expansion and upsell | 30 days | Click then seat or plan increase | Influenced, never sourced, if CS was involved |
| Winback | 45 days | Click then reactivation | Email sourced, flagged as low confidence |
| Newsletter | None | None | No revenue claim, measured on retention of engagement |
Notice what is missing. The newsletter gets no revenue line. Teams hate this, and it is the right call, because a newsletter that a customer reads for eight months before renewing is doing something real that no window based rule can capture without inventing a number. Measure it on subscriber retention and reply quality, and defend it as a brand investment rather than pretending it closed deals. The same discipline applies on the ads side, and we walk through it in paid media attribution for SaaS.
Editable CSV worksheet
Save your marketing measurement plan
Keep a worksheet for your inputs, assumptions and next actions. You can also print the calculation directly from your browser.
The four numbers to put in front of a board
Boards do not want seven email metrics. They want to know whether the program pays for itself and what would break if it stopped. Four numbers do that job.
Incremental MRR. The holdout derived difference in revenue per enrolled account, multiplied by the treated population, stated monthly. One number, one method, restated every quarter with the confidence interval attached. If the interval crosses zero, say so.
Activation lift. The percentage point difference in accounts reaching your activation milestone between treated and held out groups. This is where onboarding email earns its keep, and it usually shows a clearer effect than revenue does because it happens sooner and closer to the message.
Involuntary churn recovered. Dollars of failed payments rescued by dunning, net of the holdout baseline recovery that would have happened through in product retries anyway. Most teams report gross recovery here and overstate it by 30 to 50 percent, since some customers fix their own cards without being asked.
Expansion MRR influenced. Clearly labelled as influenced. This is the number a CFO will push on hardest, so give it the weakest claim and the clearest definition.
We reported a 12 percent email attributed revenue share for two years. When we finally ran a holdout the true incremental number was closer to 4 percent. That was a bad quarter, and the best thing we ever did for the program’s credibility.
Model those four against cost before you present them. Our SaaS email revenue calculator takes list size, click to delivered, conversion and ACV and returns incremental revenue and program payback, which is a more useful shape for a board than a funnel chart.
What this costs, and the three ways it fails
The honest cost is time, not tooling. Designing a clean test, tagging holdouts across your platform and CRM, suppressing parallel campaigns and writing the analysis takes a competent marketing operations person roughly two days per sequence up front and half a day a month after that. At a loaded rate that is a few thousand dollars a year, against a program usually worth six figures.
Three failure modes account for most of the disappointments we see.
The first is underpowered tests read as wins. A team at 300 trials a month runs a four week test, sees 19.4 percent against 18.1 percent, and ships the sequence. That difference is noise at that sample size, and the team now believes something false. If your volume cannot support the test, say the sequence is unmeasured rather than dressing up a coin flip.
This second is tooling that cannot hold a group out cleanly. List based platforms often struggle to keep a stable random assignment across sequence entries and exits, so the holdout drifts and the comparison stops being random. Event based platforms handle it natively through a random bucket attribute set at account creation, which is one of several reasons we favour them in our comparison of lifecycle email platforms, and a specific point of difference in HubSpot versus Customer.io. If you run Mautic, you will be assigning the bucket yourself in the database, which works but has to be built.
The third is organisational. Someone senior asks why 8 percent of trials are not getting onboarding email, and the holdout quietly gets switched off. Pre-empt it. Write down the annual cost of the holdout in dollars, put it next to the cost of the last silent sequence failure, and get the decision made once at the right level.
Start here in the next two weeks
Pick the sequence with the most obvious revenue outcome, which is almost always trial expiry or dunning, and put a holdout on it before you touch a single word of copy. Then get the four board numbers defined and agreed with sales while the test runs, because that conversation takes longer than the test does.
Before you call your email program measured
0 of 7 done
If you are earlier than that and still building the first sequences rather than measuring them, start with the lifecycle email fundamentals and then ship your first sequence with the holdout built in from day one. Retrofitting a control group onto a sequence that has been running for a year is possible, but you lose the clean baseline and spend the first month arguing about what changed.
Editable CSV worksheet
SaaS Email Marketing planning worksheet
A practical email planning worksheet: decisions, owners, evidence and next actions.
Frequently asked questions
How do you measure email marketing ROI in SaaS?
Run a holdout. Randomly exclude a slice of the eligible audience from the sequence, wait one full conversion cycle, then compare paid conversion rate and revenue per user across the two groups. Multiply the difference by the treated population to get incremental revenue, divide by the fully loaded program cost, and report that. Anything derived from last touch is a ceiling, not a measurement.
What size should an email holdout group be?
For a one off test, split the audience evenly, because a balanced design needs the fewest total users to reach significance. For ongoing monitoring after a sequence has proven itself, hold back 5 to 10 percent permanently. Below about 3 percent the control arm is too small to detect anything short of a catastrophic regression, and you are paying the cost of a holdout without the information.
Is open rate still useful for anything?
It has one surviving use. A sudden collapse in opens across an entire domain is a deliverability alarm worth investigating, because machine opens are consistent enough that their disappearance means something changed at the mailbox provider. It cannot judge subject lines, cannot drive re-engagement segments, and cannot appear in a revenue calculation. Use click to delivered for creative decisions.
How long should an email incrementality test run?
Long enough to cover the full conversion window plus the time needed to accumulate sample. If your trial is 14 days and most upgrades land by day 21, the earliest honest read is 21 days after the last user enters the test. Most B2B programs need six to twelve weeks of enrolment on top of that. Stopping early because the lift looks good is how teams ship sequences that do nothing.
Should email get credit if sales also touched the account?
No, not as sourced revenue. Mark it as influenced and keep it in a separate column. The rule that survives an argument with a sales leader is simple. If an account executive logged a call, meeting or manual email inside the attribution window, the deal counts as sales sourced and email influenced. Double counting across two teams is the fastest way to lose credibility with a CFO.
What is a switchback test for email?
Instead of splitting people, you split time. The sequence runs for a week, is paused for a week, and the cycle repeats for several months. You then compare cohorts that entered during on weeks against cohorts that entered during off weeks. It is useful when user level randomisation leaks, for example when several people on the same account share a workspace and talk to each other.
Can you hold out dunning emails?
Carefully, and rarely. Payment failure notices are close to a service message, and in some contracts and jurisdictions you are obliged to send them. The safer design is a staggered start rather than a true holdout. Send the recovery sequence to everyone, but delay it by 48 hours for a random slice, then measure recovery rate in that window. You get a clean read without withholding the message.
The saas-marketing.net editorial team Research and editorial
We research, write and maintain every page on this site. The library explains marketing decisions through practical frameworks, explicit assumptions and references. Corrections can be requested through the contact page.
Published September 11, 2026. Last updated .