Get the working resource ↓
SaaS Email Marketing Guide 10 min read

Measuring Lifecycle Email in SaaS

How to prove email revenue with holdout groups, incrementality tests, per sequence attribution rules, and the four board metrics that survive real scrutiny.

On this page 9 sections
  1. Why open rate stopped being a measurement input in September 2021
  2. What replaces open rate, and what each replacement actually tells you
  3. How to design a holdout when you only get 800 trials a month
  4. When a permanent holdout is worth the revenue it costs
  5. Incrementality and switchback tests for always on sequences
  6. Per sequence attribution rules that stop email double counting with sales
  7. The four numbers to put in front of a board
  8. What this costs, and the three ways it fails
  9. Start here in the next two weeks
  10. Frequently asked questions

The short answer

SaaS email attribution works through incrementality, not opens. Hold back 5 to 10 percent of the audience from every always on sequence, then compare trial conversion, activation and revenue between the held back group and the treated group. The gap is your incremental MRR. Apple Mail Privacy Protection has recorded opens for unread mail since September 2021, so report click to delivered, reply rate and downstream revenue instead. Last touch credit double counts sales activity and overstates email by a wide margin.

Key points before you start

Most lifecycle email reports open with a 44 percent open rate and close with a revenue figure lifted straight out of last touch in the CRM. Both numbers are wrong, and they are wrong in opposite directions. Opens are inflated by machines fetching tracking pixels on behalf of people who never read the message, while last touch hands email the credit for upgrades that were going to happen anyway. You do not fix that with a better attribution model. You fix it by deliberately withholding email from some people and watching what they do.

Why open rate stopped being a measurement input in September 2021

Apple shipped Mail Privacy Protection alongside iOS 15 in September 2021. When a recipient turns it on, Apple’s proxy downloads remote content in advance, including your tracking pixel, whether or not a human ever looks at the message. The open fires. Nobody read anything.

That single change broke four things at once, and most SaaS teams only rebuilt one of them. Subject line tests judged on opens became coin flips. Re-engagement segments defined as did not open in 90 days started excluding the most engaged Apple users and including people who had already forgotten you exist. Send time optimisation trained itself on proxy timestamps. And deliverability monitoring lost its cheapest proxy for inbox placement.

In the B2B lists we look at, Apple Mail typically accounts for somewhere between a third and a half of recorded opens, which means a reported 44 percent open rate might represent 25 percent real attention or 40 percent, and nothing in your platform can tell you which. Treat any open based figure in a vendor case study from 2022 onward as decorative.

MetricWhat it meant before 2021Status in 2026Use this instead
Open rateRough attention proxyContaminated by proxy fetchesClick to delivered
Subject line A/B winAttention liftNoiseClick to delivered, or downstream conversion
Did not open in 90 daysDisengagedOften an Apple user with images offNo click and no product session in 90 days
Open rate collapse on one domainDeliverability alarmStill a valid alarmKeep it, but pair with Google Postmaster Tools

The segment that quietly kills your list

Sunsetting anyone who has not opened in 90 days looks like list hygiene. Run it after 2021 and you will suppress a chunk of engaged Apple users while keeping people who have never clicked anything. Rebuild the rule on clicks plus product activity. Our common lifecycle email mistakes page has the rest of the rules that aged badly.

What replaces open rate, and what each replacement actually tells you

Four metrics survive. Click to delivered, reply rate, spam complaint rate, and downstream conversion measured against a holdout. Each answers a different question and none of them answers all three.

Click to delivered is your creative metric. Use delivered rather than sent as the denominator, because bounces are a list problem and mixing them in makes a good email look weak. It tells you whether the message earned an action, and it is the number to judge copy, offer and send time on. Click to open ratio is dead for the same reason opens are.

Reply rate matters more in sales assisted motions than most teams admit. A plain text activation email from a named human that pulls 3 percent replies is worth more than one pulling 9 percent clicks into a docs page, because a reply creates a conversation an account executive can work. It also reads as a strong engagement signal to mailbox providers.

Spam complaint rate is now a hard operating constraint rather than a vanity concern. Google and Yahoo’s bulk sender requirements, which took effect in February 2024, set a ceiling of 0.3 percent, and in practice you want to sit under 0.1 percent. Watch it per sequence, not per account, because one aggressive winback campaign can poison delivery for your onboarding mail.

0.3%

Spam complaint ceiling for bulk senders under the Google and Yahoo requirements introduced in February 2024

Google and Yahoo bulk sender requirements

Downstream conversion is the only one that carries revenue, and it is only trustworthy against a control group. That is the rest of this page. If you want the rates other teams are hitting on the first three, our SaaS email benchmarks split them by sequence type and ACV band rather than averaging a newsletter together with a password reset.

Editable CSV worksheet

SaaS benchmark evaluation worksheet

Record the source, date, cohort and metric definition before comparing your numbers with a benchmark.

We never sell your data. Your resource opens here after submission.

How to design a holdout when you only get 800 trials a month

Start from the effect you need to detect, not from the split you find comfortable. A reasonable approximation for a two proportion test at 80 percent power and 95 percent confidence is sixteen times the baseline variance divided by the squared difference you want to catch. With a baseline trial to paid rate of 18 percent, the variance term is 0.18 multiplied by 0.82, which is 0.1476.

Work an example. A product takes 800 trials a month, converts 18 percent of them, and the team wants to know whether the six email onboarding sequence does anything at all.

Lift you want to detectSplitUsers per armTotal usersMonths of enrolment at 800 trials
18% to 21% (+3 points)50/502,6245,2486.6
18% to 23% (+5 points)50/509451,8902.4
18% to 26% (+8 points)50/503707401.0
18% to 21% (+3 points)90/101,458 control, 13,122 treated14,58018.2

Read the last row twice. An unbalanced 90/10 design needs nearly three times the total population to answer the same question, because the small arm governs the standard error. That is why the split you use to test is not the split you use to monitor.

So the rule at low volume is this. Run the initial test at 50/50, accept that you can only detect large effects, and be honest that a sequence producing a two point lift is invisible to you and always will be at this volume. Once the sequence has proven itself, drop to a permanent 5 to 10 percent holdout as a regression alarm rather than a measurement instrument.

Running your first sequence holdout

  1. Pick one sequence and one outcome

    Onboarding measured on trial to paid, not on four metrics at once. You will be tempted to measure everything. Do not.

  2. Randomise at the account level

    Hash the account ID, not the user ID, so two people in the same workspace get the same treatment. Otherwise the treated user forwards the email to the control user.

  3. Freeze the sequence

    No copy edits, no new steps, no send time changes until enrolment closes. A mid test change makes the result unreadable.

  4. Enrol for a fixed window

    Calculate the weeks needed from the table above and put the end date in the calendar. Do not peek at day nine and stop.

  5. Wait one full conversion cycle

    Add the trial length plus the tail. For a 14 day trial, wait 21 to 30 days after the last enrolment before you read anything.

  6. Compare rates and revenue per user

    Conversion rate tells you if it worked. Revenue per enrolled user tells you whether it moved anyone up a plan or only pulled cheap accounts forward.

  7. Publish the result including a null

    A sequence with no measurable lift is a finding worth the same respect as a win. Turn it off or rewrite it, then say so in the channel.

When a permanent holdout is worth the revenue it costs

Once a sequence is live and proven, keep 5 to 10 percent of eligible accounts out of it permanently. That is our position, and it costs money, so here is the arithmetic.

Say your dunning sequence recovers 38 percent of failed payments and failed payments represent $50,000 of MRR a year. A permanent 5 percent holdout forgoes roughly 5 percent of the recovery, about $950 a year. In exchange you get a live control that tells you the day a deliverability problem, a broken merge tag or a platform migration quietly halves your recovery rate. Most teams discover those failures six weeks late through a revenue dip nobody can explain.

The holdout is worth it when three conditions hold. The sequence runs continuously rather than as a campaign. The outcome it drives is measurable inside 30 days. And the cost of a silent failure is larger than the cost of the withheld sends. Dunning, onboarding and trial expiry clear that bar comfortably. A quarterly product newsletter does not, and forcing a permanent holdout onto it is theatre.

Holdouts leak

A holdout only measures email if email is the only thing that differs. If your in app messages fire on the same triggers, if customer success emails manually, or if the account is in a paid retargeting audience, you are measuring the gap between two overlapping programs. Tag holdout accounts in the CRM and suppress them from parallel campaigns, or state the leak openly in the write up.

Incrementality and switchback tests for always on sequences

User level randomisation fails in two specific SaaS situations. When several people share a workspace and forward things to each other, and when the sequence is tied to a product surface you cannot hold back, such as an in app banner that fires off the same event.

Switchback testing solves both by randomising time instead of people. Turn the sequence on for a week, off the next, and repeat. Then compare the outcome for cohorts that entered during on weeks against cohorts that entered during off weeks. You need enough blocks for the comparison to mean anything, and sixteen weekly blocks is the practical floor. Fewer than that and a single bad week of seasonality dominates the result.

Two things ruin switchbacks. Carryover, where someone entering during an off week still gets touched because they were already in a related sequence, and seasonality, where your on weeks happen to land on the first week of each month. Randomise the block order rather than alternating strictly, and check that on and off weeks are balanced across the month.

DesignWhen to use itMinimum runMain weakness
Balanced user splitOne off test of a new sequence6 to 10 weeks of enrolmentLeaks inside shared workspaces
Account level splitAny B2B product with multi seat accountsSame as aboveNeeds more accounts than users
Permanent 5 to 10% holdoutProven, always on sequencesRuns foreverDetects only large regressions
Switchback by weekSequences tied to in app surfaces16 weekly blocksSeasonality and carryover
Staggered start (48h delay)Dunning and other service adjacent mail8 weeksMeasures timing, not the message itself
Pick the weakest design that answers your question. Every extra layer of rigour costs weeks.

Per sequence attribution rules that stop email double counting with sales

Set the rules before you need them, in writing, and get the sales leader to agree in the same meeting. The argument you are avoiding is the one where marketing claims 60 percent of pipeline, sales claims 70 percent, and finance stops reading either deck.

Three rules cover most cases. First, a click must precede the conversion inside the attribution window. An open never qualifies, for the reasons above. Second, if an account executive logged a call, meeting or manual email inside that same window, the deal is sales sourced and email influenced, and those go in separate columns. Third, the window varies by sequence, because a trial expiry email that works does so within 72 hours, while a winback might land three weeks later.

SequenceWindowQualifying eventCredit rule
Welcome and onboarding30 daysClick then activation milestoneActivation lift versus holdout, no revenue claim
Activation nudge7 daysClick then feature first useActivation lift only
Trial expiry72 hoursClick then upgradeEmail sourced unless an AE touched the account
Dunning14 daysClick then successful paymentInvoluntary churn recovered, always email sourced
Expansion and upsell30 daysClick then seat or plan increaseInfluenced, never sourced, if CS was involved
Winback45 daysClick then reactivationEmail sourced, flagged as low confidence
NewsletterNoneNoneNo revenue claim, measured on retention of engagement

Notice what is missing. The newsletter gets no revenue line. Teams hate this, and it is the right call, because a newsletter that a customer reads for eight months before renewing is doing something real that no window based rule can capture without inventing a number. Measure it on subscriber retention and reply quality, and defend it as a brand investment rather than pretending it closed deals. The same discipline applies on the ads side, and we walk through it in paid media attribution for SaaS.

Editable CSV worksheet

Save your marketing measurement plan

Keep a worksheet for your inputs, assumptions and next actions. You can also print the calculation directly from your browser.

We never sell your data. Your resource opens here after submission.

The four numbers to put in front of a board

Boards do not want seven email metrics. They want to know whether the program pays for itself and what would break if it stopped. Four numbers do that job.

Incremental MRR. The holdout derived difference in revenue per enrolled account, multiplied by the treated population, stated monthly. One number, one method, restated every quarter with the confidence interval attached. If the interval crosses zero, say so.

Activation lift. The percentage point difference in accounts reaching your activation milestone between treated and held out groups. This is where onboarding email earns its keep, and it usually shows a clearer effect than revenue does because it happens sooner and closer to the message.

Involuntary churn recovered. Dollars of failed payments rescued by dunning, net of the holdout baseline recovery that would have happened through in product retries anyway. Most teams report gross recovery here and overstate it by 30 to 50 percent, since some customers fix their own cards without being asked.

Expansion MRR influenced. Clearly labelled as influenced. This is the number a CFO will push on hardest, so give it the weakest claim and the clearest definition.

We reported a 12 percent email attributed revenue share for two years. When we finally ran a holdout the true incremental number was closer to 4 percent. That was a bad quarter, and the best thing we ever did for the program’s credibility.
A lifecycle lead , Series B developer tools company

Model those four against cost before you present them. Our SaaS email revenue calculator takes list size, click to delivered, conversion and ACV and returns incremental revenue and program payback, which is a more useful shape for a board than a funnel chart.

What this costs, and the three ways it fails

The honest cost is time, not tooling. Designing a clean test, tagging holdouts across your platform and CRM, suppressing parallel campaigns and writing the analysis takes a competent marketing operations person roughly two days per sequence up front and half a day a month after that. At a loaded rate that is a few thousand dollars a year, against a program usually worth six figures.

Three failure modes account for most of the disappointments we see.

The first is underpowered tests read as wins. A team at 300 trials a month runs a four week test, sees 19.4 percent against 18.1 percent, and ships the sequence. That difference is noise at that sample size, and the team now believes something false. If your volume cannot support the test, say the sequence is unmeasured rather than dressing up a coin flip.

This second is tooling that cannot hold a group out cleanly. List based platforms often struggle to keep a stable random assignment across sequence entries and exits, so the holdout drifts and the comparison stops being random. Event based platforms handle it natively through a random bucket attribute set at account creation, which is one of several reasons we favour them in our comparison of lifecycle email platforms, and a specific point of difference in HubSpot versus Customer.io. If you run Mautic, you will be assigning the bucket yourself in the database, which works but has to be built.

The third is organisational. Someone senior asks why 8 percent of trials are not getting onboarding email, and the holdout quietly gets switched off. Pre-empt it. Write down the annual cost of the holdout in dollars, put it next to the cost of the last silent sequence failure, and get the decision made once at the right level.

Start here in the next two weeks

Pick the sequence with the most obvious revenue outcome, which is almost always trial expiry or dunning, and put a holdout on it before you touch a single word of copy. Then get the four board numbers defined and agreed with sales while the test runs, because that conversation takes longer than the test does.

Before you call your email program measured

0 of 7 done

If you are earlier than that and still building the first sequences rather than measuring them, start with the lifecycle email fundamentals and then ship your first sequence with the holdout built in from day one. Retrofitting a control group onto a sequence that has been running for a year is possible, but you lose the clean baseline and spend the first month arguing about what changed.

Editable CSV worksheet

SaaS Email Marketing planning worksheet

A practical email planning worksheet: decisions, owners, evidence and next actions.

We never sell your data. Your resource opens here after submission.

Frequently asked questions

How do you measure email marketing ROI in SaaS?

Run a holdout. Randomly exclude a slice of the eligible audience from the sequence, wait one full conversion cycle, then compare paid conversion rate and revenue per user across the two groups. Multiply the difference by the treated population to get incremental revenue, divide by the fully loaded program cost, and report that. Anything derived from last touch is a ceiling, not a measurement.

What size should an email holdout group be?

For a one off test, split the audience evenly, because a balanced design needs the fewest total users to reach significance. For ongoing monitoring after a sequence has proven itself, hold back 5 to 10 percent permanently. Below about 3 percent the control arm is too small to detect anything short of a catastrophic regression, and you are paying the cost of a holdout without the information.

Is open rate still useful for anything?

It has one surviving use. A sudden collapse in opens across an entire domain is a deliverability alarm worth investigating, because machine opens are consistent enough that their disappearance means something changed at the mailbox provider. It cannot judge subject lines, cannot drive re-engagement segments, and cannot appear in a revenue calculation. Use click to delivered for creative decisions.

How long should an email incrementality test run?

Long enough to cover the full conversion window plus the time needed to accumulate sample. If your trial is 14 days and most upgrades land by day 21, the earliest honest read is 21 days after the last user enters the test. Most B2B programs need six to twelve weeks of enrolment on top of that. Stopping early because the lift looks good is how teams ship sequences that do nothing.

Should email get credit if sales also touched the account?

No, not as sourced revenue. Mark it as influenced and keep it in a separate column. The rule that survives an argument with a sales leader is simple. If an account executive logged a call, meeting or manual email inside the attribution window, the deal counts as sales sourced and email influenced. Double counting across two teams is the fastest way to lose credibility with a CFO.

What is a switchback test for email?

Instead of splitting people, you split time. The sequence runs for a week, is paused for a week, and the cycle repeats for several months. You then compare cohorts that entered during on weeks against cohorts that entered during off weeks. It is useful when user level randomisation leaks, for example when several people on the same account share a workspace and talk to each other.

Can you hold out dunning emails?

Carefully, and rarely. Payment failure notices are close to a service message, and in some contracts and jurisdictions you are obliged to send them. The safer design is a staggered start rather than a true holdout. Send the recovery sequence to everyone, but delay it by 48 hours for a random slice, then measure recovery rate in that window. You get a clean read without withholding the message.

The saas-marketing.net editorial team Research and editorial

We research, write and maintain every page on this site. The library explains marketing decisions through practical frameworks, explicit assumptions and references. Corrections can be requested through the contact page.

Published September 11, 2026. Last updated .