Measuring customer marketing
The customer marketing metric set, holdout and matched cohort designs, and how to report post sale impact to a board that does not trust attribution models.
On this page 10 sections
- Why attribution models break after the sale
- The seven metrics worth reporting
- Experiment designs that survive scrutiny
- Withholding a program from real customers
- The monthly dashboard
- The quarterly board slide
- The failure mode that ends careers
- What to measure for each program type
- What this costs
- Start here
- Frequently asked questions
The short answer
Customer marketing is measured with seven metrics: activation rate, time to first value, adoption breadth, net revenue retention contribution, save rate, advocacy influenced pipeline and reference attached win rate. Multi-touch attribution does not work after the sale because the touchpoints are not competing for credit on a single purchase decision. Use holdout groups and matched cohorts instead, publish the method beside every number, and report lift you can reproduce next quarter.
Key points before you start
Customer marketing is the first line cut in a bad quarter, and the reason is almost always the same. The function reports 340 webinar registrants and a 41% email open rate, while demand gen reports $4.2M in sourced pipeline. One of those looks like revenue. The other looks like activity, even when it is doing more for the business.
Fixing this is a measurement problem, and it’s not solved by finding a better attribution model. It’s solved by running experiments.
Why attribution models break after the sale
Multi-touch attribution answers one question: given a purchase, how should credit be split across the touches that preceded it. That question requires a conversion event to allocate against.
After the sale, that event mostly does not exist. A renewal is the absence of a decision. Expansion happens gradually across quarters, often triggered by a usage threshold nobody marketed to. Advocacy influences somebody else’s purchase entirely, which means the touch and the revenue sit on different account records.
Run a U-shaped model over renewal data and you will get numbers. They will be arithmetically correct and tell you nothing, because the model is answering a question your data was not generated by. Worse, a finance team that notices this stops trusting every number the marketing org produces, including the ones from before the sale that were fine.
The trap
Attribution software will happily produce post-sale credit splits. Being able to compute a number is not evidence that the number means anything.
The replacement is comparison. What happened to accounts that got the program versus comparable accounts that did not. That’s the only question with a defensible answer, and it’s the foundation of every design below.
The seven metrics worth reporting
Everything else is a diagnostic. Diagnostics are fine internally and poison in a board deck.
| Metric | Definition | Reportable to the board | Common failure |
|---|---|---|---|
| Activation rate | Share of new accounts reaching the defined first-value event within 30 days | Yes, with the event definition stated | Definition drifts quietly and quarters stop comparing |
| Time to first value | Median days from provisioning to that same event | Yes | Median hides a bimodal distribution, show p25 and p75 too |
| Adoption breadth | Share of accounts using two or more core capabilities | Yes | Counting logins as adoption |
| NRR contribution | Retention or expansion lift measured against a holdout | Yes, with holdout size | Claiming total NRR rather than measured lift |
| Save rate | Share of flagged at-risk accounts retained after intervention | Yes, with the flag criteria | Flagging accounts that were never really at risk |
| Advocacy influenced pipeline | Pipeline where a reference, case study or review touched the opportunity | Yes, framed as influence | Presenting influence as source |
| Reference attached win rate | Win rate of deals with a reference call versus those without | Yes, with the selection caveat | Ignoring that better deals get references |
That last caveat deserves its own sentence. Deals that get reference calls are systematically further along and better qualified than deals that do not, so the win rate gap overstates the reference effect. Say so in the footnote. A board that sees you caveat your own favourable number believes the rest of the deck.
Engagement metrics stay off the slide. Webinar attendance, email opens, community posts and NPS response rates are how you diagnose a program that stopped working. They are not outcomes, and presenting them as outcomes is precisely what taught your CFO to discount the function.
Editable CSV worksheet
SaaS benchmark evaluation worksheet
Record the source, date, cohort and metric definition before comparing your numbers with a benchmark.
Experiment designs that survive scrutiny
Three designs, in descending order of strength. Pick the strongest one your account volume supports.
Choosing a design
- Randomised holdout, the default
Randomly exclude 5 to 10% of eligible accounts from the program. Compare outcomes after a full cycle. This is the only design that answers the counterfactual cleanly. Needs enough accounts to power the comparison.
- Matched cohorts, when you cannot randomise
Match treated accounts to untreated ones on ARR band, segment, tenure, seat count and product tier. Compare outcomes. Weaker, because unobserved differences remain, but far better than nothing.
- Segment or geo split, for rollouts
Roll the program to one region or segment first and use the other as a comparison. Works well for onboarding changes. Contaminated if the segments differ in ways that matter, which they usually do somewhat.
- Before and after, last resort
Compare the same population before and after launch. Seasonality, pricing changes and product releases all confound it. Use only for large effects and label it clearly as suggestive.
- Register the design before launch
Write down the metric, the expected effect, the sample and the end date. Analysing after the fact and choosing the flattering cut is the most common way these numbers become fiction.
Sample size is where most of these designs quietly fail. Rough guidance, derived from standard two-proportion power calculations at 80% power and 95% confidence:
- Detecting a 5 percentage point difference in retention around a 90% base rate needs roughly 200 accounts per group.
- Detecting a 10 point difference in activation rate around a 50% base rate needs roughly 80 to 100 per group.
- Detecting a 2 point difference in anything needs samples most B2B SaaS companies do not have, so stop trying.
If your total eligible population is under 300 accounts, holdouts will not resolve realistic effects. Use matched cohorts, accept the weaker inference, and be honest about it. The customer marketing ROI calculator will let you plug in your own account counts and see what effect size you could actually detect before you commit to a design.
200 per side
Accounts needed to detect a five point retention difference at 80% power
saas-marketing.net model, method shown on the page
Withholding a program from real customers
The ethical objection comes up every time, and it deserves a straight answer rather than a dismissal.
You are not withholding the product or support. You’re withholding a marketing program whose value is unproven, which is exactly the thing you are trying to establish. If the program works, the holdout ends and everyone gets it. If it does not work, you have saved the excluded accounts from a set of emails nobody wanted.
Two limits worth holding to. Never hold out anything safety, security or billing related. And never hold out from accounts already flagged at risk, because the potential cost of being wrong is a churned customer rather than a slightly worse quarter.
Make the holdout permanent
Keep a standing 5% holdout across all lifecycle programs rather than creating one per test. It gives you a continuous baseline, removes the setup argument each time, and after a year you can measure the cumulative effect of the entire lifecycle motion rather than one campaign.
The monthly dashboard
Four blocks, one screen, same layout every month. Layout stability matters more than the specific charts because people learn where to look.
Block one, funnel health. Activation rate and time to first value by monthly cohort, twelve months trailing. Cohort view only. Aggregate activation rate is almost meaningless because it mixes accounts at different tenures.
- Block two, adoption. Breadth of feature use by segment, with the count of accounts using only one core capability called out separately. That single-capability group is your churn pipeline and it belongs on the dashboard permanently.
Block three, programs in flight. Each active program with its design, its holdout size, days elapsed and current directional read. Directional only. No claiming a result before the registered end date.
Block four, advocacy supply. References available, references used this month, case studies in production, and G2 or peer reviews collected. Supply constraints here show up as sales complaints three months later, so it is worth watching as an inventory number.
Pull the program list from your customer marketing plan so the dashboard and the plan stay in sync. When they diverge, the dashboard is describing a program that no longer exists.
Review request
Free SaaS marketing audit
Share your site, stage and priorities to request a review of your positioning, funnel and acquisition plan.
The quarterly board slide
One slide. Three numbers. The method printed underneath each, in the same size font as the number.
Something like this, written out:
Onboarding email sequence. Activation at day 30: 61% treated versus 52% holdout. Design: randomised 10% holdout, 940 accounts treated, 104 held out, Jan to Mar. Difference is nine points, confidence interval roughly plus or minus seven points. Directionally positive, not yet precise.
At-risk save program. 38% of flagged accounts retained through renewal versus 29% in a matched cohort. Matched on ARR band, tenure and seat count. 212 accounts per side. Estimated retained ARR from the difference: $340K.
Advocacy influenced pipeline. $6.1M of open pipeline has a reference call, case study or review touch on the opportunity. Reported as influence, not source. Deals receiving references are further along on average, so this overstates the causal effect.
That reads as less impressive than a slide claiming $6.1M of advocacy sourced pipeline. It will survive a CFO question, which the other slide will not, and surviving the question is what keeps the budget.
The general rule: report influence with holdouts rather than credit splitting, and never present a lift you cannot reproduce next quarter. If the number moves twenty points between quarters with no program change, your measurement is noise and presenting it once means presenting the reversal later.
The failure mode that ends careers
Claiming revenue customer success would have retained anyway.
It’s tempting because the numbers are large and available. Renewal value touched by a lifecycle email is a query anyone can run. Present it once and the CS leader will, correctly, point out that they ran the renewal conversation, the QBR and the escalation, and that your email arrived on a Tuesday.
You lose that argument. You should lose it. The only claim worth making is the measured difference between accounts that got the program and accounts that did not, which is a smaller number that belongs to you entirely.
The same logic applies to the retention content system. Documentation, in-product guides and adoption content plainly help retention, and the honest way to size that help is a holdout on a content-driven adoption push, not a claim on the whole retained base.
What to measure for each program type
Different programs need different designs, and matching them wrong is how good programs get killed.
| Program | Primary metric | Design | Minimum viable sample |
|---|---|---|---|
| Onboarding sequence | Day 30 activation rate | Randomised holdout | 100 accounts per side |
| Feature adoption campaign | Adoption of that feature at 60 days | Randomised holdout | 150 per side |
| At-risk intervention | Retention through next renewal | Matched cohort, never randomised | 200 per side |
| Customer webinar series | Adoption breadth change at 90 days | Matched cohort on attendance | 120 per side |
| Advocacy and reference program | Reference attached win rate | Observational with stated caveat | 60 deals with references |
| Community | Support ticket deflection and retention | Matched cohort on join date | 200 per side |
Community is the hardest of these and worth flagging. People who join a community are systematically more engaged before they join, so every community retention number in circulation overstates the effect. Match on pre-join engagement or accept that the figure is decorative. The campaign designs in customer marketing campaign ideas are worth pairing with the matching design here before you run any of them.
What this costs
An analyst’s time, mostly. Setting up a standing holdout takes a few days of data engineering once. Running a properly registered experiment adds maybe half a day of design per program and a day of analysis at the end.
The real cost is slower claims. You cannot report a result in week two, and for a function under budget pressure that delay feels dangerous. It is the right trade. One defensible number per quarter beats four impressive ones that collapse under a follow-up question, and the functions that survive budget cuts are the ones finance believes.
Start here
Pick your largest lifecycle program. Carve out a random 10% holdout this month, write the design down before you launch, and set the analysis date ninety days out. While it runs, rebuild your dashboard around the seven metrics and move everything else into a diagnostics tab. Get the vocabulary straight too, because customer advocacy and customer lifecycle marketing get used interchangeably in most orgs and that makes the reporting muddier than it needs to be.
One quarter later you will have one number with a method attached. That’s worth more than the whole previous year of engagement reporting, and the broader context for where it fits sits in SaaS customer marketing.
Editable CSV worksheet
SaaS Customer Marketing planning worksheet
A practical retention planning worksheet: decisions, owners, evidence and next actions.
Frequently asked questions
How do you measure customer marketing ROI?
Run holdouts. Withhold the program from a randomly selected 10% of eligible accounts, compare retention, expansion or activation between the two groups after a full cycle, and report the difference as lift. That number is defensible because it answers what would have happened without the program, which no attribution model can answer after the sale.
Why does multi-touch attribution fail for post-sale marketing?
Attribution models distribute credit across touches leading to one conversion event. After the sale there is no single event. Retention is the absence of a decision, expansion happens across months, and advocacy influences other people's purchases. Applying a U-shaped model to a renewal produces a number that is arithmetically valid and evidentially meaningless.
What metrics should a customer marketing team report?
Seven: activation rate, time to first value, adoption breadth, net revenue retention contribution measured against a holdout, save rate on at-risk accounts, advocacy influenced pipeline, and reference attached win rate. Report engagement metrics like webinar attendance internally as diagnostics, never to the board as outcomes.
How large does a holdout group need to be?
It depends on the effect you expect. To detect a five percentage point difference in retention between groups at typical B2B rates, you need roughly 200 accounts per side. For activation rate differences of ten points or more, 80 to 100 per side is usually enough. If your total eligible population is under 300 accounts, holdouts will not resolve small effects and you should use matched cohorts instead.
What is advocacy influenced pipeline?
Pipeline where a reference call, case study, review or customer-hosted event touched the opportunity before it closed. Track it as a flag on the opportunity record, not as a revenue split, and report it as influence rather than source. The honest framing is that these deals had advocacy involved, not that advocacy caused them.
How do you report customer marketing to a skeptical board?
One slide, three numbers, each with the method printed underneath in the same size font. Show the holdout design, the sample size and the confidence range. A board that distrusts marketing numbers distrusts them because nobody ever showed the method. Showing it, including when the result is weak, converts skeptics faster than a bigger number would.
Should customer marketing take credit for renewals?
Not for the renewal itself. Customer success owns that, and claiming it starts a turf fight you will lose. Claim the measured lift: the difference between accounts that received your program and comparable accounts that did not. That number is smaller, defensible, and worth more in a budget conversation than an inflated one.
The saas-marketing.net editorial team Research and editorial
We research, write and maintain every page on this site. The library explains marketing decisions through practical frameworks, explicit assumptions and references. Corrections can be requested through the contact page.
Published September 11, 2026. Last updated .