Lead scoring for B2B SaaS
Build a lead scoring model from fit and behaviour data, validate it against closed won history, and retire the point systems that sales quietly ignores.
On this page 8 sections
- Why workshop derived scoring models fail
- The two axis model, and why one number is not enough
- Derive weights from closed won history, with a worked example
- Negative scoring and decay, the two things most models lack
- Validate by cohort before sales sees a single score
- Wire it into routing and the SLA
- Recalibration and the signs of drift
- The honest tradeoff
- Frequently asked questions
The short answer
A working B2B SaaS lead scoring model is built backwards from closed won data, not from a workshop. Score fit and behaviour on separate axes, derive attribute weights from historical win rates rather than opinion, and route on the four quadrants instead of one threshold. Validate by cohort before sales sees a single score: if high scoring leads do not close at a materially higher rate than low scoring ones, the model is decoration and should not ship.
Key points before you start
Here is the test for whether your lead scoring model is real. Pull last quarter’s closed won deals, look up the score each lead had when it was passed to sales, and plot win rate by score band. If the line is flat, you have been routing leads at random with extra steps for however long the model has been live.
Scoring is a prediction problem. Build it like one.
Why workshop derived scoring models fail
Because the point values come from opinions and opinions are not evidence. Somebody says a demo request is worth 25 points, somebody else argues webinar attendance deserves 15, and the number that ships is the average of the room’s confidence rather than anything about your customers.
The failure is not that the weights are wrong. It is that nobody can tell whether they are wrong, so the model never improves. Sales tries it for six weeks, hits three enthusiastic 90 point leads that go nowhere, and reverts to working the list by company logo. After that the score field is furniture.
The worst outcome is not a bad model
It is a bad model nobody has checked. An unvalidated model actively hides converting leads by ranking them low, which is worse than no model at all, because with no model reps at least read every lead. Scoring concentrates attention, and concentrating attention on the wrong accounts costs more than spreading it evenly.
The two axis model, and why one number is not enough
Fit and behaviour answer different questions. Fit asks whether this account could ever be a good customer. Behaviour asks whether they are in market right now. A single composite score lets a highly engaged bad fit score the same as a quiet perfect fit, and those two leads need completely different treatment.
Score each on a 0 to 100 scale, then route by quadrant.
| Quadrant | Meaning | Routing rule | Realistic expectation |
|---|---|---|---|
| High fit, high behaviour | In market and worth winning | Route to AE within 15 minutes, Chili Piper style instant booking | Highest win rate band, should be 3x baseline |
| High fit, low behaviour | Right account, wrong time | Nurture plus outbound sequence, no SDR call blitz | Converts on a 2 to 6 month lag, do not judge weekly |
| Low fit, high behaviour | Enthusiastic and unbuyable | Self serve path only, no human time | Occasionally becomes a champion at their next company |
| Low fit, low behaviour | Noise | Newsletter, nothing else | Do not spend SDR minutes here at all |
The low fit, high behaviour quadrant is where most single threshold models leak money. A student researching a dissertation can trip every behavioural trigger you own. Under a composite score they become a 78 and land in an SDR queue. Under a two axis model they land in a self serve nurture and cost nothing.
Derive weights from closed won history, with a worked example
Export 12 months of opportunities, won and lost, with the lead attributes present at the time of creation. You need the lost ones. A model built only on wins tells you what your customers look like, not what distinguishes buyers from non buyers, and those are different questions.
Compute win rate per attribute value against your baseline. Here is a real shaped example from a mid market workflow product with a 9.2% baseline win rate on qualified opportunities.
| Attribute | Win rate | Lift vs baseline | Derived weight |
|---|---|---|---|
| 200 to 1,000 employees | 17.4% | 1.89x | +18 fit |
| Under 20 employees | 2.1% | 0.23x | -15 fit |
| Director or VP title | 14.8% | 1.61x | +14 fit |
| Uses Salesforce (from enrichment) | 16.1% | 1.75x | +16 fit |
| Requested demo | 31.2% | 3.39x | +34 behaviour |
| Viewed pricing twice in 7 days | 22.6% | 2.46x | +24 behaviour |
| Downloaded a top of funnel ebook | 7.9% | 0.86x | +2 behaviour |
| Attended a webinar | 8.8% | 0.96x | 0 behaviour |
Two things usually surprise people in this table. The ebook and the webinar, which the team had been counting as strong MQL signals, carry essentially no predictive power. And enrichment data about tech stack outperforms most self reported form fields, because people lie on forms and Clearbit does not.
Scale the weights so each axis tops out near 100. Do not add attributes that show lift below about 1.2x; they add noise and make the model harder to explain.
Editable working copy
Download this template
Save an editable working copy of the framework on this page. Add your own owners, evidence and decisions.
Negative scoring and decay, the two things most models lack
Negative scoring removes the structurally unbuyable. Free email domains where your product sells to companies, competitor domains, job seeker and student titles, countries outside your sales coverage, and careers page visits, which correlate strongly with people looking for work rather than software.
In most B2B SaaS databases somewhere between a fifth and a third of inbound leads fall into this group. Removing them from the queue is the single largest immediate improvement to lead quality available, and it requires no modelling at all.
Decay handles the other half of the problem. Behaviour is a claim about now, so it has to expire. A pricing page visit from March tells you nothing in September.
A decay rule that works
- Halve behaviour score every 30 days of inactivity
A 60 point behaviour score drops to 30 after a month of silence, 15 after two.
- Never decay fit score
Company size and tech stack do not expire. Refresh them on enrichment cadence instead, quarterly is fine.
- Reset decay on any new behavioural event
One pricing visit restores the clock, not the full previous score.
- Floor behaviour at zero
Negative behaviour scores double count with negative fit scores and distort the quadrants.
- Exempt demo requests for 14 days
A demo request deserves a full window of sales attention regardless of subsequent silence.
Validate by cohort before sales sees a single score
This is the step that separates a model from a decoration, and it takes an afternoon.
Hold out the most recent quarter. Score those leads with your new weights as if the model had been live, then plot actual win rate by score decile. You are looking for a monotonic curve where each decile converts better than the one below it, and a top quartile that closes at two to three times the bottom quartile.
2 to 3x
Win rate lift of top quartile over bottom quartile that a validated model should show on holdout data
saas-marketing.net model, method shown on the page
If the curve is flat, the attributes you chose are not predictive and adding more will not save it. That usually means one of three things: your ICP is too broad to differentiate, your form data is too thin, or the real signal lives in product usage rather than marketing behaviour. The third case is common in product led companies, and PQL scoring and routing is the right model for those teams rather than this one.
If the curve is steep but only in the top decile, you have a threshold model rather than a scoring model. That is fine, and simpler. Route the top decile and stop pretending the other nine bands mean anything.
Present the curve, not the model
When you launch to sales, show the backtest chart first and the point values second. Reps do not need to agree with the weights. They need to see that leads scoring above 70 closed at 24% last quarter while leads under 30 closed at 4%. That chart is the entire adoption argument, and it is why definitions like the marketing qualified lead only carry weight when they are tied to evidence.
Wire it into routing and the SLA
A score with no routing consequence changes nothing. Each quadrant needs a named destination, a response time and a rejection path, and all three belong in writing with sales. The sales and marketing SLA template covers the contract language; the part specific to scoring is the rejection loop.
Reps must be able to reject a routed lead with a reason code, and those reason codes feed the next recalibration. When 40% of high fit high behaviour leads come back marked “no budget”, you have learned something about a segment that no amount of attribute weighting would have surfaced.
Volume planning changes too, because a validated model means fewer MQLs of higher quality, and the lead goal arithmetic has to be redone. Run the new conversion rates through the lead goal calculator before you commit to a number, and check the economics against the cost per lead calculator and lead value calculator, since a better model usually raises cost per MQL while lowering cost per opportunity.
Review request
Free SaaS marketing audit
Share your site, stage and priorities to request a review of your positioning, funnel and acquisition plan.
Recalibration and the signs of drift
Quarterly, one hour. Plot win rate by score decile for the quarter that just closed and compare the curve to last quarter’s.
Four signals that the model has drifted: the curve flattens, the top decile’s win rate falls while the baseline holds, rep rejection rate on high scores climbs above roughly 20%, or a single attribute starts appearing in most high scores because a campaign flooded one behaviour. That last one is the sneakiest. Run a webinar to 2,000 people and suddenly a third of your database looks engaged.
Rederive weights fully once a year, or immediately after a pricing change, a new product line, or an ICP shift. Those events invalidate historical win rates by attribute, and a model trained on the old ICP will confidently misroute the new one.
The honest tradeoff
Scoring costs maintenance forever and buys you prioritisation, not more leads. A validated model with quarterly recalibration takes perhaps 15 hours a quarter of somebody’s time, and its value shows up as SDR hours redirected rather than as pipeline created. If your team is three people and everyone works every lead anyway, build the negative scoring rules and skip the rest until volume justifies it.
Where it pays is at the point where inbound exceeds capacity. Then the question stops being how many leads and starts being which ones, which is the same question lead quality for B2B SaaS works through from the demand side, and where the broader SaaS lead generation hub picks up.
Start with the backtest on your existing model. If the curve is flat, you have your answer, and the lead scoring template gives you the structure to rebuild it properly.
Editable CSV worksheet
SaaS Lead Generation planning worksheet
A practical lead gen planning worksheet: decisions, owners, evidence and next actions.
Frequently asked questions
How do you build a lead scoring model for B2B SaaS?
Export 12 months of closed won and closed lost opportunities with their originating lead attributes. Calculate win rate for each attribute value, convert the lift over your baseline win rate into weights, and split those attributes into fit and behaviour. Backtest the resulting scores against a holdout period, then route by quadrant. The whole process takes a data-literate marketer about two weeks.
What is the difference between fit score and behaviour score?
Fit describes who the account is and does not change quickly: employee count, industry, tech stack, region, job title. Behaviour describes what they have done recently: pricing page visits, demo request, product signup, doc reads, email engagement. Fit predicts whether they could ever buy. Behaviour predicts whether they are buying now. Combining them into one number destroys that distinction.
How many points should a demo request be worth?
That is the wrong question, and asking it is how bad models get built. The right answer comes from your data: if leads who request a demo close at 31% against a 9% baseline, the demo request carries roughly three and a half times the baseline weight. Every company's answer differs because the friction on that form differs.
Should we use predictive lead scoring software?
Only once you have enough closed won volume for a model to learn from, which in practice means several hundred closed opportunities a year. Below that, a transparent rules model derived from win rates outperforms a black box, mainly because sales will argue with a score they cannot inspect and arguing with the model is how it improves.
How often should a lead scoring model be recalibrated?
Quarterly as a rhythm, and immediately after any ICP change, pricing change or new product launch. The check takes an hour: plot win rate by score decile for the last quarter's closed opportunities. If the curve has flattened, the model has drifted and the weights need rederiving.
What is negative scoring and does it matter?
Negative scoring subtracts points for signals that predict non purchase: free email domains, student or job seeker titles, competitor domains, careers page visits, countries you do not sell into. It matters more than most additions, because in many B2B SaaS databases a third or more of inbound leads are structurally unbuyable and they crowd out real ones in the queue.
Why does sales ignore our lead scores?
Almost always because nobody ever showed them that high scores convert better. A model derived from a cross functional workshop has no evidence behind it, so the first time a 92 point lead goes nowhere the rep stops looking at the field. Show the backtest curve in the launch meeting and the objection usually disappears.
The saas-marketing.net editorial team Research and editorial
We research, write and maintain every page on this site. The library explains marketing decisions through practical frameworks, explicit assumptions and references. Corrections can be requested through the contact page.
Published September 11, 2026. Last updated .