Self-Reported Attribution for B2B SaaS
Where to place the how did you hear about us question, the wording and answer options that produce usable data, and how to reconcile it with CRM attribution.
On this page 8 sections
The short answer
Self-reported attribution asks buyers directly how they heard about you, usually at the point of conversion. Make the field required on demo requests and optional on trials, use open text rather than a dropdown, and code responses into a fixed taxonomy weekly. Expect 60 to 85 percent completion on required fields. Treat it as the primary signal for podcasts, communities, word of mouth and AI assistants, which set no referrer and cannot be cookied.
Key points before you start
The cheapest fix for the dark funnel is one form field, and almost everyone implements it badly. They use a dropdown, they put it behind the conversion, they never code the free text, and eighteen months later the data sits unused in a field called hdyhau__c.
Done properly it is the only read you have on podcasts, private Slack groups, word of mouth and AI assistants. All four are invisible to your analytics, and together they are where a large share of B2B buying research now happens.
Where to put the question
Placement is a trade between response rate and form conversion, and the right answer differs by form type.
| Placement | Response rate | Form conversion cost | Best for |
|---|---|---|---|
| In-form, required | 60% to 85% | Low on high-intent forms | Demo requests, contact sales |
| In-form, optional | 30% to 55% | Near zero | Trials, content downloads |
| Post-conversion thank you page | 20% to 40% | None | Any form where you fear friction |
| First in-product screen | 40% to 70% | None, but delays time to value | Self-serve signup products |
| Sales call, asked verbally | High, but inconsistently logged | None | Supplement, never the system of record |
My position: required on demo requests, optional on trials. Someone filling in a demo form has already decided to talk to a salesperson and will not abandon over one more field. Someone starting a free trial is in a fragile moment and every extra input is a real cost.
The verbal version is worth building into the discovery call script anyway, because reps hear richer answers than a form ever captures. The problem is logging. If it does not go into a structured field, it does not exist, and asking reps to summarise it in call notes produces data you cannot count.
The mistake that ruins the dataset
Putting the answer in a free-text field that nobody ever codes. Six months later you have 1,400 uncoded strings and no capacity to process them, so the field gets quietly dropped in the next form redesign. Set up the weekly coding routine before you launch the field, not after.
Wording, and why the dropdown loses
The question itself should be short and neutral. “How did you first hear about us?” works. Adding “first” matters, because otherwise a meaningful share of people answer with the channel they used ten minutes ago, which is the thing your analytics already told you.
Avoid anything that suggests an answer. “Which of our channels brought you here?” primes people toward paid and search. “Did a colleague recommend us?” produces exactly the yes-bias you would expect.
Dropdowns fail for a structural reason. You can only list channels you already know about, so the output is a ranking of your existing assumptions. The interesting answers are the ones you cannot anticipate: a specific podcast episode, a named person in a Slack group, a comparison page on someone else’s site, an answer from an AI assistant. All of those arrive as free text or not at all.
The compromise that works when volume is high: open text as the primary field, with a small optional set of chips shown after submission for the small number of people who want to pick rather than type. Keep the typed answer as the record.
Editable CSV worksheet
SaaS benchmark evaluation worksheet
Record the source, date, cohort and metric definition before comparing your numbers with a benchmark.
Coding free text at scale
Coding is the work that makes the field valuable, and it is less work than people fear. Twelve to eighteen categories, applied weekly, taking 20 to 40 minutes at a few hundred responses a month.
A taxonomy that holds up for most B2B SaaS:
- Google or general web search
- AI assistant, with the assistant named where given
- Word of mouth, colleague or friend
- Community or Slack group, with the community named
- Podcast, with the show named
- Newsletter, with the newsletter named
- Social platform, split by platform
- Event or conference
- Review site such as G2 or Capterra
- Existing customer or previous user of the product
- Partner or integration marketplace
- Press or analyst coverage
- Paid advertising, where explicitly named
- Outbound contact from our team
- Unusable or unclear
That last bucket is not optional. Roughly 10 to 20 percent of required-field answers are things like “internet”, “online” or a single character, and forcing those into a real category corrupts everything downstream. Keep them counted and separate.
Two operational rules. Keep the named detail alongside the category, because “podcast” is not actionable and “the Exit Five podcast, the pricing episode” is. And have one person own the codebook, because two people coding the same answers differently is how a taxonomy quietly stops meaning anything.
An LLM can do a first pass on coding now, and it is good at it. Have a human review a 10 percent sample weekly and correct the codebook when the model keeps making the same mistake. That gets a few hundred responses coded in minutes rather than an hour.
Reconciling with CRM last-touch data
They will disagree. That disagreement is the most useful output of the whole exercise, so build the comparison view deliberately rather than treating it as a data quality problem to resolve.
| Channel | CRM last-touch share | Self-reported share | What the gap means |
|---|---|---|---|
| Organic search | High | Moderate | Search is the final step, rarely the first influence |
| Direct | High | Near zero, nobody says ‘direct’ | Direct is a bucket for everything untracked |
| Podcast | Near zero | Moderate in podcast-heavy categories | Podcasts set no referrer and produce direct visits later |
| Community | Near zero | Moderate | Private Slack and Discord links strip referrers |
| AI assistant | Low or zero | Rising sharply through 2025 and 2026 | Many assistant visits pass no referrer at all |
| Paid | Accurate | Understated | People rarely remember or admit clicking an ad |
| Word of mouth | Zero by definition | Often the single largest category | The channel your CRM cannot see is frequently the biggest |
The pattern is consistent across the SaaS teams I have seen do this. CRM data credits the last click, which is usually a branded search, and self-reported data names the thing that caused the branded search. Both are true statements about different moments.
Use them for different decisions. CRM data for optimising trackable channels where you need click-level feedback. Self-reported data for deciding whether to sponsor the podcast, whether the community investment is working, and whether your AI citation work is producing pipeline. This is the same split covered in multi-touch attribution vs incrementality testing, where the underlying question is which measurement method suits which decision.
12 to 18
Coding categories that keep a self-reported taxonomy both specific and reportable
saas-marketing.net model, method shown on the page
Editable CSV worksheet
Get the benchmark evaluation worksheet
A worksheet for checking source dates, definitions and sample limitations before you use an industry benchmark.
Sizing dark social and AI-referred pipeline
This is the payoff. Two channel groups now carry serious B2B influence and appear in analytics as roughly nothing.
Dark social is the first: private Slack communities, WhatsApp groups, LinkedIn DMs, forwarded newsletters. Links shared in these contexts strip referrer data, so the resulting visit is direct. The measurement approach is covered in depth in measuring dark social for B2B SaaS, and the underlying concept in the dark funnel.
AI assistants are the second and newer one. Some pass a referrer, many do not, and a buyer who asked an assistant for a shortlist and then typed your name into a browser leaves no trace of the assistant at all. If your coding taxonomy has no AI assistant category, your reported share of that channel is zero regardless of its real size.
To turn coded responses into a pipeline number, multiply the category share by pipeline value for the same cohort and report it as an estimate with the response rate attached. The honest framing for a board: “Of the 312 demo requests last quarter, 214 answered the question. Community and podcast together were named by 31 percent of those, associated with 1.4 million dollars in pipeline. We cannot verify this with platform data and we are not claiming precision.”
The limits, stated plainly
Self-reported data has three real weaknesses and you should name all three before someone else does.
People misremember. A buyer who saw four touches over six months names the most recent memorable one, which biases toward vivid channels like podcasts and events and against ambient ones like display and search.
Required fields produce noise. Making the field mandatory raises completion and lowers average quality, which is why the unusable bucket matters.
And it cannot be used for credit allocation. Do not pay commission, set channel budgets to the decimal, or settle an argument between two teams using it. It sizes channels and directs attention. That is the job.
Where this gets misused
A marketing leader presents self-reported data as sourced pipeline because it makes an underperforming channel look better. Finance eventually notices that the number cannot be reconciled with anything, and then the whole dataset loses credibility including the parts that were sound. Label it as survey data every single time it appears on a slide.
Implementation, in one week
Ship it in five days
- Add the field
Open text, required on demo forms, optional on trials. One custom field on the lead and contact object, mapped so it survives conversion to opportunity.
- Write the codebook
Twelve to eighteen categories plus unusable. Name one owner. Store it somewhere the whole team can read.
- Set the weekly routine
Thirty minutes, same slot each week. LLM first pass, human review of a 10 percent sample. You know it works when the backlog never exceeds one week.
- Build the side-by-side report
CRM last-touch share next to self-reported share, by quarter. The gaps are the insight, so do not hide them behind a blended model.
- Add it to the quarterly review
One slide, labelled as survey data, with the response rate shown. Present the gap between the two datasets rather than resolving it.
Where this sits in the measurement stack
Self-reported attribution is one input among several and it is the cheapest by a wide margin. For the surrounding system, SaaS metrics and analytics covers the reporting layer, six attribution models on one dataset shows how differently the same pipeline reads depending on the model you pick, and B2B SaaS attribution tools covers the software if you decide you need more than a field and a spreadsheet.
For context on what the answers usually look like across the category, where B2B SaaS pipeline actually comes from and B2B SaaS funnel conversion benchmarks are the two reference pages worth having open when you present your first quarter of coded data. The broader method is covered in self reported attribution for B2B SaaS.
Add the field this week. The coding routine is the hard part, and it is thirty minutes.
Editable CSV worksheet
SaaS Metrics and Analytics planning worksheet
A practical metrics planning worksheet: decisions, owners, evidence and next actions.
Frequently asked questions
What is self-reported attribution?
It is asking buyers directly how they heard about you, typically through a free-text field on a demo or signup form, and coding the answers into channels. It captures the influences that tracking cannot see, including podcasts, private communities, word of mouth, events and AI assistants, all of which usually arrive in analytics as direct or organic traffic.
Where should the how did you hear about us question go?
On the demo request form itself for sales-led motions, and on the first in-product screen or a post-signup survey for self-serve. In-form placement produces the highest response rate and links the answer to the record. Post-conversion placement protects form conversion but loses roughly a third of responses because people close the confirmation page.
Does adding a how did you hear about us field hurt form conversion?
Slightly on low-intent forms and barely at all on demo requests. A single optional free-text field on a high-intent form typically costs low single-digit percentage points of conversion at most, and often nothing measurable. Test it if you have the volume, but the data value usually outweighs the cost on demo forms.
Should the answer be open text or a dropdown?
Open text. Dropdowns lead people toward the options you listed, which means you only learn about channels you already knew about. Open text surfaces specific podcast names, community names, individual people and AI assistants, which is exactly the information you cannot get anywhere else. The cost is weekly coding work.
How do you reconcile self-reported attribution with CRM data?
Store both on the record and report them side by side rather than picking a winner. Where they agree, you have confidence. Where they disagree, the pattern is usually that CRM credits the last search and self-reported names the real influence. Use CRM data for channels that can be tracked and self-reported data for channels that cannot.
How do you measure traffic from ChatGPT and other AI assistants?
Partly through referrer data, which some assistants now pass, and mainly through self-reported attribution, because a large share of assistant-influenced visits arrive with no referrer or as direct traffic. Adding explicit coding categories for AI assistants to your taxonomy is the fastest way to size a channel most analytics setups currently show as zero.
What response rate should you expect from a self-reported attribution field?
Sixty to eighty five percent on a required demo form field, and twenty to forty percent when the field is optional or placed after conversion. Quality also varies: required fields produce more low-effort answers like 'internet', which is why the coding taxonomy needs a deliberate 'unusable' bucket rather than forcing every answer into a channel.
The saas-marketing.net editorial team Research and editorial
We research, write and maintain every page on this site. The library explains marketing decisions through practical frameworks, explicit assumptions and references. Corrections can be requested through the contact page.
Published September 11, 2026. Last updated .