Running an original research program
How to run a data study that earns citations: three data sources, methodology disclosure, the release kit, an annual refresh and a way to track who quotes you.
On this page 9 sections
The short answer
Original research content is a study you run and publish yourself, built on product data, a surveyed panel, or aggregated public filings. It earns citations because it creates a number nobody else has. A working program needs a named data source, a disclosed methodology block covering sample size, collection window, definitions and exclusions, an ungated report page with embeddable charts, and an annual refresh so the figure stays current.
Key points before you start
Most SaaS teams treat a data study as a campaign. Run it once, get a spike of links, move on. That is the expensive way to do it, because the cost sits almost entirely in the first study and the return sits almost entirely in the fifth. A research program with a fixed slot on the calendar compounds. A one-off decays inside a year and takes its citations with it.
This is the build order: pick the data source you can actually defend, write the methodology before you write the headline, ship it ungated, and then spend more on distribution than you spent on the research.
Which data source can you actually defend?
Three sources work for SaaS companies, and they differ mostly in cost and in how hard they are to attack. Everything else is either repackaged vendor marketing or a survey of your own customers dressed up as market data.
| Source | Cost | Credibility risk | Best for |
|---|---|---|---|
| Anonymised product data | 40 to 80 analyst hours, no external spend | Sample is your customers, not the market | Products with 5,000+ accounts and a metric nobody else can see |
| Surveyed panel | $8,000 to $25,000 for 300 to 600 responses | Self-reported numbers drift optimistic | Attitudes, budgets, org charts and plans |
| Public filings and datasets | 30 to 60 research hours | Aggregation errors, stale filings | Small teams with no data and no budget |
Product data is the strongest when you have scale. ChartMogul built a durable reputation on SaaS retention and growth benchmarks drawn from subscription data flowing through its own platform, and the reason people quote it is that no competitor can replicate the dataset. The catch is honest and unavoidable: your customers are not the market. If your product skews toward seed stage startups, say so in the first paragraph of the methodology and segment the results by company size so a reader can find the row that matches them.
Panel surveys buy you reach into companies that will never be your customers. They also buy you a specific weakness. People misremember their own numbers. Ask a marketing leader what their content budget was last year and roughly a third will give you a figure that does not match their own finance system. That is fine for attitudes and plans. It is shaky for hard financials, and you should label the difference rather than pretending it does not exist.
Aggregating public data is underrated. SaaS Capital has run an annual survey of private B2B SaaS companies for over a decade and it is cited constantly because the method is published and the series is long. Benchmarkit does something similar with operating metrics. Neither needs a huge user base. They need discipline.
The survey nobody can use
Running a 900 respondent survey and publishing only the top-line average is the most common waste in this category. The average is the one number that applies to nobody. Segment by ARR band, motion and ACV, or you have produced an anecdote with a large n attached. This is exactly the failure that makes most circulating SaaS content marketing benchmarks unusable.
What goes in the methodology block
Everything a hostile reader needs to reproduce or attack your work. Put it on the same page as the findings, not in a linked PDF, because the crawler and the language model need to see the method next to the number to treat it as credible.
Six items, in this order:
- Sample size and base. Total responses, plus the base for any figure that used a subset. If a chart is built on 84 responses, print 84 under the chart.
- Collection window. Exact dates. “Q2 2026” is not a window, it is a vibe.
- Definitions. What you counted as an activation, a qualified lead, a published asset. Half the contradictions between competing SaaS benchmarks come from two studies defining the same word differently.
- Exclusions. Who you removed and why. Incomplete responses, straight-liners, companies under 10 employees, whatever it was.
- Margin of error and confidence level. At 300 responses against a large population you are at roughly plus or minus 5.7 points at 95 percent confidence. Say it.
- Who funded and who ran it. If a vendor paid, disclose it.
This section is the reason the whole thing works. The category is drowning in numbers with no provenance, which means a study that shows its work stands out without needing a better headline. It is the same reason a well-sourced content to pipeline conversion benchmark outranks a blog post asserting the same figure from nowhere.
±5.7 points
Margin of error at 300 completed responses, 95 percent confidence, large population
saas-marketing.net model, method shown on the page
Editable CSV worksheet
Get the benchmark evaluation worksheet
A worksheet for checking source dates, definitions and sample limitations before you use an industry benchmark.
Gated or ungated? There is no debate here
Ungated. If the goal is citation, gating destroys the asset. A form wall means no crawler reads it, no language model quotes it, and no journalist on deadline bothers. You traded the entire upside for a list of email addresses, most of which are from people who wanted the chart and will never open your next email.
The nuance worth keeping: publish the full study as an HTML page with charts inline, then gate a genuine supplement. The raw anonymised dataset as a CSV. A slide version for someone presenting internally. Those are real asks from real buyers, and the people who fill in that form are qualified in a way that a PDF downloader is not. The wider argument is laid out in gated versus ungated content, but for research specifically there is no case for a wall.
One more thing that gets forgotten. Put the figures in HTML text, not baked into images. A chart image with the number only in the PNG is invisible to the systems you are trying to reach. Every chart needs a text caption stating the finding in a sentence.
The release kit
A study is roughly 40 percent of the work. The release kit is the rest, and it is where most programs quietly fail.
Shipping a study
- Report page, ungated
Full findings in HTML with charts inline, a sentence caption under each, and the methodology block at the bottom. Check it renders and the numbers appear as selectable text.
- Chart set with embed code
Each chart as a standalone image plus a copy-paste embed snippet that links back. Writers steal charts. Make it easy and they attribute.
- Methodology appendix
The full detail, including question wording. Link it from the report page. This is what an analyst checks before citing you.
- A quotable summary
Five to seven single-sentence findings, each with the number and the base. This is the block that ends up in other people's articles verbatim.
- Journalist and newsletter list
Twenty to forty named people who have written about this topic in the last year. Not a media database dump. Pitch the finding, not the report.
- Partner distribution
Two or three companies with an adjacent audience who will share it. Offer them a cut of the data or a co-branded segment.
- Sales and CS handoff
A one-pager so reps can send the relevant chart into live deals. This is where the study pays for itself fastest.
The outreach step is the one people skip, then conclude research does not work. It works, but a study does not distribute itself. The tactics carry over almost entirely from digital PR for SaaS links, with one difference: you are pitching a number rather than a story, so the subject line should contain the number.
Pitch the surprise, not the study
“New data: 61 percent of B2B content teams publish less than four assets a month” gets opened. “Announcing our 2026 benchmark report” does not. Lead with the finding that contradicts what the recipient already believes.
What a study actually costs
Here is the real arithmetic for a mid-sized team, based on what these projects run in practice rather than an agency quote.
| Line item | Product data study | Panel survey |
|---|---|---|
| Data acquisition | 0 | $8,000 to $25,000 |
| Analyst time | 40 to 80 hrs | 25 to 40 hrs |
| Writing and editing | 20 to 30 hrs | 20 to 30 hrs |
| Design and charts | 15 to 25 hrs | 15 to 25 hrs |
| Outreach and distribution | 30 to 50 hrs | 30 to 50 hrs |
| Realistic total | $12,000 to $20,000 in loaded time | $22,000 to $45,000 |
Source: aggregated practitioner reports, saas-marketing.net estimate.
That is one to three months of a typical mid-market content budget spent on a single asset. It is defensible only if the study is designed to be refreshed, because the second edition costs roughly 40 percent of the first and earns more. If you cannot commit to running it again next year, spend the money on a topic cluster instead.
Editable CSV worksheet
SaaS benchmark evaluation worksheet
Record the source, date, cohort and metric definition before comparing your numbers with a benchmark.
The honest failure modes
Research fails in four predictable ways, and three of them are avoidable.
The finding is boring. You ran the survey, and the answer was what everyone already assumed. This happens maybe one time in four and there is no fix at the analysis stage. Mitigate by designing questions where either answer would be interesting, and by pre-registering with yourself what result would surprise you.
The sample is indefensible. You surveyed your own newsletter list and called it market data. Somebody will notice, usually in public. If the list is your only option, publish it as “a survey of 412 subscribers to our newsletter” and let the reader weight it.
Nobody distributes it. The study goes up, gets shared internally, earns eleven links from your own social accounts, and dies. This is the most common outcome and it is purely an execution failure.
The numbers age badly. You publish a 2026 figure, never refresh, and in 2028 people are citing it as current. Your credibility takes the hit, not theirs. A fixed annual slot solves this.
Never invent the sample
The single fastest way to destroy a research program is to publish a number with a sample size you did not collect. If you have practitioner experience rather than data, write it as a range and label it as an estimate. That is respectable. A fabricated n is not, and in a category where people actually check, it surfaces.
Tracking who quotes you
Backlinks and citations have split apart. A language model can restate your figure in an answer with no link and no referral session, and that is now a meaningful share of the value. Track both.
For links, Ahrefs or Semrush with an alert on the exact report URL plus the headline number as a text string, since a lot of writers cite the figure without linking and you want to find them and ask.
For answer engines, run your twelve core questions monthly across ChatGPT, Perplexity, Gemini and Google AI Mode, and log which domains get named. Do it manually in a spreadsheet for the first quarter before buying a tool, because you will learn what the engines actually do with your data and that shapes the next study. Look at whether they quote your number, your segment cuts, or your definitions. Definitions get quoted more than most people expect.
The metric that matters over a year is not link count. It is whether your figure has become the one people reach for when they need a number in this category. That is slow, and it is the only durable outcome in SaaS content marketing that competitors cannot copy in a sprint.
Where research fits alongside everything else
One study a year, at most two, for a team under $20M ARR. More than that and the distribution budget per study collapses and you get two mediocre releases instead of one good one.
It should be the top of a cluster, not an orphan. The study produces the number, and then eight or ten supporting pages reference it, which is how the 12 SaaS content marketing examples that actually worked were structured. Companies running a proper B2B SaaS content marketing program treat the annual study as the anchor and build the year’s editorial calendar around interrogating its findings. If you are segmenting by deal size, the approach in content marketing for B2B SaaS by ACV tells you which cuts your audience will actually use.
Do this next
Pick the source you can defend in one sentence to a skeptic. Write the methodology block before you collect anything, because it will expose the questions you have not thought through. Book the distribution time in the same calendar block as the research time, at roughly equal weight. And put next year’s edition on the calendar the day you publish this one, because the second study is the one that makes the first one worth it.
Editable CSV worksheet
SaaS Content Marketing planning worksheet
A practical content planning worksheet: decisions, owners, evidence and next actions.
Frequently asked questions
How much does an original research study cost for a SaaS company?
Product data studies cost mostly analyst time, roughly 40 to 80 hours plus design. Panel surveys run 8,000 to 25,000 dollars for 300 to 600 qualified B2B respondents through providers like Prolific, Dynata or Pollfish, plus incentive costs. Aggregating public filings is the cheapest route at 30 to 60 hours of research time and no external spend.
Should an original research report be gated?
No, if citations are the goal. A gated PDF cannot be crawled, cannot be quoted by an answer engine, and cannot be linked to by a journalist who will not fill in a form. Publish the full study as an HTML page with the charts inline, then offer a supplementary asset such as the raw dataset or a slide version behind a form if you want leads.
What sample size do I need for a credible B2B survey?
Around 300 completed responses gives a margin of error near 5.7 percent at 95 percent confidence for a large population, which is enough to publish headline figures. Below 150 you should present findings as directional and say so. If you plan to cut the data by segment, each segment needs its own usable base, so 300 total split six ways is not publishable at the segment level.
How do I know if my study is getting cited by ChatGPT or Perplexity?
Ask the engines directly with the queries your study answers, repeated monthly, and log which domains they name. Tools such as Profound, Peec AI and Semrush AI Toolkit automate this. Backlink counts in Ahrefs still matter but they now undercount impact, because an answer engine can quote your figure without any link existing.
How long does it take for a research study to earn links?
The first wave arrives within two to six weeks and comes almost entirely from outreach, not from search. Organic citation builds over six to eighteen months as other writers find the page while researching. A study that earns nothing in the first month usually has a distribution problem rather than a data problem.
Can I run original research without a large user base?
Yes. Aggregating public data is the route for small companies. SaaS Capital built authority on survey data before it had scale, and analysts routinely build datasets from 10-K filings, job postings, pricing pages and app store data. The constraint is rigour and repeatability, not company size.
How often should a benchmark report be updated?
Annually, on a fixed month, so the market learns when to expect it. Quarterly is only worth it for fast moving metrics such as AI citation share. A study that goes two years without a refresh starts getting cited with the wrong year attached, which damages the credibility you built.
The saas-marketing.net editorial team Research and editorial
We research, write and maintain every page on this site. The library explains marketing decisions through practical frameworks, explicit assumptions and references. Corrections can be requested through the contact page.
Published September 11, 2026. Last updated .