# Getting cited by AI answer engines

> A repeatable way to enter AI answer sets: prompt research, source pages engines prefer, third party mention seeding, schema work and citation monitoring.

Source: https://saas-marketing.net/playbooks/get-cited-by-ai-engines/
Topic: SaaS SEO
Type: playbook
Published: 2026-09-11
Last updated: 2026-09-11
Publisher: SaaS Marketing (saas-marketing.net)
License: CC BY 4.0. Quote or republish with attribution and a link to https://saas-marketing.net/playbooks/get-cited-by-ai-engines/

## Short answer

Answer engines cite sources they can crawl, parse and trust on the specific question asked. To get cited, build a set of 40 to 60 buyer prompts, record which URLs each engine currently cites for them, then classify those sources. Most will be third party listicles, review sites and community threads rather than vendor pages. Win the third party placements first, fix crawler access, and add dated, methodology backed answer blocks to your own pages.

## Key takeaways

- Around 62 percent of AI cited pages are blog posts and listicles, which is why third party inclusion beats homepage rewrites.
- Build a fixed prompt set of 40 to 60 buyer questions and re-run it weekly, or your visibility data is anecdote.
- Check your robots.txt for GPTBot, ClaudeBot, PerplexityBot and Google-Extended before doing any content work.
- Review profiles on G2 and Capterra get cited constantly, so profile completeness is a citation tactic, not a sales chore.
- Citation share moves in weeks for third party placements and in months for your own pages, so sequence accordingly.
- A citation you cannot connect to a signup is still worth tracking, but report it as reach, not as pipeline.

---

Most teams attack this backwards. They rewrite the homepage, add FAQ schema, publish a manifesto about AI search, and wait. Meanwhile the engine answering "best call recording software for sales teams" is pulling from four listicles, a G2 category page and a Reddit thread, none of which anyone at the company has read.

The work below is a 30 day sprint that starts with what the engines cite today and ends with a weekly number you can put in a board deck. It assumes you already have decent pages. If you do not, fix that first with the [answer engine optimization guide](/guides/answer-engine-optimization-saas/) and come back.

## Build the prompt set before you change anything

Your prompt set is the measurement instrument, and a bad one makes every later decision guesswork. Aim for 40 to 60 prompts that mirror how a buyer actually types into an assistant, not how they type into Google.

The difference is real. Nobody asks ChatGPT "best CRM." They ask "what CRM should a 12 person agency use if we already run everything in Slack and need it to not feel like Salesforce." Longer, more conditional, more constraint driven.

Build the set in five buckets:

| Bucket | Example prompt shape | How many |
| --- | --- | --- |
| Category selection | "What is the best [category] tool for [segment] in 2026" | 12 to 15 |
| Head to head | "Is [you] or [competitor] better for [use case]" | 8 to 10 |
| Alternatives | "What are good alternatives to [competitor] if [constraint]" | 8 to 10 |
| Problem framed | "How do I [job to be done] without [pain]" | 8 to 12 |
| Definitional | "What is [category concept] and how does it work" | 5 to 8 |

Keep them in a sheet with a fixed ID per prompt. You will run these same strings every week for a year, and changing the wording mid programme destroys the trend line. Freeze the wording the way you would freeze a survey question.

Ask the same engine the same prompt three times and you will get three slightly different source sets. That is the nature of the systems, not a bug in your process. Run each prompt at least twice per cycle and record a source as present if it appears in either run. Report a four week rolling average and never quote a single day's result to an executive.

## Record who gets cited today, then classify the sources

Run every prompt against ChatGPT, Perplexity, Google AI Mode and Gemini, and log the cited URL, the domain, the page type and if your brand is mentioned at all. This ledger is the whole point of week one. The methods for doing it without losing your weekend are in the [guide to measuring AI search visibility](/guides/measure-ai-search-visibility/).

Then classify each cited source into one of five types:

- Your own pages
- Competitor pages
- Review and directory sites, mostly G2, Capterra, TrustRadius, Software Advice
- Third party editorial, meaning listicles, category guides and roundups on media or agency sites
- Community, meaning Reddit, Stack Overflow, Hacker News, niche Slack and Discord archives

The distribution is what tells you where to spend. Published analyses through 2025 put roughly 62 percent of AI cited pages in the blog post and listicle category, which matches what most SaaS prompt sets show when you actually count. Our own [SaaS AI citation benchmark](/research/ai-citation-benchmark-saas/) breaks the split down by category, and the [social sources data](/research/social-sources-cited-in-ai-answers/) covers how much of the remainder comes from community threads.

**62.1%** Share of AI cited pages that are blog posts or listicles rather than vendor product pages

Here is the position this page takes, and it is unpopular with content teams: your fastest wins are in that third party editorial bucket, not on your own site. Getting added to four listicles that already rank and already get cited will move your mention rate in three to five weeks. Rewriting your homepage will move it in approximately never.

## Attack the gap, starting with third party listicles

Take the 15 prompts where a competitor is named and you are not. For each, open every cited source and check if you could legitimately be on that page. Most of the time you could, and nobody has ever asked.

The outreach is unglamorous and it works. A short email to the author with the three things they would need to add you: a one line product description, the pricing tier the page format uses, and one differentiator that is actually true and checkable. Offer a screenshot. Do not offer money for placement on editorial pages, both because it is against most publishers' policies and because paid inclusion tends to sit in sections the engines weight lower.

Expect roughly a 20 to 30 percent response rate on a personalised outreach list of 40, and maybe half of those to result in inclusion within six weeks. That is five to six new cited sources, which for a mid market SaaS in a defined category is usually enough to appear in answer sets you were absent from.

A 9 million ARR support automation tool ran 52 prompts in January. It appeared in 6. Of the 46 misses, 31 cited the same eight listicles, and the tool was on two of them. The team spent five weeks on outreach to the remaining six, got into four, and the appearance count went from 6 to 19 by mid March with no changes to their own website at all.

## Fix the review profiles, the entity and the community presence

Review directories show up in AI answers at a rate that surprises people who think of G2 as a sales expense. Treat profile completeness as a citation tactic.

The specific work, in order of return:

1. **G2 and Capterra completeness.** Fill every field, including the ones nobody fills: integrations, supported languages, deployment options, compliance certifications, pricing tiers with real numbers. Engines lift these attribute tables directly because they are structured and current.
2. **Review volume and recency.** Twenty reviews from the last six months beats 200 from 2022. Ask during renewal conversations, not in a quarterly blast.
3. **Wikidata entity.** Create or correct your Wikidata item with founding date, headquarters, founders, industry and official website. It takes an afternoon. It is the cheapest entity hygiene available and it feeds knowledge panels and model grounding alike.
4. **Wikipedia, carefully.** Only if you meet notability, which most Series A companies do not. Do not create a page about your own company. Do check that existing mentions of you elsewhere on Wikipedia are accurate.
5. **Reddit and community.** Not astroturfing. Genuine participation from named employees in the two or three subreddits where your buyers argue. Threads where a founder answers a technical objection in detail get cited for years.

If your site says "Acme Analytics", G2 says "Acme.io", Crunchbase says "Acme Technologies Inc" and your LinkedIn says "Acme", you have handed every model four entities to reconcile. Pick one canonical name and one canonical one line description, then propagate it everywhere including your schema markup. This is boring and it materially improves how often you get named correctly.

## The on page work that actually earns a quote

Own pages take longer to break in, but they are the only citations you fully control. The pattern that gets lifted is consistent: a narrow question, answered completely, in the first 60 to 90 words under a heading that matches the question.

**Page level citation checklist**

The methodology block deserves emphasis. When two pages make the same claim and one shows its sample size and collection window, the one showing its work gets cited noticeably more often. This is the single biggest difference between a page that gets quoted and a page that gets ignored, and it is also the cheapest to add.

Original numbers beat rewritten ones. A page reporting "we measured 340 SaaS sites in March 2026 and 61 percent blocked at least one AI crawler" becomes the source. A page repeating someone else's figure becomes a citation of a citation, which the engines mostly discard.

## Verify the crawlers can reach you, before anything else

Do this in the first hour of the sprint. Powered by Search's audit of 50 SaaS sites found 68 percent blocked at least one major AI crawler, and in almost every case nobody on the marketing team knew.

Check for these user agents in robots.txt and in your CDN bot management rules:

| Crawler | Operator | What it feeds |
| --- | --- | --- |
| GPTBot | OpenAI | Model training and browsing corpus |
| OAI-SearchBot | OpenAI | ChatGPT search results |
| ClaudeBot | Anthropic | Claude browsing and training |
| PerplexityBot | Perplexity | Perplexity answer sources |
| Google-Extended | Google | Gemini and AI Overview grounding |

Two traps. First, Cloudflare's bot fight mode and its AI crawler toggle can block these at the edge even when robots.txt allows them, so check both. Second, a WAF rule that challenges non browser user agents will silently return a JavaScript challenge that no crawler can solve. Curl each page with the user agent string set and confirm you get a 200 with real HTML.

Also confirm your pages render without JavaScript. Client side rendered marketing pages are a recurring reason a technically allowed site never gets cited, and the wider version of that problem is covered in the [B2B SaaS AI search visibility guide](/guides/b2b-saas-ai-search-visibility/).

## Where the four engines differ, and why averaging them hides the answer

They pull from different places, refresh on different clocks, and reward different things. A blended "AI visibility score" across all four is the metric most likely to make you do the wrong work, because a gain on Perplexity and a loss on ChatGPT cancel out into a flat line that describes nothing.

The practical consequence is that AI Overview presence is mostly downstream of ranking, so if you want that surface, the fix is ordinary SEO on the exact query. Perplexity rewards speed and fresh numbers. ChatGPT rewards being in the sources everyone else already cites, which is the slowest to earn and the hardest for a competitor to take back once you have it.

Report the four as four columns. Executives will ask for one number. Give them four and explain in one sentence why, because the first time a model update reshuffles one engine you will want the others as a control group.

## Weekly monitoring and what to report

Run the fixed prompt set weekly, log to the same sheet, and report three numbers: appearance rate, citation rate and sentiment of the mention. Appearance rate is the share of prompts where your brand is named at all. Citation rate is the share where one of your own URLs is the linked source. They move independently, and confusing them is how teams end up celebrating the wrong thing.

Tooling saves time once you pass about 40 prompts across four engines, which is roughly 320 manual checks a week. The [AI search visibility tracker comparison](/tools/ai-search-visibility-trackers/) covers what each tool actually samples, which matters, because several of them query a model API rather than the consumer product and get different answers as a result.

What to tell the board: appearance rate trend over four weeks, the three prompts where you gained, the three where a competitor overtook you, and the self reported attribution count from your signup form. Do not present a single week's number. Do not present citation counts as pipeline.

## The honest tradeoff

This work has a real cost and a soft return. A proper sprint runs 30 to 40 hours of setup plus about three hours a week of measurement, and the connection to revenue stays fuzzy for months. Referral traffic from assistants is a fraction of what the engagement suggests, because most people read the answer and search your brand later.

It is also less durable than SEO. A model update can reshuffle which sources it favours, and three weeks of outreach wins can flatten without explanation. Budget for it as an ongoing programme at roughly a day a week, not a project that finishes. If you need a single number to defend the spend, use the share of new signups selecting an AI assistant in your self reported attribution field, and the deeper measurement approach in our [AI citation study](/research/ai-citation-benchmark-saas/) shows what a realistic range looks like by category.

## The first week

Monday: check crawler access and CDN rules. Tuesday and Wednesday: write the 50 prompts and run the baseline against all four engines. Thursday: classify every cited source and build the gap list. Friday: send the first 15 outreach emails.

Then spend weeks two through four on the outreach list and the review profiles, and only in week five start rewriting your own pages with answer capsules and methodology blocks. That sequence is deliberate, because the third party work compounds while you wait for your own pages to be recrawled. The link and mention side of it sits alongside the rest of your [SaaS SEO](/saas-seo/) programme, and the practical version of the outreach is broken down further in the [links and citations lesson](/courses/saas-seo-sprint/04-links-and-citations/).

## Frequently asked questions

### How do you get ChatGPT to mention your SaaS product?

Get into the sources ChatGPT already reads for that question. For most software queries those are third party listicles, G2 and Capterra profiles, Reddit threads and a handful of category guides. Earning inclusion in three or four of those ranking listicles moves your mention rate faster than any change to your own website.

### Does Perplexity cite different sources than ChatGPT?

Yes, and the overlap is smaller than most teams expect. Perplexity leans heavily on freshly crawled web results and shows its sources openly, so recency and crawlability matter more. ChatGPT blends browsing results with model knowledge, which means older, widely referenced pages appear more often. Run your prompt set against both separately and never average the results.

### Do I need schema markup to be cited by AI answer engines?

Schema helps machines parse your page, but it is not what earns a citation. Organization, Product, FAQPage and Article markup with consistent entity names across your site, Wikidata and review profiles make you easier to identify. The citation itself is earned by having a clear, dated, specific answer that matches the question being asked.

### How often should I measure AI search visibility?

Weekly for a fixed prompt set, monthly for reporting. Answers vary between runs even with identical prompts, so a single check tells you almost nothing. Run the same 40 to 60 prompts on the same day each week, record which domains are cited, and report the four week rolling share rather than any individual result.

### Is blocking AI crawlers bad for SaaS marketing?

For a marketing site, yes. Powered by Search found that 68 percent of 50 SaaS sites it audited blocked at least one major AI crawler, usually by accident through an inherited robots.txt or a CDN bot rule. If your pages cannot be fetched, they cannot be cited, and the block is often invisible to everyone outside engineering.

### What content format gets cited most by AI engines?

Listicles and blog posts dominate, at roughly 62 percent of cited pages in published 2025 analyses. Comparison pages, glossary definitions and pages with original numbers also punch above their weight because they answer a narrow question completely. Product and pricing pages are cited far less often than marketing teams assume.

### Can you attribute pipeline to an AI citation?

Partially, and honestly you should say so. Referral traffic from chatgpt.com and perplexity.ai appears in analytics, but most of the influence arrives as direct or branded search later. Add a self reported attribution field to your signup form with an AI assistant option. That single field usually tells you more than any referral report.
