How to Design an AI Visibility Prompt Set Worth Tracking
Your prompt set is the instrument, not a setting
Most teams that start tracking AI visibility spend an hour picking a tool and about four minutes picking the prompts. That ratio is backwards. The prompt set is the instrument. Everything downstream — your visibility score, your AI share of voice, the chart you'll show your boss in six weeks — is just a summary of how the engines answered the questions you chose to ask. Choose shallow questions and you'll get a precise measurement of something that doesn't matter.
The failure mode I see constantly: someone exports their top 30 SEO keywords, pastes them in as prompts, and starts tracking. But nobody types "project management software" into ChatGPT the way they typed it into Google. They ask "what should a 6-person agency use to manage client projects without drowning in setup?" Keywords are compressed queries; prompts are decompressed ones. If your tracking set is built from keywords, you're monitoring a conversation nobody is having.
So before touching any tool, sit down and design the set the way you'd design a survey. Decide what you need to learn, then write the questions that would reveal it. If you're new to the measurement side of this, the AI visibility guide covers what the scores mean; this article is about what to feed them.
The intent mix — four types of AI visibility prompts
A prompt set that only asks one kind of question gives you one kind of blind spot. In practice, four intent types cover the ground, and a healthy set draws from all of them.
Best-of prompts are the obvious ones: "best CRM for a small nonprofit," "best running shoes for flat feet." These are the money questions, the ones where being named means being shortlisted. They're also the most volatile, because the engines lean heavily on retrieved listicles for them, and listicles churn. Track them, but don't let them be the whole set.
Comparison prompts pit you against named competitors: "X vs Y for freelancers," "is X worth it compared to Y?" These reveal something best-of prompts hide — how the engines describe you when you do appear. You can be present in every comparison and still lose every comparison, and only this prompt type will tell you.
Problem-first prompts never mention a product category at all: "how do I stop invoices getting lost in email threads?" This is where AI assistants differ most from classic search. A person with a problem gets walked to a category, then to a shortlist, inside one answer. If your brand only surfaces once the user already knows the category word, you're invisible at the moment the category is being chosen for them. These prompts are the hardest to write and the most valuable to track.
Brand checks ask about you directly: "is X legit?" "what do people say about X's support?" You won't win new customers in these answers, but you can absolutely lose them there. A brand check that surfaces a two-year-old complaint thread as the definitive account of your product is a fire worth knowing about.
A rough starting split that has served me well: about a third best-of, a quarter comparison, a quarter problem-first, and the remainder brand checks. Skew it toward problem-first if you're in a category buyers don't have a name for yet.
How many prompts is enough for prompt tracking?
Fewer than you fear, more than you'd like. The tension is real: every prompt you track costs money on every scan across every engine, and answers vary run to run, so each prompt needs repeated sampling before its signal means anything. A 200-prompt set scanned daily is mostly a way to set budget on fire.
Here's the way I think about it. You need enough prompts that no single answer's randomness moves your overall score, and few enough that you can actually read the per-prompt detail when the score moves. For most brands that lands somewhere between 15 and 40 prompts. A single-product company serving one audience can do honest work with 15. A brand with three product lines and two distinct buyer types needs closer to 40, because what you actually have is several small prompt sets sharing a dashboard.
The test for whether a prompt earns a slot: would a change in this answer change what you do? If the answer moved from "you're absent" to "you're cited," would anyone on your team act differently? If not, it's decoration. Cut it and spend the slot on a question where the answer would sting.
One more sizing note: depth beats breadth. Ten prompts sampled across several engines on a weekly rhythm will teach you more than fifty prompts sampled once a month, because trend is the unit of meaning here, and trends need repeated measurements to exist.
Vanity prompts — the queries that flatter and teach nothing
Every prompt set I've audited contains at least a few prompts that exist to make the dashboard green. They're worth naming so you can spot them in your own list.
The branded softball: "what is [your brand]?" Congratulations, the engines can read your homepage. This tells you your site is crawlable, which is worth confirming exactly once, not tracking weekly.
The category you've already won: prompts where you're the obvious incumbent and have been for years. Watching yourself stay first is comfortable and useless. The slot belongs to a category you're trying to enter, not one you own.
The impossibly broad head term: "best software." No specific buyer asks this, no answer to it is actionable, and your movement in it is noise. Broad prompts feel important because their search-volume cousins were important in SEO. In AI tracking they're mostly weather.
The wishful phrasing: prompts worded exactly the way your marketing describes the product, rather than the way a frustrated human describes the problem. If your positioning says "revenue intelligence platform" and your buyers say "why do my sales forecasts keep being wrong," track the second one. The gap between those phrasings is precisely what you need to see.
A quick heuristic: if a prompt could appear on a slide titled "look how visible we are" but would never appear in a real user's chat history, it's vanity. Real prompts have a person behind them, usually a mildly annoyed one.
A worked example — building a prompt set from scratch
Take a hypothetical: Ferrylight, a small SaaS that does automated invoice chasing for freelancers. One product, one audience, so a compact set is right — say 18 prompts.
Best-of (6): "best invoice chasing software for freelancers," "best way to automate late payment reminders," "top tools for getting invoices paid faster," and three variants scoped by audience — freelancers, small agencies, contractors — because the engines answer each differently. Comparison (5): Ferrylight against each of the three competitors that keep showing up in the same listicles, plus "Ferrylight vs doing it manually in email," which is the comparison most buyers are actually making. Problem-first (4): "clients keep paying my invoices late, what can I do," "how do I chase overdue invoices without sounding rude," "freelancer cash flow tips for late payers," "should I charge late fees or automate reminders?" Brand checks (3): "is Ferrylight legit," "Ferrylight reviews," "does Ferrylight work with QuickBooks?"
Notice what's absent: nothing about "accounting software" (a category Ferrylight isn't in and would drown in), no "what is Ferrylight," no head terms. Every prompt maps to a stage of one specific buyer's day. When this set moves, Ferrylight knows exactly which conversation shifted and can go read the answers to find out why.
Refresh cadence — when to change your prompt set
A prompt set is not a monument. It's also not a whiteboard. The trap on one side is never touching it, so it slowly drifts away from your market; the trap on the other side is fiddling weekly, which destroys the trend lines that were the entire point of tracking.
The rhythm that works: review quarterly, replace sparingly. Once a quarter, read your per-prompt results and ask three questions. Which prompts have gone stale — a product line you sunset, a competitor that pivoted away? Which prompts have flatlined at zero with no strategy attached — you're absent, and nothing you're doing this quarter will change that? And what new questions have appeared — a competitor launch, a feature people suddenly ask about, a new audience you're courting? Swap out the dead weight, swap in the new questions, and cap changes at around a fifth of the set so the majority of your trend history stays intact.
Two events justify off-cycle changes: a repositioning (your problem-first prompts describe problems you no longer solve) and a launch (a new product with zero tracked prompts is flying blind). Otherwise, let the set sit still and do its job. Boring is what a measurement instrument is supposed to be.
One last habit worth stealing: before you finalize any set, look at how the engines themselves expand your questions. Modern assistants fan a single prompt out into multiple sub-queries behind the scenes, and seeing that expansion tells you which phrasings actually get searched on your behalf. Run your draft prompts through the query fan-out tool and you'll usually find two or three variants you'd never have written yourself — and at least one you can delete.
This is one of eight metrics in the complete AI visibility measurement playbook.
If the concept is new, start with what prompt tracking is.