The Prompt Set Is the Strategy: How To Measure AI Visibility Across Real Buying Moments

Your AI visibility program is only as good as its prompt set. Here is the architecture, the seven prompt types, and a 20-prompt baseline you can build today.

Comfort metrics come from comfortable prompts

Most teams that start tracking AI visibility do it with prompts that feel natural to type and prove nothing. They ask "what is [our brand]," get a decent summary, and file AI search under control. Then a buyer asks "best option for a two-person marketing team under $100 a month," a competitor gets the confident recommendation, and none of it shows up in anything the team measures.

The failure is not the tool or the cadence. It is the prompt set. Ad hoc prompts fail in predictable ways:

  • Too branded. Prompts containing your name test recall, not discovery. The buyer who matters has not typed your name yet.
  • Biased toward what you already know. Teams write prompts about the positioning they wish they had, not the questions buyers actually ask.
  • Too short and too generic. "Best CRM" is not a buying moment. Real buyers attach constraints: team size, budget, industry, situation.
  • Not mapped to personas. A founder, an operator, and a growth lead ask differently. One prompt cannot represent three buyers.
  • No competitor context. If you never ask the comparison questions, you never see who is winning them.
  • Never repeated. A prompt run once is a screenshot. Answers move between runs even when nothing else changed, and shift again after model updates, so single results are noise.

The fix is to treat the prompt set as what it actually is: the measurement instrument for your entire AI visibility program. A weak prompt set does not just produce incomplete data. It produces false confidence, which is worse.

A prompt set is an architecture, not a brainstorm

Do not start by writing prompts. Start by writing down the dimensions your prompts need to cover, then generate prompts from the intersections. Six dimensions cover almost every category:

  1. Persona. Who is asking: the founder, the operator, the growth lead, the brand lead? Each has different words and different constraints.
  2. Buying moment. Where the buyer is and what is forcing the decision: just discovering the category, comparing options after a launch or a renewal, or validating a choice they have nearly made.
  3. Use case. The job to be done, in the buyer's own words, not your feature language.
  4. Objection or risk. What could stop the purchase: price, trust, switching cost, reliability, compliance.
  5. Competitor context. The alternatives the buyer would plausibly weigh, including "do nothing" and "do it manually."
  6. Geography and language. Where relevant, the same question asked in another market or language, because answers differ more than most teams expect.

You will not need every intersection. But walking the grid is what surfaces the prompts your team would never think to write, and those are usually the ones where you are silently losing.

The seven prompt types that cover a buying journey

Dimensions tell you what to vary. Prompt types tell you what to write. In practice, seven types cover the journey from first question to final validation:

Prompt typeWhat it revealsExample shape
Category discoveryWhether you exist when the buyer does not know names yet"How do I [solve job to be done] for [situation]?"
Best-for use caseWhether you win the moments closest to a decision"Best [category] for [persona] with [constraint]"
ComparisonWho gets the confident reason when options are weighed"[Competitor] vs alternatives for [use case]"
AlternativesWhether you appear when a buyer wants to leave a rival"[Competitor] alternatives for [constraint]"
Trust and proofWhat evidence the answer reaches for about you"Is [category or brand] reliable for [need]?"
Drawback and objectionWhat negatives travel with your name"What are the downsides of [brand or category]?"
Branded perceptionA control: how you are described once the name is known"What is [brand] and who is it best for?"

Weight the set toward the middle of the table. Best-for, comparison, and alternatives prompts sit closest to revenue, and they are where competitor wins concentrate. Branded prompts belong in the set, but as a control group, not the main event.

The same brand, asked two ways

The difference between a weak prompt and a strong one is the difference between testing recall and testing the market.

  • Weak: "Perceptiq." Strong: "What are the best AI visibility tools for a boutique SEO agency managing five clients?"
  • Weak: "Nordkamm backpack." Strong: "What are the most reliable backpacks for a rainy weekend hike?"
  • Weak: "best project management software." Strong: "What should a five-person design studio use to manage client projects without a dedicated PM?"

The strong versions share three properties: they do not contain a brand name, they carry a real constraint, and they are phrased the way a person actually talks to an assistant. If a prompt could not plausibly have been typed by a buyer at a decision point, it does not earn a slot in the baseline.

Split discovery from evaluation

One structural decision separates mature programs from casual ones: run discovery and evaluation as two distinct prompt sets, and read them separately.

Discovery prompts are category-shaped and unbranded. They tell you whether you enter the buyer's shortlist at all. Evaluation prompts are brand-shaped: "is [your brand] good for [use case]," "what are the weaknesses of [your brand]," "[your brand] vs [competitor]." They tell you what happens after a buyer has your name, when they turn to an assistant to validate or kill the choice.

Teams tend to obsess over discovery and ignore evaluation, which is exactly backwards for brands that already have demand. A brand can win discovery and still lose the sale in evaluation because the assistant repeats a stale price, an old limitation, or a drawback framed a year ago. Those late-stage answers are cheap to fix and expensive to ignore, and you only see them if the evaluation set exists.

How many prompts, and how often

More prompts is not more insight. The right size is the smallest set that covers your buying moments and still fits the attention you will genuinely give it.

  • Lean baseline: about 20 prompts. One brand, two or three personas, weighted toward best-for and comparison. This is enough to find your biggest gap, and it is where every team should start.
  • Growth baseline: 40 to 60 prompts. Adds evaluation prompts, objection coverage, and a second market or language if you sell in one. This is the working size for a team that reviews results weekly and ships fixes.
  • Agency baseline: audit small, manage deep. For a pitch or an initial audit, a standardized 20-to-30-prompt set is right: fast to run and comparable across prospects. Once a client is under management, the set should be at least as deep as the growth baseline, with the standardized types kept across clients and client-specific buying moments layered on. An agency running shallower measurement than its client would run in-house is not adding rigor.

On cadence: weekly is the rhythm to aim for, and it is the default Perceptiq runs on. The reason starts with something more basic than model updates: a single run is a noisy sample, not a measurement. Ask the same assistant the same question twice and you will often get a slightly different answer built on slightly different sources, because these systems sample as they generate and retrieve. We see this constantly in our own prompt runs, and independent research confirms it: only a small percentage of brands stay visible between two consecutive answers to the same prompt. On top of that baseline noise, answers genuinely drift: models update, competitors publish, new sources get indexed, and analyses of major model updates have found large shares of cited sources changing in a single release.

Short-cadence reruns are how you turn that noise into signal. Each weekly run adds another sample, and over a month the samples average into a stable, confident picture of where you actually stand, instead of four contradictory screenshots. Monthly is the floor for a fully manual setup, and either way, rerun the affected prompts a few weeks after you ship a meaningful fix. A baseline you rerun is the only way to tell drift from progress.

The 20-prompt starter baseline

If you are starting from zero, here is the allocation that produces the most insight per prompt. Write these for your own category, personas, and competitors:

  1. Four category discovery prompts, one per major job to be done, unbranded.
  2. Five best-for prompts covering your two or three core personas with their real constraints.
  3. Four comparison and alternatives prompts naming the competitors you actually lose deals to.
  4. Three trust, objection, or drawback prompts, at least one aimed at your own brand.
  5. Two evaluation prompts validating your brand for your best-fit use case.
  6. Two branded control prompts to capture how you are described once the name is known.

Then, for every answer, capture more than presence: who was named, in what order, with what stated reason, citing which sources, and whether the description matches your positioning. That is the raw material the rest of the program is built from. The audit method covers scoring and prioritizing in full.

Frequently asked questions

Should the prompt set include my brand name at all?

Yes, in two places: branded controls and evaluation prompts. The mistake is not including branded prompts. It is letting them dominate the set, because they measure recall while your losses happen in the unbranded majority of the journey.

Do I need a different prompt set for each AI assistant?

Keep one core set and run it across the assistants your buyers use, then read results per platform. The prompts should reflect buyers, not platforms. Different platforms will return different answers from the same set, and that difference is itself a finding about where each platform gets its evidence.

How do I know if my prompt set is working?

A working set regularly tells you something that changes what you ship next: a buying moment you are absent from, a competitor win with a stated reason, a stale fact in an evaluation answer. If every review meeting ends with "looks fine," your set is measuring comfort, not the market.

The set is the strategy

Two teams can buy the same tooling, track the same platforms, and report the same metrics, and one of them will still be flying blind, because everything downstream inherits the quality of the prompts. The set decides which gaps you can see, which competitor wins you catch, and which fixes look urgent.

Perceptiq starts here for exactly that reason: it generates buying-moment prompts from your personas and positioning, runs them on a schedule across platforms, and scores visibility and perception per prompt, so the measurement instrument is built into the workflow instead of improvised in a spreadsheet. Once the answers are coming in, the next step is diagnosis, and how AI decides which brands to recommend walks through exactly that: reading the answers your prompts return and turning them into fixes. Start with the 20-prompt baseline this week: write it, run it, and see which buying moment you are losing that nobody on the team has ever measured. Or let Perceptiq build your baseline and skip straight to the findings.