AI Visibility Prompt Research: Build a Benchmark That Reflects Real Buyers
AI Brand Report ·
Turn customer questions into a repeatable AI visibility benchmark, with a prompt matrix, sampling rules, and a worked example that separates discovery from brand recognition.
AI visibility prompt research is the process of turning real buyer questions into a controlled set of questions you can test repeatedly. A useful benchmark covers the decisions customers make, records the conditions of each test, and keeps brand recognition separate from unprompted discovery.
The most consequential decision in an AI visibility program often happens before the first report: deciding what to ask. If your prompt set is biased, a beautifully presented score can tell the wrong story.
A company that sells scheduling software could ask, “Why is our platform a great appointment booking tool?” The answer might mention the company every time. That establishes almost nothing about whether a buyer would discover it while researching scheduling software.
This guide offers a practical method for building a better benchmark. The examples are illustrative, not customer research or product performance claims.
Start with decisions, then write prompts
Collect questions from sales calls, support tickets, onsite search, customer interviews, and relevant search queries. Remove names, emails, account details, and other private information before turning those records into a research brief.
For each question, identify the underlying decision. “Does it integrate with my calendar?” could become a compatibility question. “Is there something less complicated?” may indicate implementation friction. A pricing objection could be about total cost rather than the monthly subscription alone.
Keep a short evidence note beside each prompt: where the question came from, which audience asks it, and what decision the answer supports. This prevents a benchmark from becoming a collection of questions that only the marketing team finds interesting.
Use keyword and prompt research together, but do not treat them as interchangeable demand data. Search keyword volume does not measure the frequency of an exact prompt in ChatGPT or another assistant.
Build a prompt matrix with explicit coverage
Here is a starter matrix for an imaginary appointment software company serving small professional practices:
| Buyer decision | Example unbranded prompt | What the test investigates |
|---|---|---|
| Define the problem | How can a small practice reduce appointment no-shows? | Whether the category is introduced appropriately |
| Discover options | Which appointment tools suit a three-person practice? | Whether relevant vendors enter the shortlist |
| Evaluate a constraint | What scheduling software supports multiple staff calendars? | Whether the answer recognizes required capabilities |
| Compare approaches | Should a small practice use scheduling software or a virtual receptionist? | Whether the product category fits the situation |
| Assess implementation | What should I check before switching appointment software? | Whether useful implementation evidence is available |
| Evaluate risk | How should I evaluate a scheduling vendor's data handling? | Whether the answer points to substantive documentation |
Create four genuinely different questions for each row to produce a 24-prompt starting set. This number is a practical workload choice, not a universal minimum or a representative sample of the market.
Do not generate the four questions by swapping synonyms. Different questions should expose different decision constraints: team size, integrations, migration effort, geography, or service model. Add a constraint only if it appears in real buyer research.
Separate three kinds of brand questions
Use distinct reporting groups for:
- Unbranded discovery: the buyer describes a need without naming vendors.
- Branded evaluation: the buyer asks about your company, product, or capabilities.
- Named comparison: the buyer supplies a shortlist and asks for tradeoffs.
These groups answer different business questions. Combining them into one mention rate can manufacture progress simply by adding more branded prompts.
Suppose a synthetic benchmark has 20 unbranded answers, with your company mentioned in 5. Another 10 answers explicitly asked about your company and all mention it. The discovery mention rate is 5/20, or 25%. Reporting 15/30, or 50%, as discovery visibility would obscure what happened.
For broader definitions, see our AI share of voice guide. In operational reports, always publish the denominator alongside the percentage.
Record the testing conditions
A prompt alone does not fully describe a test. Record the engine or product, model when exposed, date, language, location assumptions, browsing or search mode, and whether the conversation was fresh.
An API response and a consumer assistant session should not silently share the same benchmark label. They may use different tools, instructions, retrieval paths, or personalization. Document what you actually tested rather than implying that a provider name identifies one uniform experience.
Save the response and any visible citations. A summary score is useful for sorting; the original evidence is necessary for diagnosing a missed mention or an inaccurate description.
Define failures before collecting results. A timeout is not an answer with no brand mention. Track unsuccessful requests separately and show coverage so a provider outage cannot look like a sudden loss of visibility.
Use a stable panel and an exploratory panel
Keep a stable core of prompts for comparisons over time. Maintain a separate exploratory set for new products, emerging buyer questions, or a newly important competitor.
If you change the core panel, create a new version and explain the change. Run the old and new panels together once when practical. This overlap helps distinguish a real movement from a change in what you measured.
For the imaginary scheduling company, a new enterprise feature might justify an exploratory set for large teams. It should not immediately replace small-practice prompts and make the historical trend incomparable.
Decide what a finding will change
Every prompt cluster should connect to an action. If implementation questions produce vague answers, improve migration documentation. If a capability is repeatedly misdescribed, identify the conflicting sources. If a vendor does not fit the buyer's need, exclusion may be a reasonable answer rather than a visibility defect.
This distinction keeps research useful. The goal is to understand where your brand belongs in the decision process and make reliable evidence available there.
Frequently asked questions
How many prompts should an AI visibility benchmark include?
Start with enough prompts to cover your important buyer decisions. The 24-prompt example in this guide is a planning aid, not a statistically representative sample of all AI users.
Should branded prompts count toward AI discovery visibility?
Report them separately. A prompt that already names your brand tests recognition or evaluation, while an unbranded prompt tests whether the system introduces your brand.
Can keyword volume tell me how often people ask an AI prompt?
No. Search keyword volume can inform topic priorities, but it does not establish the frequency of an exact question in an AI assistant.
Put the benchmark to work
Use an AI Brand Report to start investigating how AI engines describe your brand. Then use the research method above to evaluate coverage and decide what to investigate next. For measuring changes responsibly, continue with our AI visibility experiment guide.