An aggregate AI visibility score can tell you whether your brand's presence is rising or falling. It cannot tell you which buyer questions are driving that movement, where competitors appear instead of you, or whether one engine sees your brand differently from another.
That is the role of prompt-level visibility.
Category-level and prompt-level data are not competing measurement systems. They answer different questions. Category data provides orientation: where does the brand stand across a market or topic? Prompt-level data provides diagnosis: which questions, engines, sources, and portrayals are producing that result?
The strongest AI search measurement system uses both. Category trends show where to investigate. Prompt-level evidence helps determine what to do next.
Key takeaways
- Category-level visibility shows overall competitive position; prompt-level visibility explains what is driving it.
- A prompt, an engine, and a generated response are different units of measurement and need separate denominators.
- AI answers are non-deterministic, so repeated testing is more useful than treating one response as ground truth.
- A representative prompt portfolio can model different buyer needs, but it cannot reveal every private AI conversation or an individual user's actual journey.
- Mentions, owned-domain citations, prominence, portrayal, and recommendation are related but distinct signals.
- Prompt-level data helps prioritize content and authority work; it does not independently prove revenue impact or causality.
Category scores tell you where; prompts tell you why
An AI visibility score compresses a large set of observations into one number. That is useful when you need a quick benchmark, a competitive comparison, or a trend line. It becomes risky when the summary is treated as the explanation.
Imagine two brands with the same overall visibility score. The first appears consistently across category, comparison, and buying-intent prompts. The second dominates broad informational questions but disappears when buyers ask about integrations, alternatives, or specific use cases. Their category averages may look similar, but their commercial positions are not.
Prompt-level analysis reveals that difference.
Recent industry research illustrates why both levels matter. One study of more than a thousand US categories grouped representative prompts around definitions, comparisons, alternatives, use cases, and buying questions, and found that brand visibility could vary substantially across related prompts inside the same topic. The appropriate conclusion is not that category measurement fails. It is that a category conclusion depends on the prompts underneath it.
Pixis makes the same distinction in its guide to interpreting an AI visibility score: a composite score is a directional estimate built from sampled answers, engines, prompts, and weightings. It becomes useful when you segment it, benchmark it, and follow it over time.
What prompt-level visibility measures
Prompt-level visibility examines how a brand appears across a defined portfolio of questions submitted to selected AI engines. It can capture whether the brand is mentioned, where it appears in the answer, whether it is actively recommended or merely named, how it is described, which sources or domains are cited, which competitors appear instead, and how those patterns vary by prompt cluster and engine.
This is different from traditional rank tracking. An AI answer may not present a stable ordered list, and the same prompt can produce different brands, sources, or framing on separate runs. Prompt-level visibility is therefore better understood as sampled response analysis than as a fixed ranking system.
The measurement hierarchy
Reliable interpretation starts by separating five levels, from the broadest to the most granular.
The category or topic is the overall market being evaluated, such as how visible a brand is in project-management software. Beneath it sits the prompt cluster, a set of prompts with a shared intent or use case, such as enterprise comparison prompts. Within a cluster is the unique prompt, one defined buyer-style question, such as which project-management tools support complex approval workflows. Each unique prompt then becomes a prompt-engine combination when it is evaluated on one specific engine, since Gemini and ChatGPT may answer the same question differently. Finally, the generated response is one output from one run: did the brand appear in this specific response, or not?
Confusing these levels produces misleading metrics. A brand may appear in one of three runs for a prompt on ChatGPT, all three runs on Gemini, and none on Claude or Perplexity. "Visible for the prompt" is technically true, but it hides the instability and the engine difference.
Define the metrics and their denominators
Terms such as visibility, citation rate, and share of voice can mean different things across tools, so the definitions below describe the underlying measurement concepts rather than any single platform's labels. Whichever tool you use, the discipline is the same: every report should define both the numerator and the denominator, and you should confirm how your own platform names and calculates each figure before comparing it against anything.
Response-level mention rate is the percentage of evaluated responses that name the brand: responses mentioning the brand divided by total responses evaluated. If a brand appears in 240 of 1,200 generated responses, its response-level mention rate is 20%.
Prompt coverage is the percentage of unique prompts for which the brand appears at least once across the evaluated runs and engines: unique prompts with at least one brand mention divided by total unique prompts. Prompt coverage is broader than response-level mention rate, because a brand can cover many prompts but appear inconsistently within them.
Engine-level inclusion rate is the percentage of responses from a particular engine that mention the brand: responses from the selected engine mentioning the brand divided by total responses evaluated on that engine. This prevents strong performance on one engine from masking weak performance on another.
Owned-domain citation rate is the percentage of evaluated responses that cite a URL controlled by the brand: responses citing an owned domain divided by total responses evaluated. This should not be confused with brand mention rate. An answer can recommend a brand while citing a review site, publication, community, or comparison page.
Citation share is the brand's owned-domain citations relative to a clearly defined source or competitor set. The denominator must be disclosed, because citation share could refer to all citations in the response set, citations to the tracked brands' owned domains, or citations within a particular topic cluster. Those are different measures.
Prominence estimates how visibly the brand appears in the answer. Depending on the methodology, this may reflect first mention, position in a shortlist, answer-section placement, or relative emphasis.
Portrayal estimates how the answer frames the brand, whether positive, neutral, negative, or mixed. Because portrayal classification is itself modeled, teams should review the underlying response before acting on the score. The Pixis sentiment analysis glossary provides additional context on how sentiment is interpreted.
Recommendation rate is the percentage of responses that actively recommend or include the brand in a consideration set, rather than merely mentioning it. This is more commercially meaningful than raw presence, but it still does not prove that a user saw the answer, acted on it, or converted.
How the IAB framework changes AI visibility measurement
In August 2026, IAB released Measuring Visibility in the AI Era, a framework designed to help buyers evaluate AI visibility providers and methodologies. It organizes visibility into four layers: presence (does the brand or publisher appear), prominence (where and how visibly it appears), portrayal (how accurately and favorably it is represented), and persuasion (whether that visibility connects with attention, action, or business value).
The framework also distinguishes directional measurement from decision-grade measurement and emphasizes provider disclosure. Read IAB's announcement and framework overview.
One part of the framework is particularly relevant to prompt portfolios. IAB treats measurement programs with fewer than 50 queries as exploratory rather than directional, because smaller sets may not adequately characterize a category.
That does not mean 50 queries automatically create decision-grade data. Measurement quality also depends on whether the prompt set represents the category, how prompts were generated and validated, which engines, interfaces, locations, and languages were tested, how many responses were collected for each query, how stability and reproducibility were assessed, how personalization and session state were controlled, how metrics were calculated, and whether the methodology remains comparable over time. Prompt volume is one input to rigor, not a certificate of rigor.
The IAB hierarchy also clarifies the limits of an AI visibility platform. Response analysis can measure or estimate the presence, prominence, and portrayal. Persuasion requires connecting that visibility with downstream behavior or business outcomes. Visibility alone is not ROI.
Why repeated testing matters
AI-generated answers are non-deterministic. The same engine can return different brands, sources, wording, and recommendations when a prompt is run more than once. A single response is still useful as an example, but it is weak evidence for a general conclusion.
Pixis Visibility addresses some of that variation by running each prompt 12 times across ChatGPT, Gemini, Claude, and Perplexity, with three runs per engine. This produces repeated observations rather than one snapshot and makes engine-specific differences visible. Pixis documents this approach in its guide to multi-engine testing for GEO.
Repeated sampling reduces dependence on any single output. It should not be described as automatically eliminating variance or proving statistical certainty. The more defensible use is to report the number of runs and engines, retain the underlying responses, compare patterns by engine, monitor the same prompt portfolio consistently, and investigate sustained movement rather than reacting to one answer.
Build a representative prompt portfolio
No platform can observe every private AI conversation involving a brand. Prompt-tracking tools rely on controlled libraries that represent questions buyers may ask. The quality of the output, therefore, depends heavily on the quality of the prompt set.
A useful portfolio covers distinct intent types. Branded prompts ask whether a specific brand suits a given need. Category prompts ask about the best platforms in a space. Problem-led prompts describe a pain point, such as reducing delays in creative approval. Use-case prompts ask which platforms support a specific workflow, such as multi-market campaign production. Comparison prompts weigh one brand against another. Alternative prompts ask for options beyond an incumbent. Validation prompts test suitability for a context, such as regulated industries. Buying prompts ask which platform to choose for a specific requirement.
The portfolio should also represent the business, not just the category. That means reflecting products and services, customer segments, markets and languages, high-value use cases, common objections, competitors, sales-stage questions, and priorities drawn from search, sales, customer success, and market research.
Mapping prompts against a buying journey can help ensure the set covers discovery, evaluation, comparison, validation, and selection. But the portfolio remains a model of the journey. It does not reveal the exact sequence an individual user followed in a private conversation.
How many prompts should you track?
There is no universal number.
A common starting point in the industry is around 25 prompts, enough to cover core branded, category, and competitor intent without creating an unmanageable starting set. That figure often matches the allowance in entry-level tooling, so it is best treated as starter guidance rather than an industry standard.
For formal category measurement, IAB's framework cautions that fewer than 50 queries should be considered exploratory. Larger programs should expand based on the number of categories, personas, use cases, markets, languages, and decision stages that must be represented. The correct stopping rule is coverage, not an arbitrary promise that 200 or 500 prompts will fit every business.
What prompt-level data can reveal
Engine-specific gaps. A category score averaged across engines may look healthy even when one engine rarely includes the brand. Prompt-engine reporting shows whether the weakness is broad or concentrated.
Intent gaps. A brand might appear in educational prompts and disappear in comparisons or buying questions. That difference helps teams prioritize content and third-party validation closer to the decision.
Competitive gaps. Prompt-level analysis can show which competitors enter the consideration set for specific use cases. The Pixis analysis of AI recommendations in project-management software demonstrates how visibility can be split into distinct battlegrounds within one category.
Source gaps. Citation analysis can identify whether answers rely on the brand's own pages, earned coverage, commercial comparison sites, communities, or other sources. Research on AI citation sources found that even for branded or bottom-funnel queries, 48% of cited sources came from earned media, 30% from commercial third-party content, and only 22% from the brand's own website. The finding is dataset-specific, but it illustrates why an aggregate score can conceal the source mix supporting a brand's visibility. Review the source analysis and methodology.
Portrayal gaps. A brand can be present but framed with reservations about price, reliability, audience fit, or missing features. Prompt-level response review shows which questions trigger those qualifications and which external sources may be informing them.
Turn prompt-level evidence into action
Prompt tracking is valuable only when it changes a decision.
Diagnose before creating content. If a competitor is repeatedly cited, inspect the cited pages and source types before deciding that the solution is another article. The gap may involve missing first-party product information, weak comparison coverage, unclear entity signals, insufficient third-party corroboration, inaccessible or poorly structured pages, outdated information, or a genuine product or reputation disadvantage. The Pixis guide to getting cited by ChatGPT provides a broader execution framework for content, authority, and technical readiness.
Identify the page, not an imaginary causal sentence. Citation mapping can reveal which URLs and, where available, passages are associated with an answer. It does not prove that one sentence caused the engine to select the source. Use cited pages to develop hypotheses about coverage, structure, authority, freshness, and corroboration. Then make changes and monitor whether the relevant response pattern changes over time.
Separate content quality from brand context. Improving a generated draft and improving external AI visibility are different tasks. Context supplied to a writing tool can make content more specific and on-brand, while AI search systems cite information they retrieve from the public web. Pixis explains this distinction in How Much Brand Context Does AI Content Need?.
Prioritize with business evidence. A low-inclusion prompt is not automatically a high-demand opportunity. Prioritize prompt clusters using evidence such as sales-call frequency, customer research, search demand, pipeline value, product strategy, competitive importance, and the cost of being absent from that consideration set. Prompt visibility tells you where the brand is absent. Business context tells you whether that absence matters.
What prompt-level measurement cannot tell you
A credible measurement program should state its limits clearly. Prompt-level tools cannot observe every prompt users submit privately, reconstruct an individual buyer's complete multi-turn journey, guarantee that a sampled answer was seen by a real buyer, prove that a content change caused a visibility change without a stronger design, guarantee that an external AI engine will present the brand favorably, convert visibility directly into revenue attribution, or eliminate model and interface changes outside the provider's control.
This does not make the data useless. It defines the decisions the data can support. Use prompt-level monitoring to identify patterns, compare engines, diagnose source and portrayal gaps, and prioritize work. Connect it with search performance, referral traffic, conversions, pipeline, and experiments when evaluating downstream impact.
A practical implementation framework
- Define the business scope. Choose the products, markets, personas, competitors, and decisions the program must represent.
- Build and label the prompt portfolio. Group prompts by category, intent, persona, use case, market, and funnel stage. Document how each prompt entered the library.
- Establish a repeated baseline. Run the same portfolio across the selected engines and retain response-level evidence. Avoid drawing strategic conclusions from one run.
- Report at multiple levels. Use category trends for orientation, cluster trends for prioritization, and prompt-engine-response evidence for diagnosis.
- Review mentions, sources, and portrayal together. Do not interpret a mention without reading how the brand was framed and what sources supported the answer.
- Convert gaps into hypotheses. Decide whether the likely intervention involves owned content, technical access, entity clarity, third-party authority, product messaging, or reputation work.
- Measure sustained change. Track the same portfolio consistently and investigate repeated movement. Record content and authority interventions so changes can be interpreted against a timeline.
Frequently asked questions
What is prompt-level visibility?
Prompt-level visibility measures how a brand appears across a defined portfolio of questions and sampled AI-generated responses. It can include mentions, recommendations, prominence, portrayal, citations, competitor inclusion, and engine-specific differences.
Why is a category-level score not enough?
A category score summarizes overall position but hides the prompts, intents, engines, and sources producing it. It is useful for benchmarking. Prompt-level data supplies the diagnostic evidence needed to explain the score and prioritize action.
How many prompts should a brand track?
There is no universal number. A small starter set may be useful for exploration, while category measurement requires broader representation. IAB treats programs with fewer than 50 queries as exploratory. Expand based on the categories, personas, markets, use cases, and decision stages the program needs to cover.
What is the difference between mention rate and citation rate?
Mention rate measures how often a response names the brand. Owned-domain citation rate measures how often a response cites a URL controlled by the brand. An answer can mention or recommend a brand while citing an independent source, so the gap between the measures is not a direct trust score.
Can a tool see every AI conversation about a brand?
No. Consumer AI platforms do not expose every private conversation to visibility vendors. Tools use controlled prompt libraries and sampled responses to produce a representative measurement view.
Does prompt-level visibility measure ROI?
Not independently. It measures sampled presence and representation in AI-generated responses. Assess ROI by connecting visibility work with referral traffic, conversions, pipeline, revenue, or appropriate experiments.
Use aggregates for orientation and prompts for action
Category-level visibility is useful because it shows the competitive direction of travel. Prompt-level visibility is useful because it reveals the buyer's questions, engines, sources, and portrayals behind that direction.
Pixis Visibility runs repeated prompt tests across ChatGPT, Gemini, Claude, and Perplexity and connects the resulting mention, citation, competitive, and source data with content and SEO workflows. The methodology gives teams a more stable diagnostic view than a one-response check while retaining the underlying evidence needed for review.
The Storii uses Pixis Visibility alongside traditional SEO metrics to bring prompt coverage, citation share of voice, and competitive AI visibility into its client reporting. See how The Storii applies Pixis Visibility.
Explore Pixis Visibility to build a prompt-level view of how your brand appears across AI-powered discovery.

