Tracking AI visibility through an uncategorized list of prompts produces a binary signal: you either appeared or you didn't. That list doesn't tell you whether you're improving across topics, personas, or buying stages.

Teams tend to swing to one of two extremes when building a prompt list: thousands of prompts that show only marginal gains, or a handful of prompts blended into a single visibility score that's too high-level to act on. 

The sweet spot is a carefully designed and representative prompt dataset built with KPIs and actionable insights in mind from the start. Here's how to build, track, and interpret an AI prompt dataset for your brand.

1. Organize Your Prompt Set by Intent

Prompts aren't keywords to track. While you can assign a synthetic search volume based on keyword search volume, that's just forcing an old method onto a completely new system.

Traditional search worked off a finite keyword universe plus rankings. LLMs work off a virtually infinite prompt universe, plus synthesized recommendations, citations, mentions, and comparisons. That's a layer of reasoning old tracking never had to deal with: sentiment, recommendation placement, competitor comparisons right in the chat interface.

Learn more: What Google Wants Isn’t What AI Wants

Try to track all of that complexity with one flat prompt list, and you run into the same handful of problems every time:

  • Only track citation rate, and you miss a huge chunk of text mentions, or mentions grounded in other sites instead of your own.
  • Only measure commercial, non-branded prompts, and you miss the chance to surface what an LLM actually thinks about your brand.
  • Weight everything evenly, and you're just as lost as when you were missing pieces of the picture.

There's no magic-bullet "track X metric" solution. To get actionable data as an output, you need a well-structured prompt set as an input. This means intentionally dividing your prompts into categories that each pull a different signal out of an LLM's response. That's what lets you see clearly where you already show up and where you need attention, both on-page and off.

To build a strong AI prompt set you can track, start with these four intent categories:

  • Branded: prompts that name your brand directly
  • Commercial: non-branded comparison and buying-stage prompts
  • Competitor: prompts that name or compare with competitors
  • Informational: content-driven prompts tied to your topical strategy

Step One: Identify Prompts by Type

1. Branded 

If you measure a brand's mention rate for a prompt set that mixes in non-branded prompts, that mention rate doesn't tell you much of value. Separate branded and non-branded from the start:

Non-branded:

  • "Best golf clubs for beginners"

Branded:

  • "Are [brand] golf clubs good for beginners?"

2. Commercial

Non-branded commercial prompts are the closest category to traditional keyword tracking, except ranking is no longer the finish line. When evaluating these, look beyond placement: Are you mentioned? Referenced as a source? Recommended as the top pick, or buried at the bottom?

For most businesses, this will be your largest bucket, since commercial intent isn't one kind of prompt. This category spans a few prompt "shapes":

  • Category discovery: "What are the best golf clubs for beginners?"
  • Problem/solution: "Which golf clubs are forgiving for beginners with a high handicap?"
  • Constraint-based: "Best golf club sets under $1,000"

3. Competitor Alternatives

While commercial prompts ask "what's best for this need?", this category asks something more specific: "what's an alternative to this competitor?" Naming the competitor directly lets you see whether your brand surfaces at all, whether you land among the top choices, and what pros and cons come up along the way.

Examples:

  • "Best alternatives to Callaway for beginner golf clubs"
  • "Callaway vs. [brand] for beginner golf clubs"

This category is where sentiment really starts to matter. You're not just asking "do I show up"; you're asking:

  • Am I mentioned before my competitors, or after?
  • Am I in the direct competitor set, or getting pushed down into "other alternatives"?
  • What's the sentiment when I am mentioned — favorable, neutral, or am I the brand being used as the comparison point for someone else?

This is the category that tells you what an LLM actually thinks of you relative to the field, instead of just whether you're in the field at all.

4. Informational

This one's trickier, and it needs to be dialed into your content strategy from the start. You need to be tracking your target informational prompts before you publish so you can watch whether you start surfacing once the content goes live. 

Example: "How do I choose the right golf shaft flex?"

The golden rule is to keep content useful for humans first: it might surface in traditional search, or someone might land on it through an LLM citation. Avoid pumping out hundreds of thin pages just to try to influence a model. Instead, make sure each page is genuinely helpful, quality content for people.

Step Two: Organize Prompts Into Topics

Once you've got your four categories, add a second layer of organization by tagging each prompt by topic:

  • Group prompts by the topics or entities they cover
  • Make sure each topic has enough prompts to produce meaningful data
  • Keep topics distinct enough that you can tell where visibility is strong and where it's weak

Map your topics to how your business actually categorizes its offerings, usually by product line or service area.

Example topics: beginner golf clubs, women's golf, budget golf club sets

One rule to keep straight: don't mix categories inside a topic. "Beginner golf clubs" isn't a single bucket of branded, commercial, and competitor prompts. It's three separate buckets that share a topic tag. Blend them together, and you lose the ability to tell which metric is actually moving.

Informational prompts use the same topic lens, but those topics map directly to your content strategy. This means separating by content topic, pillars, or query fan-out groups you're trying to surface for.

Step Three: Define Your Competitive Set

The final layer is defining a list of competitors to track against. For each topic and intent, a competitor list will give you a baseline to measure against rather than viewing your visibility score in a silo. For example, a 40% mention rate for commercial prompts about beginner golf clubs might look strong on its own, but competitors could be averaging much higher.

If additional competitors keep showing up in answers that aren't on your list, take note. Look into where they're getting mentioned and why, and consider whether they belong in the set. If you add new competitors down the line, make sure to mark where the change happened so you can account for it in your trend line.

2. Match KPIs to Your Prompt Groups

After you’ve organized your brand’s prompts by intent, it’s time to set tracking for each category. This is where the value of the prompt groups becomes clear: because the KPIs shift by prompt intent, your AI visibility measurement changes from a single metric to a comprehensive dataset with unique outputs.

Branded: Sentiment and Page Citations

Branded prompts surface brand-specific sentiment alongside page citations, covering both first-party assets and third-party sources like reviews or digital PR. A prompt like "what do reviews say about [brand]" reveals what the LLM says about you and where it gets its information.

Metrics to track:

  • Cited pages: which pages does the model cite? Are they owned pages or third party mentions?
  • Sentiment: what positive and negative attributes are listed?

The AI does the hard work of surfacing which sources it’s citing. Your job is to fill in the gaps (or fix the wrong information) with correct, brand-aligned content for it to find. That could mean:

  • Expanded digital PR across external review platforms and industry media
  • Detailed first-party content, such as original benchmarks, comparisons, or product breakdowns, rather than generic listicles
  • A comprehensive brand page clearly explaining who you are and what you do

Tip: The goal isn't to trick the LLM into recommending you, but instead to tell a true, consistent story across every channel. This ensures you're discoverable, recommendable, and citable no matter where the model goes looking.

Commercial: Mention Rate and Position

Non-branded commercial prompts are the comparison and buying-stage prompts: "what are the best golf clubs for beginners?" People rely on LLMs to gather this information for them, and they tend to trust what it says. So, give the model real reasons to recommend you.

Metrics to track:

  • Mention rate: are you even in the conversation, and are you in it for this specific topic?
  • Average position: once you're in the conversation, where do you land against competitors? Is that trending up or down over time?
  • AI referral traffic (from GA4): which pages are users landing on from a chat session?

Tip: Measure against your competitors, split by topic, so you can see exactly which topics need work and who's winning where.

Competitor Alternatives: Dominating the Conversation

With competitor prompts (e.g., "best alternatives to X"), you want models to pull from sources that favor you and information that represents your brand accurately in the competitive field.

Metrics to track:

  • Mention rate: are you mentioned as an alternative?
  • Sentiment: when you're placed next to a competitor, how does the model talk about you?
  • Cited pages: which pages is the model actually pulling from to make the comparison?

Informational: Citation Count and Competitor Citation Rate

This is the category where you're watching for blog or guide content to start getting picked up. 

Metrics to track:

  • Citation count per page: how many of the pages you shipped get cited, and how often?
  • Citation rate by competitor: how often your content gets cited on a topic compared to competing content on the same topic.

Tip: Consider measuring these across search-enabled platforms only, such as Google AI mode and AI overviews. LLMs often skip citations if an answer is already in their training data.

Break KPIs Down by Topic

Run every category's KPIs through the topic level too. A brand might feel like it's in a strong position overall and still be surfacing behind every competitor on the actual prompts in that category.

You won't catch that gap by looking at blended numbers. You only catch it by slicing branded, commercial, and competitor KPIs down to the topic level.

Tip: Run each prompt multiple times across multiple models, then look at the trend over a reporting window. Top cited pages across every model and every day of the month tells you a lot more than one day with one model ever will.

3. Turn the Data Into Actionable Insights

Once the categories and KPIs are in place, the actual insight work comes down to two moves.

Step One: Find the Gap

For a given topic (e.g. beginner golf clubs) are you listed as an alternative?

For the "best X" prompts in that topic, are you there at all?

If you're not, analyze the brands that are. Identify what content strategies and positioning they use, and map where the model extracts that data—whether from their domain, third-party review platforms, or affiliate comparison pages. 

Identifying the source dictates the direct strategic action: publishing targeted first-party assets, executing a digital PR campaign, or resolving brand data discrepancies across the web.

Step Two: Overlay Changes and Watch for Movement

Did you build a real, valuable comparison page? Push a piece of digital PR? Fix a brand-consistency gap?

Layer that change against the topic's visibility over time and see if it actually moved.

We tried this tactic with our own AI information page. We built this page to give models a clear, accurate base of information about our brand. 

We had the target prompts tracked before it went live, and after publishing, citation rate climbed from zero to 25-30% across Google AI Overview, Google AI Mode, and Gemini within weeks.

That's the prompt system working as intended: seed the prompts, ship the change, watch the metrics move.

Make Your Existing Prompt Set Actionable

You don't need to build this system from scratch. If you're already tracking prompts, start by sorting what you have into four categories: branded, commercial, competitor, and informational. That low-effort change can immediately make the data easier to interpret.

A mention rate, citation count, or average position means very little on its own. Once those metrics are tied to prompt intent, topic, competitors, and the changes you're making, they become much more useful for deciding what to work on next.

AI visibility measurement is still evolving, and there probably won't be one universal score that captures everything that matters. For now, a well-structured prompt set gives you something much more practical: a consistent way to understand where your brand stands, identify meaningful gaps, and track whether your work is actually improving visibility over time.