Metrics · calculator · free tracking sheet

AI Visibility: Definition, Metrics and Tracking

AI visibility is the share of relevant AI answers that mention or cite a brand. Track it by running a fixed prompt set on 6 AI engines on a fixed schedule and logging 5 values per answer: mention, citation, cited URL, position and sentiment. Calculate your mention rate, citation rate and share of model below, and download the free tracking sheet pre-filled for Google AI Overviews, AI Mode, ChatGPT, Perplexity, Gemini and Copilot.

This page covers a brand's visibility inside AI-generated answers — not visibility into AI systems as used in AI governance, and not a single product's 'AI visibility' report.

Calculate your AI visibility

Example values are pre-filled: 12 mentions in 60 answers is a 20% mention rate.

The 5 AI visibility metrics
MetricFormula
Mention rateAnswers naming the brand ÷ answers checked
Citation rateAnswers linking a brand URL ÷ answers checked
Share of modelBrand mentions ÷ mentions of all tracked brands
Answer positionAverage order of first mention (1 = first)
SentimentPositive mentions ÷ all mentions

Tracking sheet: CSV, 13 columns, pre-filled for 6 engines. Opens in Excel, Google Sheets or Numbers. Free, no sign-up.

What is AI visibility?

AI visibility is the share of relevant AI-generated answers that mention or cite a brand. It is measured per engine by running a fixed set of prompts on a fixed schedule and logging who each answer names, links and recommends.

When someone asks Google AI Overviews, AI Mode, ChatGPT search, Perplexity, Gemini or Microsoft Copilot a question, the engine writes an answer and names a few sources and brands inside it. AI visibility measures how often your brand is among them for the questions that matter to your business. It goes by several names — AI search visibility, LLM visibility and brand visibility in AI answers all describe the same quantity.

Three words in the definition carry the method. Relevant means the prompts are ones real buyers ask, not vanity queries containing the brand name. Share means the result is a rate, not a count: 12 mentions mean little until you know they came from 60 answers. Per engine means each AI product is measured separately, because each uses a different crawler, index and citation style, and an average across engines hides where the problem is.

AI visibility is the outcome that generative engine optimization works towards. Rankings measured success when a search returned ten links; an AI answer has no positions 1 to 10, only a set of claims and the sources behind them. The measurement therefore moves from “where does my page rank?” to “does the answer name me, link me and describe me well?”

The quantity is also volatile. The same prompt can produce a different answer an hour later, with different sources. A single screenshot of one answer is an anecdote; visibility is a rate across many answers, prompts and runs. Everything that follows — the 5 metrics, the tracking steps and the sizing — exists to turn that volatility into a number you can compare from one month to the next.

Which 5 metrics make up AI visibility?

AI visibility is made up of 5 metrics: mention rate, citation rate, share of model, answer position and sentiment. The first three are ratios of answers or mentions; position averages where the brand first appears; sentiment is the positive share of mentions.

The worked examples below use the calculator’s default values: 60 answers checked, 12 naming the brand, 5 linking to it, and 80 mentions of all tracked brands in total.

Mention rate

Mention rate = answers naming the brand ÷ answers checked.

With 12 of 60 answers naming the brand, the mention rate is 20%. A mention counts when the brand, product or a clear variant of the name appears in the answer text, with or without a link. Mention rate is the broadest measure of presence: it shows whether the engine associates the brand with the topic at all. Decide in advance which name variants count, write them in the sheet’s notes, and apply the rule the same way every run.

Citation rate

Citation rate = answers linking a brand URL ÷ answers checked.

With 5 of 60 answers linking to the brand’s site, the citation rate is 8.3%. A citation is a link to a page on the brand’s own domain, shown inline, in a footnote or in a source panel. It is the metric that can send traffic, and the one that shows which page the engine trusted.

A citation can appear without the brand being named in the answer text, so the two rates move independently. Engines that list numbered sources often link a page while the answer itself describes the topic generically. A high citation rate with a low mention rate means the engine uses your content but does not credit your brand; the reverse means it knows your name from other sources but does not use your pages.

Share of model

Share of model = brand mentions ÷ mentions of all tracked brands on the same prompts.

With 12 of the brand’s mentions among 80 mentions of all tracked brands, share of model is 15%. It is the AI-answer counterpart of share of voice, and it answers the competitive question: when AI engines recommend someone in this category, how often is it you? The number depends entirely on which competitors are tracked, so fix the competitor list before the first run and keep it. The share of model formula page covers how to choose that list and how to handle brands that appear only once.

Answer position

Answer position = the average order of the brand’s first mention in answers that name it (1 = first brand named).

If the brand is named first in 4 answers, second in 5 and third in 3, its average position across those 12 answers is 1.9. Position matters because readers skim: the first recommendation in a list gets more attention than the fourth. Answers that do not name the brand are excluded from this average and captured by mention rate instead, so read the two together.

Sentiment

Sentiment = positive mentions ÷ all mentions of the brand.

If 9 of the 12 mentions describe the brand favourably, sentiment is 75%. Mark each mention as positive, neutral or negative using one written rule — for example, “positive” when the answer recommends the brand or names a strength, “negative” when it names a drawback or warns against it. Sentiment catches the case where visibility rises for the wrong reason, such as a wave of answers citing a complaint.

Together the five metrics describe presence, credit, competitive standing, prominence and tone. The full set of AI search metrics adds traffic and conversion measures that sit downstream of these.

How do you track AI visibility?

Track AI visibility in 5 steps: build a fixed prompt set, fix the engines and run conditions, run each prompt more than once, log every answer in one sheet and report the 5 metrics per engine and period. Prompt wording stays unchanged between reports.

The method below is the one this site uses for its own experiments. It needs no paid software: the free tracking sheet holds every field, and the calculator turns the counts into rates.

Build a fixed prompt set from real buyer questions

A prompt is the full question a person types into an AI engine. The prompt set is the fixed list you run every time. Build it from real sources — sales calls, support tickets, site search logs, the “People also ask” questions in your market and the questions prospects send by email — rather than from keyword tools alone, because prompts are longer and more conversational than keywords.

Group the prompts by intent, so the report can show where visibility is strong and where it is missing:

  • Definition prompts ask what something is: “What is fractional CFO work?”
  • Comparison prompts weigh options: “Fractional CFO or full-time finance director for a 30-person company?”
  • Recommendation prompts ask for providers: “Best fractional CFO firms for SaaS startups in Austin.”
  • How-to prompts ask for a process: “How do I prepare my books for a Series A?”

Recommendation prompts usually matter most commercially, because the answer names brands directly. Definition and how-to prompts show whether the engine uses your content as a source. Give each prompt an ID (P001, P002 …) in the sheet, and never edit its wording once tracking starts; a reworded prompt is a new prompt.

Fix the engines and run conditions

AI answers change with who asks, where and how. Hold these conditions constant for every run:

  • Use a logged-out session or a clean profile, so past conversations and personalisation do not shape the answer.
  • Fix the country and language, because engines localise sources and brands.
  • Run in the same day-and-time window each period, to reduce drift between runs.
  • Note the engine mode — AI Overviews and AI Mode are separate Google surfaces, and some engines offer several models or search settings.

The tracking sheet is pre-filled for 6 engines: Google AI Overviews, Google AI Mode, ChatGPT search, Perplexity, Gemini and Microsoft Copilot. Drop the ones your buyers do not use, but decide before the first run and keep the list fixed.

Run every prompt more than once

A 2026 survey of 45 GEO studies (arXiv 2607.14035) reports substantial run-to-run variability in AI answers and recommends repeated measurements, paraphrased prompts and controls as the basic protocol. A prompt answered once is a sample of one.

This method runs each prompt 3 times per period and records every run. A brand named in 2 of 3 runs is a different result from one named in 3 of 3, and the difference is invisible if you check once. For prompts that matter most, add one or two paraphrases — the same question in different words — to check that visibility does not depend on exact phrasing. Log the run number (1, 2 or 3) in the notes column so repeated runs stay distinct.

Log each answer in the tracking sheet

Each row in the sheet is one answer from one engine to one prompt. The 13 columns capture everything the 5 metrics need:

The 13 columns of the AI visibility tracking sheet
ColumnWhat to enterExample
dateDate of the run2026-10-01
engineEngine and modeGoogle AI Mode
prompt_idFixed ID of the promptP001
promptExact prompt wordingBest fractional CFO firms for SaaS startups
intentIntent grouprecommendation
brand_mentioned (y/n)Brand named in the answer texty
brand_cited (y/n)Brand URL linkedn
cited_urlThe linked URL, if any—
position_in_answer (1=first)Order of the brand's first mention2
competitors_mentionedOther tracked brands namedBrand A; Brand C
competitors_citedOther tracked brands linkedBrand A
sentiment (pos/neu/neg)Tone of the brand mentionpos
notesRun number and anything unusualrun 2; answer listed 5 firms

The two competitor columns feed share of model, and the cited_url column shows which of your pages each engine trusts. Fill in every column even when the brand is absent: a row with “n” and “n” is data, not a gap.

Report per engine and per period

Count the rows, apply the formulas and report one line per engine for each period. Keep the engines separate; a single blended score hides the engine where visibility is falling.

Example figures, not client data: one engine, one period, 60 answers
EngineMention rateCitation rateShare of modelAnswer positionSentiment
Google AI Overviews20.0%8.3%15.0%1.975%

The example row uses the calculator’s defaults and the position and sentiment examples above to show the layout. A real report has one row per engine, and the comparison between periods is what matters. Add a short note under each period listing the changes made to the site since the last one, so movement in the numbers can be matched to its likely cause. The first report is the baseline everything else is compared against, and an AI visibility baseline audit produces it for you along with the checks that explain it.

How many prompts does an AI visibility tracking set need?

A 50-prompt set run weekly on 6 engines produces 300 answers per run, or 900 with 3 runs per prompt. That is the protocol used for this site's own 90-day experiment; smaller brands start at 20 prompts and keep the same schedule.

The arithmetic is simple: prompts × engines × runs. At 50 prompts, 6 engines and 3 runs, each weekly cycle produces 900 rows. That volume gives each engine 150 answers per week, enough for a change of a few mentions to show as a movement in rate rather than noise.

The 20-prompt starting point is this site’s method for smaller brands, not an industry rule. At 20 prompts on 6 engines with 3 runs, a weekly cycle is 360 rows, which one person can log by hand. Cover every intent group in the smaller set, and weight it towards recommendation prompts. Grow the set only by adding prompts; never swap old ones out mid-quarter, or the periods stop being comparable.

Experiment 01 on this site — a brand-new domain with zero backlinks — is running on the 50-prompt protocol, with a weekly log over 90 days. No results are reported yet.

What do Search Console and Bing Webmaster Tools report about AI visibility?

Search Console's Generative AI performance report shows impressions in AI Overviews and AI Mode by page, country, device and date; Bing Webmaster Tools' AI Performance report shows Copilot citations and grounding queries. Neither reports mentions without a link or competitor share.

Three first-party sources report on AI answers directly. They are free, and they come from the engines themselves, so they are the fixed reference points for any tracked figures.

First-party data on AI visibility, and what each source misses
SourceEngines coveredWhat it showsWhat it misses
Search Console Generative AI performance reportGoogle AI Overviews, AI ModeImpressions by page, country, device and date; available to all sites from 31 August 2026Clicks from AI features in this report, unlinked mentions, competitors, prompts
Bing Webmaster Tools AI PerformanceMicrosoft Copilot, AI summaries in Bing, select partnersTotal citations, average cited pages, grounding queries, page-level citation activity; public preview since 10 February 2026Unlinked mentions, competitors, sentiment
GA4 with utm_source=chatgpt.comChatGPTVisits from links clicked in ChatGPT answersAnswers that named or linked you without a click

The table shows why tracked prompts are still needed: no first-party source reports mentions without a link, competitor share or sentiment, and none covers Perplexity or Gemini. A grounding query, in Bing’s terms, is the phrase its systems used to retrieve content for an AI answer — the closest any engine comes to showing the prompts behind a citation.

Google adds a caution about the other side of the market: “No third-party tool has access to our internal ranking or AI systems.” Paid trackers observe answers from the outside, exactly as the manual method does; they do not see inside the engine. Treat Search Console and Bing figures as the ground truth for their engines, and use tracked prompts for everything they leave out.

Is AI visibility the same as AI referral traffic?

No. AI visibility counts appearances in answers; AI referral traffic counts the clicks those answers send. Many answers name a brand without a click, so referral traffic alone understates visibility, while visibility alone misses revenue.

Referral traffic is measured in web analytics. In GA4, ChatGPT visits carry the utm_source=chatgpt.com parameter, and visits from other engines appear as referrals from their domains. Those numbers show which answers send people to the site, and they connect AI search to leads and sales.

They cannot show the answers that named the brand and sent no one. A buyer who reads a recommendation in an AI answer and later searches the brand name directly arrives as branded search or direct traffic, not as an AI referral. Report both: visibility as the leading measure of presence, referral traffic as the measure of what that presence delivers.

Manual tracking works to a point; when does software take over?

When is an AI visibility tracking tool worth using?

A tracking tool is worth using once manual logging costs more hours than the tool saves. This site's threshold is one person running more than a 50-prompt set on 6 engines each week. Tools add scheduling, competitor parsing and history, but follow the same 5 metrics.

Below that threshold, the sheet and a fixed hour each week do the job, and the method stays fully transparent.

Three questions separate a useful tool from a dashboard of unexplained scores. Can you enter your own prompts, word for word, and keep them fixed? Does it state the country, language and session conditions it runs under? Does it export the raw answers, so each mention and citation can be checked by hand? A tool that answers no to any of these produces numbers you cannot reconcile with Search Console, Bing Webmaster Tools or your own sheet. Above it, software removes the logging while keeping the metrics the same — as long as the tool lets you fix the prompts, engines and run conditions yourself. The comparison of AI visibility tracking tools covers what each type of tool adds and what to check before paying for one.

Frequently asked questions

Can AI visibility be tracked for free?

Yes. A fixed prompt set, the free CSV sheet on this page and manual runs cover all 5 metrics. Search Console and Bing Webmaster Tools add free first-party data, and GA4 separates ChatGPT referrals tagged utm_source=chatgpt.com.

How often do AI answers change?

Often enough that one check proves nothing. A 2026 survey of 45 GEO studies reports substantial run-to-run variability in AI answers to the same prompt. That is why this method runs each prompt more than once and compares trends across weeks.

Related guides

Sources

  1. Google Search Central Blog — Generative AI performance reports in Search Console
  2. Search Console Help — Generative AI performance report
  3. Google Search Central — Optimizing your website for generative AI features on Google Search
  4. Bing Webmaster Blog — Introducing AI Performance in Bing Webmaster Tools (10 February 2026)
  5. OpenAI Help Center — Publishers and developers FAQ
  6. Critical survey of generative engine optimization research (arXiv 2607.14035)