On this page
← All articles
Guide·Sep 26, 2026·8 min read

How to Monitor AI Search Visibility in 2026: A Measurement Protocol

A top-down walnut desk in lamplight beside a ruled protocol for measuring AI search visibility

Here is how to monitor AI search visibility without fooling yourself: re-run a frozen prompt set against a fixed list of engines on a fixed schedule, count how often each answer mentions, cites and prefers your brand, and read the rolling average instead of any single reply. One reply is noise. SparkToro's January 28, 2026 test (600 volunteers, 12 prompts, 2,961 runs) found less than a 1-in-100 chance that ChatGPT returns the same brand list twice for the same prompt.

What Does AI Search Visibility Actually Measure?

AI search visibility is the share of AI answers, across a defined prompt set and a defined engine list, in which your brand appears. There is no results page and no position 1. Every answer is one sample from a model that may or may not search the web and rarely produces the same list twice.

Four layers stack. Presence: was the brand named. Evidence: was your domain among the sources. Position: were you the recommended pick or listed after a competitor. Tone: positive, neutral or negative framing. A brand present in 60% of answers still loses if the other 40% recommend someone else by name. For turning counts into targets, see AI search visibility metrics and KPIs; this post is about measurement.

How Many Prompts Do You Need, and How Do You Choose Them?

Start with 20 to 40 prompts per brand and market, written the way a buyer types them, and label each by intent. Conductor's July 17, 2026 analysis of 14,000 API calls across ChatGPT, Perplexity, Claude and Gemini found that intent predicts stability: comparison prompts kept the same lead brand in 91% of runs, while purchase prompts showed only 40% brand overlap between any two runs. One blended average hides that gap.

Cover five intents at minimum: comparison ("X vs Y"), recommendation ("best tool for…"), purchase, pricing and support, with 3 to 5 prompts each. Freeze the wording once the set is live; if a prompt must change, retire it and add a new one with a dated note on the chart. The argument for prompts rather than keywords as the unit is in search prompt monitoring.

Kairosy Prompt Tracking list for Figma: each buyer prompt with its platforms, last check, status and latest result

Which AI Engines Should You Lock In Before You Start Counting?

Write the engine list down, including the mode. ChatGPT with web search on and ChatGPT answering from memory are two different instruments. Kairosy fixes four (ChatGPT, Gemini, Claude, Perplexity), web-grounded on paid plans and during the trial, so each answer carries the URLs the engine leaned on.

In Kairosy's scan of Semrush (2026-08-05), the per-engine scores were ChatGPT 74/100, Gemini 58/100, Claude 59/100 and Perplexity 45/100, and the report notes: "Engine gaps like this usually mean the brand's story lives in sources one model reads and another doesn't." Never add an engine silently: a fifth engine with a lower baseline drags the blend down and reads as a decline. If ChatGPT is your priority, ChatGPT brand monitoring covers what that engine's answers are built from.

How Do You Calculate Mention Rate, Citation Rate and Share of Voice?

Mention rate

Mention rate = answers naming your brand ÷ total answers, computed per engine and per intent before blending. Worked example: 30 prompts × 4 engines × 3 runs = 360 answers. If 126 name you, mention rate is 35%. Show the per-engine split, because 35% can be 70% on Perplexity and 12% on Claude.

Citation rate

Citation rate = answers whose sources include a page on your domain ÷ answers that returned any sources. OpenAI's web search documentation separates inline citations, which "show only the most relevant references," from the sources field, which "returns the complete list of URLs the model consulted." Decide which you count: the wider list inflates the rate; inline-only is closer to what a reader sees.

Share of voice

Share of voice = your mentions ÷ mentions of every brand across the same answers. If the 360 answers contain 900 brand mentions and 126 are yours, share of voice is 14%. It moves when a competitor gains ground even while your own mention rate holds, so it belongs beside mention rate, not instead of it.

Kairosy report mid-section showing per-engine scores and answer breakdown used to compute mention rate and share of voice

How Do You Track Sentiment and Competitors in the Same Run?

Classify every answer, not every mention, with five labels: positive, neutral, negative, prefers competitor, invisible. Kairosy's Semrush report shows why the last two matter: of 24 answers, 10 were positive, 4 neutral, 4 negative, 5 preferred a competitor and 1 omitted the brand, with Ahrefs named across all four engines. The underlying answers are in the Semrush AI performance report.

Store the reason with the label; "negative: billing complaints" is actionable, "negative" is not. Kairosy attaches the reason to each answer and computes the headline score as the average of the per-engine scores of the engines that answered, so a quiet engine cannot shrink the denominator. For competitors, extract every other brand named and count per engine. The one dominating "prefers competitor" answers is your real rival in AI search, and often not the one your SEO reports show.

Kairosy public AI performance report for Semrush with per-engine scores, sentiment counts and named competitors

Should You Monitor AI Search Visibility Manually or With a Tool?

Go manual first if you have one brand, one market and a monthly cadence: a non-personalized temporary chat, N runs per prompt, answers pasted into a sheet and tagged by hand. The arithmetic is the constraint: 360 answers per cycle is roughly six hours of collection at a minute each, before classification.

Switch to a tool when you need weekly or daily cadence, several markets, real citation URLs per engine, or output someone else reads. Kairosy's honest entry: a visibility and reputation scanner, not a rank tracker. Signup is a 7-day free trial with no credit card and full reports. As of September 2026 the plans are Basic $29, Pro $99 and Growth $399 per month, with 15 / 28 / 70 tracked prompts checked daily and 3 / 6 / 20 tracking slots (one brand × market, re-scanned every Monday with an email digest). Otterly, Peec AI and Profound are alternatives; compare runs per prompt and per-engine citations. The evidence for tooling is in why teams use AI search monitoring tools; the workspace is on the AI brand monitoring page.

Kairosy Fix Strategy suggestion 4 of 8 for Figma: the schema names no author or dates, so provenance is invisible to citation ranking

How Often Should You Re-Run the Set, and How Do You Read a Trend?

Daily prompt checks catch an engine flipping its answer within a day. Weekly full scans on the same weekday produce the series you chart. Monthly reviews compare intents, engines and markets.

Read a four-week rolling mean, not the latest point. An eight-point drop on one engine in one week is inside normal variance for volatile intents; the same drop sustained three weeks is a signal. Alert on events, not levels: a new negative keyword, a competitor becoming the preferred pick, a score drop beyond tolerance, an engine changing sentiment. Annotate every prompt change and new model label as a series break. When the trend turns down, start on the pages the engines cite: run an AI-ready audit on the URLs in your citation list and confirm crawlers can reach them with the free AI crawl checker.

How to Test Prompts for LLM Search Visibility

Testing a prompt means measuring its variance before trusting its mean. Here is how to test prompts for LLM search visibility.

  1. Pilot each prompt 10 times on one engine in a clean session. Record mention rate and the spread across runs.
  2. Drop prompts that return generic advice with no brands named. They are unmeasurable, not "zero visibility."
  3. Group survivors by intent and compare spread. Expect comparison prompts to be tight and purchase prompts loose.
  4. Set runs per prompt from the spread. A prompt swinging between 20% and 70% mention rate needs 10 or more runs per cycle; one that sits at 90% every time can run 3.
  5. Log every raw answer with engine, model label, timestamp, location and session type. Aggregates can be recomputed; raw answers cannot be recovered.

The research argues for conservative run counts. Atil et al.'s "Non-Determinism of 'Deterministic' LLM Settings" (arXiv, updated April 2025) ran five models on eight tasks 10 times each and measured accuracy swings of up to 15% between runs. SparkToro's conclusion after 2,961 runs: "visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric," and any tool showing a "ranking position in AI" deserves suspicion. The protocol table fixes every variable before the first real cycle.

VariableSetting to fixSource
Runs per prompt10 in the pilot; 3–10 per cycle by measured spreadConductor: 50 runs per cell (July 2026); Atil et al.: up to 15% swings across 10 runs
Temperature / samplingLeave the engine default and record it; 0 does not mean repeatableOpenAI: range 0–2, seeded sampling is "best effort"; Anthropic: not fully deterministic at 0.0, and models after Claude Opus 4.6 reject the parameter
Session stateNon-personalized temporary chat or a dedicated API key; never your own accountOpenAI Temporary Chat FAQ: such chats "do not use memory, custom instructions, or plugins"
LocationOne fixed country (and city if you sell locally) per projectOpenAI web search: user_location with country, city, region, timezone; Perplexity API: country, region, city, latitude, longitude
Model versionRecord the model label per run; a label change is a series breakGoogle: a "-latest" alias "will get hot-swapped with every new release"; OpenAI: gpt-5.6 is an alias for gpt-5.6-sol
Prompt normalizationWording frozen; brand-matching rules (name, domain, products) written downSparkToro normalized every response into an ordered list before comparing

SparkToro research post reporting that AI tools rarely return the same brand list twice across 2,961 runs

OpenAI web search tool documentation showing the approximate user_location fields country, city, region and timezone

Why Do My Rankings Change Between Sessions in ChatGPT?

Because there is no ranking to hold still. Each session draws a fresh answer from a model that samples tokens, may or may not search, may or may not remember you, and may have been updated since yesterday. Five mechanisms explain nearly all of it.

Sampling never fully switches off

OpenAI's chat completions reference puts temperature "between 0 and 2" and says a seed makes the system "make a best effort to sample deterministically," but "determinism is not guaranteed." Thinking Machines Lab showed why in Defeating Nondeterminism in LLM Inference (September 10, 2025): 1,000 completions of one prompt at temperature 0 produced 80 unique outputs, because GPU kernels are not batch-invariant under changing server load. Anthropic's Messages API docs say it in one line: even at 0.0, results are not fully deterministic. You cannot configure this away; you can only average over it.

Session state follows you

Testing in your own ChatGPT account means memory and custom instructions shape the answer. OpenAI's Temporary Chat FAQ states that non-personalized temporary chats "do not use memory, custom instructions, or plugins." A marketer who has asked ChatGPT about their own product forty times is not a neutral observer of it.

Retrieval depends on where and when you ask

When the model searches, its sources depend on location and time. OpenAI's web search tool accepts an approximate user_location (country code, city, region, IANA timezone); Perplexity's user location guide accepts country, region, city, latitude and longitude. A tester in Austin and one in Manchester are running two different studies. Pin one location per project.

The model underneath moves

Aliases are pointers. Google's Gemini model documentation says a "-latest" alias "will get hot-swapped with every new release," while stable names "usually don't change." OpenAI's model list shows gpt-5.6 as an alias for a specific build, gpt-5.6-sol. The ChatGPT app shows a model label but no changelog on your schedule; write it down every run and start a new chart segment when it changes.

Your own counting rules drift

If "Semrush," "SEMrush," "semrush.com" and "Semrush's Site Audit" are counted inconsistently, mention rate moves without the engine doing anything. Write the matching rules once: brand name, domain, product names, misspellings, and how to treat a similarly named competitor. Apply the same rules to competitors, or share of voice compares two definitions.

Thinking Machines Lab post explaining why LLM inference is nondeterministic even at temperature zero

Google Gemini API model documentation describing stable, preview and hot-swapped latest model aliases

How to Monitor AI Search Visibility FAQs

How do you monitor brand visibility across AI search engines?

Use one prompt set for every engine and report per-engine mention rate, citation rate and sentiment side by side before blending. The per-engine view is where you learn that Perplexity cites a competitor's comparison page while ChatGPT never searched at all.

How do you monitor AI search visibility over time?

Freeze prompts, engines, location and runs per prompt; re-run on the same weekday; chart a four-week rolling mean per engine; mark prompt and model-label changes as series breaks. How to monitor AI search visibility over time is really a question of keeping the instrument constant.

How do I monitor AI search visibility if my location is the USA?

Set the country explicitly, as an API user_location parameter or a monitoring project scoped to the US market; never rely on your IP address or a VPN. Kairosy scopes each project as brand × market, with Global plus 25 countries on paid plans.

How many times should you run each prompt?

Enough that the spread is smaller than the change you want to detect. Pilot at 10 runs, measure the range, then set 3 to 10 runs per prompt per cycle.

Is there a free way to monitor visibility in AI search results?

Yes, for a small set: a non-personalized temporary chat plus a spreadsheet covers one brand and one market monthly. Kairosy's 7-day trial runs full reports with real citation URLs and no credit card, the quickest way to learn how to monitor AI search visibility for brands with several markets.

See what AI says about your brand

Run a free scan across ChatGPT, Gemini, Claude & Perplexity in about 30 seconds.

Run my free scan