Visibility runs
A visibility run asks your fixed question set (12 unbranded questions and 2 diagnostics) to ChatGPT, Claude, Gemini, Perplexity and Grok through OpenRouter, with live web search and your own key, and records who was named, recommended and cited. API answers are a proxy for the consumer apps and one run is a small sample, so confirm key questions by hand and re-run the same set every 14 days. Visibility runs and manual checks are part of Audit Week and Pro; the OpenRouter cost is separate and paid by you to OpenRouter.
What a run measures
For every question and assistant, Katman records:
- whether your brand was named;
- whether it was recommended first, and its position among the products named;
- whether one of your URLs was cited as a source;
- the other products and competitors named, and the domains cited;
- whether the model actually searched the web, a short excerpt, and the cost.
For the two diagnostic questions it also checks identity: is your brand described as the right product? A small judge model extracts the recommended products in order, with a rule-based fallback. Katman reports named, recommended and cited separately and doesn’t merge them into one score.
Every input of katman_visibility_run is listed in the tools reference. Results go to .katman/runs/<time>.json and the table in .katman/measurements.md, with changes since the previous run, the competitors and domains that come up most, and the date of the next run, 14 days later.
The question set: 12 + 2
katman_research builds the set from your research (book chapter 15). Your brand’s name appears only in the two diagnostics.
| # | Type | Template | What it tells you |
|---|---|---|---|
| 1–3 | Situation | “[Situation] but [obstacle]. Is there a site that [result]?” | The type small sites win |
| 4–5 | Tool | “Which AI tool does [job]?” | Short, generic |
| 6–7 | “Best X” | “Best [category] 2026” | Big brands dominate |
| 8–9 | Alternative | “[Big competitor] alternatives with [feature]” | List articles |
| 10–11 | Language or market | “Best site for [job] in [language]” | Local competitors |
| 12 | Problem | “How do I fix [problem]?” | Usually names no tools |
| T1–T2 | Diagnostic | “What is [brand]?” and “What is [brand] ([domain])?” | Confusion; is your domain known? |
Models and web search
Runs go through OpenRouter with your own key, using one of three presets:
economy(the default): the same five assistants and search indexes asfull, on smaller models.full: the flagship model of each assistant, listed below.probe: ChatGPT and Perplexity only.
You can also pass family names or OpenRouter slugs in models. The full preset asks:
| Family | Model (OpenRouter) | Search |
|---|---|---|
| ChatGPT | openai/gpt-6.1-sol |
OpenAI’s own web search |
| Claude | anthropic/claude-sonnet-5.5 |
Anthropic’s own web search |
| Gemini | google/gemini-3.8-flash |
Exa search by default (see below) |
| Perplexity | perplexity/sonar |
Perplexity’s own index |
| Grok | x-ai/grok-4.7 |
xAI’s own web search |
Four families use the provider’s own search, which is closer to what a buyer sees than a third-party index. Gemini runs on Exa search by default, because Google’s Gemini API grounding terms restrict storing and analysing grounded results. Pass allow_google_grounding: true to use Google’s own grounding instead; you’re then responsible for following Google’s terms. Katman uses the first model in each family’s list that is live on OpenRouter; if a model disappears during a run, its remaining questions are skipped and the next run picks the next model in that family.
What a run costs
Visibility runs are part of Audit Week and Pro. The OpenRouter cost is separate: you pay OpenRouter directly with your own key, and Katman adds nothing to that bill. Setting up the key is covered in Getting started.
- The dry run is the default. It lists every call, the estimated cost from OpenRouter’s live prices and your key’s remaining credit. Nothing is spent.
- Your agent asks you first. Only after you agree does it call the run with
dry_run: false. - The cap holds. A run refuses to start if the estimate is above
max_cost_usd($3 by default) and stops starting new calls when it reaches the cap.
Katman’s cost model gives these figures for one run of the 14-question set (estimates; the dry run shows yours):
| Preset | Typical cost | Range |
|---|---|---|
economy (default) |
about $2.4 | $1.2–3.7 |
full |
about $4.6 | $2.2–6.9 |
probe |
about $1.8 | $0.7–2.8 |
What drives the cost is search, not the length of your question: providers bill the search results the model reads as input. In our own runs on 5 September 2026, a ChatGPT answer with native search used 30,000–61,000 input tokens and a Grok answer 57,000–69,000, while a Perplexity answer used about 60 (own measurement). If an estimate is above your cap, raise max_cost_usd, use probe, or ask fewer questions with question_ids. The samples option (1–5) repeats each question to check stability, and multiplies the cost. Manual checks in the apps have no API cost.
API answers are a proxy
A run asks the providers’ models through their APIs. That isn’t exactly what people see in the ChatGPT or Claude apps, which add their own instructions, memory and personalisation:
- In a Surfer study of 1,000 prompts, the brands named through the APIs overlapped with the apps’ answers by only about 15–32% (Surfer, updated 25 September 2026; third-party study, by a vendor that sells app-based tracking).
- In our own small comparison in September 2026, API answers and logged-out ChatGPT agreed on whether the brand was present in 3 of 3 cases, but named the same brand first in only 1 of 3 (own measurement).
Read presence as fairly reliable and rank as unreliable. For the questions that matter most, check the apps by hand. Katman never automates the apps themselves; Permissions explains why.
Manual checks in the apps
katman_manual_check (Audit Week and Pro) has no API cost and stays within the apps’ terms, because you do the asking:
- Call it without
answer_text. It returns the exact questions and instructions. - Open the app logged out, in a temporary chat or private window, with web search on. ChatGPT, Claude, Gemini, Perplexity, Copilot, Grok and Google AI Mode are supported.
- Ask one question and paste the whole answer, with the source links, into
katman_manual_check. - If a logged-out session stops you after a few questions, open a fresh one.
Why logged out: signed-in chats remember you. On 4 September 2026, a signed-in ChatGPT account placed one of our own sites second for a buyer question, while a temporary chat didn’t name it at all (own measurement). Your customers don’t share your history.
One run is a small sample
Assistants rarely give the same list twice. In SparkToro’s research, there was less than a 1 in 100 chance of ChatGPT or Google’s AI giving the same brand list in two of 100 answers, although the top brands for a question still appeared in 55–77% of the answers (SparkToro, 27 January 2026; third-party study). So:
- measure how often you’re named across the whole set, not your position in one answer;
- label a single run “low sample”;
- use
samples(up to 5) when you need to know how stable an answer is; - compare runs only when the questions, models and temperature are the same.
Also look at the official reports that count real citations and clicks: Search Console’s generative AI performance report for Google, and the AI Performance report in Bing Webmaster Tools for Copilot.
Re-run every 14 days, automatically
katman_setup_monitoring (Pro) writes a GitHub Actions workflow that runs the check, and optionally the live audit, on a schedule, using your OPENROUTER_API_KEY and KATMAN_KEY repository secrets. It commits the .katman/ changes and opens an issue when your mention rate drops.
{ "schedule": "0 6 1,15 * *", "models": "economy", "write": true }
The default schedule is weekly (Mondays, 06:00 UTC); 0 6 1,15 * * matches the book’s 14-day rhythm and halves the cost. The tool shows the estimated cost per run and per month before you write the file. What to do with the results is covered in the Katman method, and the usual reasons for a zero are on Why doesn’t ChatGPT recommend my brand?.
Frequently asked questions
How much does a visibility run cost?
About $2.4 for the default economy panel of 14 questions across five assistants (range $1.2–3.7), paid to OpenRouter with your own key and separate from what you pay for Audit Week or Pro. Katman shows the estimate before anything is spent and won’t start a run above your cap, $3 by default.
Are API answers the same as what ChatGPT users see?
No, they’re a proxy. A Surfer study updated on 25 September 2026 found only about 15–32% brand overlap between API and app answers, so confirm your most important questions in the apps with katman_manual_check.
Why does Gemini use Exa search by default?
Because Google’s Gemini API terms restrict storing and analysing grounded results. You can switch to Google’s own grounding with allow_google_grounding: true, and you’re then responsible for following those terms.
Why re-run every 14 days instead of every day?
Because answers vary from run to run and site changes take time to be crawled. Comparing the same questions and models every 14 days shows the direction without paying to measure noise.
What counts as recommended?
Katman records separately whether you were named, whether you were the first product recommended, and whether one of your URLs was cited; a small judge model reads each answer to list the recommended products in order.
Can I measure without an OpenRouter key?
Yes, by hand. katman_manual_check (Audit Week and Pro) gives you the questions and records the answers you paste from the apps, with no API cost. Without a Katman key, you can ask the same questions yourself, logged out, and keep a dated table.