AI visibility audit

How every figure in the report is produced.

The audit’s honesty is the product. This page publishes the engines and models, the formulas, the research basis including where studies contradict each other, and exactly what happens to your data. If you haven’t run the audit yet, it’s free and takes about three minutes.

Engines, models and run counts

Three engines, each through its official API with search or grounding enabled on every call. Fifteen prompts per audit, three runs per prompt per engine — up to 135 grounded responses. Any call over 60 seconds is dropped and disclosed; an engine whose calls all fail is named as excluded in the report, never silently averaged in. Your report footer names the exact model version used on your run.

EngineAPIDefault model
ChatGPT (OpenAI)Responses API with the web_search tool required on every callgpt-4o (or the configured production model, named in your report footer)
Gemini (Google)generateContent with Google Search grounding on every callgemini-2.5-flash (named in your report footer)
Claude (Anthropic)Messages API with the web_search tool on every callclaude-sonnet-4-5 (named in your report footer)

How the prompts are built

Prompts are generated per company from what the site actually offers — never from a fixed list. The composition is fixed and validated before anything runs: 6 discovery, 3 provider selection, 3 constraint-based, 2 comparison, 1 brand-direct. The brand name appears in nothing except the single brand-direct prompt, because injecting the brand into a question guarantees a mention and turns the audit into a memory test. Every prompt is shown in the report with its demand weight.

The formulas

Language models extract structure from responses; every reported figure is then computed deterministically in code from that structure. No model ever produces a number you read in the report.

Recommendation rate

responses where the brand appears as a positive recommendation ÷ all successful responses

A response counts only when the extraction classifies the brand as recommended. A mention inside “we don’t suggest X” or a not-stocked list never counts. Failed or timed-out engine calls are excluded from the denominator, never counted as absence.

The ± range

(highest per-run rate − lowest per-run rate) ÷ 2, across the 3 runs

Every prompt runs three times per engine because AI answers vary. The ± range is the observed spread, not a confidence interval — we show variance rather than pretending it away.

Demand weight per prompt

monthly search volume of the prompt's topic ÷ total volume across all 15 prompts

Volumes come from third-party Google search volume data (UK). Topics with no vendor data take the smallest known volume as a floor, so every prompt carries a stated, non-zero weight. Weights are shares of measured demand — never a revenue or pound figure.

Demand-weighted absence

sum of demand weights of prompts where no engine recommended the brand in any run

“You are absent from prompts carrying 60% of the measured demand” weighs absence by how many people actually ask each question.

Source ranking

count of citations per domain across successful responses to category prompts (brand-direct prompt excluded)

Every citation URL is reduced to its domain. The brand-direct prompt is excluded from this ranking: it sends engines to the company's own site, which would hand own-domain a free top rank. Presence in each top source is then resolved for you and up to three competitors, with exactly three states: present, absent, undetermined. Undetermined is printed as undetermined — never guessed.

Accuracy verdicts

numeric claims: |claimed − published| ≤ 2% of the larger value → correct

The language model only locates the published value on your pages; the verdict on numbers is computed in code with a 2% tolerance, because price and date drift below that is noise, not hallucination. Below 8 checkable claims, claims are shown individually and the summary count is withheld.

Why there is no score

Your visibility is different on each engine, different across prompt types, and varies between runs. A composite score would average away exactly the detail that tells you what to fix. The report gives you rates with denominators, a ± range, and findings each carrying their own evidence — everything needed to check our work.

Where the research disagrees with itself

The evidence base behind AI visibility is young and partly contradictory. Rather than citing only the studies that flatter the audit, here are the contradictions we know about and how the audit handles each.

Structured data lifts citations by ~30%

Google and Microsoft state LLMs do not directly consume schema at generation time. The lift is real in correlational studies but the mechanism is debated: schema likely helps the search layer that feeds the engines, not the generation itself. We report schema completeness as a fact, not as a guaranteed lever.

llms.txt is an AEO best practice

About 97% of llms.txt files received zero crawler requests in a 300,000-domain study, and no major engine has committed to reading it. The audit reports its presence as information and will never flag its absence as a defect.

Backlinks drive AI visibility

Ahrefs measured the correlation between backlinks and AI mentions at ~0.218 across 75,000 brands — among the weakest factors. Branded web mentions (0.664) and YouTube presence (0.737) measure far stronger. That is why backlinks and domain authority do not appear anywhere in this audit.

One AI visibility score can summarise a brand

Visibility differs materially by engine, by prompt type and by run. A single score hides exactly the information that makes the audit actionable, so this audit publishes rates with denominators and a ± range instead.

What happens to your data

  • The audit runs in a single job. Raw engine responses, crawler tests and lookups exist only inside that job and are discarded when it completes.
  • The finished report travels to your browser as an encrypted token and unlocks only after you verify your work email. It is not stored server-side — the downloadable HTML file is the permanent copy.
  • What we keep: your email address and a short diagnostic summary (recommendation rate, competitor rank, source-map counts, top findings) in our CRM, plus an encrypted run record readable only by us.
  • Spend controls: every run is metered against a monthly cap, and repeated audits of the same domain are limited over a 30-day window using a one-way hash of the domain.

Data handling is covered in full in our Privacy Policy.

Run the audit

Free, about three minutes, and everything above applies to your report.

Get your free AI visibility audit