How AI Site View measures a website
Every number on your report comes from somewhere, and this page is that somewhere. What we test, what each check is worth, what counts as evidence, and where the method stops. If a result ever disagrees with this page, the page is the thing that is wrong.
Version and changelog
A measurement that changes silently is not a measurement. Observed results carry their scan time in UTC and this version number, so a result can always be traced back to the rules that produced it. Your report reads, for example, Observed 31 Aug 2026, 14:32 UTC · Methodology v1.1.
Scores produced under different methodology versions are not directly comparable. Every stored scan keeps the version that produced it, so history stays honest rather than tidy.
Changelog
The two-component model. Technical Access 60 and Machine Readability 40 are now the only things that produce the headline number. Model-training crawlers are still probed and reported, but no longer scored. The softer signals became diagnostic indicators that explain the result instead of quietly moving it.
The original five-dimension model, retroactively numbered v1.0 when v1.1 replaced it.
The AI Site View Score, 0 to 100
Two components. Nothing else contributes to the headline number, and no component is estimated to make the panel look complete.
Technical Access
60 points · all testedCan machines retrieve the important server-delivered content of this website? Only technically observable access and retrievability conditions count here.
Machine Readability
40 points · extractedCan an automated system reliably establish who you are, what you offer, who you serve and where you operate from the information your website exposes? Every point traces to a named, deterministic evidence source.
Structured data: valid JSON-LD carrying a business-relevant type scores 12. Page-level types only, such as WebSite or FAQPage, or structured data that is present but invalid, scores 6. None scores 0.
Identity establishability is the sum of three established facts: name (clear 6, partial 3, unknown 0) plus description (clear 4, partial 2, unknown 0) plus business type (clear 2, partial 1, unknown 0).
Your report shows both components natively, never rescaled into each other: 52 of 60 plus 31 of 40 is 83 of 100. If you can't add the parts up to the total on your own report, the report is lying to you.
When a probe does not answer
A crawler probe that times out twice reports "could not be determined during this scan". It leaves the denominator entirely and the remaining scored weights scale back up. A timeout is never counted as a block, because a slow network is not a policy decision. Crawlers blocked by robots.txt or denied at the server score zero of their weight, because that is a real barrier.
Two gates that sit outside the arithmetic
- No score at all. If a control request cannot reach the site, or the site blocks our prober outright, there is no score. A failed measurement is reported as a failed measurement, never as a low number.
- Capped at 30. A homepage carrying noindex is capped at 30. The cap is applied after the two components are composed, never folded inside one of them, so you can still see which part of the site is healthy underneath the barrier.
Grade letters A to F are bands over the same number and are defined once, in the methodology module the whole product reads from.
Every result is a point-in-time observation. It describes what happened when we asked, from where we asked, on the date stamped on your report.
The four diagnostic states
The state leads your report, because "your crawlers are blocked" is more useful than "you scored 41". The number stays; it just stops being the first thing you read.
| State | What it means | When it applies |
|---|---|---|
| Critical | We could not evaluate the site at all, so we give you no score. | The control request was unreachable, or the site blocked our prober. |
| Blocked | A major machine-access barrier was observed, and it is named before anything else. | The homepage carries noindex, or all six scored crawlers are blocked, or the blocked scored crawlers carry at least half of the 42 point crawler pool between them. |
| Limited | Machines can get in, but there is not enough there for them to establish your business. | Machine Readability below 20 of 40, or the server-delivered content check failed. |
| Ready | No barrier observed and your business is establishable from what the server returns. | None of the above applies. |
The crawler matrix
Nine crawler identities are examined on every scan. Six of them decide part of your score. The other three are reported because you should know about them, and scored at nothing because they answer a different question.
| Crawler | Role | How we test it | Scored | Weight |
|---|---|---|---|---|
| OAI-SearchBot OpenAI |
ChatGPT Search retrieval and indexing | Live probe, published user agent | Yes | 9 |
| PerplexityBot Perplexity |
Perplexity index | Live probe, published user agent | Yes | 8 |
| Googlebot |
Google Search, and the grounding behind AI Overviews | Live probe, published user agent | Yes | 8 |
| Claude-SearchBot Anthropic |
Claude search indexing | Live probe, published user agent | Yes | 7 |
| Bingbot Microsoft |
Bing, and the grounding behind Copilot | Live probe, published user agent | Yes | 6 |
| ChatGPT-User OpenAI |
The live fetch a person triggers by asking about you | Live probe, published user agent | Yes | 4 |
| GPTBot OpenAI |
Model-training data collection | Live probe, published user agent | Informational | 0 |
| ClaudeBot Anthropic |
Model-training data collection | Live probe, published user agent | Informational | 0 |
| Google-Extended |
Robots control token for Gemini training and grounding preference | robots.txt only, never probed | Informational | 0 |
The six scored weights sum to 42, which is the crawler share of Technical Access.
Why training crawlers do not move your score
GPTBot, ClaudeBot and Google-Extended are presented together as a training-access preference, with each one marked allowed, blocked or unknown.
A publisher who blocks model training and welcomes search retrieval has made a coherent decision, and a tool that marks them down for it is scoring its own opinion.
Why Google-Extended is never probed
Google-Extended is a robots control token, not a crawler identity we can meaningfully impersonate through an HTTP probe. We evaluate it through robots.txt policy only.
What a probe can and cannot prove
We never report a successful probe as proof that a crawler is able to reach you. A probe result is evidence about a probe, not a guarantee about the genuine crawler's full production behaviour. So every crawler result resolves into one of four honest verdicts.
Only one of those four is definitive about policy: a robots.txt disallow is a statement you published, so we can report it as fact. The rest are observations, and they are worded as observations.
Every crawler result opens an evidence drawer carrying the probe identity, the HTTP status code, the robots.txt decision with the matched rule where one applies, the test origin, the observed timestamp, and this limitation:
Server-delivered content
We check whether important content exists in the HTML returned directly by the server, without requiring browser execution. This is a durable machine-readable baseline regardless of how individual crawlers render JavaScript.
That framing is deliberate. Rendering behaviour varies between crawlers and changes over time, and many automated crawlers rely heavily on server-delivered HTML. We measure the thing that stays true either way, rather than a claim about any particular crawler's JavaScript engine on any particular day.
Why this decides how we build fixes
It is the reason there will never be an AI Site View "magic script tag" that injects your structured data in the browser. Anything a fix adds after page load is invisible to anything that reads only what the server sent, and a fix that cannot be seen by the thing it was meant to convince is not a fix. Our WordPress plugin outputs your schema server-side, in wp_head, for exactly this reason.
How every finding is labelled
Four labels, and the label is part of the finding rather than a disclaimer at the bottom of the page. You should always be able to see how hard the evidence behind a statement actually is.
The UNKNOWN discipline
Missing evidence is never converted into an assertion about your business. This is the single rule that most separates this report from a confident-sounding one, and it holds for every fact and every indicator.
UNKNOWN. No explicit pricing was established in the homepage HTML, structured data or linked evidence inspected by this scan.
This company doesn't publish pricing.
The difference is not politeness. The first sentence is about our scan and is true. The second is about your company and might be completely false, because the pricing could be on a page we did not inspect.
What counts as established
Eleven business facts are read from what the server returned. Each one lands on clear, partial or unknown, and the qualifying evidence for each is fixed in advance rather than decided in the moment. The extraction is fully deterministic: there is no language model anywhere in this path, because a language model will always produce a confident sentence, including about an empty page.
Open any fact below for the acceptance criteria. The criteria are declared in api/_methodology.js, implemented by the extractor in api/_profile.js, and pinned by a fixture suite across fourteen site archetypes so a loosened rule breaks a test rather than quietly starts lying.
Business name
Identity
Business type
Identity
Description
Identity
Products and services
Offering
Audience
Who it is for
Location
Where you operate
Pricing
What it costs
People and leadership
Who is behind it
Trust and proof
Why anyone should believe you
Contact
How to reach you
Policies
Privacy and terms
The ten diagnostic indicators
Each indicator is derived deterministically from the facts above, lands on clear, partial or unknown, and carries the same provenance and evidence that facts do. Cross-source consistency can additionally land on conflict.
- Business identity completeness. Name, type and description considered together.
- Offering clarity. Whether specific products or services are established.
- Audience clarity. Whether who you serve is established.
- Location clarity. Whether where you operate is established.
- Pricing clarity. Whether prices connect to specific things you sell.
- People and leadership clarity. Whether anyone is named.
- Trust and proof clarity. Whether claims are checkable.
- Contact clarity. Whether a machine-readable way to reach you exists.
- Policy clarity. Whether policy documents are findable.
- Cross-source consistency. Declared names compared across sources. A conflict leads your blind spots, because two different names for one business is the fastest way to become two half-businesses to a machine.
llms.txt is detected, and never scored
If your site publishes an llms.txt, we find it and report it. It earns you nothing, and that is on purpose. There is no measured evidence today that its presence changes how machines read a website, so awarding points for it would be scoring a fashion and inviting you to spend an afternoon on a file that may do nothing.
Being told what does not count is part of what you are here for. A checker that quietly rewarded every new convention would look generous and be useless.
What AI Site View will not pretend to know
The limits are not a disclaimer. They are the product. Every one of these is something a competitor could show you as a confident number today.
- We do not claim to know what ChatGPT, Claude or Gemini privately "thinks" about your company.
- We do not claim that allowing a crawler guarantees inclusion or ranking.
- We do not convert UNKNOWN evidence into confident conclusions.
- We do not award points for technical fashions without sufficient justification. llms.txt is detected and reported, never scored.
- We do not sell fake AI authority percentages. External authority measurement is not included in this scan.
- We do not call a browser-injected fix successful when crawlers cannot see it.
- We distinguish policy permission, which is what robots.txt states, from observed HTTP behaviour.
- Every important finding is traceable to evidence.
One line covers the rest of it: AI Site View reports what it can prove, labels what it infers, and says UNKNOWN when the evidence stops.
Free, no signup. Everything on this page, applied to your homepage, with the evidence for every line of it.