Can its public pages be reached and understood?
We check redirects, access rules, useful content, structure, and consistency across a small set of pages.
One point for each applicable check. Public evidence determines what counts, what passes, and what still needs to be verified.
From the first public page to the final score.
Start with the public website and declared capabilities. Must-have checks count even when evidence is missing.
Every applicable check has the same weight. More pages provide evidence, not extra points.
Divide passed checks by counted checks, then multiply by 100. Apply the strictest unmet tier gate to the overall score.
A timeout is not a failed observation. Unassessed checks stay out of the score unless marked must-have. A tier gate still requires a passing result.
Every public website is checked for reliable access and useful, understandable content. Additional capability checks normally apply when the site advertises or publishes that capability. Must-have checks and tier gates are explicit exceptions.
We check redirects, access rules, useful content, structure, and consistency across a small set of pages.
An undeclared feature is neutral unless one of its checks is marked must-have or carries a tier gate.
Each tier describes a band of final scores. Must-have checks count missing evidence as zero; unmet tier gates cap the overall score.
Scores are displayed as whole numbers, rounded down. Tier boundaries are product conventions, not guarantees of agent success.
These examples show how scope, must-have checks and tier gates affect the result.
All 20 counted checks have enough evidence to reach a conclusion, with no unmet tier gates.
18 ÷ 20 × 100. Each failed check has a concrete finding and, when supported, a fix.
Eighteen pass, one fails, and one non-mandatory check could not be verified. No tier gate applies.
18 ÷ 19 × 100, rounded down. The unassessed check is excluded from the score and listed in scan coverage.
It publishes useful pages without claiming to offer those capabilities. None of their checks is must-have or a tier gate.
Undeclared capabilities stay out of the score. Relevant ideas may appear in Go to next tier, separately.
Eighteen checks pass, one fails and one must-have check is N/A.
18 ÷ 20 × 100. The missing must-have check earns 0/1 instead of leaving the denominator.
A check is configured with maximum tier B and is failed or unassessed.
An illustrative raw score of 98 is capped at 84, the highest whole-number B score. Go to next tier puts the modules blocking tier A first.
See what each check looks for, why it matters, and which evidence supports it. Discover, Access and Use regroup the same checks; they never add extra weight.
The score describes the evidence we could inspect. It is not a probability that an agent will complete a task.