The Scanner rankings
The rankings cover roughly 28,000 public websites drawn from popular-site lists across two dozen industries. A saved scan adds a new site immediately, or replaces that site's earlier score.
One fast scanner model scores every site, roughly 28,000 of them, calibrated against an AI judge panel study.
Ranked sites answer instantly. Anything else is captured fresh, desktop and mobile, with the same repeatable protocol behind every Scanner score.
A scanner model estimates all nine dimensions from the capture in about a minute. The roughly 28,000 ranked sites were scored by exactly this model on the same protocol, so a fresh scan and a ranked row are directly comparable.
To teach and calibrate that model, a panel of eight AI judges scored 926 public websites under one frozen capture and prompt protocol. The model is trained on their per-dimension consensus, and every Scanner score is mapped onto that scale, so a 70 means what a 70 meant to the panel. No humans rate the sites.
Blind side-by-side votes on the compare page check whether the ranking order matches human preference, without showing voters any scores first.
Nautilus is an AI-generated ranking, not an audit or a certification. The scores are model estimates, calibrated but still approximate, and a single number cannot capture everything about an experience; they are most useful for comparing a site against its peers, not as an absolute grade. Sites whose homepage answers with a block page are excluded rather than scored, and captures with thin evidence are held back. For a detailed, human-actionable review of your own site, that is what FuguUX user testing is for.
The rankings cover roughly 28,000 public websites drawn from popular-site lists across two dozen industries. A saved scan adds a new site immediately, or replaces that site's earlier score.
Any other site can be scanned on demand and gets the same nine-dimension score in about a minute. Once saved, it enters the Scanner rankings automatically.
All nine are scored the same way and weighted equally in the overall score.
Inclusive interaction, semantic structure, contrast, keyboard reach, and audit evidence.
Performance as experienced by a visitor, combining lab signals with visible latency friction.
Credibility signals, policy clarity, security cues, and confidence in the next action.
Findability, menu clarity, path confidence, and task flow through public pages.
Task clarity, readable hierarchy, useful copy, and whether the site explains itself quickly.
Responsive layout, touch ergonomics, viewport stability, and mobile-first task completion.
The immediate quality signal from above-the-fold composition and visual confidence.
Spatial organization, alignment, density, scan paths, and control placement.
Color system quality, contrast intent, accent use, and visual coherence.
A site that is meaningfully ahead of most sites on the leaderboard.
Usable and credible, with visible issues that still leave room to pull ahead.
Functional but inconsistent across several high-impact UX dimensions.
Major friction, missing signals, or broken experience paths affect core tasks.