Skip to main content

How Nautilus works

One fast scanner model scores every site, roughly 28,000 of them, calibrated against an AI judge panel study.

  1. You enter a URL

    Ranked sites answer instantly. Anything else is captured fresh, desktop and mobile, with the same repeatable protocol behind every Scanner score.

  2. One fast model scores every site

    A scanner model estimates all nine dimensions from the capture in about a minute. The roughly 28,000 ranked sites were scored by exactly this model on the same protocol, so a fresh scan and a ranked row are directly comparable.

  3. A judge panel study anchors the scale

    To teach and calibrate that model, a panel of eight AI judges scored 926 public websites under one frozen capture and prompt protocol. The model is trained on their per-dimension consensus, and every Scanner score is mapped onto that scale, so a 70 means what a 70 meant to the panel. No humans rate the sites.

  4. People keep it honest

    Blind side-by-side votes on the compare page check whether the ranking order matches human preference, without showing voters any scores first.

What these scores are, and are not

Nautilus is an AI-generated ranking, not an audit or a certification. The scores are model estimates, calibrated but still approximate, and a single number cannot capture everything about an experience; they are most useful for comparing a site against its peers, not as an absolute grade. Sites whose homepage answers with a block page are excluded rather than scored, and captures with thin evidence are held back. For a detailed, human-actionable review of your own site, that is what FuguUX user testing is for.

Where the sites come from

The Scanner rankings

The rankings cover roughly 28,000 public websites drawn from popular-site lists across two dozen industries. A saved scan adds a new site immediately, or replaces that site's earlier score.

Fresh scans

Any other site can be scanned on demand and gets the same nine-dimension score in about a minute. Once saved, it enters the Scanner rankings automatically.

Nine scored dimensions

All nine are scored the same way and weighted equally in the overall score.

Accessibility

Inclusive interaction, semantic structure, contrast, keyboard reach, and audit evidence.

Speed

Performance as experienced by a visitor, combining lab signals with visible latency friction.

Trust

Credibility signals, policy clarity, security cues, and confidence in the next action.

Navigation

Findability, menu clarity, path confidence, and task flow through public pages.

Content

Task clarity, readable hierarchy, useful copy, and whether the site explains itself quickly.

Mobile

Responsive layout, touch ergonomics, viewport stability, and mobile-first task completion.

First Impression

The immediate quality signal from above-the-fold composition and visual confidence.

Layout

Spatial organization, alignment, density, scan paths, and control placement.

Color

Color system quality, contrast intent, accent use, and visual coherence.

How results are read

75-100
Strong

A site that is meaningfully ahead of most sites on the leaderboard.

60-74
Good

Usable and credible, with visible issues that still leave room to pull ahead.

45-59
Average

Functional but inconsistent across several high-impact UX dimensions.

Below 45
Weak

Major friction, missing signals, or broken experience paths affect core tasks.