Public verification
Leaderboard
30 curated sites scored by the same deterministic 40-check engine that scores designesy.org. No LLM, no paywall, no pay-to-remove.
Scores reflect what the engine measures on the live fetched surface — token architecture, motion hygiene, accessibility primitives, typography discipline. A site can look world-class and still score low if it doesn’t ship the contract primitives at :root. That is the point.
Submit a site
Enter a URL to score it against the same 40-check engine. Submissions are scored instantly and curated into the seed list on the next weekly batch. No paywall, no pay-to-remove.
Cohort snapshot
Policy. Curated seed (30 sites) + open submission. Scores are deterministic — 40 checks, no LLM. Sites scoring below 50 are flagged "needs work", not hidden. No paywall, no pay-to-remove. Scores re-run weekly. The seed list is curated across five tiers (reference, competitors, design-system exemplars, inspiration, high-traffic). Open submission is a follow-up — for now, mail hello@designesy.org.
Score distribution
How the 30 scored sites distribute across grade bands. The histogram shows the shape of the cohort — not a bell curve.
One A-grade site in a cohort of 30. The contract is demanding — most sites land in D or F because they don’t ship the primitives (token systems, reduced-motion blocks, font-synthesis rules) at :root. See the methodology page for what each check measures and why.
Ranking
Ranked by total score. The top site is the only A-grade site in the cohort. Select any row to re-score it live at /score.
| # | Site | Grade | Score | Checks | Actions |
|---|---|---|---|---|---|
| 01 | A | 100.0%↑ 0.8 | 36p · 0f · 0w · 1s | ||
| 02 | C | 77.7%↑ 11.4 | 15p · 4f · 16w · 2s | ||
| 03 | C | 77.6%• | 26p · 5f · 5w · 1s | ||
| 04 | C | 73.8%↑ 5.0 | 16p · 4f · 15w · 2s | ||
| 05 | C | 73.6%↑ 7.2 | 20p · 6f · 10w · 1s | ||
| 06 | C | 73.1%↑ 3.0 | 14p · 2f · 18w · 3s | ||
| 07 | C | 70.2%↑ 5.1 | 14p · 3f · 19w · 1s | ||
| 08 | D | 68.6%↑ 1.4 | 18p · 5f · 13w · 1s | ||
| 09 | D | 68.5%↑ 3.0 | 14p · 5f · 16w · 2s | ||
| 10 | D | 68.1%↓ 5.0 | 17p · 4f · 14w · 2s | ||
| 11 | D | 67.8%↑ 3.0 | 10p · 1f · 23w · 3s | ||
| 12 | D | 66.0%↑ 5.1 | 11p · 4f · 20w · 2s | ||
| 13 | D | 66.0%↓ 9.0 | 16p · 2f · 17w · 2s | ||
| 14 | D | 64.1%↑ 5.1 | 12p · 5f · 18w · 2s | ||
| 15 | D | 62.6%↑ 1.2 | 13p · 5f · 17w · 2s | ||
| 16 | D | 62.2%↑ 5.1 | 10p · 4f · 20w · 3s | ||
| 17 | D | 61.1%↓ 7.7 | 18p · 3f · 14w · 2s | ||
| 18 | D | 60.7%↓ 2.6 | 15p · 5f · 15w · 2s | ||
| 19 | F | 59.3%↓ 7.8 | 16p · 3f · 16w · 2s | ||
| 20 | F | 59.0%↑ 3.0 | 6p · 4f · 22w · 5s | ||
| 21 | F | 58.1%↓ 11.9 | 21p · 6f · 8w · 2s | ||
| 22 | F | 52.6%• | 6p · 3f · 23w · 5s | ||
| 23 | F | 52.5%↓ 4.0 | 7p · 3f · 23w · 4s | ||
| 24 | F | 51.0%• | 7p · 4f · 23w · 3s | ||
| 25 | F | 50.8%↓ 9.8 | 14p · 6f · 15w · 2s | ||
| 26 | F | 50.3%↓ 4.0 | 7p · 3f · 21w · 6s | ||
| 27 | F | 49.4%↓ 4.0 | 7p · 4f · 21w · 5s | ||
| 28 | F | 41.1%↓ 12.0 | 8p · 5f · 21w · 3s | ||
| 29 | F | 41.1%↓ 16.0 | 11p · 5f · 19w · 2s | ||
| 30 | F | 37.9%↓ 6.9 | 6p · 5f · 24w · 2s |
How to read this
What the engine measures
40 deterministic checks across 14 weighted categories — token architecture, motion hygiene, accessibility primitives, typography discipline, reduced-motion handling, AI disclosure, forced-colors readiness, UX copywriting, Unicode security. Plus 12 anti-slop rules (up to -20pts) and 7 originality signals (up to +8pts) so taste is part of the number. Not an LLM impression. Not a roast. The same engine scores designesy.org itself, in public, at /score?url=designesy.org.
The engine measures what is shipped, not what is documented. A design-system site can publish a rich token taxonomy in storybook and still score low if the marketing surface doesn’t expose those tokens at :root. That gap — between documented and shipped — is exactly what the leaderboard surfaces. For the full scoring methodology — every check, its category weight, the scoring math, and the accessibility floor — see the methodology page.