Evidence-informed scoring. No black box. No fake precision.
We don't claim to know Google's algorithm weights. Nobody outside Google does, and anyone who says otherwise is selling something.
We measure two things:
The overall score combines both, weighted toward traditional SEO because it still drives the majority of search traffic:
13 categories, weighted by evidence strength
| Category | Weight | What We Check | Evidence |
|---|---|---|---|
| On-page Content Quality | 22% | Word count, paragraphs, readability, lists, emphasis | Page-level content assessment only. Does not include backlinks, authority, or external signals (require external APIs). |
| Meta & Social Tags | 12% | Title, description, canonical, viewport, OG, Twitter | Title affects CTR. Keyword-in-title: near-zero ranking correlation (Backlinko 11.8M study) |
| Technical | 12% | HTTPS, response time, compression, mobile, JS rendering | Google-confirmed prerequisites |
| Headings Structure | 10% | H1, H2, H3 hierarchy, heading depth | Content understanding signal |
| Links | 10% | Internal links, anchor text, link count | Discovery + PageRank flow. Backlinks (r=0.38) require external API |
| Robots & Crawlability | 8% | robots.txt, sitemap, crawl directives | Indexability = prerequisite. Blocking = kills SEO |
| Images | 6% | Alt text, image count, dimensions | Accessibility + understanding, not primary ranking factor |
| Performance (CWV) | 5% | PageSpeed (when API available) | "Little/no correlation" — Backlinko. "Not guaranteed" — Google |
| Security & Trust Headers | 4% | HTTPS / TLS certificate, HSTS, X-Frame-Options, X-Content-Type-Options, CSP | Page-level security signals. No email-level checks (SPF/DMARC are not page security). |
| URL & Keywords | 4% | URL length, structure | Near-zero ranking correlation (Backlinko) |
| HTML Validation | 3% | W3C validation | Clean rendering, not direct ranking factor |
| Schema / Structured Data | 4% | JSON-LD presence, types | Rich result eligibility + indirect CTR. NOT a ranking factor (Ahrefs controlled study 2026, Lighthouse: unscored) |
8 categories measuring diagnostic citation-readiness. These are extraction and formatting signals, not ranking factors.
| Category | Weight | What We Check | Research Source |
|---|---|---|---|
| Brand & Entity Clarity | 16% | On-page brand consistency, sameAs, schema, social links | On-page entity signals. Off-site brand authority requires external APIs |
| Content Structure & Citability | 16% | Headings, semantic HTML, passages, lists | U.Tokyo Mar 2026: +17.3% citation rate (A-) |
| Evidence, Data & Citations | 15% | Statistics, external links, quotation patterns | Princeton KDD 2024: quotations +41%, stats +32% (A) |
| Trust & Authorship Signals | 12% | Author, credentials, reviews, dates | Google quality guidelines. Machine-readable trust proxies (author schema, credentials, linked profiles), not subjective quality |
| Answer-First (BLUF) | 11% | First 100 words, FAQ, definitions, summary | Clear answer-first sections improve extractability. AirOps: only 15% of retrieved pages get cited |
| Content Freshness | 9% | Dates, copyright, Last-Modified, blog | Ahrefs: AI-cited content 25.7% fresher (B+) |
| Technical AI Access | 14% | AI bot access, noindex, JS rendering, schema | Blocking key AI crawlers may prevent discovery. Access does not guarantee citation. Gate: all blocked = score 0 |
| Competitive Context | 7% | Positioning, pricing, process, case studies | Product heuristic (low confidence). No direct research — included for decision-oriented queries |
Hard caps that prevent meaningless high scores on broken pages
| Gate | Condition | Effect |
|---|---|---|
| Technical AI Access gate | All key AI bots blocked in robots.txt (ChatGPT-User, ClaudeBot, PerplexityBot, GPTBot, Google-Extended) | Technical AI Access = 0; Overall GEO capped at 40 |
| Content quality gate | <50 words | GEO max 15 |
| <100 words | GEO max 30 | |
| <200 words | GEO max 45 | |
| <300 words | GEO max 60 | |
| SPA gate | Heavy JS (script >5x text) + 0 headings | GEO capped at 50 |
| llms.txt | llms.txt file present | Detected & reported, not scored (no proven impact) |
Additional checks when scanning multiple pages. These detect cross-page issues that single-page audits cannot find.
| Check | What It Detects | Penalty |
|---|---|---|
| Near-Duplicate Content | Pages with 75%+ identical body text (e.g., location pages with only city name swapped). Uses SimHash fingerprinting on main content, excluding navigation and footer. | 3–5 pages: −10 • 6–15: −20 • 16+: −30 |
| Content Uniqueness Ratio | Percentage of pages with genuinely unique content. Low uniqueness (<50%) signals template-heavy sites that Google may deindex. | Reported as metric (included in template penalty above) |
| Index Bloat | Too many pages for your business type. A local painter with 137 pages dilutes authority. Thresholds are calibrated per industry. | Warning: −5 • Critical: −10 |
| Internal Broken Links | Links within your site that lead to 404 pages. These waste crawl budget and create dead ends for users and search engines. | −3 per link (max −15, combined with external) |
| Canonical Conflicts | Multiple pages pointing to the same canonical URL, or canonical targets that return 404. Confuses Google about which page to index. | −2 per issue (max −10) |
| Duplicate Titles & Descriptions | Pages sharing identical or near-identical (>85% similar) title tags or meta descriptions. | Titles: −2/group • Descriptions: −1/group |
| Duplicate FAQ Questions | Identical FAQ schema questions appearing on 3+ pages. Google may ignore FAQ rich results from pages with duplicate questions. Each page should have unique, topic-specific FAQs. | −1 per question on 3+ pages (max −5) |
| Orphan Pages | Pages with zero internal links pointing to them. Search engines may not discover or prioritize these pages. | −2 per orphan (max −8) |
Honest limitations of a free, crawl-based audit tool.
Claims we investigated and found unsupported or overstated.