Back to geoseo.today

Methodology
How GeoSEO.today Scores Work

Evidence-informed scoring. No black box. No fake precision.

Our Approach

We don't claim to know Google's algorithm weights. Nobody outside Google does, and anyone who says otherwise is selling something.

We measure two things:

  • SEO score — "Can Google crawl, index, and understand your page?"
  • GEO score — "How citation-ready is your content for AI search?"

The overall score combines both, weighted toward traditional SEO because it still drives the majority of search traffic:

Overall = SEO × 0.7  +  GEO × 0.3
This is a product-level diagnostic heuristic, not a claim about Google, OpenAI, or Perplexity algorithms. The 70/30 split reflects that traditional search indexation often acts as an important retrieval gateway for search-integrated AI systems — one 2026 AIO study found ~70% of cited domains overlapped with organic top-10, while ~30% did not.

SEO Scoring

13 categories, weighted by evidence strength

Category weights are diagnostic product weights, not Google ranking factor weights. They should not be interpreted as a ranking model. Because this audit cannot access backlink data or field Core Web Vitals, the SEO score measures on-page and technical readiness — not full ranking potential.
Category Weight What We Check Evidence
On-page Content Quality 22% Word count, paragraphs, readability, lists, emphasis Page-level content assessment only. Does not include backlinks, authority, or external signals (require external APIs).
Meta & Social Tags 12% Title, description, canonical, viewport, OG, Twitter Title affects CTR. Keyword-in-title: near-zero ranking correlation (Backlinko 11.8M study)
Technical 12% HTTPS, response time, compression, mobile, JS rendering Google-confirmed prerequisites
Headings Structure 10% H1, H2, H3 hierarchy, heading depth Content understanding signal
Links 10% Internal links, anchor text, link count Discovery + PageRank flow. Backlinks (r=0.38) require external API
Robots & Crawlability 8% robots.txt, sitemap, crawl directives Indexability = prerequisite. Blocking = kills SEO
Images 6% Alt text, image count, dimensions Accessibility + understanding, not primary ranking factor
Performance (CWV) 5% PageSpeed (when API available) "Little/no correlation" — Backlinko. "Not guaranteed" — Google
Security & Trust Headers 4% HTTPS / TLS certificate, HSTS, X-Frame-Options, X-Content-Type-Options, CSP Page-level security signals. No email-level checks (SPF/DMARC are not page security).
URL & Keywords 4% URL length, structure Near-zero ranking correlation (Backlinko)
HTML Validation 3% W3C validation Clean rendering, not direct ranking factor
Schema / Structured Data 4% JSON-LD presence, types Rich result eligibility + indirect CTR. NOT a ranking factor (Ahrefs controlled study 2026, Lighthouse: unscored)
No credible source provides exact Google ranking factor weights. These weights reflect relative importance based on available evidence.

GEO Scoring

8 categories measuring diagnostic citation-readiness. These are extraction and formatting signals, not ranking factors.

Category Weight What We Check Research Source
Brand & Entity Clarity 16% On-page brand consistency, sameAs, schema, social links On-page entity signals. Off-site brand authority requires external APIs
Content Structure & Citability 16% Headings, semantic HTML, passages, lists U.Tokyo Mar 2026: +17.3% citation rate (A-)
Evidence, Data & Citations 15% Statistics, external links, quotation patterns Princeton KDD 2024: quotations +41%, stats +32% (A)
Trust & Authorship Signals 12% Author, credentials, reviews, dates Google quality guidelines. Machine-readable trust proxies (author schema, credentials, linked profiles), not subjective quality
Answer-First (BLUF) 11% First 100 words, FAQ, definitions, summary Clear answer-first sections improve extractability. AirOps: only 15% of retrieved pages get cited
Content Freshness 9% Dates, copyright, Last-Modified, blog Ahrefs: AI-cited content 25.7% fresher (B+)
Technical AI Access 14% AI bot access, noindex, JS rendering, schema Blocking key AI crawlers may prevent discovery. Access does not guarantee citation. Gate: all blocked = score 0
Competitive Context 7% Positioning, pricing, process, case studies Product heuristic (low confidence). No direct research — included for decision-oriented queries
Reported uplift figures (e.g. +41% quotations, +17.3% structure) come from specific studies and benchmarks. They inform our diagnostic checks but do not promise equivalent gains for any page. Actual results depend on content, competition, and platform behavior.

Score Gates & Adjustments

Hard caps that prevent meaningless high scores on broken pages

Gate Condition Effect
Technical AI Access gate All key AI bots blocked in robots.txt (ChatGPT-User, ClaudeBot, PerplexityBot, GPTBot, Google-Extended) Technical AI Access = 0; Overall GEO capped at 40
Content quality gate <50 words GEO max 15
<100 words GEO max 30
<200 words GEO max 45
<300 words GEO max 60
SPA gate Heavy JS (script >5x text) + 0 headings GEO capped at 50
llms.txt llms.txt file present Detected & reported, not scored (no proven impact)
Gates apply after the weighted score calculation. A page blocking all AI crawlers cannot have direct access (Technical AI Access = 0) but may still be cited through cached versions or third-party references. Therefore, the GEO cap at 40 reflects structural limitation, not impossibility.

Site-Wide Analysis

Additional checks when scanning multiple pages. These detect cross-page issues that single-page audits cannot find.

Check What It Detects Penalty
Near-Duplicate Content Pages with 75%+ identical body text (e.g., location pages with only city name swapped). Uses SimHash fingerprinting on main content, excluding navigation and footer. 3–5 pages: −10 • 6–15: −20 • 16+: −30
Content Uniqueness Ratio Percentage of pages with genuinely unique content. Low uniqueness (<50%) signals template-heavy sites that Google may deindex. Reported as metric (included in template penalty above)
Index Bloat Too many pages for your business type. A local painter with 137 pages dilutes authority. Thresholds are calibrated per industry. Warning: −5 • Critical: −10
Internal Broken Links Links within your site that lead to 404 pages. These waste crawl budget and create dead ends for users and search engines. −3 per link (max −15, combined with external)
Canonical Conflicts Multiple pages pointing to the same canonical URL, or canonical targets that return 404. Confuses Google about which page to index. −2 per issue (max −10)
Duplicate Titles & Descriptions Pages sharing identical or near-identical (>85% similar) title tags or meta descriptions. Titles: −2/group • Descriptions: −1/group
Duplicate FAQ Questions Identical FAQ schema questions appearing on 3+ pages. Google may ignore FAQ rich results from pages with duplicate questions. Each page should have unique, topic-specific FAQs. −1 per question on 3+ pages (max −5)
Orphan Pages Pages with zero internal links pointing to them. Search engines may not discover or prioritize these pages. −2 per orphan (max −8)
Why this matters: A site can score 95+ on every individual page but still lose all organic traffic if Google detects mass template content. The site-wide analysis catches exactly these cross-page issues that per-page audits miss.

What We CAN'T Check

Honest limitations of a free, crawl-based audit tool.

Backlinks — the strongest known SEO correlation (r=0.38). Requires Moz or Ahrefs API access.
Real Core Web Vitals field data — requires Chrome UX Report (CrUX) API for actual user metrics.
Actual AI citation rate — would require querying ChatGPT, Perplexity, and Claude for your brand.
Brand mentions across the web — requires a web-scale index we don't have.
User engagement metrics — click-through rates, dwell time, and pogo-sticking are invisible to external tools.
We measure readiness, not actual visibility. A high score means your site is well-prepared for search engines and AI systems to find, understand, and cite — but it doesn't guarantee rankings or citations.

Research Sources

View full Evidence Library

Academic Papers

Industry Data

Experiments & Official

  • INDEPENDENT EXPERIMENT
    OtterlyAI / BrightonSEO (Apr 2026) — llms.txt = 0.1% AI bot traffic, 3x worse than HTML
  • INDEPENDENT EXPERIMENT
    searchVIU Schema Parser Audit — AI engines ignore JSON-LD during real-time fetch
  • VENDOR REPORT
    AirOps (Mar 2026) — ChatGPT cites only 15% of retrieved pages. Limitation: ChatGPT-specific
  • OFFICIAL GUIDANCE
    Google — "No special optimization needed for AI features"
  • OFFICIAL GUIDANCE
    Google — "CWV does not guarantee high rankings"

Debunked Myths

Claims we investigated and found unsupported or overstated.

llms.txt boosts AI visibility
OtterlyAI / BrightonSEO found llms.txt accounts for just 0.1% of AI bot traffic, and markdown versions generated 0 visits. We detect and report llms.txt but do not score it.
Schema markup directly boosts Google rankings
Schema (structured data) helps Google understand your content and qualifies you for rich results. But it is not a direct ranking boost. Content quality, relevance, and authority matter more. Lighthouse doesn't score structured data; it only checks for syntax errors.
Adding schema markup boosts AI citations
Ahrefs' controlled study (1,885 vs 4,000 pages) found no significant AI citation uplift after adding schema. This is because AI systems fetch pages over HTTP(S) and parse HTML/text. JSON-LD structured data is not accessed during real-time fetch — only during web indexing.
"3.2x freshness in 30 days"
Circular blog citations. The original source for this specific number doesn't exist. Real data from Ahrefs shows AI-cited content is 25.7% fresher — meaningful, but not 3.2x.
"4.1x data tables boost citations"
Unsourced claim originating from quolity.ai. No peer-reviewed or reproducible study supports this specific multiplier.
FirstPageSage ranking factor weights
No published methodology. Based on 116 websites studied in 2016. Widely cited but not credible as a source for current Google ranking weights.