All articles
Article··10 min read

Most GEO advice is a checklist with no numbers behind it. When a major UK insurer asked how AI-ready their knowledge centre was, I built a 10-criterion weighted rubric, checked every weight against published citation research, double-coded the judgement calls and stress-tested the weights. The average article scored 45.6 out of 100, 43% of sections buried the answer, and question headings made it worse. Here is the rubric, the method, the eight habits and what a fix looks like.

Most GEO advice you’ll read this year is a checklist with no numbers behind it. “Use question headings.” “Answer first.” “Add stats.” All sensible, and almost none of it measured.

So when a major UK insurer asked me how AI-ready their home insurance knowledge centre was, I didn’t want to hand over a list of opinions. I wanted a score they could trust, a way to see exactly where the points were being lost, and proof the findings would hold up if someone challenged them.

The result is a 10-criterion weighted GEO rubric, with each weight checked against published citation research and the scoring tested with the same methods a researcher would use: double-coding, inter-rater agreement, significance tests and sensitivity analysis. I’ve since turned it into a repeatable process I can run on any content library. Here’s how it works and what it found.

The rubric: 10 criteria, 100 points

The rubric asks four questions of every page. Does it answer the question? Can an AI lift a passage out cleanly? Is there evidence behind the claims? Is it current?

Each criterion gets a floor (where it scores zero) and a target (where it hits full marks), with a straight line in between. That matters more than it sounds. A page with 79% answer-first sections and one with 31% shouldn’t both be a “partial pass”, and a linear scale shows the difference.

The first version of the weights was my judgement. Before publishing, I checked each one against the biggest citation studies available: Cyrus Shepard’s Zyppy meta-analysis, which scores 23 citation factors out of 10 based on 54 studies, Ahrefs’ analysis of nearly 17 million AI citations, and the Princeton GEO paper (Aggarwal et al., KDD 2024). Where the evidence was strong, a criterion gained weight. Where it was thin, it lost some.

CriterionWeight (v1 → v2)How it’s scoredEvidence behind the weight
Answer-first sections
Answerability
20 → 20% of sections whose first sentence answers the heading. 0 at 30%, full marks at 80%Zyppy: answer near the top 8.8/10, query-answer match 9.2/10
Question-led headings
Answerability
10 → 6% of H2s and H3s written as questions. Full marks at 60%No direct study. Inferred from query-answer match, so weighted down
Summary at the top
Answerability
7 → 7A key points or summary block is presentZyppy: answer near the top 8.8/10
Self-contained passages
Extractability
8 → 12% of section openings that make sense lifted out on their own. 0 at 70%, full marks at 95%Zyppy: self-contained passages 8.0/10
Short sentences
Extractability
10 → 7Average sentence length. Full marks at 14 words or fewer, 0 at 22 or moreNot among the Zyppy factors. Kept for extractability, at a lower weight
Lists and tables
Extractability
12 → 14List blocks per 1,000 words (full marks at 5), plus at least one tableZyppy: AI-ready structure (headings and tables) 8.6/10
Statistics and specifics
Evidence
10 → 10Numbers, percentages or currency figures. Full marks at 5Princeton: adding statistics lifted visibility by up to 40%
Named sources
Evidence
7 → 10Mentions of a named source or data provider. Full marks at 2Zyppy: cites sources 8.0/10. Princeton: citing sources lifted visibility by up to 30%
Named expertise
Evidence
6 → 6An author or reviewer is shown, or a named expert is quotedPrinceton: adding quotations lifted visibility by up to 37%
Recently updated
Freshness
10 → 8Months since last update. Full marks at 3 or fewer, 0 at 12 or moreZyppy: freshness 7.0/10. Ahrefs: AI-cited content is 25.7% fresher than organic results, but AI Overviews cite slightly older pages
The ten criteria and their weights out of 100, version 2, with the original weight marked where it changed Weight, points out of 100. Bar: v2. Tick: v1 where it changed. 05101520 Answer-first sections Question-led headings Summary at the top Self-contained passages Short sentences Lists and tables Statistics and specifics Named sources Named expertise Recently updated Answerability Answerability Answerability Extractability Extractability Extractability Evidence Evidence Evidence Freshness 20 6, was 10 7 12, was 8 7, was 10 14, was 12 10 10, was 7 6 8, was 10
Six of the ten weights moved once they were checked against the research. Answer-first kept the top weight. Freshness and question headings lost the most.

Answer-first keeps the top weight, and the evidence now backs it: answering near the top and closely matching the question are two of the highest-scoring content factors in the Zyppy analysis. Freshness is the biggest correction. The effect is real but moderate, and Google’s AI Overviews actually cite pages about 16 days older than the organic results do. A refresh calendar matters less than whether each section answers its question.

The weights are still a practitioner’s call. The difference is that every one of them now points to a source, and the ones without strong evidence carry less weight.

Why I trust the numbers

A rubric is only as good as its consistency. If two people score the same page and get different answers, the score means nothing. So I borrowed the toolkit researchers use to make a judgement call defensible.

  1. A real corpus. 32 articles from one knowledge centre, 284 sections and around 36,000 words, captured verbatim on the same two days.
  2. Two layers of measurement. Some criteria are computed the same way for every page (sentence length, lists, tables, figures, sources, dates). The judgement-based ones are hand-coded: for every one of the 284 section openings, does the first sentence answer the heading, defer it (“it depends”), or lead in with something else? And does it make sense lifted out on its own?
  3. A second coder. A separate coder scored a random sample of 40 sections without seeing the first set. They agreed 98% of the time on how sections open (Cohen’s kappa 0.95) and 95% on whether they stand alone (kappa 0.86). Both sit in the “almost perfect” band.
  4. Significance tests. Fisher’s exact test for question vs statement headings, Kruskal-Wallis for differences between content types, Spearman correlation for which criteria move together, and a 10,000-sample bootstrap for the confidence interval on the average score.
  5. Stress-testing the weights. Even anchored to research, turning evidence into points is a judgement call, so I tried to break the weights. With equal weights, the article ranking still correlates at 0.93 with the weighted one. Across 1,000 runs with every weight randomly shifted by up to 50%, 95% of runs still correlated at 0.91 or above. Dropping any single criterion kept it at 0.85 or above.

That last point is the one I’d want a sceptical client to see. You can argue with my weights, and the pages that score badly still score badly.

Stress testWhat changedRank correlation with the weighted ranking
Equal weightsEvery criterion worth 10 points0.93
Random perturbation1,000 runs, every weight shifted by up to 50%0.91 or above in 95% of runs
Leave one outEach criterion dropped in turn0.85 or above
v1 to v2 weightsThe evidence-based reweighting described above0.96, average score unchanged at 45.6

The numbers that follow use the v2 weights. Switching from my original weights to the evidence-based ones left the average score exactly where it was, at 45.6, and barely changed the order of the articles (rank correlation 0.96). That’s the stress test doing its job on a real change.

What it found

The average article scored 45.6 out of 100 (95% confidence interval 41.7 to 49.6), and 19 of the 32 scored below 50. This is a well-resourced brand with a serious content team, which is exactly why the patterns are worth sharing.

43% of sections don’t answer their heading in the first sentence. 31 of the 32 articles do it at least once.

Question headings made it worse. Sections under a question heading answered directly only 38% of the time, against 67% for statement headings (p < 0.0001). The team had adopted question headings, the most-repeated piece of GEO advice going, without the answer-first writing that makes them work. A question with no answer underneath is arguably worse for AI retrieval than a plain statement, because it promises something the passage doesn’t deliver.

Share of sections that answer in the first sentence: 67% under statement headings, 38% under question headings Sections that answer their heading in the first sentence Statement headings 67% Question headings 38% 284 sections, 32 articles. Fisher’s exact test, p < 0.0001.
The most-repeated piece of GEO advice, adopted without the writing habit that makes it work.

Three criteria cost 47% of all lost points: answer-first sections, lists and tables, and statistics and specifics. That’s where the effort should go first.

Evidence was the weakest dimension. 20 of 32 articles never named a source, and none showed an author or reviewer. Four quoted a named in-house expert, which is a strength to build on.

Content type made no significant difference. Whatever job the article was doing, it scored about the same. That tells you these are house-style habits running across the whole team, and house-style habits can be fixed through briefing, training and QA.

The eight habits that bury the answer

Naming a problem is easy. Telling writers exactly what to stop doing is harder. So every section that didn’t answer in its first sentence was sorted, by fixed rules, into one of eight opening habits. Because the rules are fixed, the counts can be repeated on any other content and compared like for like.

HabitWhat it looks likeWhat to do instead
DeferralOpens with “It depends”, “That’s up to you” or a bare “Yes”Give the most common answer and its condition in one sentence
List set-upAnnounces a list instead of summarising itSummarise the list in one sentence, then give the list
Question restatedAsks or repeats the heading’s questionAnswer it
Advice preambleSays something is important, or to check your policy, without saying whatState the thing that’s important
Promotional teaserA “find out” or “discover” line where the answer should beMove the sales line to the end of the section
Reason before actionAn instruction heading that opens with why it mattersInstruction first, reason second
Scene-settingA general or emotional warm-up lineDelete it. The second sentence is usually the right first sentence
Adjacent factSomething true and related, but not the answerAnswer the heading first, then add the supporting fact

Scene-setting is the one I see everywhere, in every sector. It’s how most of us were taught to write for the web, and it pushes the answer out of the first sentence, which is where the evidence says it needs to be.

What a fix looks like

Every rewrite in the study followed one rule: use only facts already on the page. No new figures, claims or conditions. Each before and after was then checked by a separate reviewer who hadn’t seen the reasoning, with the job of flagging anything in the “after” that wasn’t in the source.

One page explained the difference between buildings and contents cover in running prose. Same facts, rebuilt as a comparison an AI can lift in one go:

Buildings coverContents cover
What it protectsThe structure of your homeYour belongings
ExamplesRoof, walls, windows, fitted kitchen, garage, conservatoryFurniture, white goods, gadgets, curtains

Nothing was added and the meaning is unchanged, but the passage is now far easier to extract and quote.

One caution for regulated sectors: a few rewrites added “usually” or “most” where the original page already hedged. In insurance, that’s a financial promotions question, so those went to compliance before going anywhere near a live page.

What it doesn’t tell you (yet)

I’d rather be upfront about the limits than have someone else point them out.

  • It measures how a page is written, not whether it’s cited. A high score means a page is built to be extracted. Whether AI models actually pick it up is a separate measurement.
  • The weights are judgement-based. The stress tests show the rankings hold when the weights change, but they’re still a practitioner’s call.
  • The eight habits were sorted by rules, not double-coded. The section openings were; the habit buckets were checked by hand.

The biggest limit is scope. The rubric only scores what’s on the page, and the highest-evidence factors in the Zyppy analysis sit outside it: whether a page can be crawled (9.5/10), how it ranks in classic search (9.4/10), and brand mentions across the web, which Ahrefs found correlate about three times more strongly with AI Overview visibility than backlinks. A perfectly written page that can’t be crawled can’t be cited.

The on-page work still earns its place. Yext’s study of 6.8 million AI citations found 86% came from sources brands control, with first-party websites the biggest single source at 44%. In financial services, 48.2% of citations went to brand-owned websites. For an insurer, the knowledge centre is one of the assets most likely to be cited, which is why it’s worth scoring.

The next step is the interesting one: rescoring pages after rewrites and training to measure the change, and linking scores to AI visibility data to find out which criteria actually move citations. That’s when a rubric stops being a good idea and starts being evidence.

Want your content scored?

The whole process is now repeatable. I can take 25 to 40 articles from any content library, score them against the same rubric, and hand back the scores, the charts, the habits your writers need to change and before-and-after rewrites they can learn from. Because the rubric stays fixed, your scores are comparable over time and across sections of your site.

If you’d like to know where your content is losing points in AI search, book a strategy call.

Sources

Jessica Redman

Jessica Redman

GEO, SEO and AI enablement consultant. Ten years across insurance, SaaS, eCommerce and regulated B2B.

Keep reading

The newsletter

GEO tactics, AI workflows and automation builds, from someone doing the work with enterprise clients

Every issue breaks down what's actually driving AI visibility and content performance right now: the research, the builds, the results, and the stuff most people aren't sharing.

  • GEO tactics and AI visibility strategies tested with real clients, not recycled theory
  • AI workflow and automation breakdowns (Claude, AirOps, n8n) with the exact steps to replicate them
  • Content frameworks built for how AI models find, evaluate and recommend content
  • Experiments and results from the field, shared openly
  • Early access to tools, skills and templates before they go public

Published weekly, to about 1,000 marketers

For content marketers and SEOs who want to work smarter with AI, marketing leaders who need the data to back their strategy, and practitioners who want to see the full build and replicate it.

Tell me where your growth is stuck.

A conversation costs you thirty minutes and no obligation.