Most GEO advice is a checklist with no numbers behind it. When a major UK insurer asked how AI-ready their knowledge centre was, I built a 10-criterion weighted rubric, checked every weight against published citation research, double-coded the judgement calls and stress-tested the weights. The average article scored 45.6 out of 100, 43% of sections buried the answer, and question headings made it worse. Here is the rubric, the method, the eight habits and what a fix looks like.
Most GEO advice you’ll read this year is a checklist with no numbers behind it. “Use question headings.” “Answer first.” “Add stats.” All sensible, and almost none of it measured.
So when a major UK insurer asked me how AI-ready their home insurance knowledge centre was, I didn’t want to hand over a list of opinions. I wanted a score they could trust, a way to see exactly where the points were being lost, and proof the findings would hold up if someone challenged them.
The result is a 10-criterion weighted GEO rubric, with each weight checked against published citation research and the scoring tested with the same methods a researcher would use: double-coding, inter-rater agreement, significance tests and sensitivity analysis. I’ve since turned it into a repeatable process I can run on any content library. Here’s how it works and what it found.
The rubric: 10 criteria, 100 points
The rubric asks four questions of every page. Does it answer the question? Can an AI lift a passage out cleanly? Is there evidence behind the claims? Is it current?
Each criterion gets a floor (where it scores zero) and a target (where it hits full marks), with a straight line in between. That matters more than it sounds. A page with 79% answer-first sections and one with 31% shouldn’t both be a “partial pass”, and a linear scale shows the difference.
The first version of the weights was my judgement. Before publishing, I checked each one against the biggest citation studies available: Cyrus Shepard’s Zyppy meta-analysis, which scores 23 citation factors out of 10 based on 54 studies, Ahrefs’ analysis of nearly 17 million AI citations, and the Princeton GEO paper (Aggarwal et al., KDD 2024). Where the evidence was strong, a criterion gained weight. Where it was thin, it lost some.
| Criterion | Weight (v1 → v2) | How it’s scored | Evidence behind the weight |
|---|---|---|---|
| Answer-first sections Answerability | 20 → 20 | % of sections whose first sentence answers the heading. 0 at 30%, full marks at 80% | Zyppy: answer near the top 8.8/10, query-answer match 9.2/10 |
| Question-led headings Answerability | 10 → 6 | % of H2s and H3s written as questions. Full marks at 60% | No direct study. Inferred from query-answer match, so weighted down |
| Summary at the top Answerability | 7 → 7 | A key points or summary block is present | Zyppy: answer near the top 8.8/10 |
| Self-contained passages Extractability | 8 → 12 | % of section openings that make sense lifted out on their own. 0 at 70%, full marks at 95% | Zyppy: self-contained passages 8.0/10 |
| Short sentences Extractability | 10 → 7 | Average sentence length. Full marks at 14 words or fewer, 0 at 22 or more | Not among the Zyppy factors. Kept for extractability, at a lower weight |
| Lists and tables Extractability | 12 → 14 | List blocks per 1,000 words (full marks at 5), plus at least one table | Zyppy: AI-ready structure (headings and tables) 8.6/10 |
| Statistics and specifics Evidence | 10 → 10 | Numbers, percentages or currency figures. Full marks at 5 | Princeton: adding statistics lifted visibility by up to 40% |
| Named sources Evidence | 7 → 10 | Mentions of a named source or data provider. Full marks at 2 | Zyppy: cites sources 8.0/10. Princeton: citing sources lifted visibility by up to 30% |
| Named expertise Evidence | 6 → 6 | An author or reviewer is shown, or a named expert is quoted | Princeton: adding quotations lifted visibility by up to 37% |
| Recently updated Freshness | 10 → 8 | Months since last update. Full marks at 3 or fewer, 0 at 12 or more | Zyppy: freshness 7.0/10. Ahrefs: AI-cited content is 25.7% fresher than organic results, but AI Overviews cite slightly older pages |
Answer-first keeps the top weight, and the evidence now backs it: answering near the top and closely matching the question are two of the highest-scoring content factors in the Zyppy analysis. Freshness is the biggest correction. The effect is real but moderate, and Google’s AI Overviews actually cite pages about 16 days older than the organic results do. A refresh calendar matters less than whether each section answers its question.
The weights are still a practitioner’s call. The difference is that every one of them now points to a source, and the ones without strong evidence carry less weight.
Why I trust the numbers
A rubric is only as good as its consistency. If two people score the same page and get different answers, the score means nothing. So I borrowed the toolkit researchers use to make a judgement call defensible.
- A real corpus. 32 articles from one knowledge centre, 284 sections and around 36,000 words, captured verbatim on the same two days.
- Two layers of measurement. Some criteria are computed the same way for every page (sentence length, lists, tables, figures, sources, dates). The judgement-based ones are hand-coded: for every one of the 284 section openings, does the first sentence answer the heading, defer it (“it depends”), or lead in with something else? And does it make sense lifted out on its own?
- A second coder. A separate coder scored a random sample of 40 sections without seeing the first set. They agreed 98% of the time on how sections open (Cohen’s kappa 0.95) and 95% on whether they stand alone (kappa 0.86). Both sit in the “almost perfect” band.
- Significance tests. Fisher’s exact test for question vs statement headings, Kruskal-Wallis for differences between content types, Spearman correlation for which criteria move together, and a 10,000-sample bootstrap for the confidence interval on the average score.
- Stress-testing the weights. Even anchored to research, turning evidence into points is a judgement call, so I tried to break the weights. With equal weights, the article ranking still correlates at 0.93 with the weighted one. Across 1,000 runs with every weight randomly shifted by up to 50%, 95% of runs still correlated at 0.91 or above. Dropping any single criterion kept it at 0.85 or above.
That last point is the one I’d want a sceptical client to see. You can argue with my weights, and the pages that score badly still score badly.
| Stress test | What changed | Rank correlation with the weighted ranking |
|---|---|---|
| Equal weights | Every criterion worth 10 points | 0.93 |
| Random perturbation | 1,000 runs, every weight shifted by up to 50% | 0.91 or above in 95% of runs |
| Leave one out | Each criterion dropped in turn | 0.85 or above |
| v1 to v2 weights | The evidence-based reweighting described above | 0.96, average score unchanged at 45.6 |
The numbers that follow use the v2 weights. Switching from my original weights to the evidence-based ones left the average score exactly where it was, at 45.6, and barely changed the order of the articles (rank correlation 0.96). That’s the stress test doing its job on a real change.
What it found
The average article scored 45.6 out of 100 (95% confidence interval 41.7 to 49.6), and 19 of the 32 scored below 50. This is a well-resourced brand with a serious content team, which is exactly why the patterns are worth sharing.
43% of sections don’t answer their heading in the first sentence. 31 of the 32 articles do it at least once.
Question headings made it worse. Sections under a question heading answered directly only 38% of the time, against 67% for statement headings (p < 0.0001). The team had adopted question headings, the most-repeated piece of GEO advice going, without the answer-first writing that makes them work. A question with no answer underneath is arguably worse for AI retrieval than a plain statement, because it promises something the passage doesn’t deliver.
Three criteria cost 47% of all lost points: answer-first sections, lists and tables, and statistics and specifics. That’s where the effort should go first.
Evidence was the weakest dimension. 20 of 32 articles never named a source, and none showed an author or reviewer. Four quoted a named in-house expert, which is a strength to build on.
Content type made no significant difference. Whatever job the article was doing, it scored about the same. That tells you these are house-style habits running across the whole team, and house-style habits can be fixed through briefing, training and QA.
The eight habits that bury the answer
Naming a problem is easy. Telling writers exactly what to stop doing is harder. So every section that didn’t answer in its first sentence was sorted, by fixed rules, into one of eight opening habits. Because the rules are fixed, the counts can be repeated on any other content and compared like for like.
| Habit | What it looks like | What to do instead |
|---|---|---|
| Deferral | Opens with “It depends”, “That’s up to you” or a bare “Yes” | Give the most common answer and its condition in one sentence |
| List set-up | Announces a list instead of summarising it | Summarise the list in one sentence, then give the list |
| Question restated | Asks or repeats the heading’s question | Answer it |
| Advice preamble | Says something is important, or to check your policy, without saying what | State the thing that’s important |
| Promotional teaser | A “find out” or “discover” line where the answer should be | Move the sales line to the end of the section |
| Reason before action | An instruction heading that opens with why it matters | Instruction first, reason second |
| Scene-setting | A general or emotional warm-up line | Delete it. The second sentence is usually the right first sentence |
| Adjacent fact | Something true and related, but not the answer | Answer the heading first, then add the supporting fact |
Scene-setting is the one I see everywhere, in every sector. It’s how most of us were taught to write for the web, and it pushes the answer out of the first sentence, which is where the evidence says it needs to be.
What a fix looks like
Every rewrite in the study followed one rule: use only facts already on the page. No new figures, claims or conditions. Each before and after was then checked by a separate reviewer who hadn’t seen the reasoning, with the job of flagging anything in the “after” that wasn’t in the source.
One page explained the difference between buildings and contents cover in running prose. Same facts, rebuilt as a comparison an AI can lift in one go:
| Buildings cover | Contents cover | |
|---|---|---|
| What it protects | The structure of your home | Your belongings |
| Examples | Roof, walls, windows, fitted kitchen, garage, conservatory | Furniture, white goods, gadgets, curtains |
Nothing was added and the meaning is unchanged, but the passage is now far easier to extract and quote.
One caution for regulated sectors: a few rewrites added “usually” or “most” where the original page already hedged. In insurance, that’s a financial promotions question, so those went to compliance before going anywhere near a live page.
What it doesn’t tell you (yet)
I’d rather be upfront about the limits than have someone else point them out.
- It measures how a page is written, not whether it’s cited. A high score means a page is built to be extracted. Whether AI models actually pick it up is a separate measurement.
- The weights are judgement-based. The stress tests show the rankings hold when the weights change, but they’re still a practitioner’s call.
- The eight habits were sorted by rules, not double-coded. The section openings were; the habit buckets were checked by hand.
The biggest limit is scope. The rubric only scores what’s on the page, and the highest-evidence factors in the Zyppy analysis sit outside it: whether a page can be crawled (9.5/10), how it ranks in classic search (9.4/10), and brand mentions across the web, which Ahrefs found correlate about three times more strongly with AI Overview visibility than backlinks. A perfectly written page that can’t be crawled can’t be cited.
The on-page work still earns its place. Yext’s study of 6.8 million AI citations found 86% came from sources brands control, with first-party websites the biggest single source at 44%. In financial services, 48.2% of citations went to brand-owned websites. For an insurer, the knowledge centre is one of the assets most likely to be cited, which is why it’s worth scoring.
The next step is the interesting one: rescoring pages after rewrites and training to measure the change, and linking scores to AI visibility data to find out which criteria actually move citations. That’s when a rubric stops being a good idea and starts being evidence.
Want your content scored?
The whole process is now repeatable. I can take 25 to 40 articles from any content library, score them against the same rubric, and hand back the scores, the charts, the habits your writers need to change and before-and-after rewrites they can learn from. Because the rubric stays fixed, your scores are comparable over time and across sections of your site.
If you’d like to know where your content is losing points in AI search, book a strategy call.
Sources
- AI Citation Ranking Factors, Cyrus Shepard, Zyppy Signal (May 2026): 23 factors scored out of 10 from 54 studies, experiments and patents. Summarised with Ahrefs’ brand-correlation data in What actually gets you cited in AI search (2026 data), Digital Applied
- AI assistants prefer to cite “fresher” content (17 million citations analysed), Ryan Law, Ahrefs
- 86% of AI citations come from brand-managed sources, Yext (October 2025)
- GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024 (Princeton)

Jessica Redman
GEO, SEO and AI enablement consultant. Ten years across insurance, SaaS, eCommerce and regulated B2B.
Keep reading