marketing September 16, 2026 · Mintec

The GEO paper says 40% more visibility. Here's what it actually measured.

The Princeton paper that defined GEO claims 40% more visibility. But the metric isn't traffic or clicks, and the paper has limitations the industry ignores. We break down what it measured, what it didn't, and what 2026 follow-up papers found.

Your SEO agency says it applies "GEO optimizations" and promises 40% more AI visibility. Before you pay, read what the paper actually says.

In November 2023, researchers at Princeton and Georgia Tech published a paper that coined the term GEO: Generative Engine Optimization. It was presented at KDD 2024, ACM's premier knowledge discovery conference, and has since become the foundation for almost every agency selling GEO services.

The number that gets repeated is "40% more visibility." It shows up in sales pitches, agency blogs, conference presentations. What almost nobody mentions is that the 40% is not traffic, not clicks, and not Google rankings.

What the paper actually measured

The paper built a benchmark called GEO-bench with 10,000 queries. For each query, it took the top five Google results, cleaned the text, and fed those sources to a language model to write a response. Then it rewrote one source at a time and measured what changed.

The change was measured with two metrics:

  1. Position-adjusted word count. Counts how many words in the AI response are attributed to a source, with an exponential decay by position. If your source appears early in the response, it counts more than if it appears near the end.

  2. Subjective impression. Seven scores (relevance, influence, uniqueness, diversity, subjective position, subjective count, and click probability) evaluated by G-Eval, which is another language model. It's a model judging another model's output.

The best results: evidence-adding strategies (citing sources, adding quotations, adding statistics) beat the unoptimized baseline by 41% on position-adjusted word count and 28% on subjective impression.

The limitations nobody mentions

The paper has explicit limitations the industry ignores:

It did not measure Google rankings. The authors state word for word: "owing to the black-box nature of search engine algorithms, we did not evaluate how GEO methods affect search rankings." They guess the changes are unlikely to affect rankings since they modify text rather than backlinks or metadata, but it's a guess, not a result.

All sources were already top 5. Every source in the experiment already ranked in Google's top five results for its query. The paper does not explain how to get there. It measures what happens after you're already in the top 5, not how to get there in the first place.

The same edit can hurt. In the paper's results, adding citations to a fifth-placed source lifted it by 115%, but cut a first-placed source by 30.3%. GEO, as the paper defines it, is a redistribution mechanism among sources that already made the shortlist.

The metric is invented. Position-adjusted word count is not a metric used by Google, ChatGPT, or any platform. The authors created it for the paper.

What 2026 follow-up papers found

Three important follow-ups have tested the original findings:

SAGEO Arena (KDD 2026). Researchers at Yonsei and Konkuk argued that existing benchmarks operate "on pre-determined candidate documents," abstracting away the retrieval and reranking stages that exist in a real pipeline. When they rebuilt the experiment including those stages, existing GEO optimizations frequently degraded visibility.

What Gets Cited (SIGIR 2026). A Sprinklr team ran 252,000 paired trials across six models, changing one content factor at a time. The two strongest drivers of which source gets cited first: topical relevance and list position. Explicit price information and a recent timestamp helped consistently. Formatting-only edits had little impact.

The travel paper (preprint 2025). A group trained a model on 1,905 travel content pairs rewritten with citations, statistical evidence, and improved fluency. They reported a 30.96% gain on position-adjusted word count. But it covers a single vertical, uses synthetic training pairs, and tests a single open-weight model.

What this means in practice

Most GEO optimization being sold is about rewriting text: adding citations, inserting statistics, improving fluency. These tactics can help, but only if your page is already among the sources the model retrieves.

What the 2026 research suggests is that the two most important factors are largely outside your direct control:

  1. Topical relevance. That your content is the best answer for the query. This is classic SEO: topical authority, comprehensive content, clear entity signals.

  2. External associations. That other sites mention you alongside the relevant topic. A Search Engine Land experiment in September 2026 found that 85.8% of citations came from third-party listicles, not owned content.

This connects directly to what Google published in its official GEO guidance: optimizing for AI search is still SEO, with an added layer of extractability and structured data.

The Three Jobs Framework

At Mintec (detailed framework) we've been saying that AI search has three separate jobs that most people conflate:

  1. Retrieval. That the engine finds you. This is crawlability, indexation, site structure.
  2. Citation. That the engine cites you in the response. This depends on evidence, authority, and external associations.
  3. Conversion. That the click that arrives converts. This is page experience, clear value proposition.

The GEO paper only measures job 2, and only within a scenario where job 1 is already solved.

What to do with this information

If an agency promises you "40% more AI visibility" based on the Princeton paper, ask them:

  • Is that 40% traffic or position-adjusted word count?
  • Does the paper test content rewriting alone or also retrieval and ranking?
  • What happens when the source is already in first position?

The paper's tactics (adding evidence, citations, statistics) work as a layer on top of a solid foundation. But if your site doesn't rank well organically, lacks topical authority, and nobody mentions you externally, rewriting text won't get you into AI Overviews.

For a real audit of your AI search visibility, check our 30-minute GEO audit guide and the metrics that actually matter for measuring GEO.

The questions that matter

Is GEO a separate discipline from SEO? Google says no, and 2026 papers confirm that content optimizations only work within a pipeline that includes classic retrieval.

Do you need special GEO tools? The paper doesn't test tools. The tools Mintec uses are the same as for SEO: Search Console, analytics, crawlers.

What actually moves the needle? Content that answers questions directly, built topical authority, and third-party mentions. The paper confirmed this without intending to: sources with better positions in the original listing are the ones that get cited most. Ranking matters more than rewriting.

Frequently Asked Questions

What is the original GEO paper?

It's 'GEO: Generative Engine Optimization' by researchers at Princeton and Georgia Tech, published at KDD 2024. It coined the term GEO and measured how certain content changes affect visibility in AI-generated responses.

Is the 40% improvement real traffic?

No. The 40% is a gain in position-adjusted word count, a metric that counts how many words in an AI response are attributed to a source. It's not traffic, not clicks, not Google rankings. The paper explicitly states it did not evaluate how its methods affect rankings.

What do 2026 papers say about GEO?

SAGEO Arena (KDD 2026) found that GEO optimizations can degrade visibility when retrieval and reranking stages are included. What Gets Cited (SIGIR 2026) found that topical relevance and list position are the strongest drivers of citation, not content formatting.

Related Articles