The Generative Engine Optimization Paper Decoded: What the KDD 2024 Research Actually Found — and What Practitioners Get Wrong
The Generative Engine Optimization Paper Decoded: What the KDD 2024 Research Actually Found — and What Practitioners Get Wrong
September 16, 2026

The Generative Engine Optimization Paper Decoded: What the KDD 2024 Research Actually Found and What Practitioners Get Wrong
Introduction: The Paper Everyone Cites and Almost No One Has Read
A curious pattern has emerged in the world of AI search optimization. A single academic paper gets cited constantly across blog posts, LinkedIn threads, and agency pitch decks, yet the most consequential findings inside that paper are routinely misrepresented, oversimplified, or omitted entirely. The paper is GEO: Generative Engine Optimization, authored by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande, with affiliations spanning Princeton University, IIT Delhi, the Allen Institute for AI, and Georgia Tech.
The paper carries genuine academic weight. It was first released on arXiv on November 16, 2023, then formally peer-reviewed and published in the proceedings of the 30th ACM SIGKDD Conference (KDD 2024) in Barcelona, one of the most prestigious venues in data mining and applied machine learning. This is not marketing fluff dressed up as science. It is real research, published in a real venue.
The problem is how it has been read. This article corrects three specific misreadings that dominate practitioner coverage: the 40% figure misrepresented as a universal average, the near-total omission of the equalizer effect for lower-ranked pages, and the ignored domain-specificity caveat that makes blanket GEO recommendations unreliable. This is not a summary; it is an interrogation of how the paper has been read, misread, and what its findings actually mean for practitioners operating in 2026.
The stakes have never been higher. AI search traffic has grown 16x since 2024, ChatGPT now processes 2 billion queries daily, and 68% of Google searches end without a click. Accurate interpretation of GEO research is no longer academic. It is competitive.
What the GEO Paper Actually Set Out to Do
The paper’s foundational framing is often skipped: it treats GEO as a black-box optimization problem. The internal workings of generative engines are unknown, so optimization must be inferred entirely from observable outputs. This framing matters because it defines the limits of what can be concluded.
The authors introduced a three-stakeholder model almost entirely absent from practitioner coverage: users who query the engine, generative engines that synthesize answers, and content creators who want their sources cited. Understanding these competing interests is essential to understanding what the paper measures.
Crucially, the paper set out to establish whether content modifications could measurably influence citation rates in AI-generated responses. It did not set out to prove durable traffic effects or organic discoverability gains. That distinction gets lost constantly.
The paper’s most enduring contribution was GEO-bench: a benchmark of 10,000 diverse user queries drawn from nine datasets including MS MARCO, ORCAS-1, Natural Questions, AllSouls, and LIMA, each tagged by domain, difficulty, and query intent. The benchmark is publicly available on Hugging Face, enabling reproducible research. It was the first standardized evaluation framework for the problem, and it coined the very term “Generative Engine Optimization.” Everything practitioners call GEO today traces back to this work.
The Experimental Setup: What the Paper Actually Tested
The experiment used a two-stage pipeline. In Stage 1, Google Search retrieved the top-5 sources per query. In Stage 2, GPT-3.5-turbo synthesized answers with citations drawn from those retrieved sources. This design closely mimics how commercial generative engines operate.
Understanding this pipeline is essential for interpreting the results correctly. The experiment tested whether content modifications increased citation prominence among pages already in the retrieval context. It did not test whether modifications caused pages to enter that context in the first place. That is a downstream test, not an upstream one.
The paper tested nine optimization methods: Statistics Addition, Quotation Addition, Cite Sources, Fluency Optimization, Keyword Stuffing, Authoritative Tone, Unique Content, Technical Terminology, and Simplicity. Each method was applied to source documents, and the resulting AI-generated answers were measured for changes in visibility.
To confirm the findings were not confined to a controlled setup, the authors validated results on Perplexity.ai, a live generative engine, demonstrating visibility improvements of up to 37%.
One limitation most practitioner content ignores is significant. The 2026 critical survey by Martinez explicitly notes that the 40% gains are “valid within its experimental setting but conditional on a source already being present in a fixed context.” They do not establish organic discoverability or durable traffic effects.
The Two Visibility Metrics: A Distinction That Changes Everything
Most practitioner content cites a single number from the paper without explaining that the paper used two fundamentally different ways of measuring visibility. Conflating them produces misleading interpretations.
Metric 1: Objective Impression (Position-Adjusted Word Count, or PAWC). This counts how many words from a source appear in the AI-generated answer, weighted by position. Earlier citations receive higher weight.
Metric 2: Subjective Impression. This is a human-rated assessment of how prominently a source is cited, capturing qualitative citation quality that raw word count cannot measure.
The practical implication is significant. A strategy can score well on PAWC (many words cited) yet poorly on Subjective Impression (cited in a peripheral or dismissive way), or the reverse. Practitioners need to know which metric any cited statistic is drawn from.
This is where the top strategies become instructive. Statistics Addition produced +41% PAWC and +37% Subjective Impression, making it one of the few methods that performed strongly on both dimensions simultaneously. For real-world application, the metric matters: if the goal is AI Overview inclusion (a binary presence question), PAWC is the more relevant proxy. If the goal is brand authority within a citation, Subjective Impression is more relevant.
Misreading #1: The 40% Figure Is a Maximum, Not an Average
The dominant practitioner narrative presents the 40% visibility boost as a typical or average result achievable by applying GEO methods. This is factually incorrect.
The 40% figure is an upper bound. It represents the best-performing strategy (Statistics Addition) on the Position-Adjusted Word Count metric under controlled experimental conditions. The three top performers were Statistics Addition (+41% PAWC, +37% Subjective Impression), Quotation Addition, and Cite Sources. Not all nine methods performed comparably.
One finding directly contradicts a common assumption: Keyword Stuffing showed little to no improvement in generative engine responses. Traditional SEO tactics do not automatically transfer to GEO.
The Martinez survey delivers the corrective plainly. Across 45 GEO studies, it concluded that no technique shows a stable, cross-platform causal effect on organic discoverability. The 40% figure is conditional, not generalizable.
The consequence for practitioners is direct. Anyone expecting a 40% visibility lift from adding statistics to any page in any domain is operating on a false premise. The actual effect depends on domain, query type, retrieval position, and the specific generative engine in use.
Misreading #2: The Equalizer Effect That Practitioners Ignore
One of the paper’s most actionable discoveries appears in a result almost no practitioner content covers: the equalizer effect for lower-ranked pages.
The specific finding is striking. “Cite Sources” produced a +115.1% visibility lift for pages ranked 5th in SERPs, more than double the lift observed for top-ranked pages using the same method.
Structurally, this means GEO methods do not simply amplify existing advantages for high-ranking pages. They can disproportionately benefit lower-ranked sources, partially closing the gap with top-ranked competitors.
For most brands, this is the single most strategically important finding in the entire paper. The majority of businesses do not rank first for their target queries. The equalizer effect suggests GEO offers asymmetric upside precisely for the brands that need it most.
Yet most practitioner content frames GEO as a tool for maintaining dominance among already-visible pages, rather than a mechanism for competitive disruption. The nuance still applies: the equalizer effect was observed for a specific method under specific conditions and should not be over-generalized without domain-specific testing.
Misreading #3: The Domain-Specificity Caveat That Makes Blanket Recommendations Unreliable
The paper explicitly showed that GEO strategy effectiveness varies significantly by domain. This finding is almost universally omitted from practitioner summaries.
Within the paper’s scope, the nine optimization methods did not produce uniform results across the nine content categories in GEO-bench. What worked in one domain did not necessarily work in another.
This creates a real problem. The most common practitioner advice (“add statistics,” “use authoritative tone,” “cite sources”) is presented as universally applicable, when the paper’s own evidence shows these are domain-conditional recommendations. GEO-bench was deliberately tagged by domain, difficulty, and query intent to enable exactly this level of analysis. Ignoring that dimension means discarding the paper’s most nuanced data.
The academic field recognized the problem. Researchers built E-GEO, a dedicated e-commerce GEO benchmark, precisely because the original findings did not transfer cleanly to product recommendation contexts.
The practical takeaway is a better question. Before applying any GEO method, practitioners should ask: “What domain does my content operate in, and what does the evidence say about this method in that specific domain?” Not simply: “Which method had the highest headline number?”
The Upstream Discoverability Problem: What the Paper Does Not Solve
The GEO paper’s experimental design assumes the source page is already in the retrieval context. It tests what happens after a page is retrieved, not whether a page gets retrieved at all.
This matters enormously in practice. If a page is not being retrieved by the AI engine’s search layer, no amount of content optimization will improve its citation rate. The optimization problem is upstream, not downstream.
A large-scale 2026 factorial experiment by Vishwakarma et al., comprising 252,000 trials across six LLMs and 18 factors, found that relevance and position within the retrieval context are the primary determinants of citation. This redirects GEO strategy toward upstream discoverability rather than content rewriting alone.
The decoupling is confirmed elsewhere. Only 12% of AI-cited URLs rank in Google’s top 10 for the same query, meaning traditional ranking is neither sufficient nor necessary for AI citation, but retrieval presence is.
The implication is a matter of sequencing. Content-level GEO modifications are a second-order optimization. They matter only after the first-order problem (getting into the retrieval set) is solved. Discoverability, including topical authority, structured data, crawlability, and entity recognition, must come first for the paper’s findings to apply at all.
How the Academic Field Has Evolved Since KDD 2024
The pace of development has been rapid. The July 2026 Martinez critical survey reviewed 45 GEO studies published between November 2023 and July 2026. The original paper is now one data point in a much larger body of evidence.
Key follow-on work includes E-GEO (e-commerce benchmark, 2025), IF-GEO (multi-query conflict-aware optimization, 2026), AgenticGEO (a self-evolving agentic GEO system, 2026), SAGEO Arena (realistic search-augmented evaluation, 2026), and a citation failure diagnosis paper.
IF-GEO is particularly significant. It addresses a limitation the original paper never tackled: optimizing a single document for multiple conflicting queries simultaneously, which is the norm for most commercial content.
Governance research has also emerged. A 2026 position paper identified three systemic risks: concentrated influence from low contestability, undisclosed commercial influence in AI answers, and academic-industry evaluation gaps from offline versus deployed testing. Meanwhile, SafeGEO found that GEO attacks can increase the rate at which flawed products enter AI recommendation sets by up to 83.2%, reframing GEO as both an optimization tool and a potential information integrity risk.
The takeaway is clear. Practitioners relying solely on the 2023/2024 paper are working from an incomplete picture.
What the Paper’s Findings Actually Transfer to Real-World Execution in 2026
Given everything the paper found, what it did not find, and how the field has evolved, the following represents what practitioners can confidently act on.
Transferable finding #1: Statistics and data citations are the most robust strategy. Statistics Addition was the top performer on both visibility metrics and has been replicated across follow-on studies. Adding verifiable, specific data points is the most evidence-backed GEO tactic available.
Transferable finding #2: Traditional SEO keyword tactics do not transfer. Keyword Stuffing’s failure is durable. Generative engines evaluate semantic relevance and source quality, not keyword density.
Transferable finding #3: Lower-ranked pages have asymmetric upside. The equalizer effect should reshape how brands with mid-to-lower SERP positions prioritize GEO investment.
Conditional finding: Domain matters before method. Any tactic should be evaluated against domain-specific evidence, not applied universally.
Structural prerequisite: Retrieval presence must come first. Content-level GEO only works if the page is already being retrieved. Topical authority and technical discoverability are the necessary preconditions.
The broader context reinforces the stakes. AI-sourced traffic converts at significantly higher rates than traditional organic traffic, and 83% of AI Overview queries end without a click. The reward for accurate execution is high, and so is the cost of acting on misread research.
A Framework for Reading GEO Research Without Getting Misled
Because GEO research is expanding rapidly and misreadings are common, practitioners need a structured approach to evaluating any study or claim. Six questions do most of the work:
- What was the experimental setup? Was it a controlled offline experiment or a live deployed system? Offline results may not transfer due to stochasticity, engine updates, and retrieval variability.
- What visibility metric was used? PAWC and Subjective Impression measure different things. Every percentage should be traced back to its metric definition.
- Is the cited figure a maximum, average, or median? The 40% figure is a maximum for the best method. Always determine whether a headline number reflects typical or peak performance.
- What domain was the study conducted in? Results from one content category should not be assumed to transfer to another.
- Does the finding address retrieval or citation prominence? These are separate problems. Retrieval presence is upstream of citation prominence.
- Has the finding been replicated or challenged? The Martinez 2026 survey is the most comprehensive audit of GEO replication to date. Any “proven” tactic should be checked against its conclusions.
Conclusion: Respecting the Research Means Reading It Accurately
The GEO paper is legitimate, peer-reviewed, foundational research. The practitioner ecosystem, however, has developed a simplified and often misleading version of its findings that does a disservice to both the research and the people acting on it.
Three corrections stand out. The 40% figure is a conditional maximum, not a universal average. The equalizer effect for lower-ranked pages is the paper’s most actionable finding for most brands and is almost universally ignored. Domain-specificity means blanket GEO recommendations are an oversimplification the paper itself does not support.
None of this diminishes the paper’s genuine value. It established GEO as a legitimate discipline, created the first standardized benchmark, demonstrated that content modifications can influence AI citation rates, and showed that traditional SEO tactics do not transfer. Those are durable contributions. With 45+ follow-on studies produced in under three years, however, the original paper is the foundation, not the ceiling.
In a search environment where 68% of queries end without a click, AI citation converts at significantly higher rates than organic traffic, and only 12% of AI-cited URLs rank in Google’s top 10, the difference between accurate and inaccurate GEO execution is measurable in revenue. Reading the research carefully is not an academic exercise. It is a competitive advantage.
Ready to Build a GEO Strategy Grounded in What the Research Actually Says?
Most GEO advice applies blanket tactics based on misread statistics. KOZEC takes a different approach, building content ecosystems structured around the evidence-backed principles the paper actually supports: topical authority, structured data, verifiable citations, and domain-appropriate optimization.
KOZEC’s SCO (Search Compliance Optimization) framework addresses the upstream discoverability problem first. It builds the interconnected, authoritative content infrastructure that gets pages into the AI retrieval context before content-level GEO optimization is applied. That sequencing reflects exactly what the 2026 research confirms: retrieval presence must come before citation prominence.
The platform structures content for AI Overview citation, Google AI Mode, and chat assistants, not just traditional rankings, reflecting the documented decoupling of SEO rank and AI citation. Setup takes days, not months, with no long-term contracts, so practitioners can act on accurate GEO insights without organizational delays.
To see how research-grounded GEO principles translate into automated, scalable content execution, schedule a demo at kozec.ai/schedule-a-demo/ or contact the team at (888) 545-7090.
Stay In The Loop
Subscribe to our free newsletter.
Stop Managing SEO - Start Scaling It
Let KOZEC handle strategy, content, and execution - so you can focus on growth.
Automated SEO content for growing agencies.
KOZEC helps agencies, consultants, and growing brands publish high-quality SEO content on autopilot — so your site ranks higher and converts more visitors.
Managing SEO content for many client websites doesn’t scale with traditional methods. Writers are expensive and inconsistent, keyword research is time-consuming, and publishing requires multiple manual steps. As agencies grow, maintaining both quality and consistency becomes increasingly difficult. KOZEC (Keyword Optimized Zero Effort Content) solves this by automating analysis, keyword discovery, content creation, and publishing—so your clients get reliable SEO content while your team focuses on growth.
Increase organic traffic without manual content creation
Publish keyword-optimized posts automatically to WordPress
Turn SEO into a predictable, scalable growth channel

