GEO Generative Engine Optimization arXiv:2311.09735: The Academic-to-Implementation Bridge for 2026
GEO Generative Engine Optimization arXiv:2311.09735: The Academic-to-Implementation Bridge for 2026
August 25, 2026

GEO Generative Engine Optimization arXiv:2311.09735: The Academic-to-Implementation Bridge for 2026
Introduction: Why the Original GEO Paper Still Matters in 2026
The way people find information has fundamentally changed. According to Similarweb’s 2026 Generative AI Brand Visibility Index, 35% of US consumers now use AI at the product discovery stage, compared to only 13.6% who still rely on traditional search. AI-referred web sessions grew 500% in 2025. The click-driven internet that classical SEO was built to serve is quietly being replaced by a system where answers arrive fully formed, and users rarely visit the sources behind them.
Yet most Generative Engine Optimization (GEO) content available today reads like a list of superstitions. “Add statistics.” “Cite sources.” “Use an authoritative tone.” These guides rarely explain why those tactics work, how they were measured, or what the underlying research actually proved. They present conclusions without evidence.
Those conclusions all trace back to a single document: arXiv:2311.09735, “GEO: Generative Engine Optimization”. This is the paper that first coined the term, defined the methodology, built the benchmarks, and produced the evidence base the entire field now depends on. Every practitioner tip about statistics and citations is a downstream echo of experiments run in this study.
This article treats that paper as a primary source. It walks technically minded readers through the actual experimental design, the strategy-by-strategy performance data, and the nuanced limitations that surface-level guides ignore. It then bridges those findings to production-ready content strategy, with KOZEC serving as the translator that converts peer-reviewed academic theory into automated, business-ready GEO execution.
One finding anchors everything that follows: keyword stuffing, the cornerstone of classical SEO, showed near-zero measurable lift in generative engine visibility. The old playbook is not just outdated. In this new context, it is structurally irrelevant.
The Research Behind GEO: Who Wrote It, Where It Was Published, and Why It Was Needed
The foundational paper was authored by six researchers: Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande. This was not a single-lab effort. It emerged from a collaboration between Princeton University, IIT Delhi, the Allen Institute for AI, and Georgia Tech, which is why the field often refers to it as “the Princeton GEO paper.”
That multi-institution pedigree matters. GEO did not arrive as a marketing agency’s whitepaper or a vendor’s thought-leadership piece. It arrived through rigorous academic channels. Submitted to arXiv in November 2023, the paper was later formally presented at KDD 2024, the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining held in Barcelona, appearing on pages 5 through 16 of the proceedings.
The research filled a specific and urgent gap. As generative engines like ChatGPT, Perplexity, and Google AI Overviews began handling queries end to end, the entire SEO playbook built around click-through traffic became structurally misaligned with how AI systems surface information. Optimizing for a ranking position is meaningless if the user never sees a list of links.
The paper framed this as a “three-stakeholder problem.” Generative engines benefit users, who receive better and faster answers, and they benefit the engine itself, which captures more traffic and engagement. However, they severely disadvantage content creators, who lose organic traffic because users no longer need to click through to source websites.
This is not abstract. Ahrefs data shows AI Overviews reduced click-through rates for top-ranking Google content by 58%. In May 2025, 69% of news-related Google searches resolved without a single click to any website. The academic problem the paper identified is now a measurable business crisis.
The Formal Definition of GEO: What the Math Actually Means for Content Teams
The paper defines GEO with precision that practitioner guides never bother to reproduce. GEO is a transformation, expressed as f: W → W′, where original website content W is converted into optimized content W′ that carries a higher likelihood of visibility in generative engine responses.
Crucially, the authors framed this as a black-box optimization problem. The internal ranking logic of generative engines is unknown. Unlike Google’s partially documented ranking signals, no one outside the model developers knows exactly how an LLM decides what to cite. GEO therefore had to be approached empirically, through controlled experiments, rather than by reverse-engineering a published algorithm.
This contrasts sharply with classical SEO, which operates on known and often documented signals: backlinks, keyword density, page speed, and mobile-friendliness. GEO instead operates on probabilistic content characteristics that influence LLM citation behavior. Content teams are no longer optimizing for a crawler’s index; they are optimizing for an LLM’s in-context retrieval and synthesis process, a fundamentally different target.
The terminology landscape remains messy. Related terms include Answer Engine Optimization (AEO), Large Language Model Optimization (LLMO), Artificial Intelligence Optimization (AIO), and AI SEO. As Wikipedia notes, no consensus definition distinguishing these terms had been established in academic literature as of early 2026. The formal definition matters because it clarifies that GEO is not a rebranding of SEO. It is a structurally distinct optimization problem requiring different measurement frameworks and different content interventions. For a deeper look at what is generative engine optimization and how it differs from traditional approaches, the distinction between crawler-based and LLM-based optimization is the essential starting point.
GEO-Bench: Understanding the Experimental Infrastructure
The paper’s core experimental contribution was GEO-bench, a large-scale benchmark of 10,000 queries drawn from nine diverse sources: Bing, Google, Oxford All Souls exam questions, the LIMA reasoning dataset, debate questions, trending Perplexity.ai queries, Reddit, and GPT-4-generated queries.
The source diversity reveals the intended scope. This benchmark was not built to test simple factual lookups alone. By including Oxford examination questions and reasoning datasets, the researchers designed it to span everything from straightforward informational searches to complex reasoning tasks. The queries covered 25 domains, tagged across 7 categories using GPT-4. That domain breadth is what gives the findings their validity across real-world use cases.
The dataset is publicly available on Hugging Face as GEO-Optim/geo-bench, which means independent researchers can reproduce and validate the results. Each query was paired with the top-5 Google search results, creating a realistic retrieval context that mirrors how generative engines actually access source content.
That structure, however, contains the single most important limitation of the entire study, one the 2026 critical survey later flagged. The experimental setup assumes the source is already present in the retrieval context. In other words, the reported gains measure optimization within a fixed context; they do not measure organic discoverability from scratch. This distinction reshapes how the findings should be applied, and it is a point most practitioner content ignores entirely.
The Nine GEO Strategies: What the Paper Actually Tested
The paper tested nine specific, named content optimization strategies. This precision is what separates the research from the vague “best practices” that dominate marketing blogs. Each strategy was evaluated independently against an unoptimized baseline.
The nine strategies were:
- Authoritative Tone
- Statistics Addition
- Keyword Stuffing
- Cite Sources
- Quotation Addition
- Easy-to-Understand
- Fluency Optimization
- Unique Words
- Technical Terms
To evaluate each one, the researchers developed a measurement framework centered on Position-Adjusted Word Count (PAWC). PAWC weights word contributions by their position in the AI-generated answer, giving more weight to content cited earlier in the response. The logic is intuitive: being cited in the first sentence of an AI answer is worth far more than being buried in a closing clause.
The framework also included a Subjective Impression Score, combining word count, citation position, and GPT-3.5-based quality assessments, along with G-Eval. These metrics have since been critiqued for lacking direct behavioral or economic interpretation, a fair criticism the field is still working to resolve. Importantly, validation was not purely theoretical. The findings were tested on Perplexity.ai, a commercially deployed generative engine with millions of active users.
Strategy Performance Data: The Numbers Practitioners Are Missing
The practitioner guides consistently omit the actual performance numbers.
The top-performing strategies were Statistics Addition, Cite Sources, and Quotation Addition, each achieving 30 to 40% relative improvements in PAWC visibility metrics versus unoptimized baseline content. The reason these three outperform is coherent: generative engines are trained to synthesize authoritative, verifiable information, and content that provides quantified claims, attributed sources, and expert quotations signals exactly the type of information LLMs are optimized to surface.
Mid-tier performers were the stylistic strategies. Fluency Optimization and Easy-to-Understand produced meaningful but smaller gains of 15 to 30%. This suggests generative engines value both the substance of content and the quality of its presentation, though substance wins.
Then there is the critical finding: Keyword Stuffing showed almost no measurable lift. This is a direct inversion of its role in classical SEO, and it carries profound implications for content teams still applying traditional logic to generative contexts.
The combination effect was equally striking. Fluency Optimization combined with Statistics Addition outperformed any single GEO strategy by more than 5.5%, as detailed in analyses like those from The GEO Community. Strategy combination produces superadditive effects. Real-world validation on Perplexity.ai confirmed visibility improvements of up to 37%.
One caveat deserves emphasis: GEO effectiveness varies significantly by domain. The paper explicitly states that no single universal strategy works across all query types, directly contradicting the universal advice most guides offer.
The Keyword Stuffing Finding: Why Classical SEO Logic Fails in Generative Contexts
The keyword stuffing result deserves focused attention because it is the most counterintuitive and the most consequential for anyone trained in traditional SEO.
The mechanism explains everything. In classical SEO, keyword stuffing signals relevance to a crawler’s index. The crawler counts terms, assesses density, and infers topical relevance. In generative contexts, however, the LLM is not indexing the content; it is synthesizing it. Keyword density is irrelevant to whether the model finds a passage credible, authoritative, or citable.
Generative engines evaluate content for informational value and source credibility signals, not keyword frequency. This means the entire keyword-density optimization paradigm is structurally misaligned with how LLMs process and cite content. The business implication is uncomfortable: organizations that invested heavily in keyword-optimized content may possess a large library that is well-indexed by traditional search but poorly positioned for generative citation.
Follow-up research reinforced this. A 2026 factorial experiment by Vishwakarma et al., spanning 252,000 trials across six LLMs and eighteen factors, found that relevance and position within the retrieval context are the primary determinants of first citation. This redirects GEO strategy toward upstream discoverability stages rather than keyword density.
The takeaway is not that keywords are dead. It is that keyword optimized content generation must now serve topical relevance and genuine content quality rather than density metrics.
What the Paper Reveals About Who Benefits Most from GEO
One of the paper’s most encouraging findings for smaller players: GEO is especially beneficial for lower-ranked websites, which see disproportionately larger visibility gains from optimization compared to already high-ranking sources.
The mechanism is democratizing. In a generative context, a well-optimized piece of content from a lower-authority domain can be cited alongside, or even instead of, a high-authority domain, provided its content structure better matches what the LLM is synthesizing. This dynamic is largely absent from traditional PageRank-based SEO, where domain authority creates entrenched advantages.
For growth-stage companies, this is a structural opportunity. Organizations that cannot compete on domain authority in traditional search can compete in generative search if they optimize content for citation rather than ranking. The three-stakeholder problem cuts both ways here: content creators who do not adapt face accelerating traffic loss as generative engines absorb query resolution, making GEO effectively non-optional for content-dependent businesses.
The scale of the opportunity, and the cost of inaction, is substantial. Perplexity AI processed 780 million queries in May 2025 alone, a 239% increase from August 2024. AI-sourced traffic surged 527% year over year.
The 2026 Critical Survey: How the Field Has Evolved Since the Foundational Paper
In July 2026, a critical survey published as arXiv:2607.14035 reviewed 45 GEO studies published between November 2023 and July 2026. It stands as the most comprehensive academic synthesis of the field to date.
Its headline conclusion about field maturity is sobering: GEO has expanded rapidly, but terminology, metrics, and evidence standards remain heterogeneous. Practitioners cannot assume all GEO advice is grounded in comparable evidence, because it is not.
The survey’s most valuable contribution was contextualizing the foundational paper’s famous 40% gain figure. That number is valid within its experimental setting, where the source is already present in a fixed retrieval context. It does not establish organic discoverability or durable traffic effects. The 40% gain answers one question: “If my content is retrieved, will the optimized version be cited more?” It does not answer the more important question: “Will my content be retrieved in the first place?”
That distinction redirects strategy toward upstream indexing and retrieval. The survey also identified the field’s open problems: standardized benchmarks, causal evidence of traffic effects, and domain-specific optimization frameworks all remain active research frontiers. Acknowledging these limitations, rather than papering over them, is precisely what separates sophisticated GEO strategy from surface-level advice.
Follow-Up Research: How Academic GEO Has Advanced Beyond the Original Paper
The original paper was a foundation, not a ceiling. The follow-up research directly addresses real-world implementation challenges the 2023 study did not fully solve.
IF-GEO: Solving the Multi-Query Optimization Problem
IF-GEO (arXiv:2601.13938, January 2026) tackles the most practically significant limitation of the original work: a single page must rank for hundreds of conflicting queries simultaneously, not just one query in a fixed context.
Its “diverge-then-converge” framework first identifies conflicting optimization signals across multiple queries (diverge), then synthesizes a unified content strategy that resolves those conflicts without sacrificing visibility for any individual query (converge). Consider the business reality: a SaaS product page must be cited when users ask about pricing, features, comparisons, use cases, and integrations. Each query type may favor a different content structure, creating genuine optimization conflicts. IF-GEO, developed at the University of Science and Technology of China, confirms that GEO research is now a global academic field.
E-GEO: Red-Teaming and the E-Commerce Testbed
E-GEO (arXiv:2511.20867, November 2025, updated July 2026) is the first systematic GEO framework designed specifically for e-commerce.
Its most intellectually interesting dimension is red-teaming: E-GEO tests manipulation tactics against genuine content improvement. This adversarial evaluation helps distinguish durable GEO gains from short-term exploits that generative engines will eventually penalize. This matters because e-commerce product pages have fundamentally different structures than editorial content. Strategies that work for blog posts may not transfer to product descriptions or category pages. E-GEO adds a domain-specific benchmark to the ecosystem, directly answering the original paper’s call for domain-specific optimization methods. For brands exploring SEO automation for ecommerce, the E-GEO findings underscore why domain-specific frameworks matter more than generic optimization checklists.
Role-Augmented Intent-Driven GEO and Multi-Agent Frameworks
Role-Augmented Intent-Driven GEO (arXiv:2508.11158, August 2025, updated March 2026) extends GEO into LLM-powered and RAG-based generative search architectures. Its core idea is that different query intents (informational, navigational, transactional, and investigational) require different content roles, and content structure should be matched to the intent role the generative engine is fulfilling.
Beyond that, a multi-agent GEO framework (arXiv:2604.19516) represents the field’s movement toward automated, system-level GEO execution rather than manual content interventions. The trajectory is clear: the field is moving from “what content changes improve citation?” toward “how do we build systems that continuously optimize content for generative citation at scale?” That shift validates the automated, agentic approach.
From Academic Theory to Business Execution: The Implementation Bridge
The research has direct, actionable implications. The paper’s top three strategies map cleanly to content production practices:
- Statistics Addition: Every substantive claim should be quantified wherever possible.
- Cite Sources: Every quantified claim should be attributed to a named source.
- Quotation Addition: Expert perspectives should be embedded as direct quotations rather than paraphrased summaries.
The combined-strategy finding becomes an operational standard. Because Fluency Optimization plus Statistics Addition outperforms any single strategy, every piece of content should pass both a readability review and a quantification audit before publication.
The domain-specificity finding rules out one-size-fits-all templates. Query category analysis should precede content strategy, with different optimization profiles applied across different topic domains. The upstream discoverability gap, surfaced by the 2026 factorial experiment, means GEO must begin at the indexing and retrieval stage, ensuring content is structured for topical authority and internal linking rather than isolated page optimization. Building a content ecosystem rather than individual blog posts is precisely the structural approach the upstream retrieval evidence supports.
Measurement remains a challenge. PAWC and Impression Score are research metrics, not standard dashboard figures. Practitioners need proxy metrics: AI Overview citation tracking, AI-referred traffic in GA4, and brand mention monitoring in LLM outputs. The context is now mainstream. AI Overviews appear on 48% of Google queries, up from 31% in February 2025. GEO-aligned structure is no longer an advanced tactic; it is a baseline requirement.
Why Automated GEO Execution Is the Only Scalable Response to These Findings
The research demands applying multiple optimization strategies across every piece of content, at publishing cadences manual teams cannot sustain. AI content platforms already produce 4.6x more content per marketer per month, and teams at Level 3 AI maturity produce 5 to 10x more content at 75 to 85% lower cost per article. Manual GEO application is simply not competitive.
The IF-GEO multi-query finding compounds this challenge for manual teams. Optimizing a single page for hundreds of conflicting queries requires systematic conflict identification and resolution that exceeds human content-team capacity. Because generative engines evolve continuously, content must also be re-evaluated and updated on an ongoing basis, a process that demands automated monitoring and revision workflows.
This is precisely where KOZEC operates. As an AI-powered content automation platform with built-in GEO optimization, KOZEC translates the academic findings, including Statistics Addition, Fluency Optimization, source citation, authoritative tone, and topical ecosystem building, into automated production workflows that execute at scale. KOZEC’s reported performance metrics reflect this: +386% AI Overview Citation Growth and +621% Keyword Visibility Increase. Its agentic AI makes strategic decisions autonomously rather than requiring manual prompting at each step, aligning directly with the multi-agent GEO research direction that represents the field’s forward trajectory.
What Most GEO Guides Are Getting Wrong: A Research-Based Critique
Most GEO guides list tactics without explaining the experimental evidence behind them, the conditions under which they were validated, or their known limitations.
The most common misrepresentation is citing the 40% visibility gain as a universal average, when both the paper and the 2026 survey clarify it is a maximum under favorable conditions where the source is already in the retrieval context. Some rare exceptions, such as the critical breakdown from Blck Alpaca, get this right, but they remain outliers.
The most consequential omission is the failure to explain that keyword stuffing actively underperforms in generative contexts. Practitioners applying traditional SEO logic to GEO are not merely missing opportunity; they may be actively misallocating content resources.
Other gaps compound the problem. Universal GEO advice contradicts the paper’s explicit domain-specificity finding. Most guides never address how to measure GEO performance, leaving teams unable to distinguish real gains from organic fluctuation. Nearly all ignore the upstream retrieval stage, despite the factorial experiment showing position within retrieval context is a primary citation determinant. Strategy built on incomplete evidence will produce diminishing returns as generative engines grow more sophisticated. Understanding how to evaluate AI SEO software through the lens of these research standards is a practical first step toward separating evidence-based tools from those built on surface-level assumptions.
Conclusion: The Academic Foundation Is the Competitive Advantage
arXiv:2311.09735 is not merely a historical document. It is the evidence base that explains why specific GEO strategies work, under what conditions they produce measurable gains, and where their limitations lie.
The findings practitioners must internalize are precise: Statistics Addition, Cite Sources, and Quotation Addition produce 30 to 40% visibility gains; Fluency Optimization plus Statistics Addition outperforms any single strategy by 5.5%; keyword stuffing produces near-zero lift; domain-specific optimization is necessary; and the 40% figure is a ceiling under favorable conditions, not an average.
The 2026 critical survey of 45 studies confirms these findings while identifying the next frontier: organic discoverability, durable traffic effects, and standardized measurement. With AI Overviews on 48% of Google queries and 35% of US consumers using AI at product discovery, GEO-aligned content is a baseline requirement for content-dependent businesses.
Organizations that understand the research, not just the tactics, will build GEO strategies that are durable, measurable, and adaptable. That academic-to-implementation bridge is exactly where KOZEC operates: converting peer-reviewed findings into automated workflows that apply evidence-based strategies at a scale and consistency manual teams cannot match.
Ready to Apply the Research? See How KOZEC Automates GEO at Scale
Understanding the research is the first step. Applying it consistently across every published page is the harder one. KOZEC’s agentic AI platform implements the GEO strategies validated in arXiv:2311.09735, including Statistics Addition, Fluency Optimization, source citation, authoritative tone, and topical ecosystem building, as automated, production-ready content workflows.
The value proposition is straightforward: KOZEC delivers 15 to 60+ content pieces per month at $600 to $1,500/month, with GEO and SCO (Search Compliance Optimization) built into every piece. That means translating academic research into business results without requiring a content team to manually apply each optimization strategy.
Early users are seeing measurable organic traffic growth within 60 to 90 days, with setup measured in days rather than months. There are no long-term contracts and organizations can cancel anytime, which lowers the barrier for those ready to move from academic understanding to practical execution.
Schedule a demo at kozec.ai/schedule-a-demo/ to see it in action. Prefer to speak directly with the team? Call (888) 545-7090 or visit kozec.ai.
Stay In The Loop
Subscribe to our free newsletter.
Stop Managing SEO - Start Scaling It
Let KOZEC handle strategy, content, and execution - so you can focus on growth.
Automated SEO content for growing agencies.
KOZEC helps agencies, consultants, and growing brands publish high-quality SEO content on autopilot — so your site ranks higher and converts more visitors.
Managing SEO content for many client websites doesn’t scale with traditional methods. Writers are expensive and inconsistent, keyword research is time-consuming, and publishing requires multiple manual steps. As agencies grow, maintaining both quality and consistency becomes increasingly difficult. KOZEC (Keyword Optimized Zero Effort Content) solves this by automating analysis, keyword discovery, content creation, and publishing—so your clients get reliable SEO content while your team focuses on growth.
Increase organic traffic without manual content creation
Publish keyword-optimized posts automatically to WordPress
Turn SEO into a predictable, scalable growth channel

