Published on: May 30, 2026
โ€ข
7 min read

AEO, GEO, LLMO Explained – Their Origin and what researchers say

Written By

Ayush Verma

Talk to an Expert โ†’

One thing before the definitions.

Somebody is going to sell you one of these this quarter. They’ll open with how everything changed in the last eighteen months, and how you need a new discipline with a new retainer attached. I’ve been on the other side of that pitch and I’ve also written it, years ago, about something else. So this piece is mostly definitions and sources. Where I have an opinion I’ll say so and you can ignore it.

– Ayush

Three acronyms. One practice. Let’s do the definitions first, then look at what the research says, because only one of these has research behind it at all.

What is AEO?

Answer Engine Optimization. It came out of the SEO community around 2018, with no founding paper and no formal launch.

At the time it meant featured snippets and voice assistants. Position zero was worth real money. Structuring content so a machine could lift a direct answer out of it was sensible work.

The term has since been repurposed for AI search, which is fair enough. The underlying problem barely changed: your content has to survive being summarised by something else.

What is GEO?

Generative Engine Optimization. Formally defined in November 2023 by researchers at Princeton, IIT Delhi, Georgia Tech and the Allen Institute for AI.

This is the only one of the three with an academic origin. The paper was presented at KDD 2024, the ACM SIGKDD conference. Authors are Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande.

I’ll come back to what it found, because it’s the most useful thing in this whole conversation and almost nobody selling GEO has read past the abstract.

What is LLMO?

Large Language Model Optimization. Coined by Jina.ai in December 2022, a few weeks after ChatGPT launched, in a post titled “SEO is Dead, Long Live LLMO.”

The obituary was premature. The term stuck anyway.

You’ll also run into AIO, GAIO, AISO and AI SEO around the edges. Treat those as variations, not distinctions.

What’s the difference between AEO, GEO and LLMO?

AEOGEOLLMO
Stands forAnswer Engine OptimizationGenerative Engine OptimizationLarge Language Model Optimization
OriginSEO community, around 2018Princeton et al., Nov 2023Jina.ai, Dec 2022
Founding paperNoneAggarwal et al., KDD 2024None
Originally aboutFeatured snippets, voice assistantsCitation inside generated answersVisibility in LLM responses
Honest scope todayThe same practiceThe same practiceThe same practice

Look at the bottom row. That’s the article in three words.

The differences are historical. They tell you who coined the term, not what the work involves.

So why does one practice have three names?

Because a term that sounds new justifies a new line item on an invoice. A term that sounds like a subset of SEO does not.

Watch who adopted which one. Enterprise platforms and venture-backed startups went heavily for GEO, because it signals novelty and borrows credibility from a real paper. Agencies picked whichever term was distinct enough to put on a service page.

That’s the whole mechanism. It explains why the vocabulary keeps churning while the work changes far more slowly.

I should say I’m not neutral here. I work in this space and I have a commercial interest in it. Which is exactly why I’d rather point you at the primary research than at my framing of it.

What the GEO research found

The setup matters before you quote anything from it. The team built GEO-bench, a benchmark of diverse queries across multiple domains. Each query was paired with the sources a generative engine would draw on. Then they tested nine content modifications and measured which ones moved citation visibility.

What worked. Adding statistics. Adding quotations from credible sources. Citing sources. Improving fluency.

The best methods improved on baseline by 41% on Position-Adjusted Word Count and 28% on Subjective Impression. Cite Sources and Fluency Optimization both landed around 28% on the first metric.

Fluency is worth pausing on. It adds no new information at all. Clearer prose is simply easier for a model to parse, summarise and attribute.

What failed, and this is the more interesting half.

Keyword stuffing didn’t just fail to help. It scored 8% below the unmodified baseline, and 10% below when validated against Perplexity.

Two decades of an industry’s founding tactic. The machine now marks you down for it.

An authoritative tone did nothing. The researchers expected a confident, persuasive voice to lift visibility. It didn’t move. Sounding certain is free, and the machine has priced that in.

Read the working list again. Cite your sources. Quote experts. Use real numbers. Write clearly.

That reads like a description of competent writing. I quoted it that way for months before I worked out why that reading is too comfortable.

The part the study didn’t test

It measured whether adding citations, quotations and statistics increases how often a source gets pulled into an answer.

It did not measure whether any of those citations or statistics were true.

The paper’s own examples make this hard to miss. One showcase inserts a per-capita consumption figure attributed to a research body I cannot locate. Another adds a large percentage claim with no source at all.

The caption says the methods increase visibility without adding substantial new information.

So the finding isn’t that good writing wins. It’s that generative engines reward the surface features of rigour. Numerical specificity. Quotation marks. Citation formatting.

Nothing in the study shows an engine can tell those features apart from their genuine counterparts. That’s a darker result and it’s the one the paper supports.

What other researchers say about it

The GEO paper gets quoted constantly. The work criticising it almost never does.

A 2026 benchmarking paper puts the objection plainly. The original evaluation used a single task and a single metric, word count. It never tested competitive multi-actor settings.

Which is to say, it never tested what happens when everybody optimises at once.

That matters, because it’s the question every buyer should be asking. A tactic that lifts you 40% while you’re alone in doing it is one proposition. The same tactic once your whole category has adopted it is another.

The same body of work covers adversarial ranking manipulation, where optimised token sequences can push a product to the top of a recommendation. Take it seriously. And note the framing: the researchers describe a vulnerability, not a strategy. Vulnerabilities get closed.

Why hold all of this loosely

The GEO study ran on a model generation that’s now three years old. The pipeline was built to imitate the search products of the time.

The direction is more durable than the numbers. Content that’s specific, well-sourced and clearly written gets cited more than content that isn’t. That was true before generative engines and it’ll outlast whatever replaces them.

The exact figures are a snapshot of one system at one moment. Anyone presenting them as current operating parameters either hasn’t read the methodology, or is hoping you won’t.

What to do on Monday

Win narrow questions, not broad ones. “Best CRM” is unwinnable and mostly worthless. “CRM for freight brokerages under fifty seats” is winnable. And whoever types that is much closer to buying.

Make your claims liftable. Specific numbers beat vague ones. Sourced claims beat asserted ones. A model has to extract a discrete claim and attribute it. One long persuasive argument gives it nothing to lift.

Don’t buy a service on the strength of its acronym. The terminology tells you nothing about the method. Ask what the work involves. And ask whether the person selling it can name one piece of research that isn’t their own case study.

And notice which of the working tactics cost nothing. Fluent prose and the appearance of sourcing are both things a language model produces for free. A signal that costs nothing to fake stops carrying information once enough people fake it.

Real citations to real sources, and numbers from your own data. Those are the two on that list that need someone to do the work.

Those are the half with a future.

In 2018 the answer was to structure your content so it could be extracted as a direct answer. In 2026 the answer is to structure your content so it can be extracted as a direct answer.

The acronym changed. The invoice went up. The work didn’t move nearly as much as anybody wanted it to.

For anything else, you know how to reach me.