Tested in public · data published at the end of each test

GEO Experiments

GEO experiments are controlled tests of how ChatGPT, Perplexity, Google AI Overviews, AI Mode and Gemini choose and cite sources, run in public on this site. 6 experiments are listed: 1 running, 5 planned. Each states a hypothesis, setup, start date and engines, and its raw data is published when it ends, including null results. Filter the list by engine or status, and see the research papers each test replicates.

These are practitioner field tests on live websites. For the academic GEO literature, see the research papers section below.

Which GEO experiments are running or planned?

Six GEO experiments are listed: one running and five planned. The running test tracks whether a brand-new domain gets cited by 6 AI engines within 90 days; the planned tests cover statistics and quotations, llms.txt, answer-first passages, structured data and Reddit mentions.

Engine
Status

Showing 6 of 6 experiments

  • RunningExp 01

    Can a brand-new domain get cited by AI engines in 90 days?

    Started on launch day with zero backlinks. 50 fixed prompts checked weekly on 6 engines. Weekly log published.

    Engines: all 6 · Started: launch day · Ends: day 90

  • PlannedExp 02

    Do statistics and quotations raise citation rate?

    Replicates the 2023 GEO paper's top tactics on 20 live pages, with 20 unchanged control pages.

    Engines: ChatGPT, Perplexity · Duration: 8 weeks

  • PlannedExp 03

    Does llms.txt change how AI crawlers behave?

    Server logs 30 days before and 30 days after publishing llms.txt, split by user agent.

    Crawlers: GPTBot, ClaudeBot, PerplexityBot, others · Duration: 60 days

  • PlannedExp 04

    Answer-first vs story-first: which passage do AI Overviews lift?

    Paired pages with identical facts; one leads with the answer, one with context.

    Engines: AI Overviews, AI Mode · Duration: 6 weeks

  • PlannedExp 05

    Does structured data change citation rate?

    Article and FAQ markup added to half of a matched page set; citations compared across 8 weeks.

    Engines: AI Overviews, Gemini · Duration: 8 weeks

  • PlannedExp 06

    Do Reddit mentions change ChatGPT recommendations?

    Tracks recommendation prompts before and after a brand is discussed in relevant threads. Observes organic discussion only — no seeded or undisclosed posts.

    Engines: ChatGPT, Perplexity · Duration: 12 weeks

A GEO experiment is a controlled test of how AI engines choose, quote and cite sources, with its hypothesis, setup, start date and engines stated before it begins. Experiment 01 is the only test running; the other five start in sequence, and each card’s status changes to “Running” on its start date. No experiment on this page has a result yet, so none is claimed.

How is each GEO experiment run?

Each GEO experiment runs on fixed prompts, a fixed schedule and repeated runs, with a control group wherever the design allows. The hypothesis, setup, start date and engines are published before the first run, and the full dataset is published when the test ends.

The protocol has seven parts, and every experiment states each one:

  1. Hypothesis — the single change being tested and the effect expected, written before the first run.
  2. Setup — the pages, prompts and changes involved, described precisely enough to repeat.
  3. Start date — fixed in advance, so the before and after periods are defined.
  4. Engines — the AI engines or crawlers observed, named on each card.
  5. Runs — every prompt run more than once per check under fixed conditions: logged-out session, fixed country and language, same time window.
  6. Control — unchanged pages or periods compared with the changed ones, wherever the design allows.
  7. Data — every logged answer published at the end of the test, including null results.

The answer logging follows the same prompt tracking method used for AI visibility baselines: each answer is recorded with the brand named, the brand linked, the cited URL, the competitors named and the position of the first mention.

Repeated runs and controls are not optional extras. A 2026 critical survey of 45 GEO studies reports substantial run-to-run variability in AI answers and recommends repeated measurements, paraphrases, controls, human validation, and attention to multi-actor interference as the basic protocol. A test that checks each prompt once and has no control group cannot separate a real effect from an engine’s normal drift.

What has published research found about generative engine optimization?

Published research finds that rewriting content changes how often already-retrieved sources are cited, by up to 40% in the original 2023 benchmark. A 2026 survey of 45 studies found no technique with a proven, lasting effect on being discovered in the first place.

Seven papers define the current state of the field. Each is listed with its arXiv identifier, first posting date and the finding stated in its abstract.

  • arXiv 2311.09735 · 16 Nov 2023 · KDD 2024

    GEO: Generative Engine Optimization

    Aggarwal et al. (6 authors). Introduces the term and the GEO-bench benchmark; content rewrites raise source visibility by up to 40%.

  • arXiv 2510.11438 · 13 Oct 2025 · preprint

    What Generative Search Engines Like and How to Optimize Web Content Cooperatively

    Wu et al. (4 authors). AutoGEO learns engine preference rules and rewrites content while preserving search utility.

  • arXiv 2602.02961 · 3 Feb 2026 · industry deployment

    Generative Engine Optimization: A VLM and Agent Framework for Pinterest Acquisition Growth

    Zhang et al. (6 authors). Pinterest's production system reports 20% organic traffic growth.

  • arXiv 2603.20213 · 2 Mar 2026 · preprint

    AgenticGEO: A Self-Evolving Agentic System for Generative Engine Optimization

    Yuan et al. (6 authors). Outperforms 14 baselines across 3 datasets on 2 engines.

  • arXiv 2603.09296 · 10 Mar 2026 · preprint

    Diagnosing and Repairing Citation Failures in Generative Engine Optimization

    Tian et al. (5 authors). Classifies citation failure modes; AgentGEO gains over 40% relative citation rate while changing about 5% of content, and generic optimization can harm long-tail content.

  • arXiv 2604.19516 · 21 Apr 2026 · ACL 2026 Findings

    From Experience to Skill: Multi-Agent Generative Engine Optimization via Reusable Strategy Learning

    Wu et al. (10 authors). MAGEO and the MSME-GEO-Bench benchmark; engine-specific preference modelling drives gains on 3 engines.

  • arXiv 2607.14035 · 15 Jul 2026 · survey of 45 studies

    Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)

    Martinez. Gains depend on the source already being in the model's context; answers show substantial run-to-run variability; no technique has a stable cross-platform effect on discoverability.

Every study in this list runs on fixed benchmarks or a single company’s platform. The experiments on this page carry the same questions to live websites and public AI engines. For the depth behind the original paper — its setup, its 9 methods and why its gains reversed for top-ranked sources — see GEO research put into practice.

Which research question does each experiment test?

Each experiment carries one question from lab research or official guidance onto live AI engines. Experiment 02 replicates the 2023 paper's statistics and quotation tactics; experiments 03 and 05 test statements in Google's own AI search guidance; experiment 01 tests discoverability.

The source each experiment tests
ExperimentQuestionSource it tests
Exp 01Can a new domain get cited within 90 days?2026 survey: no technique shows a stable, cross-platform effect on discoverability
Exp 02Do statistics and quotations raise citation rate?2023 GEO paper: Statistics Addition and Quotation Addition among the top methods
Exp 03Does llms.txt change crawler behaviour?Google: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search."
Exp 04Do AI Overviews lift answer-first passages more often?Common GEO advice that each section should open with its answer, untested on live AI Overviews
Exp 05Does structured data change citation rate?Google: structured data isn't required for its generative AI features
Exp 06Do Reddit mentions change ChatGPT recommendations?Google: AI features reflect forum discussions; inauthentic mentions are not helpful

The table shows that two experiments replicate research findings, three test statements from Google’s guidance, and one tests advice that is widely repeated but unmeasured on live engines. Experiment 01 exists because the survey’s gap is the question most site owners ask: whether a site with no history can be found at all.

The same protocol runs on client sites.

How do these experiments connect to client audits?

The experiments and the client audit use the same measurement protocol: fixed prompts, repeated runs on 6 engines and a logged baseline. Findings from finished experiments feed the fix plans delivered in audits.

When a site owner requests the AI visibility audit service, the baseline is built exactly as it is here, so the client’s numbers and the experiment data can be compared directly. Until an experiment finishes, its question is treated in audits as open, not settled.

Frequently asked questions

Are failed experiments published?

Yes. Null and negative results are published with the same detail as positive ones: the hypothesis, the setup, the dates, the engines and the full dataset. A test that shows no effect is still evidence, and leaving it out would bias the record.

When does the first result appear?

The first result appears when Experiment 01 ends, 90 days after the site's launch. Its weekly log is published while it runs, but no conclusion is drawn before day 90. Planned experiments publish results when their own test periods end.

Related guides

Sources

  1. Aggarwal et al. — GEO: Generative Engine Optimization (arXiv 2311.09735; KDD 2024)
  2. Wu et al. — What Generative Search Engines Like and How to Optimize Web Content Cooperatively (arXiv 2510.11438)
  3. Zhang et al. — Generative Engine Optimization: A VLM and Agent Framework for Pinterest Acquisition Growth (arXiv 2602.02961)
  4. Yuan et al. — AgenticGEO: A Self-Evolving Agentic System for Generative Engine Optimization (arXiv 2603.20213)
  5. Tian et al. — Diagnosing and Repairing Citation Failures in Generative Engine Optimization (arXiv 2603.09296)
  6. Wu et al. — From Experience to Skill: Multi-Agent Generative Engine Optimization via Reusable Strategy Learning (arXiv 2604.19516)
  7. Martinez — Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (arXiv 2607.14035)
  8. Google Search Central — Optimizing your website for generative AI features on Google Search