What makes content easier for generative search to cite?

A visual explanation of the relationship between perplexity, semantic relevance, source position, AI polishing, citation diversity, and user outcomes.

By dotSuper Research DeskPublished Sep 1, 2026Reviewed Sep 1, 202611 min read
Search & discoveryWebsite analysis, RAG experiment, and randomized trial from arXiv:2509.14436Updated Sep 1, 2026

/ THE SHORT ANSWER

No. The study found that lower-perplexity source text and stronger semantic similarity to the query were associated with higher citation probability, and controlled AI polishing increased citation breadth. The defensible practice is to make original evidence clearer and easier to retrieve, not to flatten every page into generic AI prose.

Key takeaways
  • 01A one-standard-deviation decrease in perplexity moved predicted citation probability from 47% to 56% in the website-level model.
  • 02In controlled RAG tests, semantic similarity had a positive relationship with citation while source position showed a negative relationship as content moved deeper in the document.
  • 03General polishing added 1.0744 citations per query and citation-directed polishing added 2.1059 in the paper’s controlled experiment.
  • 04The study does not justify scaled AI content or bland prose. Google warns that low-value scaled pages may violate spam policies.

/ dotSuper point of view

Preserve the insight. Reduce the retrieval friction around it.

The observed citation-probability shift

The paper uses perplexity as a measure of how predictable a text is to a language model. Across 98,477 unique websites, lower perplexity correlated with greater likelihood of appearing in Google AI Overview citations. A decrease of one standard deviation, equal to 9.52 perplexity points in the sample, moved predicted citation probability from 47% to 56%.

The authors found no equivalent perplexity effect in their conventional organic-ranking robustness analysis, which suggests the mechanism may be more specific to generative retrieval than to classic ranking.

PREDICTED PROBABILITY

Nine percentage points separate the two modeled states.

Website-level predicted citation probability at the sample mean and after one standard deviation lower perplexity.
Sample mean47%
Mean website state
1 SD lower perplexity56%
9.52 points lower in the sample

This is an observational model result. It does not mean any nine-point copy edit will create a nine-point citation lift.

View the chart data
Modeled citation-probability comparison
StatePredicted probabilityChange
Sample mean47%Reference
One standard deviation lower perplexity56%+9 percentage points

What happened in the controlled polishing experiment

The authors created three versions of source content for each of 4,060 queries: original, generally polished, and polished with a citation objective. The controlled RAG pipeline then measured citation count and similarity among selected sources.

Both treatments expanded the number of citations. Citation-directed polishing produced the larger increase and the larger decrease in similarity among cited sources. The result suggests that clearer source material can broaden the eligible citation pool inside that experimental system.

TREATMENT EFFECTS

More citations, with greater source diversity.

Regression coefficients relative to original source content in the controlled RAG experiment.
General polish+1.0744

Additional citations per query

Citation-directed+2.1059

Additional citations per query

General similarity−0.0319

Change among cited sources

Directed similarity−0.0722

Change among cited sources

All four reported effects were statistically significant at p < 0.01 in the paper. The setting was a controlled Gemini RAG pipeline.

View the chart data
Content-polishing treatment effects
TreatmentCitation count effectSimilarity effect among cited sources
General AI polishing+1.0744−0.0319
Citation-directed AI polishing+2.1059−0.0722

The user experiment adds an important second-order effect

The randomized trial retained 147 participants. Participants using the polished source set produced submissions with higher information diversity overall and completed the task 1.5942 minutes faster on average. The education subgroup results differed: participants without college degrees showed the larger information-diversity effect, while college-educated participants showed the larger time reduction.

RANDOMIZED TRIAL

Different users benefited in different ways.

Treatment coefficients reported for the full sample and education subgroups.
All participants+0.3554

Information-diversity effect

All participants−1.5942

Minutes spent on task

No college degree+0.6123

Information-diversity effect

College degree−2.9186

Minutes spent on task

The task concerned a school smartphone-ban policy and used 65 source websites. External validity to commercial search tasks is not established.

View the chart data
Selected randomized-trial treatment effects
GroupInformation effectTime effect
All participants, n=147+0.3554−1.5942 minutes
College degree, n=68+0.0690, not significant−2.9186 minutes
No college degree, n=79+0.6123−0.3858 minutes, not significant

A safer editorial practice for dotSuper

Use AI as an editor around human-owned evidence. Begin with a first-hand observation, dataset, interview, operating artifact, or explicit point of view. Then improve the answer structure, definitions, headings, tables, source placement, and internal linking without removing the specificity that makes the page worth citing.

Important claims should appear in text, not only inside motion graphics or images. The direct answer should arrive early, while supporting context, limitations, and methodology remain available for readers who need depth.

  • Lead with the answer and state the scope in the same screen.
  • Place the evidence close to the sentence it supports.
  • Use comparison tables when the decision depends on several explicit criteria.
  • Keep original examples, numbers, and limitations intact during editing.
  • Reject pages that merely restate public information without a dotSuper contribution.

What this page cannot conclude

  • 01Both papers are preprints. Their findings should be treated as evidence to test, not as a settled ranking formula.
  • 02The GEO engine study collected data in August 2025. Models, retrieval systems, interfaces, and citation behavior can change quickly.
  • 03Observed citation patterns do not prove that changing one page element will cause an engine to cite that page.
  • 04Google states that there is no special schema, file, or content format required for its generative AI features. Core SEO and useful, original content remain the foundation.
  • 05Perplexity depends on the model and tokenization used to calculate it. It is not a universal readability score.
  • 06The controlled polishing results came from a specific Gemini RAG setup and should not be assumed to transfer at the same magnitude to every engine.

Sources

  1. 01When Content is Goliath and Algorithm is David: The Style and Semantic Effects of Generative Search EnginearXiv · accessed Sep 1, 2026
  2. 02Optimizing your website for generative AI features on Google SearchGoogle Search Central · accessed Sep 1, 2026
  3. 03Google Search's guidance on using generative AI content on your websiteGoogle Search Central · accessed Sep 1, 2026
BUILD A MEASURABLE DISCOVERY SYSTEMWhat makes content easier for generative search to cite?

/ APPLY THE THINKING

Turn evidence into discoverable demand.

dotSuper connects original research, search foundations, AI visibility monitoring, content operations, and qualified lead routing as one accountable system.

Question for the working sessionShould a business rewrite content to sound more like an LLM?

/ Topic-led working session · What makes content easier for generative search to cite?

Turn this question\ninto a useful first move.

Bring how this question currently shows up in your business: “Should a business rewrite content to sound more like an LLM?” We’ll test the page’s evidence against your context and define the smallest useful next move.

Live availability from ceo@dotsuper.net Your time zone · Local time
  1. 01Bring the contextWhere this issue shows up in the work.
  2. 02Test the relevanceUse the evidence against your reality.
  3. 03Choose the next moveOne accountable action, clearly owned.
Live availability
  1. Date
  2. Time
  3. Booked

Syncing live times