← Knowledge LibraryChapter 14 · Experimental Controls

Experimental Design for GEORepeated Runs, Paraphrases, and Factorial Controls

Why generative variability demands repeated observations, intent-preserving paraphrases, and structured factorial designs.

Why generative variability demands repeated observations, intent-preserving paraphrases, and structured factorial designs.

Simple before-and-after designs confound treatment with time

The difference between one pre-treatment week and one post-treatment week contains the intervention, platform drift, competition, and random noise. More observations improve precision but do not repair a design that lacks a credible comparison.

Repeated runs estimate a distribution

One prompt can yield different sources, brands, citations, and ordering across executions. Repeated runs reveal average outcomes, variance, flip rates, and uncertainty. The required number depends on baseline variance, target effect size, clustering, and the cost of false conclusions.

Prompt paraphrases are experimental controls

Several formulations of the same information need test whether an effect survives wording changes. “Best CRM,” “recommend a CRM,” and “which CRM should I choose?” should be treated as related variants—not pooled as independent evidence without preserving their intent group.

Factorial designs expose interactions

A factorial design can cross treatment with engine, prompt family, paraphrase, and time. It estimates not only an average effect but whether the intervention behaves differently across platforms or information needs.

Example design

Treatment × Engine × Prompt family × Paraphrase × Run

Frequently asked questions

Questions about this topic

Why are repeated runs necessary?+

Generated answers are stochastic, so one execution cannot estimate a stable rate or its uncertainty.

How many repetitions are enough?+

There is no universal number; it depends on variability, desired precision, effect size, clustering, and budget.

Why use paraphrases?+

They test whether the effect belongs to the underlying intent instead of one favorable wording.

What is a factorial GEO experiment?+

A design that crosses controlled factors such as treatment, engine, prompt type, paraphrase, and time to estimate main effects and interactions.

Source notes

References

Sources are listed in APA 7 style. Preprints are identified as such and should not be treated as peer-reviewed findings unless separately published.

  1. Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14035
  2. Kumar, P. (2026). Generative engine optimization at scale: Measuring brand visibility across AI search engines [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.20065