← Knowledge LibraryChapter 14 · Statistical Inference

Experimental Design for GEONull Outputs, Missingness, Bootstrap, and Clustering

Why failed and search-off observations must remain visible, and how resampling and cluster-aware uncertainty prevent false precision.

Why failed and search-off observations must remain visible, and how resampling and cluster-aware uncertainty prevent false precision.

Null and search-off outputs are experimental outcomes

A valid response with no citation, a search-off answer, a refusal, an API timeout, and a parser failure mean different things. Removing null or zero-citation observations can inflate conditional performance; treating system failures as zero can create a false decline.

Missingness can be systematic

If one engine fails more often on complex prompts or one treatment produces longer responses that time out, missing records depend on the experimental condition. Report missingness by arm, engine, prompt family, and period, and investigate whether results change under defensible handling rules.

Bootstrap confidence intervals estimate uncertainty

Resampling can approximate the sampling distribution when analytical formulas are awkward. But the resampled unit must preserve dependence: repeatedly sampling individual rows as if they were independent can understate uncertainty.

Recognize clustered observations

Runs sharing the same prompt, paraphrase family, page, engine, or date are correlated. Cluster-aware bootstrap or hierarchical models should reflect the level at which treatment was assigned and the structure that generates repeated outcomes.

More rows do not necessarily mean more independent evidence.

Frequently asked questions

Questions about this topic

Should no-citation responses be deleted?+

No. They are part of end-to-end performance and removing them changes the estimand and denominator.

Is an API timeout the same as zero visibility?+

No. It is a failed or missing observation rather than a valid response with no mention.

What should a GEO bootstrap resample?+

The independent or clustered unit appropriate to assignment and dependence, such as prompts, pages, or prompt–engine cells.

Why are repeated runs clustered?+

They share prompts, sources, engines, dates, or treatment units and therefore may be correlated.

Source notes

References

Sources are listed in APA 7 style. Preprints are identified as such and should not be treated as peer-reviewed findings unless separately published.

  1. Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14035
  2. Kumar, P. (2026). Generative engine optimization at scale: Measuring brand visibility across AI search engines [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.20065