← Knowledge LibraryChapter 4 · Measurement Methods

Discoverability and RetrievalMeasuring Discoverability and Retrieval

How to measure two hidden upstream stages without confusing visible citations for the complete candidate set.

How to measure two hidden upstream stages without confusing visible citations for the complete candidate set.

Potential access and realized selection are different

Discoverability asks whether a source or entity is technically and semantically eligible to be found. Retrieval asks whether the system actually selected it for a particular query.

A technically discoverable source that never appears for a prompt family and one that repeatedly enters candidate sets represent very different strategic situations.

Discoverability is measured through partial signals

Commercial engines rarely expose their complete index, so practical indicators include crawler accessibility, observable index presence, named-brand recognition, domain recurrence, unbranded mention rate, and visible-source presence.

Each captures only part of the construct. None should be described as direct access to a hidden retrieval probability unless the engine exposes it.

Open pipelines reveal more than commercial systems

In a controlled RAG system, researchers can log candidate sets, retrieval scores, rank, and top-k inclusion. Commercial systems may expose only citations or source cards, which occur downstream of retrieval.

Citation absence does not prove retrieval absence.

A source may have been retrieved, excluded from final context, used without attribution, or ignored during generation.

Locate the largest upstream drop-off

The retrieval funnel moves from the accessible web to crawlable pages, indexed and eligible pages, query-relevant candidates, the retrieved set, reranked sources, and top-k context. Measurement should identify where the largest loss occurs before proposing an intervention.

Frequently asked questions

Questions about this topic

How is discoverability different from retrieval?+

Discoverability describes the potential to be found; retrieval is the realized selection of a source for a specific information need.

Can commercial AI retrieval be measured directly?+

Usually only partially. Visible citations and source cards are downstream proxies, while full candidate sets and scores are commonly hidden.

Does no citation mean the page was not retrieved?+

No. It may have been retrieved but dropped during reranking or context selection, used without visible attribution, or ignored during generation.

Source notes

References

Sources are listed in APA 7 style. Preprints are identified as such and should not be treated as peer-reviewed findings unless separately published.

  1. Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14035
  2. Zhang, K., He, X., & Yao, J. (2026). From citation selection to citation absorption: A measurement framework for generative engine optimization across AI search platforms [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.25707
  3. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., & Narasimhan, K. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5–16). Association for Computing Machinery. https://doi.org/10.1145/3637528.3671900