Why being retrieved is not enough: limited context, passage selection, and source position shape citation opportunity.
Only part of the candidate pool moves forward
Top-k refers to the subset of retrieved or reranked material retained for downstream generation. Even with large context windows, systems still decide which sources and passages deserve limited processing and answer space.
A page outside this retained set is technically retrieved but functionally absent from the generator.
Context allocation controls exposure
Context allocation describes how many tokens, which passages, and how much material from each source are supplied to the model. Two retrieved pages can therefore receive radically different opportunities to influence the final answer.
Martinez treats context exposure as its own visibility component because retrieval and meaningful context allocation are distinct outcomes.
Position inside context matters
Controlled research indicates that query-document relevance and context position are among the more consistently supported determinants of citation behavior. Information placed in a high-priority part of the context may receive more attention than equally valid evidence buried among other passages.
Retrieval and context selection are upstream constraints on citation opportunity.
This is why post-retrieval writing improvements cannot fully compensate for never entering the high-priority context.
Questions about this topic
What does top-k mean?+
It is the limited subset of the highest-priority retrieved or reranked candidates retained for later processing.
What is context allocation?+
It is the amount and selection of source material—often measured in passages or tokens—made available to the generative model.
Why does context position matter?+
Models may not use every part of a long context equally. Controlled studies suggest that where relevant evidence appears can affect its likelihood of contributing to a citation or answer.
References
Sources are listed in APA 7 style. Preprints are identified as such and should not be treated as peer-reviewed findings unless separately published.
- Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14035
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., & Narasimhan, K. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5–16). Association for Computing Machinery. https://doi.org/10.1145/3637528.3671900
- Zhang, K., He, X., & Yao, J. (2026). From citation selection to citation absorption: A measurement framework for generative engine optimization across AI search platforms [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.25707