← Knowledge LibraryChapter 7 · Influence Measurement

Citation AbsorptionMeasuring Citation Absorption

How observable influence proxies estimate source contribution without pretending to expose a model’s hidden reasoning.

How observable influence proxies estimate source contribution without pretending to expose a model’s hidden reasoning.

Internal source influence is not directly observable

Commercial engines do not expose attention states, hidden retrieval scores, or exact source-to-token causal traces. Researchers therefore estimate answer-level correspondence from outputs that can actually be observed.

The correct claim is that a page shows higher observed answer influence—not that it received a known percentage of internal reasoning weight.

Influence combines several observable dimensions

Zhang, He, and Yao’s influence score combines reference count, first-position ratio, paragraph coverage, TF-IDF cosine similarity, and bigram or trigram overlap. Together they represent repetition, prominence, answer coverage, and textual correspondence.

The weighting is a research construct and should not be treated as a universal engine metric.

Absorption extends earlier impression metrics

Aggarwal and colleagues measured the share of answer words associated with a source, adjusted visibility for position, and used an evaluator to estimate subjective impression. These measures established that citations occupy unequal amounts of answer real estate.

Absorption goes further by asking how much factual and semantic substance a cited page contributes.

Use a bundle of proxies and preserve uncertainty

Reference repetition, early position, coverage, semantic similarity, lexical overlap, claim support, and paraphrase matching each capture a different observable signal. None alone establishes causality.

An absorption score is an observational estimate of answer correspondence—not a window into hidden model attention.

Frequently asked questions

Questions about this topic

Can citation absorption be measured directly?+

Not from normal commercial interfaces. It is estimated through observable answer–source correspondence and prominence signals.

What signals can estimate absorption?+

Reference count, first citation position, paragraph coverage, semantic similarity, lexical overlap, and claim-level support are useful signals.

What does TF-IDF similarity measure here?+

It estimates lexical correspondence between source and answer, but does not by itself prove causal use.

Is an influence score a documented platform metric?+

No. It is a research proxy constructed from observable features rather than a score exposed by commercial engines.

Source notes

References

Sources are listed in APA 7 style. Preprints are identified as such and should not be treated as peer-reviewed findings unless separately published.

  1. Zhang, K., He, X., & Yao, J. (2026). From citation selection to citation absorption: A measurement framework for generative engine optimization across AI search platforms [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.25707
  2. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., & Narasimhan, K. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5–16). Association for Computing Machinery. https://doi.org/10.1145/3637528.3671900
  3. Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14035