L-03 · Measurement

How to measure source diversity in AI answers

Primary intent
Define source-diversity measures that remain interpretable across repeated AI-answer observations.
Evidence state
Source-grounded reference
Review owner
Matthias Ramahi · independent review not claimed
Last reviewed
2026-08-22
Direct answer

Source diversity

Source diversity describes how broadly visible citations are distributed across distinct URLs, domains or declared source classes within a defined observation set. A valid result names the unit, normalization rule, sample, denominator and missing-data treatment; a count of unique domains alone is not enough.

Use this method when a study needs to compare the breadth or concentration of visible sources across questions, surfaces or dates.

Choose the diversity unit first

URL diversity, registered-domain diversity and source-type diversity answer different questions. Ten cited URLs from one publisher can be diverse at page level and concentrated at domain level. A study should report the unit explicitly rather than using “sources” as an undefined count.

Normalize hosts and URLs before counting. Decide how to handle www aliases, tracking parameters, fragments, syndicated copies, subdomains and redirects. Preserve the original visible URL beside the normalized value so the transformation remains reviewable.

  • Unique visible URLs: breadth at document level.
  • Unique registered domains: publisher concentration.
  • Declared source classes: mix of documentation, research, editorial or other categories.
  • Recurrence share: how much of the sample is occupied by repeating sources.

Publish the denominator

A result such as “24 domains appeared” cannot be interpreted without the number of questions, answers returned, answers with citations and total visible citation slots. The denominator also changes when an interface returns no answer or no links.

Report coverage before diversity: eligible observations, captured answers, answers with at least one visible source, and total source occurrences. Then report unique values and concentration. This prevents a sparse surface from looking diverse simply because only a few observations contained links.

Diversity is descriptive, not automatically good

A higher unique-domain count may reflect broader sourcing, noisy citations, repeated one-off domains or a different question mix. A lower count may reflect appropriate reliance on primary documentation. Diversity should therefore be interpreted beside question class, source relevance and source role.

Do not turn the measure into an unvalidated quality score. If quality, authority or correctness matters, define and review those constructs separately.

Comparing two windows

Use the same control questions, surface, locale, route, timing rule and normalization code. Compare both the set overlap and the distribution of occurrences. A stable unique count can hide a complete turnover in which domains appeared.

Annotate interface or provider changes. If a change affects citation availability, the safest result may be a comparability break rather than a trend line.

S

Source notes

These sources support the definitions, standards or project boundaries named in this reference. They do not prove that a public observation dataset exists.

  1. portfolio-dossier
    Canonical ai-fanout.com domain dossier

    Confirmed ownership, accepted public Evidence Lab purpose, named Research Owner, indexable website launch and separately gated provider research.

    Owner record
  2. nist-ai-rmf-genai
    NIST AI RMF Generative AI Profile

    Supports explicit measurement, documentation, monitoring and limitations for generative-AI evaluations.

    Open
  3. w3c-prov-o
    PROV-O: The PROV Ontology

    Provides provenance concepts for entities, activities, agents, derivations, sources and versions.

    Open
  4. fair-principles
    The FAIR Data Principles

    Supports reusable research data with metadata, provenance and clear usage licenses.

    Open