L-12 · Field guide

How to audit sources in public AI answers

Primary intent
Run a bounded, reviewable audit of visible AI-answer sources.
Evidence state
Source-grounded reference
Review owner
Matthias Ramahi · independent review not claimed
Last reviewed
2026-08-22
Direct answer

Audit AI-answer sources

To audit AI-answer sources, freeze the question and observation conditions, capture the public answer and every visible source, preserve original URLs, normalize copies, classify missing cases, verify a bounded sample of target pages, and report counts with denominators and limitations. The audit covers visible citations, not hidden retrieval behavior.

Use this guide for a one-off evidence audit or as the collection procedure inside a registered study.

1. Freeze the audit scope

Write the research question, public surface, route, locale, account or session rule, observation window and question list before capture. Define what counts as a visible source: inline citation, linked source panel, expandable source or another interface element.

State what the audit cannot see. It does not recover private query fan-out, retrieval candidates, ranking weights or model reasoning.

2. Capture the public record

Assign an observation ID, record an offset timestamp, preserve the exact submitted question and capture the visible answer state. Store every visible source URL as shown, with its location or visibility tier when that distinction is reliable.

When the answer or source panel is unavailable, record the missing reason. Do not rerun only failed cases until a preferred result appears unless the retry rule was defined in advance.

3. Normalize and verify without erasing provenance

Derive normalized URLs and registered domains while preserving the original values. Document parameter removal, redirect handling and canonical selection. Then open a bounded, stated sample of cited pages to check status, title, topic match and whether the page is actually accessible.

A cited URL is not automatically evidence for the answer’s claim. The audit can distinguish source presence from source support, but claim-level verification needs a separate coding guide and reviewer process.

4. Report coverage before conclusions

Publish scheduled and completed observations, answers with visible citations, total citation occurrences, unique URLs and domains, missing reasons and the verification sample size. Then report source recurrence, concentration or category mix if those measures were preregistered.

Include the source list or a rights-safe representation when permitted. Name capture limitations, possible personalization, interface changes and any pages that could not be verified.

  • Keep raw capture and analysis tables separate.
  • Use one rule for every eligible observation.
  • Show denominators beside percentages.
  • Treat common ownership and conflicts explicitly.
  • Publish corrections without deleting the old record.
S

Source notes

These sources support the definitions, standards or project boundaries named in this reference. They do not prove that a public observation dataset exists.

  1. portfolio-dossier
    Canonical ai-fanout.com domain dossier

    Confirmed ownership, accepted public Evidence Lab purpose, named Research Owner, indexable website launch and separately gated provider research.

    Owner record
  2. google-query-fanout
    AI features and your website

    Google describes query fan-out publicly without exposing a general private-query inspection interface.

    Open
  3. w3c-prov-o
    PROV-O: The PROV Ontology

    Provides provenance concepts for entities, activities, agents, derivations, sources and versions.

    Open
  4. rfc-3339
    RFC 3339: Date and Time on the Internet

    Supports an interoperable timestamp representation tied to UTC.

    Open
  5. fair-principles
    The FAIR Data Principles

    Supports reusable research data with metadata, provenance and clear usage licenses.

    Open