Provenance and Knowledge Updating in Large Language Models

Determine effective methods for large language models to provide verifiable provenance for generated outputs and to update their internal world knowledge efficiently and reliably, enabling transparent decision-making and overcoming limitations of static parametric memory.

Background

The survey highlights fundamental limitations of LLMs, including hallucinations and outdated knowledge due to reliance on static training data. It stresses that beyond accuracy, systems must attribute their outputs to precise sources and be able to update internal knowledge to stay current.

This problem motivates Retrieval-Augmented Generation and multimodal extensions, but the paper emphasizes that provenance and updating remain unresolved at the core of LLMs.

References

Moreover, providing provenance for their decisions and updating their world knowledge remain critical open problems .

— Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation  (2502.08826 - Abootorabi et al., 12 Feb 2025) in Section 1, Introduction (Background)

How a text that has passed through more than one provider's model is to be checked is an open question, and a text watermark and the signed provenance metadata that can accompany a generated file need not stay in step once the file is edited.

— Watermarks Without Verification: AI Text Watermarking After the EU AI Act  (2609.09604 - Nemecek et al., 9 Sep 2026) in Section “What Remains Open,” subsection on coordination across providers

We do not identify a reliable black-box method when the objective is to reconstruct one specific tool, because functionally identical tools provide little observable evidence for determining data provenance.

— Extracting Knowledge from Tools in LLM Agents  (2608.30288 - Zang et al., 31 Aug 2026) in Appendix I, “Functionally Identical Tools with Disjoint Data”