Uncertain data provenance risk for AI

Transparency Icon representing transparency risks.
Transparency
Training data risks
Amplified by generative AI
Amplified by synthetic data

Description

Data provenance refers to the traceability of data (including synthetic data), which includes its ownership, origin, transformations, and generation. Proving that the data is the same as the original source with correct usage terms is difficult without standardized methods for verifying data sources or generation.

Why is uncertain data provenance a concern for foundation models?

Not all data sources are trustworthy. Data might be unethically collected, manipulated, or falsified. Verifying that data provenance is challenging due to factors such as data volume, data complexity, data source varieties, poor data management, and synthetic data generation methods. Using such data can result in undesirable behaviors in the model.

Parent topic: AI risk atlas

We provide examples covered by the press to help explain many of the foundation models' risks. Many of these events covered by the press are either still evolving or have been resolved, and referencing them can help the reader understand the potential risks and work toward mitigations. Highlighting these examples are for illustrative purposes only.