Unrepresentative data risk for AI
Description
Unrepresentative data occurs when the training or fine-tuning data is not sufficiently representative of the underlying population or does not measure the phenomenon of interest. Synthetic data might not fully capture the complexity and nuances of real-world data. Causes include possible limitations in the seed data quality, biases in generation methods, or inadequate domain knowledge. Thus, AI models might struggle to generalize effectively to real-world scenarios.
Why is unrepresentative data a concern for foundation models?
If the data is not representative, then the model will not work as intended.
Autonomous Vehicle Simulation
Bai et al. explicitly study the 'domain gap' between synthetic and real data. They show that models trained purely or heavily on synthetic data often perform worse on real data, especially in conditions or environments not well represented in the synthetic training set.
Parent topic: AI risk atlas
We provide examples covered by the press to help explain many of the foundation models' risks. Many of these events covered by the press are either still evolving or have been resolved, and referencing them can help the reader understand the potential risks and work toward mitigations. Highlighting these examples are for illustrative purposes only.