Language models can generate large amounts of example data quickly: questions, conversations, labelled texts, edge cases.
Useful Applications
- Expanding test sets with variations of real cases.
- Bootstrapping a classifier when labelled data is scarce.
- Edge cases that rarely appear in real data but must be handled.
- Privacy-preserving examples that resemble real records without containing real people's data.
- Data for fine-tuning smaller models on a narrow task.
Risks
- Tidiness: generated examples tend to be cleaner, more fluent and more uniform than real data, so models trained or evaluated on them may disappoint in practice.
- Bias and blind spots: the generator's own biases and gaps carry into the data.
- Errors: labels or facts in generated data can be wrong.
- Licensing: check the model provider's terms on using outputs to train other models.
Good Practice
- Seed generation with real examples and real distributions.
- Ask for diversity explicitly: different lengths, styles, difficulty levels.
- Review samples by hand and filter low-quality items.
- Keep real, human-labelled data for final evaluation.
- Label synthetic data clearly so it isn't mistaken for real data later.
Measure the Effect
Compare model performance on real held-out data with and without the synthetic additions. If real-data performance doesn't improve, the synthetic data isn't helping.