Language models can quickly produce test data: sample customer messages, edge-case inputs, fake records and adversarial prompts.
Be Specific About Variety
Ask for variation along defined dimensions: length, tone, language, typos, topic, difficulty. "Generate 30 support messages: 10 billing, 10 technical, 10 account; vary length from one line to three paragraphs; include some with spelling mistakes and some in mixed languages."
Include Edge Cases
Request empty inputs, extremely long inputs, unusual characters, ambiguous requests, off-topic messages and attempts to misuse the system.
Realistic Records
For fake personal records, specify formats and plausible distributions, and make sure generated data doesn't accidentally match real people.
Label as You Generate
Ask for expected outputs or labels alongside each input, but verify them — generated labels can be wrong.
Review Samples
Generated data tends to be cleaner and more uniform than real data. Read samples and supplement with real (anonymised) inputs.
Structured Output
Request JSON or CSV so data loads directly into test frameworks.
Keep It Separate
Mark synthetic test data clearly, and keep final evaluation anchored in real data where possible.