Tweaking a prompt until one example looks good often breaks others. Systematic iteration avoids this.
Start With Test Cases
Collect 10–50 representative inputs, including difficult ones, with notes on what good output looks like.
Establish a Baseline
Run the current prompt on all test cases and record results.
Change One Thing at a Time
Adjust a single element — an instruction, an example, the format — and re-run all cases. Otherwise you can't tell which change helped or hurt.
Look at Failures
Categorise what goes wrong: missing information, wrong format, hallucination, tone. Each category suggests a specific fix.
Keep Records
Version prompts in source control or a document, with the date, change and results. You'll want to go back.
Automate Checks
Where possible, score outputs automatically: format validity, required fields, length, keyword presence. Use human or model grading for quality.
Know When to Stop
When remaining failures are rare, acceptable or better handled elsewhere (validation, human review), stop tuning.
Re-Test When Models Change
A prompt tuned for one model may behave differently on another or on an updated version. Re-run tests before switching.