Most AI features begin the same way: someone writes a prompt, tries three examples, and likes what they see. Then the feature ships, and the first real users find the cases nobody tried.
An eval set fixes that. It is a small collection of real inputs, each with a note on what a good answer looks like. You run the feature against it every time something changes.
What goes in it
- Typical inputs, taken from real usage or from the people who will use it.
- Awkward inputs: very short, very long, misspelled, in another language.
- Inputs where the right answer is “I don’t know” or “I can’t help with that”.
- Every bug you have already fixed, so it stays fixed.
Twenty to fifty cases is plenty to start. You can grow it as you find new failures.
How to score it
Use the simplest check that works. For structured output, check the fields. For short answers, check for the key facts. For open-ended text, a person reading ten answers beats a clever metric you don’t trust. A second model can grade answers too, but spot-check its grades by hand.
Why it pays off
With an eval set, a model upgrade is a question you can answer in minutes: run the set, compare, decide. Without one, every change is a guess, and you find out from your users.
By Foldox. All posts