It is tempting to pick one model and use it everywhere. It keeps the code simple. But the tasks inside a product are rarely equal.
Three questions to ask per task
- How hard is it? Sorting a message into a category is not the same as planning a multi-step job.
- How fast must it be? A chat reply that streams in a second feels different from a report that takes a minute.
- What does a mistake cost? A wrong tag is cheap. A wrong answer about money or health is not.
Small, fast models suit classification, extraction and routing. Larger models suit open-ended writing, reasoning and tool use. Specialised models suit images, video and speech.
Keep the model behind one function
Whatever you choose, wrap the call in one place in your code. Pass in the task, get back the result. When a better or cheaper model appears, you change one line and re-run your tests.
Don’t lock yourself in
Providers change prices, limits and models. A gateway or a thin adapter layer lets you switch, or fall back to a second provider when one has an outage.
By Foldox. All posts