There is no single best AI model. There is the model that does your specific task well enough, fast enough, cheaply enough and within your data rules. A model that tops public leaderboards may be overkill for sorting support tickets, while a small, fast model may fall apart on a long legal document. Choosing well means defining the job first and measuring candidates against it.
This guide gives you a decision framework you can apply to any provider. It covers task fit, quality, latency, how to estimate cost per task using tokens, context length, privacy, the trade-offs between open-weight and hosted models, multimodality and vendor lock-in. It finishes with the most important step: running a small evaluation on your own examples.
Start with task fit
Write a one-paragraph task card before you compare anything. It should answer:
- What goes in and what comes out? For example, "a customer email in, a JSON object with category and a draft reply out".
- How hard is the reasoning? Classification and extraction are usually easy; multi-step analysis, maths and code are harder.
- What does a mistake cost? A mislabelled newsletter is cheap; a wrong figure in a financial summary is not.
- How much volume? Ten requests a day and ten thousand an hour lead to very different choices.
- Is a person in the loop? Human review tolerates a lower accuracy rate than full automation.
Most providers offer a family of models in tiers: large, capable and slower; medium; small, fast and cheap. A common pattern is to test the largest model first to see whether the task is achievable at all, then try smaller tiers to see how far down you can go before quality drops below your bar. If you are new to how these systems work, how large language models work explains the basics.
Quality and latency
Quality means accuracy on your task, judged against your criteria: correct facts, correct format, appropriate tone, following instructions. Public benchmarks are a rough filter at best; they test general skills on someone else's questions, and providers can tune for them.
Latency has two parts that matter differently:
- Time to first token: how quickly output starts appearing. This matters for chat interfaces where people watch the response stream in.
- Total generation time: how long until the full answer is ready. This matters for background jobs and for steps where the next action waits on the complete output.
Set a latency budget, such as "under two seconds for an interactive reply" or "under a minute per document overnight". Models with extended reasoning or "thinking" modes often improve quality on hard problems but add time and cost, so test both with and without them.
Want this working in your business, not just on paper? Get a free, written AI starting plan.
Get my free AI planEstimating cost per task
Hosted models are typically priced per million tokens, with separate rates for input (what you send) and output (what the model writes). A token is a chunk of text; in English, a rough rule of thumb is about three-quarters of a word per token, so 1,000 words is roughly 1,300 tokens. Other languages and code can use more tokens per word.
The formula is:
Cost per task = (input tokens × input price per token) + (output tokens × output price per token)
Worked example (hypothetical prices)
The prices below are invented for illustration only. Check each provider's current pricing page for real figures.
- Task: summarise a support conversation into a structured note.
- Input: a 600-token system prompt with instructions and examples, plus a conversation averaging 1,400 tokens = 2,000 input tokens.
- Output: a structured summary averaging 300 tokens.
- Model A (large): hypothetical $3 per million input tokens and $15 per million output tokens. Cost = 2,000 × $0.000003 + 300 × $0.000015 = $0.006 + $0.0045 = $0.0105 per task.
- Model B (small): hypothetical $0.25 per million input and $1.25 per million output. Cost = $0.0005 + $0.000375 ≈ $0.0009 per task.
- At 50,000 tasks a month: Model A ≈ $525; Model B ≈ $44.
If Model B meets your quality bar, the choice is easy. If it does not, the question becomes whether Model A's extra accuracy is worth about $480 a month, which you can compare with the cost of a person correcting errors.
Context length
The context window is how much text, in tokens, a model can consider in one request, covering your instructions, any documents and the conversation so far. Large windows let you include whole contracts or codebases, but bigger is not automatically better. Cost rises with every token you send, responses can slow down, and models do not always use information buried in the middle of a very long input equally well.
If your source material is larger than you can sensibly send each time, retrieval is usually better: store documents in a searchable index and send only the relevant passages. Retrieval-augmented generation explained covers how this works.
Privacy and data handling
For many organisations, data rules eliminate options before quality is even considered. Check these points for every provider:
- Is data sent through the business API or enterprise plan used to train models? Is that the default, and can it be turned off?
- How long are prompts and outputs retained, and for what purposes?
- In which regions is data processed and stored?
- Which security certifications and contractual commitments, such as a data processing agreement, are available?
- Can you control who in your organisation has access to logs?
Consumer chat apps and business APIs often have different terms, so read the ones that apply to how you will actually use the model. Frameworks such as the NIST AI Risk Management Framework are useful for structuring a review. For regulated data, involve a qualified legal or compliance professional; this guide is not legal advice. Our AI privacy and security checklist goes further.
Open-weight vs hosted models
Hosted models run on the provider's infrastructure and are accessed through an API. Open-weight models publish their trained weights so you can run them on your own hardware or with a hosting provider of your choice. Licences vary, so read the terms before commercial use.
| Factor | Hosted (API) | Open-weight (self-run or third-party hosted) |
|---|---|---|
| Top-end capability | Usually the strongest models available | Often close behind; strong for many focused tasks |
| Setup effort | Minimal: an API key and a few lines of code | Higher: hardware or hosting, serving software, updates |
| Cost model | Pay per token; no idle cost | Pay for compute whether busy or idle; can be cheaper at steady high volume |
| Data control | Governed by provider terms | Data can stay entirely within your environment |
| Customisation | Prompting; fine-tuning on some models | Full fine-tuning and modification, subject to licence |
| Version stability | Provider may update or retire models on its schedule | You decide when, or whether, to change versions |
| Scaling | Handled by provider, within rate limits | Your responsibility |
A practical approach for many teams: prototype on a hosted model to prove the task works, then evaluate whether an open-weight model meets the bar if cost, privacy or control becomes a priority.
Multimodality and vendor lock-in
Multimodality. If your inputs include images, scanned documents, charts, audio or video, test models on those inputs specifically. Support varies: some models read images well but not handwriting; some accept audio directly; others need a separate transcription step. Image and video generation is usually a separate model family, covered in our AI image and video generation guide.
Vendor lock-in. Switching models is easier if you plan for it:
- Keep all model calls behind one internal function or service, so changing provider is a configuration change, not a rewrite.
- Store prompts as versioned files rather than scattering them through code.
- Avoid depending heavily on one provider's unique features unless the benefit is clear.
- Keep your evaluation test set ready, so you can check a new model in an afternoon.
Run your own evaluation
A small, careful test on your own data beats any amount of reading reviews. Here is a process that works for most teams:
- Collect 30 to 100 real examples that represent your actual inputs, including difficult and edge cases. Remove sensitive details if needed.
- Define what "good" means. For extraction and classification, write the correct answer. For open-ended writing, write a short scoring rubric, for example accuracy, completeness, tone and format, each scored 1 to 3.
- Pick three or four candidates across tiers and providers, plus your current approach as a baseline.
- Use the same prompt for all, then allow one round of prompt tuning per model, since different models respond to slightly different instructions.
- Run every example through every model and record output, latency and token counts.
- Score blind where possible. Hide which model produced which output from the person scoring, to reduce bias.
- Compare in a table: accuracy or average score, failure types, median and slowest latency, and cost per task at your expected volume.
- Choose the cheapest, fastest option that clears your quality bar, not the one with the highest score regardless of cost.
Using another model to grade outputs can speed up scoring for large sets, but check its grades against a human-scored sample first. Rerun the evaluation whenever a provider updates a model or you change your prompt. If you are wiring the chosen model into automated processes, see AI automation workflows.
Frequently asked questions
Should I always use the most powerful model available?
No. Larger models cost more and are usually slower. Many tasks, such as classification, extraction and short drafting, work well on smaller models. Test from the top down and stop at the smallest model that meets your quality bar.
How many test examples do I need?
Thirty well-chosen examples will reveal major differences between models. For decisions with higher stakes or subtle quality differences, aim for 100 or more, and make sure edge cases are represented.
Can I use different models for different steps?
Yes, and it is common. A small model might classify incoming requests while a larger model handles the complex ones. Routing by difficulty is one of the most effective ways to control cost.
Are open-weight models safe for business use?
They can be, and they offer strong data control when run in your own environment. You take on responsibility for security, updates and safety filtering, and you must follow the model's licence terms.
This guide is general information, not professional advice. Spotted an error? Tell us.