LLM guide
Choosing a model
There is no single best model. The right one is the cheapest model that does your task well enough, on infrastructure your data rules allow. The only reliable way to find it is to test candidates on your own examples.
Start fromYour task, not a leaderboard
Test on100-300 real examples
CompareQuality, speed, cost
RevisitEvery few months
The main choices
| Question | Options | How to decide |
|---|---|---|
| Commercial or open-weight? | Commercial models via API; open models (Llama, Mistral, Qwen and others) you can run yourself | Data rules and volume first. Open models are good enough for many enterprise tasks today. |
| How big? | Small (around 8B parameters), medium (around 70B), very large | Use the smallest that passes your test set. Small models are far cheaper and faster. |
| General or specialised? | General chat models, code models, embedding models for search | Use specialised models where they exist, especially for search embeddings. |
| Which languages? | English-first models, multilingual models | If users write in Hindi or other Indian languages, test that explicitly. Quality varies a lot. |
| Licence | Apache 2.0 and similar, or custom community licences | Read the licence. Some restrict use above a user count or in certain products. |
How we test models
- Collect real examples. Questions or documents from the actual workflow, with the answers a good employee would give.
- Define what good means. Correct facts, right format, right tone, says "I don't know" when it should.
- Run each candidate. Same prompts, same examples, recorded outputs.
- Score. Automatic checks where possible, plus review by the people who do the work.
- Weigh cost and speed. A model 3% better but 10 times more expensive is rarely the right choice.
Keep the test set. When a new model version arrives, you can re-run it in an afternoon and decide with evidence.
Common mistakes
- Choosing from public leaderboards. General benchmarks say little about your documents and your users.
- Testing only on easy examples. Include the messy, ambiguous and edge cases. That is where models differ.
- Locking in one model forever. Put a gateway in front so the model behind it can change without changing every application.
More LLM guides: Private hosting · Fine-tuning · Governance and LLMOps · LLM FAQ · Use case: claims summaries
Starting with LLMs, or stuck after a pilot?
Tell us the task you want AI to help with and any rules about where your data can go. We will come back with a plain recommendation: which kind of model, where to run it, roughly what it costs, and how to know if it is working.