ChatGPT, Claude, Gemini, Qwen … Sol, Fable, Sonnet, Flash, Nano Banana, Copilot. I don't know about everyone else, but the number of AI providers and models seems to grow by the week — and it can be hard to know which are right for your business and your current needs. Hopefully this article can shed some light on what it all means, and give you a feel for when to use what.
Before we begin, I'd like to emphasise that I've tried most of the following across different use cases — software development, spreadsheets, website design — and they can all be powerful in their own ways. At AIbility we've settled on Claude: over time it has proven the best at following instructions for our workflow, the best fit for our tooling, and the tool whose updates most closely match our direction. But that is personal preference. It doesn't mean it's right for your needs — and by the end of this article, you should be able to judge that for yourself.
First, the providers
Let's clear up the biggest source of confusion straight away: the difference between a provider and a model. If you've read What is AI? you'll know I'm fond of a car analogy, so here's another one. The provider is the manufacturer; the model is the car in the showroom. Nobody drives a Volkswagen called "Volkswagen" — you drive a Polo or a Golf. It's exactly the same here: nobody actually uses a model called "ChatGPT". That's the showroom. The cars inside have names like GPT-5.6.
So, the main manufacturers you'll hear about. OpenAI make ChatGPT — the household name that started the frenzy, and for many people the only name they know. Anthropic make Claude — the one we use at AIbility. Google make Gemini, which is steadily being baked into everything they own: Search, Gmail, Android, the lot. And Alibaba make Qwen, the best known of the open models — ones you can download and run on your own hardware, free of charge.
Then there are two names that get lumped in with the providers but aren't quite the same thing. Ollama isn't a manufacturer at all — it's more like a garage: a free tool that lets you run open models such as Qwen on your own machine, no cloud required. And Copilot is Microsoft's assistant built into Windows and Office — a very nice badge on the bonnet, but for the most part it's OpenAI's engine underneath.
Then, the models
Now for the part that makes the whole menu far less intimidating: every provider's range has roughly the same shape. The names change; the tiers don't. Each provider ships a small family of models, and almost every family breaks down into three sizes.
The flagship — the biggest and deepest thinker in the range, and the most expensive to run. It will happily chew through complex analysis, large documents and genuinely difficult problems, but it takes its time doing it. For Anthropic that's Fable; for OpenAI it's GPT-5.6 Sol; for Google it's Gemini 3 Pro.
The all-rounder — the sensible middle of the range, and the default in most of the chat apps. Fast enough for conversation, clever enough for real work. Claude Sonnet, GPT-5.6 Terra and Gemini Flash live here.
The small and speedy — light, quick and very cheap. Ideal for simple, repetitive jobs at volume. Claude Haiku, GPT-5.6 Luna and Gemini's Flash-Lite are the runabouts of the fleet.
And alongside the main range sit the specialists — models built for one job. The delightfully named Nano Banana is Google's family of image generation and editing models, and every provider has similar specialists for images, video and voice.
As for the numbers — Claude 5, GPT-5.6, Gemini 3 — they're simply the generation, like a car's registration year. Higher means newer, and newer generally means better. You don't need to memorise them.
So when do you use what?
With the tiers in mind, choosing becomes surprisingly simple: match the size of the model to the size of the job.
For everyday work — drafting emails, summarising documents, answering questions — the all-rounder is all you need, and it's what the chat apps give you by default anyway. When the problem is genuinely hard — analysis across a stack of contracts, planning a project, untangling tricky code — step up to the flagship: slower and pricier, but when the quality of the answer matters, it more than pays for itself. And when you're automating something high-volume — categorising every incoming email, extracting data from thousands of invoices — the small and speedy tier is the one. At that scale the cost difference is enormous, and the little models are far more capable than you'd think.
Two special cases. If the job is visual, reach for the image specialists like Nano Banana. And if your data is too sensitive to leave the building, that's where the open models come in — Qwen running through Ollama on your own hardware trades away some convenience and polish, but nothing ever leaves your premises.
Pick a provider you trust, start in the middle of the range, and move up or down as the work demands.
And if all of that still feels like a lot — the reassuring truth is that most of the chat apps now pick the model for you, quietly routing simple questions to the small models and hard ones to the big ones. You mostly need to care when you're paying per use, automating at volume, or building AI into your business. Which, funnily enough, is where we come in.
You don't have to pick just one
Here's the part that surprises most people: the tiers aren't rivals — they're colleagues. The most powerful pattern we use at AIbility is putting a flagship in charge of a team of smaller models. Fable or Opus plays the foreman: it reads the brief, breaks the job into pieces, and hands each piece to a Sonnet or a Haiku to carry out — several of them at once, working in parallel — then checks the results before reporting back.
Why bother? Cost and speed. The flagship does the thinking; the smaller models do the legwork. You get flagship-quality judgement across the whole job without paying flagship prices for every keystroke — and because the work runs in parallel, it finishes far sooner too. If you read Agentic Software Development, this is the same idea: one developer working like a team of ten. The tiers are how that team is staffed — an architect directing, and a crew of quick, capable juniors doing the building.
So which provider, then?
Honestly? The gap between the top models is far smaller than the marketing wars would have you believe, and they leapfrog each other every few months. Chasing whichever model tops this week's leaderboard is a full-time job with very little payoff.
What actually compounds is consistency: learning one assistant's quirks, refining your instructions, and building your workflows and tooling around it. That's why we chose Claude — instruction following and tooling were what mattered for how we work. Your deciding factor might be completely different: if your business lives in Google Workspace, Gemini is right there; if it lives in Office, so is Copilot. Pick for fit, settle in, and review it once or twice a year — not every time a new model trends on LinkedIn.
How we can help you
Working out where AI pays off in a business — and which provider, models and tooling to build on — is exactly the journey we've been on ourselves at AIbility, and it's the one we now guide others through. Whether you want the confidence to choose and explore for yourself, or you'd rather leave the tooling to us, we'll help you find the fit that's right for you.
If that sounds like the kind of help your team could use, get in touch — we'd love to start you on your journey.