Kept Organised

7 August 2026

All writing

AI has too many models now

GPT-this, Claude-that, a dozen version numbers with no clear guidance on which one actually fits your task. A plain way to think about it.

At some point in the last couple of years, picking an AI model became its own small research project. There’s a fast cheap one, a slow expensive one, one that’s good at code, one that’s good at writing, a couple with numbers after their names that don’t obviously mean anything, and a new option roughly every few months that supposedly makes all the previous ones look outdated. Nobody explains this well, because most of the companies making these models would rather you just used theirs.

If you’ve ever opened a chat interface, seen a dropdown with six or seven model names, and picked whichever one you used last time out of sheer inertia, you’re not missing some obvious piece of public knowledge. The information genuinely isn’t presented anywhere in a useful way. It’s scattered across benchmark leaderboards that measure things irrelevant to your actual task, marketing pages that call everything “our most capable model yet,” and forum threads full of people arguing about their own narrow use case.

A simpler way to think about it

Ignore the version numbers for a moment and think about three questions instead: how much does this task actually need to reason through something complicated, how much does it matter if the answer comes back in two seconds versus twenty, and how much does a mistake actually cost you.

Quick factual lookups, simple rewrites, short summaries: these don’t need the most powerful model available. A faster, cheaper model handles them just as well, and the extra reasoning power of a top-tier model mostly goes to waste. This is most of what people actually use AI for day to day, and it’s the case most people over-serve by defaulting to whatever the flagship model is.

Genuinely difficult reasoning, like working through a multi-step problem, writing code that needs to actually run correctly, or untangling a messy legal or technical document: this is where the gap between models actually shows up. A weaker model will confidently give you a wrong answer here just as fast as a strong one gives you a right one, and you often won’t be able to tell the difference just by reading it.

High-stakes accuracy, where being wrong is expensive or embarrassing: this is worth slowing down for regardless of which model you use. Even the best current models get things wrong with total confidence, so for anything you’re going to act on without checking, a second pass, a different model for comparison, or a human review is worth the extra step.

The part nobody tells you

Different models are also just better at different kinds of work, independent of how “powerful” they are. Some are noticeably better at following a specific formatting instruction. Some write more naturally and some write more stiffly. Some are much faster at code and comparatively weak at nuanced writing, and vice versa. This isn’t really captured by any single ranking, because rankings tend to average across tasks that have nothing to do with each other.

The practical result is that there’s rarely one correct model to use for everything, and the fact that this isn’t obvious from the outside is a real design failure across the industry, not something you’re missing.

This is the exact problem Oversight is being built to solve: routing your request to whichever model actually fits it, without you needing to keep track of which one is good at what this month. It’s still in build, but the frustration behind it is one we’ve felt plenty ourselves.

Tags

Tool Updates

15 September 2026 · Windfall

Receipt capture is live

Photograph a receipt, get an HMRC-reasoned category back in seconds, no manual entry.

1 September 2026 · The Insight Distillery

Playbooks are live

Save screenshots and links, and Distillery groups them into playbooks by outcome, with actions cited back to your own material.

Stay in touch for launches and updates