Life = Content

Writing

Benchmarks Lie: How I Actually Pick a Model for the Job

updated 2026-07-29

Every model launch comes with a leaderboard victory lap. Fine — I still have to ship something by Friday.

Benchmarks and day-to-day usefulness measure different things. A model can look like it beat the field on paper and still feel exhausting in a real session: over-apologizing, stopping early, asking you to confirm tiny fixes that didn't need a ceremony. That friction is a real cost. It just never shows up in ARC scores.

What I actually optimize for

  • Does it stay out of my way? If I'm fighting the personality more than the problem, it's not an upgrade — chart or no chart.
  • Is it better than what I already have? Most of us aren't freely shopping across labs. Ask "is this better than the tool I'm allowed to use today?" not "did it beat the internet's favorite model this week?"
  • Can I turn the dial down? Effort / thinking settings are rarely linear. I've seen peaks below max, and cheaper settings that barely lose quality. Try mid before max and keep whichever wins on *your* task.
  • Is the harness cleaner than last time? Sometimes the win isn't the weights — it's cutting a bloated system prompt. Over-constrained instructions can hurt more than they help. Cut one instruction, re-run the same task, and keep the trim if quality holds.

A 15-minute bake-off

When a new model drops and you're tempted by the chart, do this instead:

  1. Pick one real task you already know well (a PR review, a refactor, a draft rewrite).
  2. Run it on your current default and the contender, same prompt, same context.
  3. Score only what you feel: time-to-useful-output, how often you had to steer, and whether you'd happily do it again tomorrow.
  4. Keep the winner as default for a week. Revisit if the work changes.

If a model solves something with a strategy nobody else tried, that gets my attention more than another #1 badge. Leaderboards tell you who won the exam. A weird new approach tells you it might actually be exploring.

What I'm not doing

I'm not pretending capital-markets headlines or lab drama help me pick a model for a ticket. Fun to watch. Wrong input for the decision.

And if multi-model routers keep winning, "which model should I use" content has a shelf life anyway. The durable skill is knowing what good looks like in *your* workflow — speed, reliability, tone, cost — and swapping tools when those signals change.

Benchmarks aren't useless. They're just not the job. The job is getting useful output without wanting to throw your laptop.

← back to writing