Twelve models in one month, one rule for staying sane
February 2026 shipped a dozen serious models. I tried to evaluate all of them, failed, and found a better rule.
February was the month I started a spreadsheet I would like to formally apologize for.
The releases came in like weather. Gemini 3.1 Pro. Claude Opus 4.6 and Sonnet 4.6. GPT-5.3 Codex. Qwen 3.5, GLM-5, MiniMax M2.5, ByteDance’s Seed 2.0 in two sizes, something called Mercury 2 that generates text with diffusion instead of the usual left-to-right way, which I still find slightly unsettling. Twelve significant models in four weeks, give or take how you count.
My spreadsheet had columns for benchmark scores, context windows, price per million tokens, and a “vibes” column that I am not proud of. I spent two full evenings maintaining it. By the third week it was already wrong in six places because prices had changed and two models had been updated in place. I was doing unpaid QA for an industry that ships faster than I can type.
The spreadsheet died on a Sunday and got replaced by a rule that has held up since: one daily driver, one challenger, ignore the rest.
The daily driver is the model I actually use, all day, for everything: code, writing, thinking out loud, explaining my own automation scripts back to me. Familiarity compounds. After a month with one model I knew its habits, where it pads answers, where it goes confidently wrong (units and off-by-one date math, since you ask), and when a weird answer means I should rephrase. That knowledge is worth more than five benchmark points, because I can correct for it.
The challenger is whichever new model genuinely seems interesting this month. It gets one honest week on my real work, not on benchmark puzzles. Same tasks, same prompts. If it clearly beats the daily driver at things I actually do, it takes the seat. In February it did not, though it was close enough to be annoying.
Everything else I let pass by, on the theory that the train leaves every three weeks anyway.
Two things I did take from the February flood, beyond the rule.
First, the open-weight models stopped being a curiosity. Qwen 3.5 and GLM-5 were not toys; they were a real tier, close behind the frontier, and you could run them on hardware a company like mine actually owns. I filed that away under: probably matters sooner than I think.
Second, benchmark numbers and daily usefulness have visibly drifted apart. Every one of those twelve models was “state of the art” on something. The differences between the top models were smaller than the difference between a good and a bad prompt on any one of them. At some point the honest conclusion is that model choice is no longer the bottleneck for someone like me. I am the bottleneck. That was less flattering but more actionable.
A uni colleague asked me around then which model is the best, in the tone of someone asking which car to buy. I gave them the answer I wish someone had given me earlier: pick any of the top four, use it every day for a month, and the ranking will matter less than what you learned about driving. They looked disappointed. The spreadsheet version of me from early February would have looked disappointed too.
The corner of the industry I am probably not watching closely enough is none of the model releases. It is the quiet stuff about agent frameworks moving into production, which started in January and keeps growing. The models get the headlines; I have a creeping suspicion the plumbing around them is becoming the actual story. For now it is only a suspicion.
If the flood taught me one thing, it is this: familiarity with one model beats headlines about twelve.