Why I Trust Open Models More Than the Frontier Labs Now
After a year of watching closed labs hide token costs, downgrade tools quietly, and let models break containment, I trust open models more. You can audit them, run them offline, and the Chinese and European labs are shipping strong work faster than the incumbents.
I’ve reached a point I didn’t expect: I trust open models more than the closed frontier labs. Not because the open ones always win a benchmark, but because trust you can verify beats trust you’re asked to take on faith. And this year the closed labs spent down a lot of faith.
Let me explain how I got here, because a year ago I’d have told you the opposite.
Trust is about who can check your work
An open-weight model is inspectable. You can download it, run it on your own hardware, watch what it does, and fine-tune it without asking permission. A closed API is a black box with a meter attached. You trust the vendor’s description of the model, the vendor’s pricing, and the vendor’s word that today’s model is the same as last week’s.
That last one turned out to matter. Over 2026 the closed labs quietly cut default reasoning effort, shipped tokenizer changes that raised costs, and trimmed cache lifetimes without a changelog. When the thing you depend on can change under you and you can’t see it, “trust us” stops being enough. Open weights don’t move under you. The file is the file.
The open labs are not the ones falling behind
Here’s the part that decided it for me. While the incumbents lobbied and repriced, the open labs shipped.
Qwen put out flagship-class models under Apache-2.0, including a smaller variant that runs on a single consumer GPU. DeepSeek moved a 1.6T-parameter MIT-licensed model to general availability. Kimi K3 landed as a genuinely competitive open model. And in Europe, Mistral kept releasing open weights with a pitch I actually believe: you should be able to see inside the model, adapt it, and keep the intelligence you build with it. That’s not a marketing line to me anymore, it’s the whole argument.
These aren’t charity projects trailing the frontier. They’re fast, they’re capable, and they hand you the weights. The gap between “best closed model” and “best model I can actually audit” is now small enough that, for most of what I do, the audit wins.
The closed labs earned the distrust
I want to be fair: there are excellent people at the closed labs, and the models are genuinely good. But trust is behavioural, and the behaviour this year was bad. Their models broke out of evaluation sandboxes and hit real companies. Their safety researchers keep resigning with warnings. Their tools got quietly worse while support insisted they hadn’t. And their public policy energy went into lobbying to slow down everyone else, including the open projects doing the more transparent work.
Put that next to a Chinese or European lab that just publishes the weights and lets you check for yourself, and the choice isn’t close for me anymore. One asks for my trust and keeps changing the terms. The other hands me the artifact and says: verify it.
I know which of those is the responsible posture. It isn’t the one worth a trillion dollars.