Chinese Flash models are the Coup de grâce for US AI Lab investor fantasies
Chinese Flash models are the Coup de grâce for US AI Lab investor fantasies, or “Just because something you invented is valuable doesn’t mean you are going to extract the value”
I’ve been playing around with Qwen-3.8-Flash-Next and GLM-5.3-Flash and it is apparent that these models represent another, Deepseek Level bomb into the heart of the imaginary token economy US labs and hyperscalers have been selling their investors.
First off, the Qwen Model is really good. Good enough for a surprising number of tasks, including coding, reverse engineering, forensic analysis, proxmox cluster optimisation, Kubernetes configuration, and more. It’s the first truly local model that feels like it has no major tradeoffs, running lightning fast with 80-130 tp/s, 256k context, multimodal video and audio understanding, excellent tool use.
If you can get 94GB of VRAM (RTX6000) for the NV4 quant and enough additional host memory for n-gram offload, it’s the best local AI you can get today. And GLM-Flash ... oh my, here goes the neighbourhood.
Is it “Fable Level”? I dunno because what does that even mean, but it doesn’t randomly refuse, doesn’t lecture, doesn’t cook up seam-loadbearing-jargon babble, and is dirt cheap on the API and feels more than good enough for serious work.
We have to be clear: Cost savings are not a reason to use local AI. Buying, at current hardware prices, the rig and GPU to run would cost the equivalent of an army of VC subsidised Claude Code or Codex subscriptions and is certain to never amortise.
These same models will be hosted in the cloud at bargain rate and, especially on current hardware prices, there’s no break-even point against cloud.
Yes privacy, yes business continuity, yes control, but no cost savings. But exactly this mechanism breaks AI investors dreams of massive ROI - these cheap models will take over the vast majority of inference work, or simply put, will destroy the value of the labor they provide by giving them to everyone for pennies.
Jevon’s law will kick in, yes, AI as far as the eye can see, but at bargain bin prices for “Good enough”.
Big AI now desperately needs to find use cases that use so much compute that their hold on capacity matters again before time runs out. I’m not holding my breath for investors here, human society only moves that fast.
It’s the calculator all over again: A valuable invention but commoditised quickly.
And yes, economies of scale favour platform centralisation, even with powerful decentralisation options available, as the internet and the rise of Google/Meta have shown us.
But still, it’s not the same world anymore: The cost of centralisation is increasingly becoming visible as the primary beneficiary is turning against the rest of the world trying to extract value.
These flash models rip out the 80% Pareto bottom from the market, good enough for 80+% of the work and so cheap nobody (except the bros who pocketed the investments) is getting rich.
Usecases that need more-than-this AI are not able to satisfy investors.
Even expensive Apple boxes for local AI are temporary, it’s only a matter of time before we see cheap and powerful edge boxes, likely from China crashing the party.
Chinese Flash models are the Coup de grâce for US AI Lab investor fantasies, or "Just because something you invented is valuable doesn't mean you are going to extract the value" I've been playing… | Georg Zoeller | 26 comments
Chinese Flash models are the Coup de grâce for US AI Lab investor fantasies, or “Just because something you invented is valuable doesn’t mean you are going to extract the value” I’ve been playing around with Qwen-3.8-Flash-Next and GLM-5.3-Flash and it is apparent that these models represent another, Deepseek Level bomb into the heart of the imaginary token economy US labs and hyperscalers have been selling their investors. First off, the Qwen Model is really good. Good enough for a surprising number of tasks, including coding, reverse engineering, forensic analysis, proxmox cluster optimisation, Kubernetes configuration, and more. It’s the first truly local model that feels like it has no major tradeoffs, running lightning fast with 80-130 tp/s, 256k context, multimodal video and audio understanding, excellent tool use. If you can get 94GB of VRAM (RTX6000) for the NV4 quant and enough additional host memory for n-gram offload, it’s the best local AI you can get today. And GLM-Flash ... oh my, here goes the neighbourhood. Is it “Fable Level”? I dunno because what does that even mean, but it doesn’t randomly refuse, doesn’t lecture, doesn’t cook up seam-loadbearing-jargon babble, and is dirt cheap on the API and feels more than good enough for serious work. We have to be clear: Cost savings are not a reason to use local AI. Buying, at current hardware prices, the rig and GPU to run would cost the equivalent of an army of VC subsidised Claude Code or Codex subscriptions and is certain to never amortise. These same models will be hosted in the cloud at bargain rate and, especially on current hardware prices, there’s no break-even point against cloud. Yes privacy, yes business continuity, yes control, but no cost savings. But exactly this mechanism breaks AI investors dreams of massive ROI - these cheap models will take over the vast majority of inference work, or simply put, will destroy the value of the labor they provide by giving them to everyone for pennies. Jevon’s law will kick in, yes, AI as far as the eye can see, but at bargain bin prices for “Good enough”. Big AI now desperately needs to find use cases that use so much compute that their hold on capacity matters again before time runs out. I’m not holding my breath for investors here, human society only moves that fast. It’s the calculator all over again: A valuable invention but commoditised quickly. And yes, economies of scale favour platform centralisation, even with powerful decentralisation options available, as the internet and the rise of Google/Meta have shown us. But still, it’s not the same world anymore: The cost of centralisation is increasingly becoming visible as the primary beneficiary is turning against the rest of the world trying to extract value. These flash models rip out the 80% Pareto bottom from the market, good enough for 80+% of the work and so cheap nobody (except the bros who pocketed the investments) is getting rich. Usecases that need more-than-this AI are not able to satisfy investors. Even expensive Apple boxes for local AI are temporary, it’s only a matter of time before we see cheap and powerful edge boxes, likely from China crashing the party.