linkedin.com
Qwen-3.8-Flash-Next is a big deal and a herald of things to come. ↗
Flash models are models distilled down from full size models, in this case from Qwen-3.8-Max, a 2.4T model requiring some 8 datacenters GPUs to run. The Flash model instead features 125B parameters, roughly 5% of the original model of which again only 6B are active per token 4%....
