Qwen 3.8 Model with 2.4T Parameters Open-Sourced, Domestic Large Model Architectures Converge
Half a month after Kimi K3 was open-sourced, Alibaba's Qwen released its first Max-scale weights, Qwen3.8-2.4T, on August 13. The model features 92 layers, 2.4T parameters, 95B activated…
Half a month after Kimi K3 was open-sourced, Alibaba's Qwen released its first Max-scale weights, Qwen3.8-2.4T, on August 13. The model features 92 layers, 2.4T parameters, 95B activated parameters, and employs 512 experts with a 3:1 hybrid attention architecture. Post-training focuses on agent capabilities while retaining reasoning states. The open-source release uses a custom commercial license, with the overall architecture converging with Kimi K3.
Original: https://wallstreetcn.com/articles/3779379
insigtX content is informational and educational, not investment advice.