Huawei published a 505-billion-parameter model, and the stack behind it
Huawei released openPangu-2.0-Pro on 4 August, publishing the model weights, inference code and a technical report through GitCode's Ascend Tribe community and Huawei Cloud's ModelArts Studio. It is a mixture-of-experts model with 505 billion total parameters and 18 billion active per token, a 512,000-token context window, and roughly 34 trillion training tokens behind it. The reported architecture and training stack include the Muon optimiser, multi-head latent attention, decoupled sparse attention, sliding-window attention, and a three-stage post-training process using online policy distillation.
The detail drawing attention is the hardware: it was trained on Ascend 910B NPUs, without Nvidia GPUs.
Two things follow for the cost of intelligence. Published weights put a ceiling on what comparable capability can be charged for, because anyone can run the model on rented hardware. And a second complete training stack means the world's supply of frontier-scale models is not gated by one vendor's allocation queue. Both push the same direction, and a technical report detailed enough to build against is a larger contribution than a model card.
The honest limits are substantial. Huawei's performance claims have not been checked by any independent benchmark organisation, so nothing about how good this model is has been established here — only that it exists and can be downloaded. The release does not create a fully domestic Chinese hardware supply chain either: earlier Ascend chips used TSMC 7-nanometre compute dies and Samsung memory, with SMIC and CXMT expected to supply future production. The licence terms are not stated in the available account. Community work to run the model on Nvidia GPUs is already under way.
Source: Open Source For You
MANY MINDED