transformers v5.16.1

v5.16.1

Release v5.16.1

Added 1
  • Add support for GLM-5.3-Flash, a natively multimodal model with 320B total parameters and 18B active parameters
Fixed 2
  • Restore backward compatibility for the tensor-parallel API
  • Fix kernel commit and repo paths for ESMFold2

From transformers

This is a special release as we include GLM! (and a few small fixes)

GLM-5.3-Flash

GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.

Links: Documentation

  • [Glm 5.3 Flash] GLM 5.3 Flash Support (#48342) by @Dovis01 in #48342
Small patch fixes

Mainly BC behavior for TP and pinning a hf kernel for security reasons :hugs:

  • Restore BC for the tensor-parallel API (#48300) by @ArthurZucker
  • Fix kernel commit and repo paths for ESMFold2 (#48186) by @Rocketknight1

Full Changelog: https://github.com/huggingface/transformers/compare/v5.16.0...v5.16.1

View original

Upgraded? How did it go?

Discussion