Alibaba Releases "Qwen2.5" — Open-Sourcing More Than 100 Models at Once
On September 19, 2024, Alibaba Cloud announced its new model family “Qwen2.5” at its annual Apsara Conference. Alongside general-purpose models, the release included the code-specialized “Qwen2.5-Coder” and math-specialized “Qwen2.5-Math,” and — counting quantized versions — more than 100 models were released at once to the global open source community. Alibaba described it as “perhaps the largest open source release in history.”
Details
- Three model families, 100+ models: General-purpose “Qwen2.5” (0.5B to 72B), code-specialized “Qwen2.5-Coder” (1.5B, 7B, 32B), and math-specialized “Qwen2.5-Math” (1.5B, 7B, 72B)
- Training scale: Pretrained on up to 18 trillion tokens, scoring above 85 on MMLU, above 85 on HumanEval, and above 80 on the MATH benchmark
- License: Released under the Apache 2.0 license, except for the 3B and 72B models
- Expanded capabilities: Context length of up to 128K tokens, generation of up to 8K tokens, and support for 29 or more languages
- Related announcements: Alibaba simultaneously announced a text-to-video model, an enhanced large vision-language model, and an improved version of its top-tier model “Qwen-Max”
What happened next
Qwen2.5 continued to rank near the top of various leaderboards, marking a turning point where Qwen established itself as the leading name in the open source camp. The following January, this trajectory extended to the release of the closed model “Qwen2.5-Max.”