Library

Research Library

Whitepaper2024

Qwen2.5 Technical Report

Qwen, :, An Yang, Baosong Yang et al. · Alibaba Qwen (Technical Report) · 2024

Abstract

Alibaba's report on a model family spanning half a billion to seventy-two billion parameters, trained on roughly eighteen trillion tokens. The interesting material is the data work — how the pretraining mixture was filtered and rebalanced — and the fact that the same recipe is carried across a wide size range, which is what makes the family useful as a controlled comparison rather than a single checkpoint.

Why it matters

The most practical open-weight family to build on, because a consistent recipe across sizes means you can prototype small and scale up without changing behaviour underneath you.

open weightstrainingpretrainingsmall models
Read the source

https://arxiv.org/abs/2412.15115