PaLM 2 Technical Report
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson et al. · Google (Technical Report) · 2023
Abstract
Google's report on a model trained with a rebalanced compute budget — smaller than its predecessor but fed considerably more data, following the compute-optimal argument rather than raw parameter count. It also leans much harder on multilingual and reasoning-heavy data, and reports results across translation and code as well as English benchmarks.
Why it matters
A useful marker of when the industry stopped equating scale with parameters. Read it next to the Chinchilla paper here — this is that argument applied at production scale.
https://arxiv.org/abs/2305.10403