Library

Research Library

Whitepaper2023

PaLM 2 Technical Report

Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson et al. · Google (Technical Report) · 2023

Abstract

Google's report on a model trained with a rebalanced compute budget — smaller than its predecessor but fed considerably more data, following the compute-optimal argument rather than raw parameter count. It also leans much harder on multilingual and reasoning-heavy data, and reports results across translation and code as well as English benchmarks.

Why it matters

A useful marker of when the industry stopped equating scale with parameters. Read it next to the Chinchilla paper here — this is that argument applied at production scale.

pretrainingscaling lawsmultilingualtraining
Read the source

https://arxiv.org/abs/2305.10403