Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac et al. · Google DeepMind (Technical Report) · 2023
Abstract
Google DeepMind's report on a model family built multimodal from the start — trained jointly on text, images, audio and video rather than assembling a language model and a vision encoder afterwards. It spans three sizes aimed at datacentre, general and on-device use, and reports across all four modalities.
Why it matters
The clearest statement of the native-multimodal bet: that joint training beats bolting encoders onto a language model. Worth pairing with the scaling study in this collection, which examines whether that bet pays off under a fixed compute budget.
https://arxiv.org/abs/2312.11805