Library

Research Library

Whitepaper2023

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac et al. · Google DeepMind (Technical Report) · 2023

Abstract

Google DeepMind's report on a model family built multimodal from the start — trained jointly on text, images, audio and video rather than assembling a language model and a vision encoder afterwards. It spans three sizes aimed at datacentre, general and on-device use, and reports across all four modalities.

Why it matters

The clearest statement of the native-multimodal bet: that joint training beats bolting encoders onto a language model. Worth pairing with the scaling study in this collection, which examines whether that bet pays off under a fixed compute budget.

multimodalarchitectureevalslong context
Read the source

https://arxiv.org/abs/2312.11805