Research Library

Library

Papers, patents & whitepapers, annotated

Whitepaper2024

DeepSeek-V3 Technical Report

DeepSeek-AI, Aixin Liu, Bei Feng et al. · DeepSeek (Technical Report) · 2024

Read it for the training economics rather than the benchmarks. It is the clearest public account of how a frontier-scale run is actually made affordable, and much of what followed in open-weight modelling traces back to choices documented here.

Whitepaper2024

Qwen2.5 Technical Report

Qwen, :, An Yang et al. · Alibaba Qwen (Technical Report) · 2024

The most practical open-weight family to build on, because a consistent recipe across sizes means you can prototype small and scale up without changing behaviour underneath you.

Whitepaper2024

The Llama 3 Herd of Models

Meta AI · Meta (Technical Report) · 2024

A rare open look at how a frontier-scale model is actually built end to end. Read it when you want the engineering reality behind the benchmarks.

Whitepaper2024

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Marah Abdin, Jyoti Aneja, Hany Awadalla et al. · Microsoft (Technical Report) · 2024

The strongest evidence that data curation substitutes for parameters at the small end. Relevant if you are deciding between a hosted frontier model and something you can run on-device.

Whitepaper2024

StarCoder 2 and The Stack v2: The Next Generation

Anton Lozhkov, Raymond Li, Loubna Ben Allal et al. · BigCode (Technical Report) · 2024

Notable for treating training-data provenance as a design problem with an actual mechanism rather than a disclaimer. If you care where code models get their material, this is the reference implementation.

Whitepaper2023

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud et al. · Google DeepMind (Technical Report) · 2023

The clearest statement of the native-multimodal bet: that joint training beats bolting encoders onto a language model. Worth pairing with the scaling study in this collection, which examines whether that bet pays off under a fixed compute budget.

Whitepaper2023

Mistral 7B

Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch et al. · Mistral AI (Technical Report) · 2023

The paper that made small open models credible. Both attention tricks are now near-universal, so it doubles as the clearest short explanation of why modern models serve cheaply.

Whitepaper2023

Llama 2: Open Foundation and Fine-Tuned Chat Models

Hugo Touvron, Louis Martin, Kevin Stone et al. · Meta (Technical Report) · 2023

For years this was the reference text on how a chat model is actually aligned, because it showed the process rather than just the result. The two-reward-model split is the detail most worth carrying away.

Whitepaper2023

PaLM 2 Technical Report

Rohan Anil, Andrew M. Dai, Orhan Firat et al. · Google (Technical Report) · 2023

A useful marker of when the industry stopped equating scale with parameters. Read it next to the Chinchilla paper here — this is that argument applied at production scale.

Whitepaper2023

GPT-4 Technical Report

OpenAI, Josh Achiam, Steven Adler et al. · OpenAI (Technical Report) · 2023

The moment frontier reports stopped being reproducible science and became capability-and-safety disclosures. Worth reading alongside an open report like DeepSeek's or OLMo's to see exactly which questions each one refuses to answer.

Whitepaper2022

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al. · Anthropic (Technical Report) · 2022

Notable for making the values explicit and auditable rather than implicit in whichever labels the annotators happened to produce. The self-critique loop is also the ancestor of the LLM-as-judge pipelines now used far beyond safety work.

Whitepaper2022

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BigScience Workshop, :, Teven Le Scao et al. · BigScience (Technical Report) · 2022

The first serious demonstration that a frontier-scale model could be built outside a private lab. Its language coverage remains unusual — most open models are still overwhelmingly English.

Showing 112 of 12 sources · newest first