Research Library
Library
Papers, patents & whitepapers, annotated
DeepSeek-V3 Technical Report
DeepSeek-AI, Aixin Liu, Bei Feng et al. · DeepSeek (Technical Report) · 2024
Read it for the training economics rather than the benchmarks. It is the clearest public account of how a frontier-scale run is actually made affordable, and much of what followed in open-weight modelling traces back to choices documented here.
Qwen2.5 Technical Report
Qwen, :, An Yang et al. · Alibaba Qwen (Technical Report) · 2024
The most practical open-weight family to build on, because a consistent recipe across sizes means you can prototype small and scale up without changing behaviour underneath you.
The Llama 3 Herd of Models
Meta AI · Meta (Technical Report) · 2024
A rare open look at how a frontier-scale model is actually built end to end. Read it when you want the engineering reality behind the benchmarks.
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Marah Abdin, Jyoti Aneja, Hany Awadalla et al. · Microsoft (Technical Report) · 2024
The strongest evidence that data curation substitutes for parameters at the small end. Relevant if you are deciding between a hosted frontier model and something you can run on-device.
StarCoder 2 and The Stack v2: The Next Generation
Anton Lozhkov, Raymond Li, Loubna Ben Allal et al. · BigCode (Technical Report) · 2024
Notable for treating training-data provenance as a design problem with an actual mechanism rather than a disclaimer. If you care where code models get their material, this is the reference implementation.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud et al. · Google DeepMind (Technical Report) · 2023
The clearest statement of the native-multimodal bet: that joint training beats bolting encoders onto a language model. Worth pairing with the scaling study in this collection, which examines whether that bet pays off under a fixed compute budget.
Mistral 7B
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch et al. · Mistral AI (Technical Report) · 2023
The paper that made small open models credible. Both attention tricks are now near-universal, so it doubles as the clearest short explanation of why modern models serve cheaply.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone et al. · Meta (Technical Report) · 2023
For years this was the reference text on how a chat model is actually aligned, because it showed the process rather than just the result. The two-reward-model split is the detail most worth carrying away.
PaLM 2 Technical Report
Rohan Anil, Andrew M. Dai, Orhan Firat et al. · Google (Technical Report) · 2023
A useful marker of when the industry stopped equating scale with parameters. Read it next to the Chinchilla paper here — this is that argument applied at production scale.
GPT-4 Technical Report
OpenAI, Josh Achiam, Steven Adler et al. · OpenAI (Technical Report) · 2023
The moment frontier reports stopped being reproducible science and became capability-and-safety disclosures. Worth reading alongside an open report like DeepSeek's or OLMo's to see exactly which questions each one refuses to answer.
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al. · Anthropic (Technical Report) · 2022
Notable for making the values explicit and auditable rather than implicit in whichever labels the annotators happened to produce. The self-critique loop is also the ancestor of the LLM-as-judge pipelines now used far beyond safety work.
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BigScience Workshop, :, Teven Le Scao et al. · BigScience (Technical Report) · 2022
The first serious demonstration that a frontier-scale model could be built outside a private lab. Its language coverage remains unusual — most open models are still overwhelmingly English.
Showing 1–12 of 12 sources · newest first
About Nybble™
The AI space moves fast.
Nybble™ is how you keep up — and stay sharp.
What happened. In two minutes.
The AI news cycle moves at a pace no one can keep up with. Snack distills what launched, what shipped, and what matters — every day, without the filler.
Go to SnackThe concepts behind the headlines.
News tells you what. Stack tells you why and how. From RAG architectures to agentic evals, these are the ideas that will shape what you build next.
Go to StackProve you actually get it.
Reading about LangChain is not the same as knowing it. Hack challenges you with production-grade questions, then shows you the references that make the answer stick.
Go to Hack