Research Library

Library

Papers, patents & whitepapers, annotated

Paper2015

Deep Residual Learning for Image Recognition

Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · CVPR · 2015

Residual connections are in every transformer block you will ever use, which makes this required background even though the paper is about images. The framing is the lesson: the obstacle to depth was trainability, not capacity.

Patent2015

Computing numeric representations of words in a high-dimensional space

Google LLC · USPTO — US9037464B1 · 2015

The word-embedding idea as granted intellectual property. Useful for seeing how a technique that became foundational infrastructure was framed in claim language — and a reminder that the paper and the patent are separate artefacts with different purposes.

Paper2014

Adam: A Method for Stochastic Optimization

Diederik P. Kingma, Jimmy Ba · ICLR · 2014

Still the default optimiser for essentially every model on this shelf, more than a decade on. Worth reading precisely because it is the piece of the stack most people never look inside.

Paper2014

Sequence to Sequence Learning with Neural Networks

Ilya Sutskever, Oriol Vinyals, Quoc V. Le · NIPS · 2014

The structural ancestor of every generative model you use. Its weakness — squeezing a whole sentence through one vector — is the specific problem attention was invented to solve three years later.

Preprint2014

Generative Adversarial Networks

Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza et al. · arXiv · 2014

The adversarial framing dominated generative modelling for most of a decade before diffusion displaced it. Worth reading for the idea that a learned critic can replace a hand-specified loss — a pattern that reappears throughout modern alignment work.

Paper2013

Efficient Estimation of Word Representations in Vector Space

Tomas Mikolov, Kai Chen, Greg Corrado et al. · ICLR · 2013

Where the idea that meaning can live in a vector became practical, and the direct ancestor of every embedding in your retrieval stack. Also the most approachable paper here — the model is simple enough to read in one sitting.

Showing 6166 of 66 sources · newest first