Multi-task learning using knowledge distillation
Google LLC · USPTO — US10635977B2 · 2020
Abstract
Covers training a smaller network to reproduce the outputs of a larger one across several tasks at once, using the larger model's full output distribution as the training signal rather than hard labels alone. The claims cover the multi-task arrangement and the transfer of knowledge between the two networks.
Why it matters
Distillation as claimed intellectual property, and the technique behind nearly every small model that punches above its size — including several on this shelf. Read it next to the on-policy distillation work here to see how far the idea has since travelled.
https://patents.google.com/patent/US10635977B2/en