StarCoder 2 and The Stack v2: The Next Generation
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano et al. · BigCode (Technical Report) · 2024
Abstract
An open code model together with the corpus it was trained on, drawn from over six hundred programming languages plus issues, notebooks and commits. The data pipeline is the substance — license filtering, deduplication, and an opt-out process for developers who did not want their code included.
Why it matters
Notable for treating training-data provenance as a design problem with an actual mechanism rather than a disclaimer. If you care where code models get their material, this is the reference implementation.
https://arxiv.org/abs/2402.19173