DataArticle
The Stack v3 is the largest open code dataset publicly released, with 114 TB raw, 224M repositories, 44B files, 770 languages, and roughly 5T deduplicated/filtered tokens.
Anton Lozhkov announced The Stack v3, now the largest open code dataset publicly released, with significant gains over v2 including a jump from ~550B to ~5T filtered tokens and large per-language increases in C++, TypeScript, Rust, and Python. ✦ AI generated
AINews / Latent.Space · Latent Space · 2026-07-24 · original ↗
@anton_lozhkov announced The Stack v3, now the largest open code dataset publicly released: 114 TB raw, 224M repositories, 44B files, 770 languages, and roughly 5T deduplicated/filtered tokens. Relative to v2, the filtered corpus jumps from ~550B to ~5T tokens, with especially large gains in C++ (x15), TypeScript (x7.5), Rust (x7), and Python (x4.8).
Read full article ↗excerpt · fair-use quotation
Around this claim