❌

Normal view

  • ✇InfoQ
  • Hugging Face Releases FinePDFs: a 3-Trillion-Token Dataset Built from PDFs Robert Krzaczyński
    Hugging Face has unveiled FinePDFs, the largest publicly available corpus built entirely from PDFs. The dataset spans 475 million documents in 1,733 languages, totaling roughly 3 trillion tokens. At 3.65 terabytes in size, FinePDFs introduces a new dimension to open training datasets by tapping into a resource long considered too complex and expensive to process. By Robert Krzaczyński
     

Hugging Face Releases FinePDFs: a 3-Trillion-Token Dataset Built from PDFs

15 September 2025 at 16:55

Hugging Face has unveiled FinePDFs, the largest publicly available corpus built entirely from PDFs. The dataset spans 475 million documents in 1,733 languages, totaling roughly 3 trillion tokens. At 3.65 terabytes in size, FinePDFs introduces a new dimension to open training datasets by tapping into a resource long considered too complex and expensive to process.

By Robert Krzaczyński
❌