Running on CPU Upgrade
282
The Synthetic Data Playbook: Generating Trillions of the Finest Tokens
π
Explore synthetic data benchmarks with an interactive bookshelf
We release large pre-training datasets to accelerate open LLM development. Part of the Hugging Face Science team (hf.co/science)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Explore synthetic data benchmarks with an interactive bookshelf
Viewer to explore the finewiki dataset
Explore and download the FineWeb webβscale text dataset
Evaluate multilingual models using FineTasks
Explore and analyze experiment results
Launch an interactive demo interface