# Hummingbird-V2 training data Hummingbird-V2 was trained from scratch on the 10,000,000,000-token packed FlightMix Balanced training split. Separate 25,000,000-token validation and 25,000,000-token held-out splits were not used as training records. The selected checkpoint is the 10B-token checkpoint at step 31,250. | Source | Hub repository | Training tokens | Declared source license | |---|---|---:|---| | FineWeb-Edu | `HuggingFaceFW/fineweb-edu` | 2,050,000,000 | ODC-By 1.0 | | FineWeb-Edu-Dedup | `HuggingFaceTB/smollm-corpus` | 3,450,000,000 | ODC-By 1.0 | | Cosmopedia v2 | `HuggingFaceTB/smollm-corpus` | 1,500,000,000 | ODC-By 1.0 | | FineMath 4+ | `HuggingFaceTB/finemath` | 1,000,000,000 | ODC-By 1.0 | | DCLM baseline 1.0 | `mlfoundations/dclm-baseline-1.0` | 1,000,000,000 | CC-BY 4.0 | | Science exposition | `allenai/dolma3_dolmino_mix-100B-1025` | 400,000,000 | ODC-By 1.0 | | Grounded QA | `allenai/dolma3_dolmino_mix-100B-1025` | 200,000,000 | ODC-By 1.0 | | FinePDFs-Edu | `HuggingFaceFW/finepdfs-edu` | 300,000,000 | ODC-By 1.0 | | Educational code | `allenai/dolma3_dolmino_mix-10B-1025` | 100,000,000 | ODC-By 1.0 | The bundled `training/corpus_contract.yaml` records pinned upstream revisions, source selections, quality filters, deduplication, and split assignment. The bundled `training/packed_metadata.json` records the realized packed token and document totals. The pipeline used a benchmark protection index for rendered HellaSwag, ARC-Easy, ARC-Challenge, PIQA, and ArithMark-3 prompts and choices. Protected matching reduces direct overlap but cannot prove complete absence of benchmark-related content in upstream data. Public task scores across checkpoints and comparisons with other candidate runs informed the release choice. This selection bias is disclosed in README.md. Dataset licenses and notices remain applicable to the upstream data; the Apache-2.0 license for this model package does not replace them.