Hummingbird-V2 / TRAINING_DATA.md
juinron's picture
Publish Hummingbird-V2 10B base model and model card
ebd2f40 verified
|
Raw History Blame Contribute Delete
1.95 kB

Hummingbird-V2 training data

Hummingbird-V2 was trained from scratch on the 10,000,000,000-token packed FlightMix Balanced training split. Separate 25,000,000-token validation and 25,000,000-token held-out splits were not used as training records. The selected checkpoint is the 10B-token checkpoint at step 31,250.

Source Hub repository Training tokens Declared source license
FineWeb-Edu HuggingFaceFW/fineweb-edu 2,050,000,000 ODC-By 1.0
FineWeb-Edu-Dedup HuggingFaceTB/smollm-corpus 3,450,000,000 ODC-By 1.0
Cosmopedia v2 HuggingFaceTB/smollm-corpus 1,500,000,000 ODC-By 1.0
FineMath 4+ HuggingFaceTB/finemath 1,000,000,000 ODC-By 1.0
DCLM baseline 1.0 mlfoundations/dclm-baseline-1.0 1,000,000,000 CC-BY 4.0
Science exposition allenai/dolma3_dolmino_mix-100B-1025 400,000,000 ODC-By 1.0
Grounded QA allenai/dolma3_dolmino_mix-100B-1025 200,000,000 ODC-By 1.0
FinePDFs-Edu HuggingFaceFW/finepdfs-edu 300,000,000 ODC-By 1.0
Educational code allenai/dolma3_dolmino_mix-10B-1025 100,000,000 ODC-By 1.0

The bundled training/corpus_contract.yaml records pinned upstream revisions, source selections, quality filters, deduplication, and split assignment. The bundled training/packed_metadata.json records the realized packed token and document totals. The pipeline used a benchmark protection index for rendered HellaSwag, ARC-Easy, ARC-Challenge, PIQA, and ArithMark-3 prompts and choices. Protected matching reduces direct overlap but cannot prove complete absence of benchmark-related content in upstream data.

Public task scores across checkpoints and comparisons with other candidate runs informed the release choice. This selection bias is disclosed in README.md. Dataset licenses and notices remain applicable to the upstream data; the Apache-2.0 license for this model package does not replace them.