Instructions to use juinron/Hummingbird-V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use juinron/Hummingbird-V2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="juinron/Hummingbird-V2", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("juinron/Hummingbird-V2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use juinron/Hummingbird-V2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "juinron/Hummingbird-V2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juinron/Hummingbird-V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/juinron/Hummingbird-V2
- SGLang
How to use juinron/Hummingbird-V2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "juinron/Hummingbird-V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juinron/Hummingbird-V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "juinron/Hummingbird-V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juinron/Hummingbird-V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use juinron/Hummingbird-V2 with Docker Model Runner:
docker model run hf.co/juinron/Hummingbird-V2
Download TRAINING_DATA.md from juinron/Hummingbird-V2: direct link, hf CLI and curl.
- Browser
- Download file 1.95 kB
-
https://huggingface.co/juinron/Hummingbird-V2/resolve/main/TRAINING_DATA.md
- Command line
-
hf download hf://juinron/Hummingbird-V2/TRAINING_DATA.md
-
curl -L -o TRAINING_DATA.md https://huggingface.co/juinron/Hummingbird-V2/resolve/main/TRAINING_DATA.md
Hummingbird-V2 training data
Hummingbird-V2 was trained from scratch on the 10,000,000,000-token packed FlightMix Balanced training split. Separate 25,000,000-token validation and 25,000,000-token held-out splits were not used as training records. The selected checkpoint is the 10B-token checkpoint at step 31,250.
| Source | Hub repository | Training tokens | Declared source license |
|---|---|---|---|
| FineWeb-Edu | HuggingFaceFW/fineweb-edu |
2,050,000,000 | ODC-By 1.0 |
| FineWeb-Edu-Dedup | HuggingFaceTB/smollm-corpus |
3,450,000,000 | ODC-By 1.0 |
| Cosmopedia v2 | HuggingFaceTB/smollm-corpus |
1,500,000,000 | ODC-By 1.0 |
| FineMath 4+ | HuggingFaceTB/finemath |
1,000,000,000 | ODC-By 1.0 |
| DCLM baseline 1.0 | mlfoundations/dclm-baseline-1.0 |
1,000,000,000 | CC-BY 4.0 |
| Science exposition | allenai/dolma3_dolmino_mix-100B-1025 |
400,000,000 | ODC-By 1.0 |
| Grounded QA | allenai/dolma3_dolmino_mix-100B-1025 |
200,000,000 | ODC-By 1.0 |
| FinePDFs-Edu | HuggingFaceFW/finepdfs-edu |
300,000,000 | ODC-By 1.0 |
| Educational code | allenai/dolma3_dolmino_mix-10B-1025 |
100,000,000 | ODC-By 1.0 |
The bundled training/corpus_contract.yaml records pinned upstream revisions,
source selections, quality filters, deduplication, and split assignment. The
bundled training/packed_metadata.json records the realized packed token and
document totals. The pipeline used a benchmark protection index for rendered
HellaSwag, ARC-Easy, ARC-Challenge, PIQA, and ArithMark-3 prompts and choices.
Protected matching reduces direct overlap but cannot prove complete absence
of benchmark-related content in upstream data.
Public task scores across checkpoints and comparisons with other candidate runs informed the release choice. This selection bias is disclosed in README.md. Dataset licenses and notices remain applicable to the upstream data; the Apache-2.0 license for this model package does not replace them.