Instructions to use TrustLLMeu/trustllm-v2-8b-sft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TrustLLMeu/trustllm-v2-8b-sft with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="TrustLLMeu/trustllm-v2-8b-sft", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("TrustLLMeu/trustllm-v2-8b-sft", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use TrustLLMeu/trustllm-v2-8b-sft with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TrustLLMeu/trustllm-v2-8b-sft" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TrustLLMeu/trustllm-v2-8b-sft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/TrustLLMeu/trustllm-v2-8b-sft
- SGLang
How to use TrustLLMeu/trustllm-v2-8b-sft with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "TrustLLMeu/trustllm-v2-8b-sft" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TrustLLMeu/trustllm-v2-8b-sft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "TrustLLMeu/trustllm-v2-8b-sft" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TrustLLMeu/trustllm-v2-8b-sft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use TrustLLMeu/trustllm-v2-8b-sft with Docker Model Runner:
docker model run hf.co/TrustLLMeu/trustllm-v2-8b-sft
TrustLLM v2 8B SFT
Supervised fine-tuned checkpoint of TrustLLMeu/trustllm-v2-8b, prepared by the TrustLLM project.
Revision note (27 September 2026).
mainnow holds the M3 release candidate: the same nine-language SFT data, trained from a midtraining continuation of the base model with a corrected schedule. It gains 12–17 points on grade-school math in English, German and Icelandic, and matches or slightly beats the previous release on the multiple-choice tasks (evaluation below).Earlier revisions stay available:
Research-only; not for redistribution or commercial use. Access requests are reviewed manually by the TrustLLM project, under the same access conditions as the base model.
Model and training
The model has approximately 7.19 billion total parameters, 24 layers, 64 routed experts with top-8 routing, and one shared expert. The repository contains BF16 weights, the tokenizer, chat template, generation configuration, and the custom OptMoE model code.
Starting point. Step 2500 of a TrustLLM midtraining run on the final step-120000 base checkpoint: about 84B tokens of mostly English continued pretraining, with Icelandic about 0.4% of tokens.
SFT. Full-parameter training with the TorchTitan fork and DiSCO optimizer on 16 H100 GPUs at MareNostrum 5.
| Setting | Value |
|---|---|
| Context length | 4,096 tokens |
| Global batch size | 128 packed sequences |
| Optimizer | DiSCO, learning rate 0.1: 64 warmup steps, constant, then linear decay to 0 over the last 50% of steps |
| Training steps | 2,048 (this revision is the final step) |
| Precision | FP32 parameters and gradient reduction; BF16 computation and exported weights |
| Objective | Assistant-only loss with conversation packing |
Checkpoint selection. A selection rule was frozen before training. Among checkpoints 1024/1536/1792/2048 that pass two gates, it picks the highest mean of ARC-IS, Belebele-IS, WinoGrande-IS, ARC-EN and MGSM-EN. The gates are fewer than 1 Scandinavian-suffix non-word per 1,000 Icelandic words, and at most 10% degenerate outputs. It selected step 2048. An independent second seed selected the same step (rule score 38.72, against 39.09 for this seed).
Data
Same SFT data and mix as the earlier revisions: an export of allenai/Dolci-Instruct-SFT and project translations into Icelandic, Faroese, Norwegian Bokmål, Norwegian Nynorsk, Swedish, Danish, Dutch and German. Sampling targeted about 50% English-partition tokens and equal shares of the eight translation languages. Source IDs and their translations were grouped across train, validation and test splits.
The included chat template defaults to non-thinking generation. Tool schemas and function-call strings were preserved verbatim; environment messages were treated as tool responses and excluded from the loss.
Evaluation
Development evaluation, the same protocol for both revisions. Multiple-choice tasks use option-text log-likelihood under the chat template. Math is greedy step-by-step generation, scored on the last number in the answer.
| this revision (M3) | previous main (step 1792) |
|
|---|---|---|
| ARC-Challenge, English (acc_norm, n = 1,172) | 0.359 | 0.347 |
| ARC-Challenge, Icelandic translation (acc_norm) | 0.324 | 0.317 |
| Belebele, Icelandic (acc_norm, 300-item subset) | 0.343 | 0.340 |
| WinoGrande, Icelandic (partial scoring, n = 1,088) | 0.536 | 0.535 |
| MGSM, English (n = 250) | 0.392 | 0.224 |
| MGSM, German (n = 250) | 0.332 | 0.200 |
| MGSM, Icelandic (hafsteinn/mgsm-is, machine-translated, n = 250) | 0.256 | 0.164 |
| GSM8K test, English (n = 1,319) | 0.375 | 0.221 |
| GSM8K test, Icelandic translation (n = 1,319) | 0.243 | 0.157 |
| Icelandic answers: non-words with Scandinavian endings per 1,000 words (200 open questions) | 0.08 | 0.11 |
What is and is not established:
- The math gain is the robust result. Across 3 seeds of a matched 256-step comparison, SFT from the midtraining checkpoint gains +11.9 (MGSM-EN), +5.6 (MGSM-DE) and +8.4 (MGSM-IS) points, and every seed-paired 95% interval excludes zero.
- The multiple-choice differences are within noise.
- Midtraining did not add Icelandic factual knowledge in our probes.
- Absolute scores are modest: this is a short multilingual SFT, not a tuned assistant.
- The non-word count is a heuristic: word forms absent from the BÍN inflection database that end in a mainland-Scandinavian suffix.
Translated training and evaluation data may contain translation artifacts, and the Icelandic math sets have not been manually verified. No safety evaluation has been performed.
Loading notes
Access approval and Hugging Face authentication are required to download the weights. The architecture uses the included custom Transformers code, so loading through Transformers requires trust_remote_code=True. Use the supplied chat template with enable_thinking=False. The generation configuration stops on both the end-of-text and end-of-message token IDs.
Evaluation used Transformers with the included remote code (BF16, eager attention, batch size 1). Batched left-padded generation did not match single-sequence generation in our tests, so use batch size 1 or verify batching. The repository does not claim compatibility with a particular vLLM release.
- Downloads last month
- 23
Model tree for TrustLLMeu/trustllm-v2-8b-sft
Base model
TrustLLMeu/trustllm-v2-8b