Instructions to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="local-inference-lab/Qwen3.8-Flash-Next-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("local-inference-lab/Qwen3.8-Flash-Next-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("local-inference-lab/Qwen3.8-Flash-Next-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "local-inference-lab/Qwen3.8-Flash-Next-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/Qwen3.8-Flash-Next-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/local-inference-lab/Qwen3.8-Flash-Next-NVFP4
- SGLang
How to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "local-inference-lab/Qwen3.8-Flash-Next-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/Qwen3.8-Flash-Next-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "local-inference-lab/Qwen3.8-Flash-Next-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/Qwen3.8-Flash-Next-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with Docker Model Runner:
docker model run hf.co/local-inference-lab/Qwen3.8-Flash-Next-NVFP4
Model Description
local-inference-lab/Qwen3.8-Flash-Next-NVFP4 is a mixed-precision model distilled from Qwen/Qwen3.8-Flash-Next using quantization-aware distillation (QAD). The student is trained against the original BF16 teacher with quantized weights in its forward pass, learning to compensate for quantization error rather than relying on post-training quantization alone.
The architecture is unchanged: 48 decoder layers, 512 routed experts per layer with 10 active per token, hybrid Gated DeltaNet/Qwen Sparse Attention (QSA), and n-gram embedding tables. Compression comes from lower-precision weights, not fewer layers or experts. The checkpoint occupies approximately 98 GiB on disk and is particularly suited to run on a single RTX 6000 (with PLE offload), or a DGX Spark (with or without PLE offload).
What's quantized
| Component | Weight format | Distillation |
|---|---|---|
| Text routed experts: gate, up and down projections | NVFP4 | Trained |
| N-gram embedding tables (PLE) | NVFP4 | Trained |
| Text shared experts: gate, up and down projections | MXFP8 | Trained |
| Text attention projections, including QSA indexers | MXFP8 | Frozen |
| Text routers, residual-stream mixing and PLE projections/convolution | BF16 | Trained |
| Text token embeddings, LM head, attention norms and recurrent parameters | BF16 | Frozen |
| Vision encoder | MXFP8 attention/FC1/merger projections, NVFP4 FC2; remaining parameters BF16 | Unchanged |
| Multi-token prediction (MTP) module | NVFP4 routed experts; remaining parameters BF16 | Unchanged |
NVFP4 stores 4-bit E2M1 values with FP8 block scales per 16 elements and an FP32 global scale. MXFP8 stores 8-bit E4M3 values with power-of-two block scales per 32 elements. Residual-mixing and PLE normalization weights are trained in FP32 and exported in BF16. Shared-expert scalar gates remain frozen in BF16.
Vision and MTP retain their weights. They are included in the release but were not part of text distillation.
Quantization-aware distillation
The BF16 teacher generates responses and supplies token-level probability and hidden-state targets. The trainable NVFP4 and MXFP8 weights are quantized and reconstructed for each forward pass. Gradients update the underlying weights so their low-precision representations better match the teacher.
The objective combines next-token probability matching using total variation, hidden-state matching, and separate losses for end-of-response and end-of-thinking behavior. These boundary losses target the teacher's decisions to continue, finish reasoning, and stop responding.
Training uses the Quatrain distillation trainer: 2,500 trunk updates followed by 1,500 joint-refinement updates, which also train the n-gram embedding tables and simulate MXFP8 shared-expert weights. Attention remains frozen, using its MXFP8 weights during joint refinement.
Training data
The distillation corpus contains 200,004,844 prompt-and-response tokens across 53,573 documents, with sequences up to 8,192 tokens. Prompts cover coding, reasoning, general instruction following and multilingual conversations. Assistant responses are generated by the original BF16 Qwen teacher using its native chat template and thinking mode; responses from other models are not used as training targets.
Activation calibration
Distillation uses BF16 activations. Separately, routed-expert activation scales are calibrated over the 200M-token chat corpus and an additional 104M raw-text tokens, using natural routing. The raw-text corpus is used for calibration, not gradient updates.
The serving configuration uses static NVFP4 activation quantization for text routed experts and dynamic MXFP8 activation quantization for MXFP8 projections. PLE tables, MTP routed experts and vision FC2 use weight-only NVFP4. Activation quantization is therefore separate from the weight quantization simulated during distillation.
Requirements
Use a runtime that supports Qwen3.8-Flash-Next and this mixed NVFP4/MXFP8 ModelOpt layout, including NVFP4 PLE tables. The disk size is not a runtime memory estimate: KV cache and runtime buffers require additional memory. The tokenizer, chat template and vision assets are preserved from the mixed-precision base checkpoint.
Evaluation
| Model | AA-LCR | GPQA Diamond | IFBench | Tool Eval |
|---|---|---|---|---|
| NVIDIA NVFP4 | 74.1 | 91.5 | 81.0 | โ |
| local-inference-lab QAD | 79.4 | 91.41 | 82.31 | 95 |
- Downloads last month
- 53,918
Model tree for local-inference-lab/Qwen3.8-Flash-Next-NVFP4
Base model
Qwen/Qwen3.8-Flash-Next