worthant commited on
Commit
227d5b1
·
verified ·
1 Parent(s): 0bdffb7

forge: regenerate the model card

Browse files
Files changed (1) hide show
  1. README.md +33 -34
README.md CHANGED
@@ -10,12 +10,11 @@ pipeline_tag: text-generation
10
  library_name: gguf
11
  tags:
12
  - atomic-chat
 
13
  - qwen
14
- - qwen3
15
  - gguf
16
- - imatrix
17
- - quantized
18
  - llama.cpp
 
19
  ---
20
 
21
  <center>
@@ -28,23 +27,24 @@ tags:
28
 
29
  <br/>
30
 
31
- <img src="https://huggingface.co/AtomicChat/qwen36-27b-GGUF/resolve/main/hero.png" alt="Qwen3.6 27B" style="width:420px; max-width:100%; height:auto; margin-bottom:0.6em;"/>
32
 
33
  <div style="display:flex; justify-content:center; gap:0.5em;">
34
  <a href="https://huggingface.co/Qwen/Qwen3.6-27B"><strong>Base model: Qwen/Qwen3.6-27B</strong></a>
35
  </div>
36
  </center>
37
 
38
- **Qwen3.6 27B**, self-quantized to GGUF by [Atomic Chat](https://atomic.chat). Built straight from Qwen's original weights with a per-tensor importance matrix. Runs fully offline.
39
 
40
  ## Highlights
41
 
42
- - **First open-weight Qwen3.6 variant**, following the Qwen3.5 series, with a focus on stability and real-world utility.
43
- - **Agentic coding** that handles frontend workflows and repository-level reasoning with greater fluency and precision.
44
- - **Thinking Preservation**, a new option to retain reasoning context from historical messages to streamline iterative development.
45
- - **Base model is multimodal** (vision encoder); these GGUF quants cover the text path.
46
- - **262,144-token native context**, extensible up to ~1,010,000 tokens.
47
- - **Full quant ladder** with an importance matrix on every quant over [`calibration_datav3`](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8).
 
48
 
49
  > [!NOTE]
50
  > These GGUFs are **self-quantized from the original weights**, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
@@ -57,32 +57,33 @@ tags:
57
  | Property | Value |
58
  |---|---|
59
  | Base model | `Qwen/Qwen3.6-27B` |
60
- | Total parameters | 27B |
61
  | Layers | 64 |
62
- | Context length | 262,144 native, extensible up to ~1,010,000 |
63
- | Architecture | Causal LM with vision encoder (Gated DeltaNet + Gated Attention) |
64
- | This repo | GGUF quants (imatrix), text path |
 
 
65
 
66
  <img src="https://huggingface.co/AtomicChat/qwen36-27b-GGUF/resolve/main/benchmark.png" alt="Qwen3.6 27B benchmark scores" style="width:100%; max-width:900px;"/>
67
 
68
- Scores are Qwen's published results for the base `Qwen/Qwen3.6-27B`. Quantization preserves the large majority of this; `Q4_K_M` and up sit within a point or two of full precision.
69
-
70
 
71
  ## Choosing a quant
72
 
73
  | Quant | Size | Notes |
74
  |---|---|---|
75
- | `Q2_K` | 10.7 GB | Smallest. Minimal RAM, clear quality drop. |
76
- | `IQ3_M` | 12.6 GB | Beats Q3 at similar size thanks to imatrix. Best low-RAM pick. |
77
  | `Q3_K_M` | 13.3 GB | Low quality but usable. |
78
  | `Q3_K_L` | 14.3 GB | A step above Q3_K_M. |
79
  | `IQ4_XS` | 15.1 GB | Excellent quality for size. Recommended low-bit. |
80
- | `Q4_K_S` | 15.6 GB | Compact Q4, fast. |
81
  | **`Q4_K_M`** | 16.5 GB | **Recommended default. Best balance of size, speed and quality.** |
82
- | **`UD-Q4_K_XL`** | 17.5 GB | **Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint.** |
83
- | `Q5_K_S` | 18.7 GB | Higher quality. |
84
  | `Q5_K_M` | 19.2 GB | Higher quality, low loss. |
85
- | `Q6_K` | 22.1 GB | Near lossless. |
86
  | `Q8_0` | 28.6 GB | Effectively lossless, reference quality. |
87
 
88
  > [!TIP]
@@ -101,38 +102,36 @@ Run Qwen3.6 27B locally with:
101
 
102
  | Parameter | Value |
103
  |---|---|
104
- | temperature | 0.7 |
105
- | top_p | 0.8 |
106
  | top_k | 20 |
107
  | min_p | 0.0 |
108
- | presence_penalty | 1.5 |
109
  | repetition_penalty | 1.0 |
110
 
111
- Qwen's recommended Instruct (non-thinking) settings. Thinking mode for general tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0.
112
 
113
  ## Run in llama.cpp
114
 
115
  ```bash
116
- git clone https://github.com/ggerganov/llama.cpp
117
  cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
118
  cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
119
  ```
120
 
121
  ```bash
122
  ./llama.cpp/build/bin/llama-server \
123
- -hf AtomicChat/qwen36-27b-GGUF:UD-Q4_K_XL \
124
  --jinja -ngl 99 -c 8192 -fa on
125
  ```
126
 
127
  ## How these were made
128
 
129
  1. Download `Qwen/Qwen3.6-27B` (original weights).
130
- 2. Convert to f16 GGUF with [llama.cpp](https://github.com/ggerganov/llama.cpp).
131
- 3. Build an importance matrix over `calibration_datav3` (100 chunks).
132
- 4. Quantize the full ladder with `--imatrix`.
133
  5. `UD-Q4_K_XL` additionally pins the token-embedding and output tensors to `Q8_0`.
134
 
135
  ## License
136
 
137
- Released by Qwen under the Apache 2.0 license. Quantized by Atomic Chat.
138
-
 
10
  library_name: gguf
11
  tags:
12
  - atomic-chat
13
+ - qwen3.6
14
  - qwen
 
15
  - gguf
 
 
16
  - llama.cpp
17
+ - quantized
18
  ---
19
 
20
  <center>
 
27
 
28
  <br/>
29
 
30
+ <img src="https://huggingface.co/AtomicChat/qwen36-27b-GGUF/resolve/main/hero.png" alt="Qwen3.6 27B" style="width:100%; max-width:100%; height:auto; margin-bottom:0.6em;"/>
31
 
32
  <div style="display:flex; justify-content:center; gap:0.5em;">
33
  <a href="https://huggingface.co/Qwen/Qwen3.6-27B"><strong>Base model: Qwen/Qwen3.6-27B</strong></a>
34
  </div>
35
  </center>
36
 
37
+ **Qwen3.6 27B**, self-quantized to GGUF by [Atomic Chat](https://atomic.chat). Built straight from Qwen's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
38
 
39
  ## Highlights
40
 
41
+ - **27.8B parameters**: the weights this repo quantizes.
42
+ - **Context length**: 262,144 tokens (256K), as published by Qwen.
43
+ - **64 layers**: Dense decoder.
44
+ - **Modalities**: the base model handles Text, Image; this repo ships text-only quants, it carries no vision projector.
45
+ - **Full imatrix ladder**: every quant is calibrated with an importance matrix.
46
+ - **Agentic Coding:**: the model now handles frontend workflows and repository-level reasoning with greater fluency and precision.
47
+ - **Thinking Preservation:**: we've introduced a new option to retain reasoning context from historical messages, streamlining iterative development and reducing overhead.
48
 
49
  > [!NOTE]
50
  > These GGUFs are **self-quantized from the original weights**, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
 
57
  | Property | Value |
58
  |---|---|
59
  | Base model | `Qwen/Qwen3.6-27B` |
60
+ | Parameters | 27.8B |
61
  | Layers | 64 |
62
+ | Context length | 262,144 tokens (256K) |
63
+ | Vocabulary | 248,320 |
64
+ | Modalities | Text, Image in the base model; text only in this repo, it ships no vision projector |
65
+ | Architecture | Dense decoder, 24 attention heads over 4 KV heads, `Qwen3_5ForConditionalGeneration` |
66
+ | This repo | GGUF quants (imatrix). Quants: `Q2_K`, `IQ3_M`, `Q3_K_M`, `Q3_K_L`, `IQ4_XS`, `Q4_K_S`, `Q4_K_M`, `UD-Q4_K_XL`, `Q5_K_S`, `Q5_K_M`, `Q6_K`, `Q8_0` |
67
 
68
  <img src="https://huggingface.co/AtomicChat/qwen36-27b-GGUF/resolve/main/benchmark.png" alt="Qwen3.6 27B benchmark scores" style="width:100%; max-width:900px;"/>
69
 
70
+ Scores are Qwen's published results for the base `Qwen/Qwen3.6-27B`, not our own measurements. Quantization preserves the large majority of this; `Q4_K_M` and up stay close to full precision.
 
71
 
72
  ## Choosing a quant
73
 
74
  | Quant | Size | Notes |
75
  |---|---|---|
76
+ | `Q2_K` | 10.7 GB | Smallest K-quant. Minimal RAM, clear quality drop. |
77
+ | `IQ3_M` | 12.6 GB | Beats Q3 at a similar size thanks to imatrix. Best low-RAM pick. |
78
  | `Q3_K_M` | 13.3 GB | Low quality but usable. |
79
  | `Q3_K_L` | 14.3 GB | A step above Q3_K_M. |
80
  | `IQ4_XS` | 15.1 GB | Excellent quality for size. Recommended low-bit. |
81
+ | `Q4_K_S` | 15.6 GB | Compact 4-bit, fast. |
82
  | **`Q4_K_M`** | 16.5 GB | **Recommended default. Best balance of size, speed and quality.** |
83
+ | `UD-Q4_K_XL` | 17.5 GB | Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint. |
84
+ | `Q5_K_S` | 18.7 GB | Higher quality, slightly more compact than Q5_K_M. |
85
  | `Q5_K_M` | 19.2 GB | Higher quality, low loss. |
86
+ | `Q6_K` | 22.1 GB | Near lossless, noticeably lighter than Q8_0. |
87
  | `Q8_0` | 28.6 GB | Effectively lossless, reference quality. |
88
 
89
  > [!TIP]
 
102
 
103
  | Parameter | Value |
104
  |---|---|
105
+ | temperature | 1.0 |
106
+ | top_p | 0.95 |
107
  | top_k | 20 |
108
  | min_p | 0.0 |
 
109
  | repetition_penalty | 1.0 |
110
 
111
+ Qwen's recommended sampling configuration for `Qwen/Qwen3.6-27B`.
112
 
113
  ## Run in llama.cpp
114
 
115
  ```bash
116
+ git clone https://github.com/ggml-org/llama.cpp
117
  cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
118
  cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
119
  ```
120
 
121
  ```bash
122
  ./llama.cpp/build/bin/llama-server \
123
+ -hf AtomicChat/qwen36-27b-GGUF:Q4_K_M \
124
  --jinja -ngl 99 -c 8192 -fa on
125
  ```
126
 
127
  ## How these were made
128
 
129
  1. Download `Qwen/Qwen3.6-27B` (original weights).
130
+ 2. Convert to f16 GGUF with [llama.cpp](https://github.com/ggml-org/llama.cpp).
131
+ 3. Build an importance matrix over our calibration corpus.
132
+ 4. Quantize the ladder with `--imatrix`.
133
  5. `UD-Q4_K_XL` additionally pins the token-embedding and output tensors to `Q8_0`.
134
 
135
  ## License
136
 
137
+ Original model by Qwen, released under the Apache 2.0 license. Full terms: [Apache 2.0](https://huggingface.co/Qwen/Qwen3.6-27B/blob/main/LICENSE). Quantized by Atomic Chat.