How to use from
Docker Model Runner
docker model run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0
Quick Links

code-daemon-enrich-v1

A distilled Qwen3.5-0.8B that writes the short, structured labels of a code-intelligence pipeline: names for clusters of code entities, names for graph communities, and link picks. It is the high-volume worker of the Code-Daemon daemon — a purpose-built component, not a general assistant; outside these tasks its behaviour is undefined.

What makes it different: label stages are prefill-bound — long prompt, a few words of answer — and run once per cluster, community and link candidate of an index. A sub-billion model is enough for them if it learns the contract: name a group instead of listing it, fold a level instead of concatenating it, decline a link instead of always picking one. Distillation from a 27B teacher taught it exactly that; the stock base and the 0.6B it replaced do not do it.

The weights under this id changed on 2026-09-16. Until then it served a Qwen3-0.6B distilled from Qwen2.5-7B; that model is retired and its card is in this repository's git history.

The tasks

task answer example
cluster label, levels 0 and 1 <topic>: name1, name2, name3 DI service provider construction: BuildServiceProvider, GetService, ServiceProviderOptions
graph-community label a 2-to-5-word noun phrase Go standard library packages
link selection a candidate id c<N>, or none c3

Numbers

Held-out rows scored as the GGUF the daemon loads, through llama-server:

Q8_0 GGUF label tasks, rougeL vs teacher link selection, exact of 41 community labels that are a bare list L1 label, median words
this model 0.493 18 0 % 6
stock Qwen3.5-0.8B 0.193 14 1 % 33
the retired Qwen3-0.6B 0.220 11 38 % 33
teacher (Qwen3.8-27B) — — 0 % 7
  • +0.300 rougeL over its own base (95 % CI +0.257 … +0.343, 158 wins / 22 losses). The base writes a paragraph where a label is asked for.
  • It can decline a link. The teacher answers none to 24 of 41 link questions, this model to 6, the retired 0.6B to none — every ambiguous case became an edge.
  • Where it is worse: it paraphrases what should be copied — iter, dur can come back as iterations, elapsed. If your use needs exact identifier echo, measure that first.

Speed and memory — the trade

Laptop RTX 5060 8 GB, CUDA, Q8_0, n_ctx=8192, live inside the daemon, against the retired 0.6B the same day:

this model retired 0.6B
decode, single stream 253 tok/s 305 tok/s
decode, 8 concurrent slots 794 tok/s 1 083 tok/s
label stage end to end, prompt tokens counted 5 858 tok/s 6 265 tok/s

Decode is 17–27 % slower; the whole stage is 6.5 % slower, because these stages are prefill-bound. Plan capacity from the last row.

Memory: 2 281 MB resident with 28 parallel sequences, not the 774 MB of the file. The model is hybrid — 18 of its 24 layers are gated DeltaNet — so each sequence carries 19.3 MiB of recurrent state whatever the context length. Fit a smaller card by cutting sequences, not context: half the sequences gives back ~270 MiB.

How to use it

ChatML, one user turn, no system turn, and an empty think block opening the answer — the model was trained that way and drifts without it. Greedy decoding (temperature 0).

llama-cli -m code-daemon-enrich-v1-Q8_0.gguf -c 8192 --temp 0 \
  -p '<|im_start|>user
Write ONE line — 2 to 5 plain-English words — labelling this group. No quotes, no explanation.

Group members:
- parseArgs
- Command
- Usage
<|im_end|>
<|im_start|>assistant
<think>

</think>

'

n_ctx=8192 covers the label prompts; parallel sequences are the memory knob (above).

How it was made

  • Base: Qwen/Qwen3.5-0.8B — 24 layers, 6 of them attention and 18 gated DeltaNet.
  • Teacher: Qwen3.8-27B, answering the prompts the daemon sends in production.
  • Method: sequence-level knowledge distillation, merged into the base and exported to GGUF without the base's multi-token-prediction block. Q8_0 keeps a small model's logits crisp for short, single-pick outputs.

Files

file size what it is
code-daemon-enrich-v1-Q8_0.gguf 774 MB the model; needs llama.cpp with the qwen35 architecture (b10809 or newer)

License & attribution

Apache-2.0, matching the Qwen/Qwen3.5-0.8B base. Not legal advice — check the base and teacher model cards before redistributing. Base and teacher © the Qwen team; please also honour their cards.

Downloads last month
49
GGUF
Model size
0.8B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for faxenoff/code-daemon-enrich-v1

Quantized
(317)
this model