Instructions to use faxenoff/code-daemon-enrich-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use faxenoff/code-daemon-enrich-v1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0 # Run inference directly in the terminal: llama cli -hf faxenoff/code-daemon-enrich-v1:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0 # Run inference directly in the terminal: llama cli -hf faxenoff/code-daemon-enrich-v1:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf faxenoff/code-daemon-enrich-v1:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf faxenoff/code-daemon-enrich-v1:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf faxenoff/code-daemon-enrich-v1:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf faxenoff/code-daemon-enrich-v1:Q8_0
Use Docker
docker model run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0
- LM Studio
- Jan
- vLLM
How to use faxenoff/code-daemon-enrich-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "faxenoff/code-daemon-enrich-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "faxenoff/code-daemon-enrich-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0
- Ollama
How to use faxenoff/code-daemon-enrich-v1 with Ollama:
ollama run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0
- Unsloth Desktop
- Pi
How to use faxenoff/code-daemon-enrich-v1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "faxenoff/code-daemon-enrich-v1:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use faxenoff/code-daemon-enrich-v1 with Docker Model Runner:
docker model run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0
- Lemonade
How to use faxenoff/code-daemon-enrich-v1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull faxenoff/code-daemon-enrich-v1:Q8_0
Run and chat with the model
lemonade run user.code-daemon-enrich-v1-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use faxenoff/code-daemon-enrich-v1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default faxenoff/code-daemon-enrich-v1:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use faxenoff/code-daemon-enrich-v1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "faxenoff/code-daemon-enrich-v1:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
code-daemon-enrich-v1
A distilled Qwen3.5-0.8B that writes the short, structured labels of a code-intelligence pipeline: names for clusters of code entities, names for graph communities, and link picks. It is the high-volume worker of the Code-Daemon daemon — a purpose-built component, not a general assistant; outside these tasks its behaviour is undefined.
What makes it different: label stages are prefill-bound — long prompt, a few words of answer — and run once per cluster, community and link candidate of an index. A sub-billion model is enough for them if it learns the contract: name a group instead of listing it, fold a level instead of concatenating it, decline a link instead of always picking one. Distillation from a 27B teacher taught it exactly that; the stock base and the 0.6B it replaced do not do it.
The weights under this id changed on 2026-09-16. Until then it served a Qwen3-0.6B distilled from Qwen2.5-7B; that model is retired and its card is in this repository's git history.
The tasks
| task | answer | example |
|---|---|---|
| cluster label, levels 0 and 1 | <topic>: name1, name2, name3 |
DI service provider construction: BuildServiceProvider, GetService, ServiceProviderOptions |
| graph-community label | a 2-to-5-word noun phrase | Go standard library packages |
| link selection | a candidate id c<N>, or none |
c3 |
Numbers
Held-out rows scored as the GGUF the daemon loads, through llama-server:
| Q8_0 GGUF | label tasks, rougeL vs teacher | link selection, exact of 41 | community labels that are a bare list | L1 label, median words |
|---|---|---|---|---|
| this model | 0.493 | 18 | 0 % | 6 |
| stock Qwen3.5-0.8B | 0.193 | 14 | 1 % | 33 |
| the retired Qwen3-0.6B | 0.220 | 11 | 38 % | 33 |
| teacher (Qwen3.8-27B) | — | — | 0 % | 7 |
- +0.300 rougeL over its own base (95 % CI +0.257 … +0.343, 158 wins / 22 losses). The base writes a paragraph where a label is asked for.
- It can decline a link. The teacher answers
noneto 24 of 41 link questions, this model to 6, the retired 0.6B to none — every ambiguous case became an edge. - Where it is worse: it paraphrases what should be copied —
iter, durcan come back asiterations, elapsed. If your use needs exact identifier echo, measure that first.
Speed and memory — the trade
Laptop RTX 5060 8 GB, CUDA, Q8_0, n_ctx=8192, live inside the daemon, against the retired 0.6B
the same day:
| this model | retired 0.6B | |
|---|---|---|
| decode, single stream | 253 tok/s | 305 tok/s |
| decode, 8 concurrent slots | 794 tok/s | 1 083 tok/s |
| label stage end to end, prompt tokens counted | 5 858 tok/s | 6 265 tok/s |
Decode is 17–27 % slower; the whole stage is 6.5 % slower, because these stages are prefill-bound. Plan capacity from the last row.
Memory: 2 281 MB resident with 28 parallel sequences, not the 774 MB of the file. The model is hybrid — 18 of its 24 layers are gated DeltaNet — so each sequence carries 19.3 MiB of recurrent state whatever the context length. Fit a smaller card by cutting sequences, not context: half the sequences gives back ~270 MiB.
How to use it
ChatML, one user turn, no system turn, and an empty think block opening the answer — the model was trained that way and drifts without it. Greedy decoding (temperature 0).
llama-cli -m code-daemon-enrich-v1-Q8_0.gguf -c 8192 --temp 0 \
-p '<|im_start|>user
Write ONE line — 2 to 5 plain-English words — labelling this group. No quotes, no explanation.
Group members:
- parseArgs
- Command
- Usage
<|im_end|>
<|im_start|>assistant
<think>
</think>
'
n_ctx=8192 covers the label prompts; parallel sequences are the memory knob (above).
How it was made
- Base:
Qwen/Qwen3.5-0.8B— 24 layers, 6 of them attention and 18 gated DeltaNet. - Teacher:
Qwen3.8-27B, answering the prompts the daemon sends in production. - Method: sequence-level knowledge distillation, merged into the base and exported to GGUF without the base's multi-token-prediction block. Q8_0 keeps a small model's logits crisp for short, single-pick outputs.
Files
| file | size | what it is |
|---|---|---|
code-daemon-enrich-v1-Q8_0.gguf |
774 MB | the model; needs llama.cpp with the qwen35 architecture (b10809 or newer) |
License & attribution
Apache-2.0, matching the Qwen/Qwen3.5-0.8B base.
Not legal advice — check the base and teacher model cards before redistributing. Base and teacher
© the Qwen team; please also honour their cards.
- Downloads last month
- 49
8-bit
docker model run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0