Instructions to use HaloWang/rwkv-weights with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RWKV
How to use HaloWang/rwkv-weights with RWKV:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use HaloWang/rwkv-weights with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf HaloWang/rwkv-weights:Q4_K_M # Run inference directly in the terminal: llama cli -hf HaloWang/rwkv-weights:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf HaloWang/rwkv-weights:Q4_K_M # Run inference directly in the terminal: llama cli -hf HaloWang/rwkv-weights:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf HaloWang/rwkv-weights:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf HaloWang/rwkv-weights:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf HaloWang/rwkv-weights:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf HaloWang/rwkv-weights:Q4_K_M
Use Docker
docker model run hf.co/HaloWang/rwkv-weights:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use HaloWang/rwkv-weights with Ollama:
ollama run hf.co/HaloWang/rwkv-weights:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use HaloWang/rwkv-weights with Docker Model Runner:
docker model run hf.co/HaloWang/rwkv-weights:Q4_K_M
- Lemonade
How to use HaloWang/rwkv-weights with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull HaloWang/rwkv-weights:Q4_K_M
Run and chat with the model
lemonade run user.rwkv-weights-Q4_K_M
List all available models
lemonade list
- Atomic Chat
RWKV Weights
Current release: G1k for RWKV Chat 4.8.5
The current release contains 40 published artifacts from the 2026-09-30 G1k checkpoints with context length 25,600. Both repositories contain the same file sizes and SHA-256 identities.
MLX
| Model | Quantization | Supported platforms | Download |
|---|---|---|---|
| 1.5B | W6 / group 64 | macOS, iOS | ModelScope 路 Hugging Face |
| 2.9B | W6 / group 64 | macOS, iOS | ModelScope 路 Hugging Face |
| 7.2B | W6 / group 64 | macOS, iOS | ModelScope 路 Hugging Face |
| 13.3B | W6 / group 64 | macOS | ModelScope 路 Hugging Face |
These MLX packages include the configuration used by the Swift runtime. Select a package matching your platform and available memory.
Complete artifact catalog
The G1k release manifest (ModelScope 路 Hugging Face) lists all 40 files, quantization recipes, platform and chip limits, sizes, hashes, and frozen mirror revisions. It covers MLX, CoreML, llama.cpp, WebRWKV, Palm, QNN and MediaTek artifacts according to the existing supported combinations.
Use index.json for discovery (ModelScope 路 Hugging Face). Its latestManifest points to the current G1k release. The App consumer catalog is recorded in the release manifest.
Historical releases
Older immutable artifacts and manifests are retained for reproducibility. The previous G1i download notes are archived (ModelScope 路 Hugging Face). Superseded G1i/G1j MLX entries are removed from the current 4.8.5 App catalog.
Repository layout
artifacts/<family>/<task>/<size>/<backend>/<artifact>
manifests/<family>/<task>/<immutable-manifest>.json
index.json
- Downloads last month
- 806
4-bit
6-bit
8-bit