MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF

MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF

GGUF quantizations of MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic for llama.cpp, Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes.

This repository provides local-deployment builds of a 2B Thinking model fine-tuned on Claude data atop openbmb/MiniCPM5-2B, with a strong focus on agentic tool calling, coding, and instruction following. MiniCPM5's native chat template is embedded in the GGUF files.

Transformers checkpoint: MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic


Files

File Quant Size Notes
MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-F16.gguf F16 ~4.7 GB full-precision conversion base
MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-Q8_0.gguf Q8_0 ~2.5 GB recommended default
MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-Q4_K_M.gguf Q4_K_M ~1.5 GB smallest, for memory-constrained devices

Q8_0 is the recommended default quant for this 2B model. Q4_K_M is available for ultra-lightweight deployment.


Quick start

llama.cpp (llama-cli)

llama-cli \
  -m MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-Q8_0.gguf \
  -p "Write a Python function to merge two sorted lists." \
  -n 512 \
  --temp 1.0 --top-p 0.95 --min-p 0.0 \
  -c 8192

The model supports up to 128K tokens (131,072) per config.json. Set -c according to your available VRAM/RAM.

llama.cpp server

llama-server \
  -m MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-Q8_0.gguf \
  -c 8192 --port 8080

LM Studio / jan / KoboldCpp

Load any .gguf file from this repository. The MiniCPM5 chat template is embedded in the GGUF metadata.


Sampling recommendations

Inherited from openbmb/MiniCPM5-2B:

Scenario Params
Default temperature=1.0, top_p=0.95, min_p=0.0
If repetitive outputs temperature=1.0, top_p=0.95, min_p=0.0, repetition_penalty=1.05

This model is Thinking-only โ€” chain-of-thought reasoning is always active.

Support for sampling parameters varies across inference frameworks โ€” check your runtime's documentation.


Capabilities

  • Agentic tool calling โ€” reliable function-calling / tool-use behavior designed for multi-step agentic workflows
  • Claude fine-tune โ€” post-trained on Claude data
  • Coding โ€” code generation, debugging, and software-engineering workflows
  • Instruction following โ€” reliable adherence to user prompts and task constraints
  • Thinking mode โ€” chain-of-thought reasoning; MiniCPM5 chat template baked into the GGUF
  • Long context โ€” up to 128K tokens (131,072 tokens per upstream config.json)

Benchmark

Scores for the Transformers checkpoint MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic:

ClawBench (Agentic Coding)

Model QwenClawBench WildClawBench
MiniCPM5-2B 42.11 23.19
MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic 44.56 (+2.45) 24.32 (+1.13)

ClawBench evaluates agentic coding ability โ€” the model's capacity to autonomously use tools, navigate codebases, and complete multi-step software engineering tasks.

More benchmarks (BFCL, SWE-bench, Tau-Bench, etc.) coming soon.


Limitations

  • Thinking outputs โ€” the model may emit reasoning blocks before the final answer
  • 2B scale โ€” lightweight local deployment; not frontier-scale
  • Runtime context โ€” actual usable context depends on your GGUF runtime and hardware limits

Provenance & licensing

Apache-2.0, inherited from MiniCPM5-2B.

Acknowledgements

Downloads last month
6,201
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF

Space using GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF 1