Instructions to use lightonai/colgrep-agent-2B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lightonai/colgrep-agent-2B-GGUF with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="lightonai/colgrep-agent-2B-GGUF")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("lightonai/colgrep-agent-2B-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use lightonai/colgrep-agent-2B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/lightonai/colgrep-agent-2B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use lightonai/colgrep-agent-2B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lightonai/colgrep-agent-2B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lightonai/colgrep-agent-2B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/lightonai/colgrep-agent-2B-GGUF:Q4_K_M
- SGLang
How to use lightonai/colgrep-agent-2B-GGUF with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "lightonai/colgrep-agent-2B-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lightonai/colgrep-agent-2B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "lightonai/colgrep-agent-2B-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lightonai/colgrep-agent-2B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use lightonai/colgrep-agent-2B-GGUF with Ollama:
ollama run hf.co/lightonai/colgrep-agent-2B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use lightonai/colgrep-agent-2B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "lightonai/colgrep-agent-2B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use lightonai/colgrep-agent-2B-GGUF with Docker Model Runner:
docker model run hf.co/lightonai/colgrep-agent-2B-GGUF:Q4_K_M
- Lemonade
How to use lightonai/colgrep-agent-2B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull lightonai/colgrep-agent-2B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.colgrep-agent-2B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use lightonai/colgrep-agent-2B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default lightonai/colgrep-agent-2B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use lightonai/colgrep-agent-2B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf lightonai/colgrep-agent-2B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "lightonai/colgrep-agent-2B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
ColGrep-Agent-2B
ColGrep-Agent-2B is a lightweight code-localization agent designed to identify the files and line ranges relevant to a GitHub issue, bug report, feature request, or natural-language question about a repository.
Given a question such as "Where are retries configured?" or "Why does the export crash on empty files?", the agent iteratively searches the repository using ColGrep, reads candidate files through a read-only terminal, and returns the precise code locations relevant to the task.
The model powers colgrep --agent and is designed to offload repository exploration from larger coding agents. Rather than consuming thousands of tokens searching through files before making a fix, a larger model can delegate code localization to ColGrep-Agent-2B and receive a focused list of relevant locations.
Its small footprint enables practical local inference, with GPU acceleration when available, without requiring Python, a separate inference server, or an API key.
🚀 Quick start
1. Install ColGrep and the model
Homebrew (macOS / Linux)
brew install lightonai/tap/colgrep && colgrep --install-agent
Shell installer (macOS / Linux)
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/lightonai/next-plaid/releases/latest/download/colgrep-installer.sh | sh && "${CARGO_HOME:-$HOME/.cargo}/bin/colgrep" --install-agent
Windows (PowerShell)
irm https://github.com/lightonai/next-plaid/releases/latest/download/colgrep-installer.ps1 | iex
colgrep --install-agent
Cargo (from source)
cargo install colgrep && colgrep --install-agent
The colgrep --install-agent command downloads the default colgrep-agent-2B-Q8_0.gguf model (approximately 2.5 GB), verifies that it loads, and prepares the prompt cache to accelerate the first inference.
Installation is required only once. No Python environment, external inference server, or API key is necessary.
2. Index a repository
cd /path/to/project
colgrep init
After the initial indexing, subsequent searches automatically detect modified files and update the index before answering.
3. Ask a question
colgrep --agent "where does the agent decide which GPU backend llama.cpp uses"
The following example shows a real session on ColGrep's own source code, running on an M3 Pro with Metal acceleration. Individual tool results have been omitted for readability.
$ colgrep --agent "where does the agent decide which GPU backend llama.cpp uses"
◆ colgrep agent · colgrep-agent-2B-Q8_0.gguf (GPU)
❯ where does the agent decide which GPU backend llama.cpp uses
● colgrep llama.cpp GPU backend
● terminal cat -n agent/src/llm/llama.rs
● terminal cat -n agent/src/llm/server.rs
● terminal cat -n agent/src/engine.rs
● terminal cat -n agent/src/llm/llama.rs | sed -n '80,150p'
● terminal cat -n agent/src/llm/server.rs | sed -n '65,130p'
● colgrep find_gpu
● terminal cat -n agent/src/llm/server.rs | sed -n '1,100p'
● finish agent/src/llm/server.rs:103-105, agent/src/llm/server.rs:65-69, …
✓ 5 locations · 9 turns · 2 searches · 6 terminal commands · 21.7s
agent/src/llm/server.rs:103-105
agent/src/llm/server.rs:65-69
agent/src/llm/llama.rs:101-101
agent/src/llm/llama.rs:115-116
agent/src/engine.rs:57-67
The output follows the path/to/file:START-END format, making it directly usable by developers, scripts, and other coding agents.
Agent progress is printed to stderr only when running in a terminal, while stdout contains only the resulting code locations. This separation makes the command easy to compose with other tools.
Additional usage examples
Scope the search to a directory and display the matching code lines:
colgrep --agent "crash when the config file is empty" ./backend -c
Return structured JSON for downstream applications:
colgrep --agent --json "where are retries configured"
🤖 Integration with coding agents
ColGrep can be integrated into existing coding-agent workflows, allowing larger models to delegate repository exploration to a smaller specialized model.
Install the corresponding hooks and skills:
colgrep --install-claude-code # Claude Code
colgrep --install-codex # Codex
colgrep --install-opencode # OpenCode
For example, once configured, Claude Code can delegate open-ended questions such as "Where is this behavior implemented?" to colgrep --agent, receiving exact file paths and line ranges instead of performing every exploratory search itself.
This division of labor allows larger coding agents to reserve more of their context and computation for understanding issues, planning changes, and generating patches.
🔎 ColGrep
ColGrep is a semantic code search tool for terminals and coding agents, built on the NextPlaid search engine. It combines natural-language search with regex filtering to find code by both meaning and exact matches.
ColGrep-Agent-2B extends this search functionality with an iterative search-and-read loop. Rather than relying on a single search query, the agent follows leads, inspects candidate files, refines its searches, and identifies relevant code locations over multiple turns.
🧠 Training and harness
ColGREP-Agent-2B is built on MiniCPM5-2B and trained in two stages: a short supervised fine-tuning (SFT) cold start followed by reinforcement learning (RL) using DAPO.
We construct multilingual code-localization environments from nebius/SWE-rebench, internlm/SWE-Fixer-Train-110K, and additional scraped repositories covering multiple programming languages. Each training instance pairs a repository with an issue description and a dedicated bug-fix commit. We use the pre-patch repository state as the search environment and derive gold file- and line-level localization targets from the corresponding patch. These data sources are used across both the SFT and RL stages.
We first perform a short cold-start SFT phase using traces generated and filtered with Qwen3.5-27B as a teacher model. This stage teaches tool-call formatting, command ordering, argument and flag usage, and effective repository exploration strategies, with an emphasis on semantic search through ColGrep rather than traditional grep-based search. We then train for 500 steps using DAPO in a dedicated code-localization environment, rewarding accurate file- and line-level identification while penalizing unnecessary exploration steps.
The model operates in a repository exploration environment with access to three tools:
| Tool | Purpose |
|---|---|
colgrep |
Search code semantically, with optional regex and path filters. |
terminal |
Read files and inspect the repository through read-only commands. |
finish |
Submit relevant locations as path/to/file:START-END. |
The agent has a 10-turn budget to search and inspect the repository before submitting its final prediction as a set of relevant file paths and line ranges.
📊 Evaluation
We evaluate ColGREP-Agent-2B on SWE-bench Lite and Multi-SWE-bench Flash, adapting both benchmarks into code-localization tasks following the methodology described in our paper. Rather than generating patches or resolving issues end-to-end, the model is tasked with identifying the files and lines of code relevant to each issue. This allows us to evaluate localization accuracy independently of code generation.
| Model | Harness | SWE-bench Lite | Multi-SWE-bench Flash | ||
|---|---|---|---|---|---|
| File F1 | Line F1 | File F1 | Line F1 | ||
| lightonai/colgrep-agent-2B | ColGrep | 76.18 | 19.92 | 62.57 | 15.93 |
| openbmb/MiniCPM5-2B | ColGrep | 66.72 | 9.62 | 52.01 | 11.31 |
| google/gemma-4-31B-it | ColGrep | 79.93 | 17.08 | 68.30 | 19.14 |
| Qwen/Qwen3.5-27B | ColGrep | 77.93 | 15.71 | 68.28 | 16.63 |
| google/gemma-4-26B-A4B-it | ColGrep | 76.35 | 13.34 | 63.50 | 14.75 |
| Qwen/Qwen3.8-27B | ColGrep | 60.80 | 7.70 | 59.68 | 11.50 |
| deepseek-ai/DeepSeek-V4-Flash | ColGrep | 74.04 | 11.36 | 64.26 | 13.64 |
All models use the ColGREP harness. File and line F1 measure overlap with reference files and lines, respectively. Scores are percentages averaged over five seeds (300 instances per benchmark per seed), using strict predictions without rescue, a 10-turn limit with mandatory final submission, and temperature 0.6 without reasoning for comparison models.
📚 Citation
@misc{colgrepagent_2026,
title = {ColGrep-Agent},
author = {Maxence Lasbordes and Raphael Sourty and Aarush Sinha and Amélie Chatelain},
year = {2026},
howpublished = {\url{https://huggingface.co/blog/lightonai/colgrep-agent/}}
}
- Downloads last month
- 120
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for lightonai/colgrep-agent-2B-GGUF
Base model
openbmb/MiniCPM5-2B
