Post
47
Ternary Transformers & Micro-Agent Architecture
CMSManhattan : Center Business Solutions Inc.
JiRack — Ternary Transformers & Micro-Agent Architecture
We build highly efficient large language models using 1.58-bit ternary weights {-1, 0, 1} for extreme compression and fast CPU/GPU inference.
Core focus:
JiRack Ternary Transformer Architecture — fresh Qwen base, trained on DeepSeek-style datasets, optimized for fast CPU inference (MIT License)
JiRack Micro-Agent Deployment — specialized small models + smart router for low-cost agentic systems
Production-ready ONNX Runtime & Docker inference stacks
Public Models
ModelSizeStatusJiRackUltra series (1B / 7B / 14B / 32B)—Released
CMSManhattan/JiRackUltra_1b
CMSManhattan/JiRackUltra_7b
CMSManhattan/JiRackUltra_14b
CMSManhattan/JiRackUltra_32b
JiRackTernary series1B → 10B+ReleasedJiRackPrecisionTokenizer—Released
Mission
Democratize frontier-scale language models through extreme efficiency. Train and run powerful models on accessible hardware without sacrificing quality.
Solved issues
Benefits of JiRack Micro-Agent Architecture:
Solves catastrophic forgetting during training by using small, specialized models for each domain, managed by a smart router Enables extremely cheap inference using ternary models Significantly reduces cloud inference costs while maintaining high performance In classical architecture, an expensive model has to search for MCP-agents every time, while JiRack uses a very small model and cheap router for agent tasks, saving big money right from the start Considered one of the best approaches for enterprise AI deployments
Hugging Face:
CMSManhattan
Ollama : https://ollama.com/cmsmanhattan
Docker Hub: cmsmanhattan Contact: grabko@cmsmanhattan.com
CMSManhattan : Center Business Solutions Inc.
JiRack — Ternary Transformers & Micro-Agent Architecture
We build highly efficient large language models using 1.58-bit ternary weights {-1, 0, 1} for extreme compression and fast CPU/GPU inference.
Core focus:
JiRack Ternary Transformer Architecture — fresh Qwen base, trained on DeepSeek-style datasets, optimized for fast CPU inference (MIT License)
JiRack Micro-Agent Deployment — specialized small models + smart router for low-cost agentic systems
Production-ready ONNX Runtime & Docker inference stacks
Public Models
ModelSizeStatusJiRackUltra series (1B / 7B / 14B / 32B)—Released
CMSManhattan/JiRackUltra_1b
CMSManhattan/JiRackUltra_7b
CMSManhattan/JiRackUltra_14b
CMSManhattan/JiRackUltra_32b
JiRackTernary series1B → 10B+ReleasedJiRackPrecisionTokenizer—Released
Mission
Democratize frontier-scale language models through extreme efficiency. Train and run powerful models on accessible hardware without sacrificing quality.
Solved issues
Benefits of JiRack Micro-Agent Architecture:
Solves catastrophic forgetting during training by using small, specialized models for each domain, managed by a smart router Enables extremely cheap inference using ternary models Significantly reduces cloud inference costs while maintaining high performance In classical architecture, an expensive model has to search for MCP-agents every time, while JiRack uses a very small model and cheap router for agent tasks, saving big money right from the start Considered one of the best approaches for enterprise AI deployments
Hugging Face:
Ollama : https://ollama.com/cmsmanhattan
Docker Hub: cmsmanhattan Contact: grabko@cmsmanhattan.com