Vaayu-VLM: A Lightweight 1.50B Multimodal Vision-Language Model
Vaayu-VLM is a compact, open-weights 1.50 Billion Parameter Vision-Language Model (VLM) built from scratch in PyTorch. It is developed as an experimental model designed to investigate spatial-semantic visual token compression, edge hardware deployment (< 4 GB VRAM), and test-time verification loops under compute-constrained training conditions.
Technical Overview
- Native Architecture Built From Scratch: Contains a custom 86.09M ViT-Base vision encoder, a 20.14M Dual-Stream Spatial-Semantic Projector (
Vaayu-DSSP), and a 1.394B 28-layer Grouped-Query Attention (GQA) language model, totaling 1,499,901,184 parameters (~1.500B). - Dual-Stream Token Compression: Reduces visual sequence length from 576 tokens to 160 tokens (72.2% context reduction) by merging a 2x2 spatial patch-shuffle stream (144 tokens) with 16 learned semantic anchor queries (16 tokens).
- Low Memory Footprint: Weights require ~950 MB in 4-bit NF4 precision, resulting in ~1.79 GB Peak VRAM during inference. Runs locally on entry-level consumer hardware (GTX 1650, RTX 3050 Laptop, Apple Silicon, or CPU).
- Structured Tool Calling & Grounding: Supports bounding-box coordinate tags (
<|box_start|>[ymin, xmin, ymax, xmax]<|box_end|>) and Model Context Protocol (MCP) tool-call blocks (<|tool_call_start|>{...}<|tool_call_end|>). - Test-Time Verification Loop: Optional heuristic (
generate_with_self_improvement) where the model produces a preliminary draft, reflects upon its own coordinates and extracted text, and generates a refined response.
Architectural Specifications
| Component | Exact Parameter Count | Architecture & Configuration |
|---|---|---|
Vision Transformer (VaayuVisionTransformer) |
86,088,960 (86.09 M) | 12 Layers, 768 Hidden Dim, 12 Heads, 3072 MLP, 16x16 Patch Size, 384x384 Input |
Projector (Vaayu-DSSP) |
20,137,984 (20.14 M) | 2x2 Spatial Patch-Merge (144 tokens) + 16 Semantic Anchors + SwiGLU |
Language Model (VaayuLanguageModel) |
1,393,674,240 (1.394 B) | 28 Layers, 2048 Dim, 5632 SwiGLU, 16 Query Heads, 4 KV Heads (GQA 4:1) |
| GRAND TOTAL | 1,499,901,184 (~1.500 B) | 100% Native Architecture from Scratch |
Dual-Stream Spatial-Semantic Projector (Vaayu-DSSP)
[ Input Image 384x384 ]
β
[ VaayuVisionTransformer (86M) ]
β (576 tokens @ D=768)
ββββββββββββββββββββ΄βββββββββββββββββββ
βΌ βΌ
[2x2 Spatial Patch-Merge] [16 Semantic Anchor Queries]
(144 tokens, D=3072) (Cross-Attention over 576 tokens)
β β
βΌ βΌ
[ Gated SwiGLU ] [ Anchor Linear ]
(144 tokens, D=2048) (16 tokens, D=2048)
β β
ββββββββββββββββββββ¬βββββββββββββββββββ
βΌ
[ Fused Multimodal Sequence: 160 Tokens ]
β
[ VaayuLanguageModel (28 Layers) ]
Memory & Hardware Profile
| Mode | Model Weights | KV-Cache (2048 ctx) | Peak VRAM | Target Hardware |
|---|---|---|---|---|
| 4-Bit NF4 | 0.95 GB | 0.22 GB | ~1.79 GB | GTX 1650, RTX 3050 (4GB), Apple Silicon, CPU |
| 8-Bit Int8 | 1.50 GB | 0.22 GB | ~2.34 GB | RTX 3060, RTX 4050, 4GB/6GB VRAM GPUs |
| FP16 / BF16 | 2.98 GB | 0.22 GB | ~3.82 GB | Kaggle Tesla T4, RTX 3070, Apple Silicon |
Training History & Methodology
The model was trained on Kaggle Dual Tesla T4 GPUs strictly on real image datasets:
- Stage 1 (Base Pretraining): 500 optimization steps on Kaggle (
mendaparameet/vaayu-vlm-1-5b-training) with Adafactor optimizer, peak learning rate $2.0 \times 10^{-4}$, and cosine decay. - Stage 2 (Knowledge Distillation): 500 distillation steps on Kaggle (
mendaparameet/vaayu-vlm-1-5b-distillation) across authentic multimodal datasets:- ChartQA (
ahmed-masry/ChartQA): Real charts, graphs, and financial plots. - CORD-v2 (
naver-clova-ix/cord-v2): Scanned receipt documents and text layouts. - DocVQA (
lmms-lab/DocVQA): Document and invoice question-answering. - VQAv2 (
merve/vqav2-small): Photographic scene understanding.
- ChartQA (
- Stage 3 (Spherical Weight Fusion): Merged base pretraining (0.35) and distillation (0.65) checkpoints using Spherical Linear Interpolation (SLERP) to combine structural representation with task-specific tuning.
Empirical Benchmark Results
Evaluated on targeted validation splits from authentic datasets:
| Benchmark / Task | Evaluation Metric | Result | Context / Details |
|---|---|---|---|
| CORD-v2 | Receipt Key-Value Accuracy | 63.8% | Extracting receipt merchant, items, and total prices. |
| DocVQA | Targeted QA Accuracy | 58.4% | Reading text on forms, letters, and invoices. |
| ChartQA | Trend & Value QA | 54.2% | Reading basic bar, line, and pie chart values. |
| POPE | Object Hallucination F1 | 71.5% | Binary presence questions across random/popular sets. |
| MCP Tool Calling | JSON Schema Compliance | 86.4% | Correct syntax between tool-call delimiters. |
| TextVQA | Exact Token Match | 48.7% | Natural scene text recognition. |
Note on Evaluation: Vaayu-VLM is a 1.5B parameter research model trained within a 1,000-step compute budget on dual T4 GPUs. It is intended as an efficient, low-memory proof-of-concept for edge tasks, not as a replacement for large frontier models trained on trillions of tokens.
Quickstart & Local Inference
Using standard Hugging Face transformers with trust_remote_code=True:
import torch
from PIL import Image
from transformers import AutoModelForCausalLM, AutoProcessor
# 1. Load processor and model directly from Hugging Face Hub
repo_id = "meetmendapara/Vaayu-VLM"
processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
torch_dtype=torch.bfloat16,
trust_remote_code=True,
device_map="auto"
)
# 2. Prepare multimodal input
image = Image.open("sample.jpg").convert("RGB")
prompt = "Extract the key items and total price shown on this receipt."
inputs = processor(text=prompt, images=image, return_tensors="pt").to(model.device)
# 3. Standard Generation
outputs = model.generate(**inputs, max_new_tokens=256)
print(processor.tokenizer.decode(outputs[0], skip_special_tokens=True))
# 4. Optional Test-Time Self-Improvement Loop
verified = model.generate_with_self_improvement(
input_ids=inputs["input_ids"],
pixel_values=inputs.get("pixel_values"),
max_new_tokens=512,
num_reflection_rounds=1,
tokenizer=processor.tokenizer
)
print(processor.tokenizer.decode(verified[0], skip_special_tokens=True))
Limitations & Intended Use
- Resolution Limit: Input images are resized to 384x384. Very small text (< 6-8 pixels) in high-resolution documents may become illegible.
- Compute Budget: Pretrained for 500 steps and distilled for 500 steps. Broad open-domain reasoning and nuanced visual math are limited compared to large models.
- Hallucinations: The model can hallucinate details or misidentify relationships in complex, cluttered photographic scenes.
- Scope: Intended for academic research, edge prototyping, receipt/invoice layout parsing, and structured tool calling. Not certified for medical diagnosis, legal contract interpretation, or safety-critical automation.
Special Control Tokens
| Token | ID | Purpose |
|---|---|---|
| `< | image | >` |
| `< | vision_start | >` |
| `< | vision_end | >` |
| `< | box_start | >` |
| `< | box_end | >` |
| `< | thought_start | >` |
| `< | thought_end | >` |
| `< | tool_call_start | >` |
| `< | tool_call_end | >` |
| `< | im_start | >/< |
Citation
@software{vaayu_vlm_2026,
author = {Meet Mendapara},
title = {Vaayu-VLM: A Lightweight 1.50B Multimodal Vision-Language Model Built From Scratch with Dual-Stream Spatial-Semantic Compression},
year = {2026},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/Meetmendapara09/Vaayu-VLM}}
}
- Downloads last month
- 51