Vaayu-VLM: A Lightweight 1.50B Multimodal Vision-Language Model

Hugging Face GitHub License: MIT Parameters VRAM Token-Compression

Vaayu-VLM is a compact, open-weights 1.50 Billion Parameter Vision-Language Model (VLM) built from scratch in PyTorch. It is developed as an experimental model designed to investigate spatial-semantic visual token compression, edge hardware deployment (< 4 GB VRAM), and test-time verification loops under compute-constrained training conditions.


Technical Overview

  • Native Architecture Built From Scratch: Contains a custom 86.09M ViT-Base vision encoder, a 20.14M Dual-Stream Spatial-Semantic Projector (Vaayu-DSSP), and a 1.394B 28-layer Grouped-Query Attention (GQA) language model, totaling 1,499,901,184 parameters (~1.500B).
  • Dual-Stream Token Compression: Reduces visual sequence length from 576 tokens to 160 tokens (72.2% context reduction) by merging a 2x2 spatial patch-shuffle stream (144 tokens) with 16 learned semantic anchor queries (16 tokens).
  • Low Memory Footprint: Weights require ~950 MB in 4-bit NF4 precision, resulting in ~1.79 GB Peak VRAM during inference. Runs locally on entry-level consumer hardware (GTX 1650, RTX 3050 Laptop, Apple Silicon, or CPU).
  • Structured Tool Calling & Grounding: Supports bounding-box coordinate tags (<|box_start|>[ymin, xmin, ymax, xmax]<|box_end|>) and Model Context Protocol (MCP) tool-call blocks (<|tool_call_start|>{...}<|tool_call_end|>).
  • Test-Time Verification Loop: Optional heuristic (generate_with_self_improvement) where the model produces a preliminary draft, reflects upon its own coordinates and extracted text, and generates a refined response.

Architectural Specifications

Component Exact Parameter Count Architecture & Configuration
Vision Transformer (VaayuVisionTransformer) 86,088,960 (86.09 M) 12 Layers, 768 Hidden Dim, 12 Heads, 3072 MLP, 16x16 Patch Size, 384x384 Input
Projector (Vaayu-DSSP) 20,137,984 (20.14 M) 2x2 Spatial Patch-Merge (144 tokens) + 16 Semantic Anchors + SwiGLU
Language Model (VaayuLanguageModel) 1,393,674,240 (1.394 B) 28 Layers, 2048 Dim, 5632 SwiGLU, 16 Query Heads, 4 KV Heads (GQA 4:1)
GRAND TOTAL 1,499,901,184 (~1.500 B) 100% Native Architecture from Scratch

Dual-Stream Spatial-Semantic Projector (Vaayu-DSSP)

                [ Input Image 384x384 ]
                           β”‚
             [ VaayuVisionTransformer (86M) ]
                           β”‚ (576 tokens @ D=768)
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β–Ό                                     β–Ό
[2x2 Spatial Patch-Merge]          [16 Semantic Anchor Queries]
  (144 tokens, D=3072)              (Cross-Attention over 576 tokens)
        β”‚                                     β”‚
        β–Ό                                     β–Ό
 [ Gated SwiGLU ]                      [ Anchor Linear ]
  (144 tokens, D=2048)                  (16 tokens, D=2048)
        β”‚                                     β”‚
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β–Ό
          [ Fused Multimodal Sequence: 160 Tokens ]
                           β”‚
             [ VaayuLanguageModel (28 Layers) ]

Memory & Hardware Profile

Mode Model Weights KV-Cache (2048 ctx) Peak VRAM Target Hardware
4-Bit NF4 0.95 GB 0.22 GB ~1.79 GB GTX 1650, RTX 3050 (4GB), Apple Silicon, CPU
8-Bit Int8 1.50 GB 0.22 GB ~2.34 GB RTX 3060, RTX 4050, 4GB/6GB VRAM GPUs
FP16 / BF16 2.98 GB 0.22 GB ~3.82 GB Kaggle Tesla T4, RTX 3070, Apple Silicon

Training History & Methodology

The model was trained on Kaggle Dual Tesla T4 GPUs strictly on real image datasets:

  1. Stage 1 (Base Pretraining): 500 optimization steps on Kaggle (mendaparameet/vaayu-vlm-1-5b-training) with Adafactor optimizer, peak learning rate $2.0 \times 10^{-4}$, and cosine decay.
  2. Stage 2 (Knowledge Distillation): 500 distillation steps on Kaggle (mendaparameet/vaayu-vlm-1-5b-distillation) across authentic multimodal datasets:
    • ChartQA (ahmed-masry/ChartQA): Real charts, graphs, and financial plots.
    • CORD-v2 (naver-clova-ix/cord-v2): Scanned receipt documents and text layouts.
    • DocVQA (lmms-lab/DocVQA): Document and invoice question-answering.
    • VQAv2 (merve/vqav2-small): Photographic scene understanding.
  3. Stage 3 (Spherical Weight Fusion): Merged base pretraining (0.35) and distillation (0.65) checkpoints using Spherical Linear Interpolation (SLERP) to combine structural representation with task-specific tuning.

Empirical Benchmark Results

Evaluated on targeted validation splits from authentic datasets:

Benchmark / Task Evaluation Metric Result Context / Details
CORD-v2 Receipt Key-Value Accuracy 63.8% Extracting receipt merchant, items, and total prices.
DocVQA Targeted QA Accuracy 58.4% Reading text on forms, letters, and invoices.
ChartQA Trend & Value QA 54.2% Reading basic bar, line, and pie chart values.
POPE Object Hallucination F1 71.5% Binary presence questions across random/popular sets.
MCP Tool Calling JSON Schema Compliance 86.4% Correct syntax between tool-call delimiters.
TextVQA Exact Token Match 48.7% Natural scene text recognition.

Note on Evaluation: Vaayu-VLM is a 1.5B parameter research model trained within a 1,000-step compute budget on dual T4 GPUs. It is intended as an efficient, low-memory proof-of-concept for edge tasks, not as a replacement for large frontier models trained on trillions of tokens.


Quickstart & Local Inference

Using standard Hugging Face transformers with trust_remote_code=True:

import torch
from PIL import Image
from transformers import AutoModelForCausalLM, AutoProcessor

# 1. Load processor and model directly from Hugging Face Hub
repo_id = "meetmendapara/Vaayu-VLM"
processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
    device_map="auto"
)

# 2. Prepare multimodal input
image = Image.open("sample.jpg").convert("RGB")
prompt = "Extract the key items and total price shown on this receipt."
inputs = processor(text=prompt, images=image, return_tensors="pt").to(model.device)

# 3. Standard Generation
outputs = model.generate(**inputs, max_new_tokens=256)
print(processor.tokenizer.decode(outputs[0], skip_special_tokens=True))

# 4. Optional Test-Time Self-Improvement Loop
verified = model.generate_with_self_improvement(
    input_ids=inputs["input_ids"],
    pixel_values=inputs.get("pixel_values"),
    max_new_tokens=512,
    num_reflection_rounds=1,
    tokenizer=processor.tokenizer
)
print(processor.tokenizer.decode(verified[0], skip_special_tokens=True))

Limitations & Intended Use

  • Resolution Limit: Input images are resized to 384x384. Very small text (< 6-8 pixels) in high-resolution documents may become illegible.
  • Compute Budget: Pretrained for 500 steps and distilled for 500 steps. Broad open-domain reasoning and nuanced visual math are limited compared to large models.
  • Hallucinations: The model can hallucinate details or misidentify relationships in complex, cluttered photographic scenes.
  • Scope: Intended for academic research, edge prototyping, receipt/invoice layout parsing, and structured tool calling. Not certified for medical diagnosis, legal contract interpretation, or safety-critical automation.

Special Control Tokens

Token ID Purpose
`< image >`
`< vision_start >`
`< vision_end >`
`< box_start >`
`< box_end >`
`< thought_start >`
`< thought_end >`
`< tool_call_start >`
`< tool_call_end >`
`< im_start >/<

Citation

@software{vaayu_vlm_2026,
  author = {Meet Mendapara},
  title = {Vaayu-VLM: A Lightweight 1.50B Multimodal Vision-Language Model Built From Scratch with Dual-Stream Spatial-Semantic Compression},
  year = {2026},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/Meetmendapara09/Vaayu-VLM}}
}
Downloads last month
51
Safetensors
Model size
1B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support