Distilling the 13B SpaceLLaVA VLM-as-a-Judge into a Florence-2 model to efficiently quality filter spatialVQA datasets like OpenSpaces
Salma Mayorquin PRO
salma-remyx
AI & ML interests
None yet
Recent Activity
new activity about 11 hours ago
kernels-community/README:Fused Triton kernels for DoRA — 1.24× over the PyTorch reference on A100 posted an update about 17 hours ago
I unblocked 15 improvements for VQASynth in a day!
VQASynth, the pipeline behind SpaceThinker-Qwen2.5VL-3B, SpaceLLaVA, and the SpaceThinker dataset, had a backlog spanning new annotation methods, curation pipelines, model integrations, and agent tools.
I scoped each item in a short design brief. In GitHub Actions, Outrider mapped them to the repo’s existing modules and interfaces, implemented the changes, added tests, ran the repo checks, and opened draft PRs.
The 15 PRs covered:
* object orientation and 3D bounding boxes
* SAM2 regional captioning and multi-view matching
* LLaMA-Mesh tokenization
* Qwen2.5-VL fine-tuning
* spatial-reasoning data generation
* CLIP, SigLIP, and LLM2CLIP backends
and more
The batch added 12,275 lines, nearly doubling the codebase.
SpatialAnnotator went from two tools to seven, creating combinatorially more possible annotation pipelines for each image.
Outrider handled the implementation, tests, and integration. At about $1 per branch, I spent my time reviewing 15 concrete changes, fixing what needed fixing, and deciding what to ship.
Outrider is open source!
Point it at a paper, issue, or design brief and review the PR: https://github.com/remyxai/outrider
Check out the updates on VQASynth: https://github.com/remyxai/VQASynth