ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport Paper • 2609.34899 • Published 14 days ago • 7
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 25 days ago • 228
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • Sep 3 • 121
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published Jul 13 • 78
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation Paper • 2608.10636 • Published Aug 11 • 11
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels Paper • 2607.24651 • Published Jul 27 • 4
view article Article ViDoRe V3: a comprehensive evaluation of retrieval for enterprise use-cases QuentinJG • Nov 5, 2025 • 68
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Paper • 2605.20177 • Published May 19 • 8
Encoder-Free Human Motion Understanding via Structured Motion Descriptions Paper • 2604.21668 • Published Apr 23 • 3
Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment Paper • 2604.00913 • Published Apr 1 • 3
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval Paper • 2603.12824 • Published Mar 13 • 7
view article Article NanoVDR: A 70M Text-Only Model That Retrieves Visual Documents as Well as a 2B VLM Ryenhails • Mar 16 • 3