Native Action-Prior Learning from Videos for World Action Models Paper • 2610.03391 • Published 10 days ago • 88
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training Paper • 2609.36659 • Published 13 days ago • 87
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite Paper • 2610.02826 • Published 10 days ago • 103
FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation Paper • 2609.38839 • Published 12 days ago • 125
MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation Paper • 2609.38078 • Published 13 days ago • 114
Does Learning Protein Folding Generalize to Broader Reasoning? Paper • 2609.38879 • Published 12 days ago • 62
World Action Modeling with Progressive Visual Planning Paper • 2610.02508 • Published 11 days ago • 97
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training Paper • 2609.00188 • Published Aug 31 • 54
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models Paper • 2608.27550 • Published Aug 27 • 83
VLAct Collection Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models • 12 items • Updated Aug 31 • 3
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published Aug 17 • 9
Mage Collection A family of lightweight multimodal models, including understanding and generation. • 8 items • Updated Jul 26 • 30
StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Paper • 2608.05703 • Published Aug 6 • 17
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
SciForma: Structure-Faithful Generation of Scientific Diagrams Paper • 2607.18091 • Published Jul 20 • 25
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 78
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Paper • 2605.20342 • Published May 19 • 31
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation Paper • 2605.18739 • Published May 18 • 117
FlowAnchor: Stabilizing the Editing Signal for Inversion-Free Video Editing Paper • 2604.22586 • Published Apr 24 • 15