HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 5 days ago • 123
AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning Paper • 2608.14028 • Published 8 days ago
Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments Paper • 2606.15587 • Published Jul 15
PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training Paper • 2503.06486 • Published Mar 9, 2025
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering Paper • 2604.08209 • Published Apr 9 • 27
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model Paper • 2602.09638 • Published Feb 10
World Guidance: World Modeling in Condition Space for Action Generation Paper • 2602.22010 • Published Feb 25 • 16
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert Paper • 2510.03896 • Published Oct 4, 2025
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning Paper • 2508.20840 • Published Aug 28, 2025
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model Paper • 2508.06206 • Published Aug 8, 2025 • 1
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Paper • 2507.01016 • Published Jul 1, 2025 • 1
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 5 days ago • 123
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks Paper • 2607.08768 • Published Jul 9 • 34
Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots Paper • 2606.28133 • Published Jun 26 • 40
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality Paper • 2512.07951 • Published Dec 8, 2025 • 51
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation Paper • 2510.05057 • Published Oct 6, 2025 • 13
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation Paper • 2510.05057 • Published Oct 6, 2025 • 13
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation Paper • 2510.05057 • Published Oct 6, 2025 • 13 • 3
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling Paper • 2509.12201 • Published Sep 15, 2025 • 107