LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow Paper • 2607.10623 • Published 27 days ago • 14
BadWAM: When World-Action Models Dream Right but Act Wrong Paper • 2607.15207 • Published 23 days ago • 53
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 126
One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation Paper • 2605.29429 • Published May 28 • 8
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents Paper • 2605.30723 • Published May 29 • 17
On-Policy Adversarial Flow Distillation for Autoregressive Video Generation Paper • 2605.26105 • Published May 25 • 20
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards Paper • 2605.21467 • Published May 20 • 207
Where Does Authorship Signal Emerge in Encoder-Based Language Models? Paper • 2605.19908 • Published May 19 • 5
iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance Paper • 2605.21431 • Published May 20 • 2
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining Paper • 2605.14747 • Published May 14 • 147
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization Paper • 2605.13641 • Published May 13 • 51
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers Paper • 2605.06169 • Published May 7 • 238
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents Paper • 2605.05185 • Published May 6 • 106
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction Paper • 2604.22880 • Published Apr 24 • 10
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces Paper • 2604.05172 • Published Apr 6 • 25
Adam's Law: Textual Frequency Law on Large Language Models Paper • 2604.02176 • Published Apr 2 • 510
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models Paper • 2604.08546 • Published Apr 9 • 116
Efficient and Principled Scientific Discovery through Bayesian Optimization: A Tutorial Paper • 2604.01328 • Published Apr 1 • 9