RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 10 days ago • 279
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems Paper • 2609.39050 • Published 11 days ago • 18
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 9 days ago • 33
EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling Paper • 2610.02298 • Published 10 days ago • 57
EVO-WAM: Evolving World Action Models through Video-Action Verification Paper • 2609.38057 • Published 12 days ago • 49
Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations Paper • 2609.32522 • Published 15 days ago • 71
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 15 days ago • 326
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation Paper • 2609.36601 • Published 12 days ago • 95
MaLiang-Harness: A Programmable Path to Image and Video Generation Paper • 2609.34309 • Published 13 days ago • 394
MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x Viewer • Updated about 7 hours ago • 15k • 5.17k • 297