MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 10 days ago • 33
Latent-MOPD: Latent Multi-Teacher On-Policy Distillation Paper • 2610.02381 • Published 11 days ago • 70
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 11 days ago • 283
Native Action-Prior Learning from Videos for World Action Models Paper • 2610.03391 • Published 10 days ago • 88
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 16 days ago • 326
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR Paper • 2609.37868 • Published 13 days ago • 65
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation Paper • 2609.36601 • Published 13 days ago • 95