RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 11 days ago • 282
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability Paper • 2610.08448 • Published 6 days ago • 207
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL Paper • 2609.32577 • Published 16 days ago • 106
MaLiang-Harness: A Programmable Path to Image and Video Generation Paper • 2609.34309 • Published 14 days ago • 394
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 15 days ago • 675
Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 18 days ago • 275
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published Sep 4 • 119
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 24 days ago • 138
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 28 days ago • 252
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published Aug 26 • 333
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published Aug 26 • 154
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published Aug 24 • 212
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published Aug 15 • 452
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 181
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published Aug 10 • 800