-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 83 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20
Collections
Discover the best community collections!
Collections including paper arxiv:2607.21653
-
Why Fine-Tuning Encourages Hallucinations and How to Fix It
Paper • 2604.15574 • Published • 25 -
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper • 2604.24763 • Published • 70 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 90 -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
Paper • 2604.26752 • Published • 114
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 102 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
Paper • 2605.27366 • Published • 28 -
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Paper • 2607.21653 • Published • 33 -
EvoOntology: A Self-Evolving Ontology Layer for Data Agents
Paper • 2609.15779 • Published • 158
-
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Paper • 2603.19220 • Published • 70 -
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Paper • 2605.20164 • Published • 5 -
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
Paper • 2605.19577 • Published • 58 -
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
Paper • 2605.18703 • Published • 49
-
AI for Auto-Research: Roadmap & User Guide
Paper • 2605.18661 • Published • 70 -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Paper • 2605.18287 • Published • 14 -
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
Paper • 2605.16865 • Published • 9 -
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Paper • 2603.28069 • Published • 8
-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 83 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20
-
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
Paper • 2605.27366 • Published • 28 -
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Paper • 2607.21653 • Published • 33 -
EvoOntology: A Self-Evolving Ontology Layer for Data Agents
Paper • 2609.15779 • Published • 158
-
Why Fine-Tuning Encourages Hallucinations and How to Fix It
Paper • 2604.15574 • Published • 25 -
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper • 2604.24763 • Published • 70 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 90 -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
Paper • 2604.26752 • Published • 114
-
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Paper • 2603.19220 • Published • 70 -
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Paper • 2605.20164 • Published • 5 -
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
Paper • 2605.19577 • Published • 58 -
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
Paper • 2605.18703 • Published • 49
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 102 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
AI for Auto-Research: Roadmap & User Guide
Paper • 2605.18661 • Published • 70 -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Paper • 2605.18287 • Published • 14 -
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
Paper • 2605.16865 • Published • 9 -
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Paper • 2603.28069 • Published • 8