07. IA agentica
updated
Paper
• 2606.01533
• Published • 8
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper
• 2606.06741
• Published • 30
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper
• 2606.07412
• Published • 13
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper
• 2606.08348
• Published • 17
SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating
Paper
• 2606.07074
• Published • 13
Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory
Paper
• 2606.06523
• Published • 8
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Paper
• 2606.10917
• Published • 77
Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts
Paper
• 2606.05922
• Published • 72
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
Paper
• 2606.11025
• Published • 42
Rethinking the Divergence Regularization in LLM RL
Paper
• 2606.09821
• Published • 34
Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval
Paper
• 2606.04391
• Published • 11
What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems
Paper
• 2606.05304
• Published • 5
Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking
Paper
• 2606.07689
• Published • 7
Paper
• 2606.10650
• Published • 10
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application
Paper
• 2606.12191
• Published • 73
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Paper
• 2606.11926
• Published • 131
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
Paper
• 2606.13120
• Published • 5
WebChallenger: A Reliable and Efficient Generalist Web Agent
Paper
• 2606.10423
• Published • 3
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Paper
• 2606.12370
• Published • 22
Large Language Models Are Overconfident in Their Own Responses
Paper
• 2606.03437
• Published • 3
HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness
Paper
• 2606.12882
• Published • 15
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment
Paper
• 2606.10747
• Published • 13
LLM Agents Can See Code Repositories
Paper
• 2606.14061
• Published • 21
Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
Paper
• 2606.06574
• Published • 26
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
Paper
• 2606.14249
• Published • 52
APPO: Agentic Procedural Policy Optimization
Paper
• 2606.12384
• Published • 79
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
Paper
• 2606.15300
• Published • 14
FastContext: Training Efficient Repository Explorer for Coding Agents
Paper
• 2606.14066
• Published • 96
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
Paper
• 2606.22388
• Published • 97
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
Paper
• 2606.21337
• Published • 76
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
Paper
• 2606.22883
• Published • 38
Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation
Paper
• 2606.18844
• Published • 20
Self-Compacting Language Model Agents
Paper
• 2606.23525
• Published • 19
Causal Discovery in the Era of Agents
Paper
• 2606.23608
• Published • 8
FastMix: Fast Data Mixture Optimization via Gradient Descent
Paper
• 2606.14971
• Published • 4
DanceOPD: On-Policy Generative Field Distillation
Paper
• 2606.27377
• Published • 84
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
Paper
• 2606.26790
• Published • 59
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
Paper
• 2606.26300
• Published • 53
PhysiFormer: Learning to Simulate Mechanics in World Space
Paper
• 2606.27364
• Published • 12
When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
Paper
• 2606.27288
• Published • 5
Are We Ready For An Agent-Native Memory System?
Paper
• 2606.24775
• Published • 138
Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents
Paper
• 2606.22953
• Published • 1
Forecasting Future Behavior as a Learning Task
Paper
• 2606.11445
• Published • 2
Qwen-AgentWorld: Language World Models for General Agents
Paper
• 2606.24597
• Published • 163
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
Paper
• 2606.24428
• Published • 53
OpenThoughts-Agent: Data Recipes for Agentic Models
Paper
• 2606.24855
• Published • 49
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Paper
• 2606.27378
• Published • 61
Towards Automating Scientific Review with Google's Paper Assistant Tool
Paper
• 2606.28277
• Published • 12
Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving
Paper
• 2606.27457
• Published • 5
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Paper
• 2607.01071
• Published • 31
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
Paper
• 2607.00461
• Published • 30
CausalMix: Data Mixture as Causal Inference for Language Model Training
Paper
• 2607.01104
• Published • 22
The State-Prediction Separation Hypothesis
Paper
• 2607.01218
• Published • 13
AutoTrainess: Teaching Language Models to Improve Language Models Autonomously
Paper
• 2606.31551
• Published • 25
Valdi: Value Diffusion World Models
Paper
• 2607.00917
• Published • 16
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Paper
• 2607.01211
• Published • 14
Autonomous Scientific Discovery via Iterative Meta-Reflection
Paper
• 2607.01131
• Published • 9
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors
Paper
• 2606.32029
• Published • 15
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
Paper
• 2606.29082
• Published • 44
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
Paper
• 2606.23127
• Published • 27
Little Brains, Big Feats: Exploring Compact Language Models
Paper
• 2606.30062
• Published • 17
Hierarchical Experimentalist Agents
Paper
• 2606.29315
• Published • 6
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Paper
• 2607.01874
• Published • 24
PACE: A Proxy for Agentic Capability Evaluation
Paper
• 2607.02032
• Published • 21
When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search
Paper
• 2606.27669
• Published • 17
Paper
• 2607.27201
• Published • 111
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper
• 2608.05987
• Published • 103
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
Paper
• 2608.05102
• Published • 69
Weak-to-Strong On-Policy Distillation
Paper
• 2607.26246
• Published • 59
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Paper
• 2608.05000
• Published • 64
Progressive Agent Skill Generation via Reinforcement Learning
Paper
• 2608.01678
• Published • 62
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
Paper
• 2608.06301
• Published • 36
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Paper
• 2608.02603
• Published • 36
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
Paper
• 2608.05013
• Published • 38
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Paper
• 2608.03764
• Published • 28
Continual Learning in Transition
Paper
• 2608.06216
• Published • 27
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
Paper
• 2608.00155
• Published • 26
CADENA: Stepwise CAD Reverse Engineering
Paper
• 2608.00799
• Published • 43
Kimi K3: Open Frontier Intelligence
Paper
• 2607.24653
• Published • 523
Program-as-Weights: A Programming Paradigm for Fuzzy Functions
Paper
• 2607.02512
• Published • 310
Metis: Memory Foundation Model
Paper
• 2607.26760
• Published • 276
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
Paper
• 2607.13285
• Published • 236
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Paper
• 2607.21461
• Published • 157
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Paper
• 2607.16617
• Published • 146
Weak-to-Strong Generalization via Direct On-Policy Distillation
Paper
• 2607.05394
• Published • 150
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
Paper
• 2607.17250
• Published • 94
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Paper
• 2607.05382
• Published • 89
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Paper
• 2607.04438
• Published • 65
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
Paper
• 2608.07222
• Published • 14
Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
Paper
• 2608.00675
• Published • 11
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
Paper
• 2608.06714
• Published • 11
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
Paper
• 2608.06501
• Published • 4
CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
Paper
• 2608.09848
• Published • 7
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
Paper
• 2608.09043
• Published • 8
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper
• 2608.09888
• Published • 783
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Paper
• 2608.07169
• Published • 51
Motif 3: Technical Report
Paper
• 2608.09119
• Published • 48
Scaling Inherently Interpretable Language Models
Paper
• 2608.07594
• Published • 22
Stealing Reasoning Traces from Proprietary LLM APIs
Paper
• 2608.09867
• Published • 119
Evo-Bench: Can Language Models Improve Agent Harness?
Paper
• 2608.09096
• Published • 19
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Paper
• 2608.10915
• Published • 195
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
Paper
• 2608.10299
• Published • 137
Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Paper
• 2608.07645
• Published • 28
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Paper
• 2608.10875
• Published • 18
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
Paper
• 2608.09900
• Published • 13
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
Paper
• 2608.10692
• Published • 13
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
Paper
• 2608.08389
• Published • 12
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
Paper
• 2608.10636
• Published • 11
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
Paper
• 2608.10366
• Published • 11
Gaze Target Estimation Anywhere with Concepts
Paper
• 2608.11367
• Published • 5
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
Paper
• 2608.05604
• Published • 80
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
Paper
• 2608.08160
• Published • 30
Self-Evolving Embodied Agents via Skill-Harness Evolution
Paper
• 2608.11350
• Published • 15
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Paper
• 2608.12307
• Published • 115
DarwinX: Evolving Agent Harnesses Through Natural Selection
Paper
• 2608.07545
• Published • 115
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Paper
• 2608.13560
• Published • 64
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Paper
• 2608.12743
• Published • 44
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
Paper
• 2608.12990
• Published • 14
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning
Paper
• 2607.29211
• Published • 14
Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Paper
• 2608.14530
• Published • 35
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Paper
• 2608.11341
• Published • 70
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Paper
• 2608.14290
• Published • 35
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
Paper
• 2608.13667
• Published • 17
Multimodal Model Diffing for Feature Discovery and Control
Paper
• 2608.09928
• Published • 11
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Paper
• 2608.03744
• Published • 6
A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
Paper
• 2608.14075
• Published • 5
UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
Paper
• 2608.10835
• Published • 4
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Paper
• 2608.15089
• Published • 447
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
Paper
• 2608.15669
• Published • 63
Demystifying Agent Skills: Why They Work-Until They Don't
Paper
• 2608.14036
• Published • 170
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Paper
• 2608.17310
• Published • 108
Agent Lightning v1.0: Towards Harnessed Agentic RL
Paper
• 2608.17528
• Published • 35
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
Paper
• 2608.18933
• Published • 13
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Paper
• 2608.17597
• Published • 10
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Paper
• 2608.17800
• Published • 11
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
Paper
• 2608.16002
• Published • 7
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
Paper
• 2608.14229
• Published • 17
LLMs Get Smarter from Targeted Synthetic Multilingual Data
Paper
• 2608.15964
• Published • 6
The Embedder's Dilemma: LLMs Are Better, but at What Cost?
Paper
• 2608.12875
• Published • 15
EnvHarness: Awakening Static Worlds for Agent Learning
Paper
• 2608.19880
• Published • 276
TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
Paper
• 2608.15767
• Published • 9
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
Paper
• 2607.21596
• Published • 21
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
Paper
• 2608.23041
• Published • 64
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
Paper
• 2608.24876
• Published • 30
CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
Paper
• 2608.24794
• Published • 6
Automata from Agent Traces: Failure and Next-Step Prediction
Paper
• 2608.23670
• Published • 5
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Paper
• 2608.23691
• Published • 3
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
Paper
• 2608.25593
• Published • 117
FrontierChallenge: Evaluating Scientific Workflow Completion
Paper
• 2608.24979
• Published • 149
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning
Paper
• 2608.23318
• Published • 32
Code World Model: Coding Agent as World Brain
Paper
• 2608.25927
• Published • 39
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Paper
• 2608.25529
• Published • 17
A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
Paper
• 2608.21140
• Published • 5
Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
Paper
• 2608.22071
• Published
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Paper
• 2608.27345
• Published • 146
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
Paper
• 2608.15763
• Published • 54
CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension
Paper
• 2608.23172
• Published • 10
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
Paper
• 2608.28281
• Published • 106
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
Paper
• 2608.28122
• Published • 67
J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
Paper
• 2608.26582
• Published • 45
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
Paper
• 2608.24804
• Published • 41
Fast Weight Attention for Continual Learning
Paper
• 2608.27763
• Published • 33
Evaluating the Hidden Costs of Personalization in Large Language Models
Paper
• 2608.28833
• Published • 30
WebWorld: The Browser as a World Model for Self-Improving Web Code
Paper
• 2608.30530
• Published • 10
The Safeguard Worked. Is the LLM System Safer?
Paper
• 2609.00519
• Published • 5
StudentSim: Training LLM-based Student Simulators
Paper
• 2609.01591
• Published • 493
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Paper
• 2609.00111
• Published • 386
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
Paper
• 2609.01572
• Published • 36
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
Paper
• 2609.01404
• Published • 28
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
Paper
• 2609.01481
• Published • 19
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
Paper
• 2608.26623
• Published • 21
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
Paper
• 2608.30730
• Published • 18
Agents in the Large: Perception-Centered Architecture for Persistent Agents
Paper
• 2608.30478
• Published • 11
Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
Paper
• 2609.00621
• Published • 9
Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
Paper
• 2608.30457
• Published • 9
Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
Paper
• 2608.31082
• Published • 5
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
Paper
• 2609.02749
• Published • 555
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
Paper
• 2609.01437
• Published • 269
Aspire: Can Models Self-Evolve from Vague Goals?
Paper
• 2608.31111
• Published • 232
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
Paper
• 2609.02783
• Published • 120
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
Paper
• 2609.02886
• Published • 152
Language Models Can Control Their Own Attention
Paper
• 2609.02737
• Published • 75
Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents
Paper
• 2608.30322
• Published • 4
FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
Paper
• 2609.00377
• Published • 4
Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations
Paper
• 2608.29692
• Published • 2
Environment Evolution for Terminal Agents
Paper
• 2609.04128
• Published • 22
Using Grounded Theory for Agent Behavior Analysis at Scale
Paper
• 2608.30391
• Published • 19
Enoki: Efficient Multi-Level Hallucination Detection
Paper
• 2609.00581
• Published • 28
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Paper
• 2609.05258
• Published • 21
Paper
• 2609.03003
• Published • 31
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
Paper
• 2609.04482
• Published • 12
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
Paper
• 2609.01281
• Published • 16
Omni Interaction Agent Technical Report
Paper
• 2609.08977
• Published • 134
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
Paper
• 2609.07398
• Published • 77
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Paper
• 2609.09153
• Published • 41
Agentic Visual Generation: From Generative Models to Agentic Control
Paper
• 2609.06758
• Published • 36
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
Paper
• 2609.06289
• Published • 33
Kalman Delta Networks: Uncertainty-aware Associative Memory
Paper
• 2609.07816
• Published • 29
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
Paper
• 2609.05663
• Published • 22
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
Paper
• 2510.08999
• Published • 21
MOLE: Detecting Insider Threats in AI Agents
Paper
• 2609.06966
• Published • 19
Graph Machine: Towards Better Pretraining via Edges
Paper
• 2609.02881
• Published • 7
Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
Paper
• 2609.06140
• Published • 7
Scaling Automatic Research Agents via World Models
Paper
• 2608.12564
• Published • 458
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
Paper
• 2609.08572
• Published • 95
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
Paper
• 2609.11042
• Published • 58
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
Paper
• 2609.08149
• Published • 28
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Paper
• 2609.06702
• Published • 24
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Paper
• 2609.10226
• Published • 21
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
Paper
• 2609.09113
• Published • 20
Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning
Paper
• 2609.06806
• Published • 10
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Paper
• 2609.09134
• Published • 10
DF26: We Cannot Tell Fake From Real Anymore
Paper
• 2609.07369
• Published • 6
SenseNova-U1.5: Towards Native Unified Visual Intelligence
Paper
• 2609.11929
• Published • 259
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem
Paper
• 2609.07064
• Published • 141
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
Paper
• 2609.05903
• Published • 62
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat
Paper
• 2609.11155
• Published • 25
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
Paper
• 2609.10712
• Published • 43