SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks Paper ⢠2410.03769 ⢠Published Oct 2, 2024
Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning Paper ⢠2506.23056 ⢠Published Jun 29, 2025
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Paper ⢠2508.07407 ⢠Published Aug 10, 2025 ⢠100
InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem Paper ⢠2602.14367 ⢠Published Feb 16 ⢠17
AgentSearchBench: A Benchmark for AI Agent Search in the Wild Paper ⢠2604.22436 ⢠Published Apr 24 ⢠14
Ref-NeuS: Ambiguity-Reduced Neural Implicit Surface Learning for Multi-View Reconstruction with Reflection Paper ⢠2303.10840 ⢠Published Mar 20, 2023 ⢠1
GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs Paper ⢠2412.11258 ⢠Published Dec 15, 2024 ⢠13
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond Paper ⢠2509.26636 ⢠Published Sep 30, 2025 ⢠1
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm Paper ⢠2512.08629 ⢠Published Dec 9, 2025 ⢠1
LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding Paper ⢠2405.17104 ⢠Published May 27, 2024
InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem Paper ⢠2602.14367 ⢠Published Feb 16 ⢠17
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents Paper ⢠2602.06855 ⢠Published Feb 6 ⢠83
Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation Paper ⢠2510.02306 ⢠Published Oct 2, 2025 ⢠4
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation Paper ⢠2406.08482 ⢠Published Jun 12, 2024
Geospatial Foundational Embedder: Top-1 Winning Solution on EarthVision Embed2Scale Challenge (CVPR 2025) Paper ⢠2509.06993 ⢠Published Sep 3, 2025
"Ask Me Anything": How Comcast Uses LLMs to Assist Agents in Real Time Paper ⢠2405.00801 ⢠Published May 1, 2024
Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation Paper ⢠2311.09569 ⢠Published Nov 16, 2023
SpeechNet: Weakly Supervised, End-to-End Speech Recognition at Industrial Scale Paper ⢠2211.11740 ⢠Published Nov 21, 2022