-
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Paper • 2604.13016 • Published • 116 -
Thinking-Space/Qwen3-1.7B-SFT
Text Generation • 2B • Updated • 236 • 4 -
Thinking-Space/Qwen3-4B-Base-GRPO
Text Generation • 4B • Updated • 961 • 3 -
Thinking-Space/OpenThought3-Qwen3-4B
Viewer • Updated • 305k • 165 • 3
Collections
Discover the best community collections!
Collections including paper arxiv:2604.13016
-
Code as Agent Harness
Paper • 2605.18747 • Published • 224 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 73 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 61
-
LLM Pruning and Distillation in Practice: The Minitron Approach
Paper • 2408.11796 • Published • 62 -
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
Paper • 2408.09174 • Published • 53 -
To Code, or Not To Code? Exploring Impact of Code in Pre-training
Paper • 2408.10914 • Published • 45 -
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
Paper • 2408.11878 • Published • 64
-
DataComp-VLM: Improved Open Datasets for Vision-Language Models
Paper • 2606.28551 • Published • 51 -
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Paper • 2607.01874 • Published • 21 -
PACE: A Proxy for Agentic Capability Evaluation
Paper • 2607.02032 • Published • 18 -
Measuring the Gap Between Human and LLM Research Ideas
Paper • 2607.01233 • Published • 19
-
Thinking-Space/Qwen3-1.7B-SFT
Text Generation • 2B • Updated • 236 • 4 -
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Paper • 2604.13016 • Published • 116 -
Thinking-Space/Qwen3-4B-Base-GRPO
Text Generation • 4B • Updated • 961 • 3 -
Thinking-Space/OpenThought3-Qwen3-4B
Viewer • Updated • 305k • 165 • 3
-
Visual Spatial Tuning
Paper • 2511.05491 • Published • 53 -
Adam's Law: Textual Frequency Law on Large Language Models
Paper • 2604.02176 • Published • 108 -
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
Paper • 2604.10098 • Published • 83 -
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Paper • 2604.13016 • Published • 116
-
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Paper • 2604.13016 • Published • 116 -
Thinking-Space/Qwen3-1.7B-SFT
Text Generation • 2B • Updated • 236 • 4 -
Thinking-Space/Qwen3-4B-Base-GRPO
Text Generation • 4B • Updated • 961 • 3 -
Thinking-Space/OpenThought3-Qwen3-4B
Viewer • Updated • 305k • 165 • 3
-
DataComp-VLM: Improved Open Datasets for Vision-Language Models
Paper • 2606.28551 • Published • 51 -
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Paper • 2607.01874 • Published • 21 -
PACE: A Proxy for Agentic Capability Evaluation
Paper • 2607.02032 • Published • 18 -
Measuring the Gap Between Human and LLM Research Ideas
Paper • 2607.01233 • Published • 19
-
Code as Agent Harness
Paper • 2605.18747 • Published • 224 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 73 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 61
-
Thinking-Space/Qwen3-1.7B-SFT
Text Generation • 2B • Updated • 236 • 4 -
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Paper • 2604.13016 • Published • 116 -
Thinking-Space/Qwen3-4B-Base-GRPO
Text Generation • 4B • Updated • 961 • 3 -
Thinking-Space/OpenThought3-Qwen3-4B
Viewer • Updated • 305k • 165 • 3
-
Visual Spatial Tuning
Paper • 2511.05491 • Published • 53 -
Adam's Law: Textual Frequency Law on Large Language Models
Paper • 2604.02176 • Published • 108 -
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
Paper • 2604.10098 • Published • 83 -
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Paper • 2604.13016 • Published • 116
-
LLM Pruning and Distillation in Practice: The Minitron Approach
Paper • 2408.11796 • Published • 62 -
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
Paper • 2408.09174 • Published • 53 -
To Code, or Not To Code? Exploring Impact of Code in Pre-training
Paper • 2408.10914 • Published • 45 -
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
Paper • 2408.11878 • Published • 64