agi-noobs/chess-sft-20k-llm-reasoning-enriched-dpo-hard-negatives-v1 Viewer • Updated Dec 30, 2025 • 1.57k • 93 • 6
ticoAg/llm-complex-reasoning-train-qwen2-72b-instruct-correct Viewer • Updated Aug 8, 2024 • 7.11k • 113 • 6
Yuhan123/vicuna-13b-self_consistency_random_var_3 Text Generation • 13B • Updated Mar 14, 2025 • 79 • 8
s-emanuilov/LLMBG-Llama-3.1-8B-BG-Reasoning-v0.1 Text Generation • 8B • Updated Feb 9, 2025 • 266 • 15
Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure Paper • 2609.23551 • Published 17 days ago • 7
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 16 days ago • 157
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay Paper • 2609.25001 • Published 16 days ago • 132
Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation Paper • 2609.20758 • Published 20 days ago • 4
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 19 days ago • 138
xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning Updated Dec 9, 2025 • 280 • 12
Yuhan123/vicuna-13b-self_consistency_neg_exp_var_5 Text Generation • 13B • Updated Mar 14, 2025 • 168 • 8
Yuhan123/vicuna-13b-self_consistency_neg_exp_var_4 Text Generation • 13B • Updated Mar 14, 2025 • 168 • 8
Learning Foresight without Explicit Trajectories for 3D Diffusion Policies Paper • 2609.20669 • Published 20 days ago • 9