rlsamplingJF/Llama-3.2-1B-finemath_part1_part2-rm-lr1e-6-constant-warmup_0.05-bs16-gc1.0-cc0.01-ls0.1-step150 1B • Updated about 22 hours ago • 15
rlsamplingJF/Llama-3.2-1B-finemath_part1_part2-rm-lr1e-6-constant-warmup_0.05-bs16-gc1.0-cc0-ls0-initial 1B • Updated about 24 hours ago • 21
rlsamplingJF/Llama-3.2-1B-finemath_part1_part2-rm-lr1e-6-constant-warmup_0.05-bs16-gc1.0-cc0-ls0-step405 1B • Updated 1 day ago • 18
rlsamplingJF/Qwen2.5-7B-Instruct-finemath_part1-rm-lr1e-6-constant-warmup_0.05-bs16-gc1.0-cc0.01-ls0.1-step30 7B • Updated 1 day ago • 28
rlsamplingJF/Llama-3.2-3B-finemath_part1_part2-rm-lr1e-5-constant-warmup_0.05-bs44-gc1.0-step105 3B • Updated 2 days ago • 17
rlsamplingJF/Llama-3.2-3B-finemath_part1_part2-rm-lr1e-6-constant-warmup_0.05-bs32-gc1.0-step195 3B • Updated 2 days ago • 27
rlsamplingJF/Qwen2.5-7B-Instruct-finemath_part1-rm-lr5e-7-constant-warmup_0.05-bs8-gc1.0-cc0.0-ls0.0-initial 7B • Updated 3 days ago • 23
rlsamplingJF/Qwen2.5-7B-Instruct-finemath_part1-rm-lr5e-7-constant-warmup_0.05-bs8-gc1.0-cc0.0-ls0.0-step105 7B • Updated 3 days ago • 23
rlsamplingJF/Qwen2.5-7B-Instruct-finemath_part1-rm-lr1e-6-constant-warmup_0.05-bs8-gc1.0-step60 7B • Updated 3 days ago • 94
rlsamplingJF/Llama-3.2-3B-finemath_part1-rm-lr1e-6-constant-warmup_0.05-bs32-gc1.0-cc0.01-ls0-initial 3B • Updated 4 days ago • 28