tt-x1_lr-lr4e6 (TaskTrove X1 lr=4e-6, global_step_30)

X1 learning-rate ablation arm (lr = 4e-6) of the TaskTrove hyperparameter sweep.

  • Base model: Qwen/Qwen3-Coder-30B-A3B-Instruct
  • Harness: terminus-2 agentic RL
  • Source: DCAgent/exp_rpt_multifile (TaskTrove, pytest verifier, pass_ratio reward shaping)
  • Geometry: 6 nodes x 4 GH200, Jupiter (JSC)

Status

TERMINATED by owner at step 36/80 after reward collapsed to exactly 0.000 at steps 33-36 (a hard collapse, not the gradual decay expected from an over-large learning rate). This is not a completed / valid X1 result: it did not reach its 80-step horizon. Selected checkpoint = global_step_30, the trailing-5 EMA peak (0.1700) among saved exports and the last healthy checkpoint before the step-33 collapse. See training_logs/report.md for the full curve.

Training Traces

Companion trial-level trace dataset: https://huggingface.co/datasets/penfever/tt-x1_lr-lr4e6

training_logs/

Per-step metric surface, the parse_skyrl_metrics.py analysis report, and the reward plot (metrics.csv, report.md, reward_plot.png).

Downloads last month
-
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for laion/tt-x1_lr-lr4e6-30-30B

Finetuned
(71)
this model