theworker02 commited on
Commit
df972ca
·
verified ·
1 Parent(s): 9b47349

Honest XL model card: 443719680 params, 120 CPU steps, NLL 5.8116

Browse files
Files changed (1) hide show
  1. README.md +18 -11
README.md CHANGED
@@ -12,21 +12,28 @@ datasets:
12
  base_model: gpt2-scratch
13
  ---
14
 
15
- # Open Reason open-reason-xl (CPU)
16
 
17
- This is an **XL** GPT-2-style causal LM (~450M parameters) trained from scratch on the Open Reason SFT split. It is larger than `theworker02/open-reason-large` and is **not** a 1B model and is **not** `theworker02/open-reason-1b`.
 
 
18
 
19
- - Parameters: 443719680
20
- - Architecture: n_layer=22 n_embd=1280 n_head=20
21
- - Steps: 120
22
- - Backend: cpu-host
23
- - CUDA used: False
24
- - Hardware: Host CPU (AMD64 Family 26 Model 68 Stepping 0, AuthenticAMD); torch 2.12.0+cpu; cuda_available=False; docker_installed=False; docker_used=False. NVIDIA CUDA was not used. AMD GPU/ROCm/DirectML were not used.
25
- - Dataset: theworker02/open-reason pipeline 1.4.0
26
- - SFT rows: 3175
27
- - Final loss: 5.811557769775391
 
 
 
 
28
 
29
  Related checkpoints (none of these is a 1B model):
 
30
  - Dataset: https://huggingface.co/datasets/theworker02/open-reason
31
  - Small: https://huggingface.co/theworker02/open-reason-small
32
  - Medium: https://huggingface.co/theworker02/open-reason-medium
 
12
  base_model: gpt2-scratch
13
  ---
14
 
15
+ # Open Reason XL (CPU)
16
 
17
+ GPT-2-style causal LM trained **from scratch** on the Open Reason SFT split.
18
+ Exact parameter count: **443,719,680**. This is **not** a 1B model and is **not**
19
+ `theworker02/open-reason-1b`.
20
 
21
+ ## Training facts
22
+
23
+ - Parameters: 443,719,680 (`sum(p.numel() for p in model.parameters())`)
24
+ - Architecture: n_layer=22, n_embd=1280, n_head=20, vocab=8192, seq=256
25
+ - Steps: 120 (batch 1, gradient accumulation 2, gradient checkpointing)
26
+ - Device: **host CPU** (AMD Ryzen 9 9950X, 32 threads)
27
+ - `torch`: 2.12.0+cpu; `cuda_available=False`; Docker not installed
28
+ - NVIDIA CUDA was not used. AMD GPU / ROCm / DirectML were not used
29
+ - Dataset: [`theworker02/open-reason`](https://huggingface.co/datasets/theworker02/open-reason) pipeline **1.4.0**
30
+ - SFT rows: **3175** (`data/release/all.jsonl`)
31
+ - Final training NLL: **5.8116**
32
+ - License: Apache-2.0
33
+ - No held-out exact-match / coding / math benchmark scores are claimed
34
 
35
  Related checkpoints (none of these is a 1B model):
36
+
37
  - Dataset: https://huggingface.co/datasets/theworker02/open-reason
38
  - Small: https://huggingface.co/theworker02/open-reason-small
39
  - Medium: https://huggingface.co/theworker02/open-reason-medium