cagliostro-v3.5

The chat version is out: cagliostro-v3.5-chat, this model after one epoch of supervised fine-tuning on smol-smoltalk. It ends its own turns, follows system prompts, and comes out level with SmolLM2-135M-Instruct on MT-Bench.

A 146M parameter decoder-only language model. It is cagliostro-v3's final checkpoint with one more short training stage: 0.39B tokens of v3's own cooldown data mixed with how-to and textbook text from Cosmopedia. Same architecture, same tokenizer, 75.39B tokens in total.

It scores 27.49 on the Open SLM Leaderboard metric, above every model listed there at the time of release.

Results

Zero-shot, measured with lm-evaluation-harness and the leaderboard's own ArithMark-3 script, on the float32 weights in this repository.

Benchmark Metric Score
HellaSwag acc_norm 43.41
ARC-Easy acc_norm 53.75
ARC-Challenge acc_norm 29.35
PIQA acc_norm 68.06
ArithMark-3 acc_norm 45.30
Open SLM Index 27.49

The Index is the leaderboard's own formula, (N(HellaSwag,25) + N(CombinedARC,25) + N(PIQA,50) + 0.65*N(ArithMark,25)) / 3.65 where N(v,c) = 100(v-c)/(100-c) and CombinedARC is the mean of ARC-Easy and ARC-Challenge.

For context, using the leaderboard's published figures for every other model:

Model Params Tokens Index
cagliostro-v3.5 146M 75.4B 27.49
SmolLM2-135M 135M 2T 27.13
cagliostro-v3 146M 75B 26.55
SmolLM-135M 135M 600B 25.74
GPT-X2.5-135M 135M 75B 25.17
Haidass1.5-143M 143M 400B 25.07

Index against the sub-150M field

Per-benchmark comparison

The lead over SmolLM2-135M is 0.36 Index on the leaderboard's figures. Re-evaluated in the same harness, SmolLM2-135M scores 27.01 and cagliostro-v3.5 is ahead by 0.48, but a paired bootstrap over benchmark items puts the 95% interval for that gap at [-0.82, +1.75]. On points it is first. On these five benchmarks the two models cannot be told apart with confidence, and we would not claim more than that.

cagliostro-v3 appears in the same-harness charts as it measures today, Index 26.31. Its leaderboard row lists 26.55.

What changed from v3

Field Value
Starting point cagliostro-v3 step 762,939, weights and AdamW state
Steps 4,000
Tokens per step 98,304
Tokens 0.39B
Learning rate warmup over 200 steps to 3e-4, then cosine to zero
Peak learning rate a tenth of v3's 3e-3
Optimizer AdamW, weight decay 0.01
Hardware one RTX PRO 6000 Blackwell
Wall clock about 41 minutes at 162,000 tokens per second
Source Share
Cosmopedia v1, WikiHow-style tutorials 30.0%
FineWeb-Edu (deduplicated) 20.35%
Cosmopedia v1, textbooks (OpenStax, Khan Academy, Stanford) 15.0%
Cosmopedia v2 13.75%
FineMath 3+ 8.25%
OpenMathInstruct-2 7.15%
SmolTalk 2.75%
DCLM-Baseline 2.75%

The bottom six rows are v3's cooldown mixture scaled to 55%. The WikiHow set is 174M tokens, so it was seen about 0.7 times. The textbook slice is the first 150M tokens of those three subsets.

What changed, benchmark by benchmark

The gain is mostly ArithMark-3 (+2.6 points) and HellaSwag (+0.9), with smaller gains on ARC-Challenge and PIQA. ARC-Easy moved the other way by about a point. That drop also shows up when v3 is trained the same way on its own cooldown mixture, so it comes from the extra low learning rate stage rather than from the new data.

The post-training run

How it was chosen

Six runs started from the same v3 checkpoint, identical except for the data. Before any of them finished we wrote down what a winner had to do: score above SmolLM2-135M's 27.13, beat SmolLM2-135M re-measured in the same harness, and beat v3 with the paired 95% interval above zero. Then a second run of the same recipe with a different data seed had to clear 27.13 as well.

All six post-training runs

Run Index
100% SmolTalk chat data 26.00
v3's own cooldown mixture 26.54
+ 30% WikiHow, seed 1 27.14
+ 30% WikiHow, seed 2 26.98
+ 30% WikiHow + 15% textbooks, seed 1 27.31
+ 30% WikiHow + 15% textbooks, seed 2 27.49

WikiHow alone passed once and then failed its second seed, so it was dropped. WikiHow with textbooks passed on both seeds. This repository holds the seed 2 checkpoint, the stronger of the two. The recipe averages 27.40 over both.

Chat data lost points, as it does for most base and instruct pairs on the leaderboard. The benchmarks score raw log-likelihood with no chat template, so text that looks like a conversation does not help.

Contamination checks

No benchmark training, validation or test split was used at any point.

Every new document went through the same 13-gram decontamination v3 used, against HellaSwag validation, ARC test, PIQA validation and ArithMark-3, dropping any document with a match. In a sample of 8,000 documents from the two Cosmopedia sources it dropped none.

The HellaSwag gain deserves a closer look, because about two thirds of HellaSwag is built from real WikiHow articles and Cosmopedia's WikiHow set was generated from real WikiHow titles.

HellaSwag gain by item source

19.3% of HellaSwag's WikiHow validation items share an article title with the Cosmopedia set. For the training split the figure is 18.5%, so the validation items are not singled out. The gain is about the same on items whose title appears in the training data (+1.45) as on items whose title does not (+1.24). ActivityNet items, which have nothing to do with WikiHow, barely moved (+0.09). That pattern is a model getting better at everyday how-to text, not one that has seen the test articles.

Sample outputs

Greedy decoding, float32, first sentence of each completion, unedited. These are all twelve prompts that were tried; the video shows four of them. The full completions are in samples.json.

Prompt Completion
To solve 47 + 68, first add the numbers together: \[ 47 + 68 = 115 \] Now, divide 115 by 2 to find the quotient: ...
A prime number is a positive integer greater than 1 that has no positive divisors other than 1 and itself.
Photosynthesis is the process by which plants convert sunlight into chemical energy.
How to boil an egg: First, you need to crack the egg into a bowl.
To tie a shoelace, start by wrapping the yarn around the hook of the shoelace.
The water cycle begins when rain falls onto the ground and turns into water droplets.
If a train travels 60 miles in 2 hours, its speed is 60 miles per hour.
Sam had 24 apples and gave away 9, so Sam now has 24 - 9 = 15 apples.
The capital of France is Paris.
Gravity is the force that pulls objects towards each other.
The area of a rectangle is found by multiplying its length by its width.
Plants need sunlight because they need to grow.

The misses are typical for this size. It adds 47 and 68 correctly and then divides by 2 for no reason, gets the train's speed wrong, starts boiling an egg by cracking it, and describes a shoelace as yarn.

Model details

Field Value
Parameters 146,352,000
Non-embedding parameters 85.7%
Layers 30
Hidden size 640
Intermediate size 1,536
Attention heads 10
Key/value heads 5
Attention Grouped query attention with cross-head subspace attenuation
Activation SwiGLU
Normalization RMSNorm, eps 1e-6
Positional encoding RoPE, theta 100,000
Context length 2,048
Vocabulary 32,768 BPE
Embeddings Tied input and output
Logit cap 15.0
Tokens seen 75.39B (75B pretraining, 0.39B post-training)
Weights float32 safetensors

The architecture is defined in this repository. trust_remote_code=True is required because CagliostroForCausalLM is not part of transformers.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "bench-labs/cagliostro-v3.5"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, dtype=torch.float32)

ids = tok("The capital of France is", return_tensors="pt")
out = model.generate(**ids, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

This is a base model with no instruction tuning and no chat template. It completes text. For chat, use cagliostro-v3.5-chat.

Fast inference

generate() now uses a KV cache, so it is fast out of the box. On a GPU, generate_fast() goes further with a static cache and CUDA graphs:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "bench-labs/cagliostro-v3.5"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, dtype=torch.float16).to("cuda")

ids = tok("The water cycle begins when", return_tensors="pt").input_ids.to("cuda")
out = model.generate_fast(ids, max_new_tokens=200, repetition_penalty=1.1)
print(tok.decode(out[0], skip_special_tokens=True))

Speed of cagliostro-v3.5-chat, which shares this architecture and this code:

Precision GPU CPU
Previous release, generate() without a cache fp32 31 tok/s (1.0x) 8 tok/s (1.0x)
generate(), now with a KV cache fp32 43 tok/s (1.4x) 22 tok/s (2.7x)
generate_fast() fp32 88 tok/s (2.8x) 23 tok/s (2.8x)
generate_fast() fp16 150 tok/s (4.8x)
generate_fast(compile=True) fp16 230 tok/s (7.4x)

Measured with the code in this repository. GPU numbers use 256 new tokens; CPU numbers use 128. Both use a 23-token prompt and repetition penalty 1.1, median of three runs. On a GPU, generate_fast() uses CUDA graphs; on a CPU it falls back to a static-cache loop.

  • The forward pass used for benchmark scoring is unchanged bit for bit, so leaderboard numbers do not move.
  • In float32, generate() and generate_fast() give exactly the same greedy text as the previous release on all 12 sample prompts.
  • float16 and bfloat16 are about 1.7x faster than float32. Greedy text can drift from float32 at a near tie, usually after a few dozen tokens (float16: 8 of 12 prompts identical over 200 tokens). float16 tracked float32 slightly better than bfloat16.
  • compile=True fuses the decode step with torch.compile for about another 1.5x. It needs a working Triton install and costs about 40 seconds the first time in each process.
  • generate_fast() takes one sequence at a time and supports temperature, top_p, seed, repetition_penalty and a transformers TextStreamer. Use generate() for batches.

Evaluate in float32 as described below. These speed options are for generation only.

Reproducing the evaluation

pip install lm-eval
python -m lm_eval --model hf \
  --model_args pretrained=bench-labs/cagliostro-v3.5,dtype=float32,trust_remote_code=True \
  --tasks hellaswag,arc_easy,arc_challenge,piqa \
  --num_fewshot 0 --batch_size 64 --device cuda:0

ArithMark-3 uses the script linked from the leaderboard, pointed at the same model id with --dtype float32. Evaluate in float32. A bfloat16 round trip moves logits enough at this logit cap to change borderline multiple-choice answers.

These exact commands, run against this repository, return the numbers in the results table.

Limitations

English only. 2,048 token context. No instruction tuning (see cagliostro-v3.5-chat for that), no safety tuning, no RLHF. At 146M parameters it confabulates freely and should not be relied on for factual questions. ARC-Easy is about a point below cagliostro-v3. The mathematics ability measured by ArithMark is arithmetic and short symbolic work, not general mathematical reasoning.

Beyond the leaderboard

The leaderboard reports five benchmarks, so we also ran a wider set against SmolLM2-135M, the model this one replaced at the top. Both sides got the same setup: lm-evaluation-harness 0.4.13, zero-shot, float32, full test sets, 2,048-token context. The first four rows reproduce the leaderboard's own numbers exactly. Differences are cagliostro-v3.5 minus SmolLM2-135M, with a 95% interval from a paired bootstrap over items.

Benchmark Metric cagliostro-v3.5 SmolLM2-135M Difference [95% CI]
HellaSwag acc_norm 43.41 43.11 +0.30 [-0.34, +1.01]
ARC-Easy acc_norm 53.75 58.59 -4.84 [-6.48, -3.03]
ARC-Challenge acc_norm 29.35 29.61 -0.26 [-2.30, +1.79]
PIQA acc_norm 68.06 68.50 -0.44 [-2.07, +1.20]
LAMBADA (OpenAI) acc 38.70 42.89 -4.19 [-5.39, -2.97]
LAMBADA (OpenAI) perplexity, lower is better 27.25 19.26
WikiText-2 bits per byte, lower is better 0.910 0.847
WikiText-2 word perplexity, lower is better 29.13 23.13
BLiMP (67 sets) acc 77.26 78.97 -1.71 [-2.02, -1.36]
BoolQ acc 61.25 60.31 +0.95 [-0.25, +2.11]
Social IQa acc 40.23 39.36 +0.87 [-0.92, +2.66]
WinoGrande acc 52.41 52.96 -0.55 [-3.87, +2.61]
OpenBookQA acc_norm 32.20 33.40 -1.20 [-3.80, +1.60]
SciQ acc_norm 75.10 78.20 -3.10 [-5.50, -0.90]
CommonsenseQA acc 19.33 19.49 -0.16 [-3.11, +2.95]
MMLU (57 subjects) acc 24.57 24.35 +0.22 [-0.78, +1.22]
ArithMark-3 official script 45.30 38.80 +6.50
ArithMark-2 (not on the board) official script 59.84 33.20 +26.64 [+24.32, +28.92]
Open SLM Index 27.49 27.01 +0.48 [-0.82, +1.75]

Bold marks an interval that excludes zero.

Outside the leaderboard, SmolLM2-135M is the stronger general language model. It is clearly ahead on plain language modelling (WikiText, LAMBADA), grammar (BLiMP), science questions (SciQ) and ARC-Easy. The two are statistically level on the other nine tasks; CommonsenseQA and MMLU are at chance for both. This model's lead on the board comes from arithmetic, which the Index weights at 0.65: 6.5 points on ArithMark-3 and 26.6 on ArithMark-2, a separate arithmetic set the board does not use. SmolLM2-135M was trained on 2 trillion tokens, this model on 75.4 billion.

The chat versions were compared too. On MT-Bench, judged pairwise by gpt-oss-120b in both answer orders, cagliostro-v3.5-chat won 66 question-turns to SmolLM2-135M-Instruct's 55, with 39 ties (sign test p = 0.36, so level). Full tables, every answer, all 320 judge verdicts and the scripts are in the chat repo's evals/ folder and card.

License

Apache-2.0. The training data is drawn from FineWeb-Edu and FineMath (ODC-By), DCLM-Baseline and OpenMathInstruct-2 (CC-BY-4.0), Cosmopedia v1, Cosmopedia v2 and SmolTalk (Apache-2.0).

Downloads last month
695
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bench-labs/cagliostro-v3.5

Finetuned
(1)
this model
Finetunes
1 model

Datasets used to train bench-labs/cagliostro-v3.5

Spaces using bench-labs/cagliostro-v3.5 2

Collection including bench-labs/cagliostro-v3.5