TrynMini v1 β a sentence transformer built from scratch (research prototype)
A compact transformer sentence encoder implemented from scratch β embeddings, sinusoidal
positional encoding, multi-head self-attention, feed-forward blocks, and masked mean pooling,
all hand-written (no transformers library) and trained on STS-B plus synthetic similarity data.
This is an early research prototype, not a production encoder.
v1 shipped with a tiny auto-generated vocabulary (~164 tokens) learned from limited, templated training text. As a result it handles the phrases it saw in training but does not generalize to arbitrary English β most real-world words fall back to
[UNK]. It's published as a learning milestone and for provenance.For a model you can actually use, see LNTTushar/trynmini-v2-static-7m-v2 β a properly trained static embedder (30k vocab, 0.715 STSB-dev, 7.8 MB, CPU-only, 5-line quickstart). A re-trained transformer v2 with a real 30k vocabulary is also in progress.
Architecture
- Custom Transformer encoder: 4 layers Β· 6 attention heads Β· 384-dim Β· mean pooling
- Parameters: 3,299,584 (~3.3M)
- Model size: ~12.6 MB
- Vocabulary: ~164 tokens (auto-generated β the main limitation; fixed in v2's 30k vocab)
- Max sequence length: 128 tokens
Results
On its own STS-B evaluation (within the limits of the 164-token vocab):
| Metric | Score |
|---|---|
| Spearman | 0.677 |
| Pearson | 0.672 |
These numbers reflect the small-vocab setup and should not be compared directly to encoders trained on a full vocabulary.
What this project is
TrynMini is a from-scratch study of how sentence embedding models work, built day by day: tokenizer β transformer β distillation β static compression β calibration. v1 (this repo) is the first working transformer stage. The line continues in the static v2 models linked above, which are the ones intended for real use.
License
Apache-2.0.
- Downloads last month
- 87