TrynMini v1 β€” a sentence transformer built from scratch (research prototype)

A compact transformer sentence encoder implemented from scratch β€” embeddings, sinusoidal positional encoding, multi-head self-attention, feed-forward blocks, and masked mean pooling, all hand-written (no transformers library) and trained on STS-B plus synthetic similarity data.

This is an early research prototype, not a production encoder.

v1 shipped with a tiny auto-generated vocabulary (~164 tokens) learned from limited, templated training text. As a result it handles the phrases it saw in training but does not generalize to arbitrary English β€” most real-world words fall back to [UNK]. It's published as a learning milestone and for provenance.

For a model you can actually use, see LNTTushar/trynmini-v2-static-7m-v2 β€” a properly trained static embedder (30k vocab, 0.715 STSB-dev, 7.8 MB, CPU-only, 5-line quickstart). A re-trained transformer v2 with a real 30k vocabulary is also in progress.

Architecture

  • Custom Transformer encoder: 4 layers Β· 6 attention heads Β· 384-dim Β· mean pooling
  • Parameters: 3,299,584 (~3.3M)
  • Model size: ~12.6 MB
  • Vocabulary: ~164 tokens (auto-generated β€” the main limitation; fixed in v2's 30k vocab)
  • Max sequence length: 128 tokens

Results

On its own STS-B evaluation (within the limits of the 164-token vocab):

Metric Score
Spearman 0.677
Pearson 0.672

These numbers reflect the small-vocab setup and should not be compared directly to encoders trained on a full vocabulary.

What this project is

TrynMini is a from-scratch study of how sentence embedding models work, built day by day: tokenizer β†’ transformer β†’ distillation β†’ static compression β†’ calibration. v1 (this repo) is the first working transformer stage. The line continues in the static v2 models linked above, which are the ones intended for real use.

License

Apache-2.0.

Downloads last month
87
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support