hv-multimodal-fsdd-16384
A 250 KB cross-modal audio-text model in the hypervector framework. Runs on CPU. NumPy only. No pretrained embeddings, no GPU, no PyTorch.
Prototype classification (FSDD): 0.967 (10 spoken digits)
Cross-modal retrieval (audio โ category): 0.917
Text-to-audio retrieval (category โ audio): 1.000
Random baseline: 0.100
Model size: ~250 KB
Dependencies: NumPy only
Model Description
A cross-modal retrieval model that binds audio hypervectors to their category labels in a single memory. Given an audio clip, retrieve the matching category. Given a category name, retrieve a matching audio clip. Both directions use the same 250 KB model.
This is not a classifier in the deep-learning sense. It has no learned parameters, no gradient descent, and no pretraining. It is a random projection followed by bundling. The accuracy comes from the geometry of the hypervector space, not from training.
Architecture
The model uses four components.
1. Feature extraction (112 dimensions). Each audio clip is processed into a fixed feature vector:
- 32 mel-band means + 32 mel-band standard deviations (64 dims)
- 12 MFCC means + 12 MFCC stds + 12 delta means + 12 delta-delta means (48 dims, C0 dropped)
C0 is the first MFCC coefficient and represents overall frame energy. It has 84ร the magnitude of the median dimension and washes out the remaining signal after L2 normalization. Dropping it and z-scoring the rest is essential.
2. Per-dimension z-scoring. After extraction, each of the 112 dimensions is standardized across the training set. This ensures no single feature dominates the projection.
3. Random projection to hypervector. A random bipolar matrix R โ {-1, +1}^{16384 ร 112} projects the feature vector to a 16384-bit bipolar hypervector. Sign of the projection becomes the bit.
4. Binding + bundling. Each training clip's hypervector is bound to its category's codebook hypervector (element-wise product), then all pairs are summed. The result is a single 16384-bit memory. Retrieval: unbind with the query, clean up against the category codebook.
Evaluation
Dataset: FSDD. Free Spoken Digit Dataset. 300 clips, 10 categories (digits 0โ9), 30 clips per digit. 240 training, 60 test (80/20 split).
| Metric | Value | Chance |
|---|---|---|
| Prototype classification | 0.967 | 0.100 |
| Cross-modal retrieval | 0.917 | 0.100 |
| Text-to-audio retrieval | 1.000 | 0.017 |
The gap between prototype (0.967) and cross-modal (0.917) is 3 clips out of 60. On FSDD, the cross-modal architecture reaches the prototype ceiling.
Dataset: ESC-50. 500 clips, 50 environmental sound categories. Not competitive. The same model reaches 0.230 prototype accuracy versus chance of 0.020. This is 11ร chance but far below the 0.85+ that pretrained audio models achieve. ESC-50 is not included in the model card claims.
Intended Use
- Cross-modal retrieval. Find audio clips matching a text label, or find labels matching audio.
- Small-footprint audio tagging. Fixed 250 KB model, CPU-only inference.
- On-device audio classification. Any device with NumPy can run the model in under 1 ms per query.
- Few-shot adaptation. Training on a new label set takes under 1 second.
Limitations
- FSDD-scale only. Works well on 10 acoustically distinct classes. Fails on 50 overlapping environmental classes (see ESC-50 result).
- Feature-bound. The 112-dim feature vector determines the ceiling. MFCC on ESC-50 saturates at 0.23 prototype. Any task harder than spoken digits requires a pretrained audio embedding.
- No semantics. The model does not understand audio content. It computes similarity in a fixed feature space.
- Random projection. The projection matrix is not learned. The accuracy depends on the feature extractor alone.
- Not a replacement for Whisper or PANNs. If you have access to a pretrained audio model, use it. This is for environments where a 250 MB model is impossible.
Comparison to Other Solutions
| Approach | Size | Latency | FSDD accuracy | Dependencies |
|---|---|---|---|---|
| hv-multimodal | 250 KB | < 1 ms | 0.967 | NumPy only |
| Whisper-tiny (39M params) | 151 MB | ~500 ms | ~0.99 | PyTorch |
| PANNs CNN14 (81M params) | 327 MB | ~50 ms | ~0.99 | PyTorch |
| all-MiniLM-L6-v2 audio head | 90 MB | ~10 ms | ~0.95 | PyTorch |
| MFCC + logistic regression | ~1 MB | ~10 ms | ~0.90 | scikit-learn |
The gap to pretrained models is small on FSDD but large on ESC-50. Use the model when size or dependency constraints rule out the alternatives.
How to Use
from hv_multimodal import HVMultimodal
model = HVMultimodal.load("hv-multimodal-fsdd-16384")
# Audio -> category
category = model.classify("path/to/recording.wav")
# -> "7"
# Category -> retrieve matching audio
top_indices = model.retrieve_audio("7", k=3)
# Cross-modal retrieval in both directions
scores = model.score_audio_against_categories("path/to/recording.wav")
# -> {'0': 0.02, '1': 0.01, ..., '7': 0.91, ...}
- Downloads last month
- 30
Evaluation results
- Prototype Classification on FSDD (Free Spoken Digit Dataset)self-reported0.967
- Cross-Modal Retrieval on FSDD (Free Spoken Digit Dataset)self-reported0.917
- Text-to-Audio Retrieval on FSDD (Free Spoken Digit Dataset)self-reported1.000
- Random Baseline on FSDD (Free Spoken Digit Dataset)self-reported0.100