hv-multimodal-fsdd-16384

A 250 KB cross-modal audio-text model in the hypervector framework. Runs on CPU. NumPy only. No pretrained embeddings, no GPU, no PyTorch.

Prototype classification (FSDD): 0.967 (10 spoken digits)

Cross-modal retrieval (audio โ†’ category): 0.917

Text-to-audio retrieval (category โ†’ audio): 1.000

Random baseline: 0.100

Model size: ~250 KB

Dependencies: NumPy only


Model Description

A cross-modal retrieval model that binds audio hypervectors to their category labels in a single memory. Given an audio clip, retrieve the matching category. Given a category name, retrieve a matching audio clip. Both directions use the same 250 KB model.

This is not a classifier in the deep-learning sense. It has no learned parameters, no gradient descent, and no pretraining. It is a random projection followed by bundling. The accuracy comes from the geometry of the hypervector space, not from training.

Architecture

The model uses four components.

1. Feature extraction (112 dimensions). Each audio clip is processed into a fixed feature vector:

  • 32 mel-band means + 32 mel-band standard deviations (64 dims)
  • 12 MFCC means + 12 MFCC stds + 12 delta means + 12 delta-delta means (48 dims, C0 dropped)

C0 is the first MFCC coefficient and represents overall frame energy. It has 84ร— the magnitude of the median dimension and washes out the remaining signal after L2 normalization. Dropping it and z-scoring the rest is essential.

2. Per-dimension z-scoring. After extraction, each of the 112 dimensions is standardized across the training set. This ensures no single feature dominates the projection.

3. Random projection to hypervector. A random bipolar matrix R โˆˆ {-1, +1}^{16384 ร— 112} projects the feature vector to a 16384-bit bipolar hypervector. Sign of the projection becomes the bit.

4. Binding + bundling. Each training clip's hypervector is bound to its category's codebook hypervector (element-wise product), then all pairs are summed. The result is a single 16384-bit memory. Retrieval: unbind with the query, clean up against the category codebook.

Evaluation

Dataset: FSDD. Free Spoken Digit Dataset. 300 clips, 10 categories (digits 0โ€“9), 30 clips per digit. 240 training, 60 test (80/20 split).

Metric Value Chance
Prototype classification 0.967 0.100
Cross-modal retrieval 0.917 0.100
Text-to-audio retrieval 1.000 0.017

The gap between prototype (0.967) and cross-modal (0.917) is 3 clips out of 60. On FSDD, the cross-modal architecture reaches the prototype ceiling.

Dataset: ESC-50. 500 clips, 50 environmental sound categories. Not competitive. The same model reaches 0.230 prototype accuracy versus chance of 0.020. This is 11ร— chance but far below the 0.85+ that pretrained audio models achieve. ESC-50 is not included in the model card claims.

Intended Use

  • Cross-modal retrieval. Find audio clips matching a text label, or find labels matching audio.
  • Small-footprint audio tagging. Fixed 250 KB model, CPU-only inference.
  • On-device audio classification. Any device with NumPy can run the model in under 1 ms per query.
  • Few-shot adaptation. Training on a new label set takes under 1 second.

Limitations

  • FSDD-scale only. Works well on 10 acoustically distinct classes. Fails on 50 overlapping environmental classes (see ESC-50 result).
  • Feature-bound. The 112-dim feature vector determines the ceiling. MFCC on ESC-50 saturates at 0.23 prototype. Any task harder than spoken digits requires a pretrained audio embedding.
  • No semantics. The model does not understand audio content. It computes similarity in a fixed feature space.
  • Random projection. The projection matrix is not learned. The accuracy depends on the feature extractor alone.
  • Not a replacement for Whisper or PANNs. If you have access to a pretrained audio model, use it. This is for environments where a 250 MB model is impossible.

Comparison to Other Solutions

Approach Size Latency FSDD accuracy Dependencies
hv-multimodal 250 KB < 1 ms 0.967 NumPy only
Whisper-tiny (39M params) 151 MB ~500 ms ~0.99 PyTorch
PANNs CNN14 (81M params) 327 MB ~50 ms ~0.99 PyTorch
all-MiniLM-L6-v2 audio head 90 MB ~10 ms ~0.95 PyTorch
MFCC + logistic regression ~1 MB ~10 ms ~0.90 scikit-learn

The gap to pretrained models is small on FSDD but large on ESC-50. Use the model when size or dependency constraints rule out the alternatives.

How to Use

from hv_multimodal import HVMultimodal

model = HVMultimodal.load("hv-multimodal-fsdd-16384")

# Audio -> category
category = model.classify("path/to/recording.wav")
# -> "7"

# Category -> retrieve matching audio
top_indices = model.retrieve_audio("7", k=3)

# Cross-modal retrieval in both directions
scores = model.score_audio_against_categories("path/to/recording.wav")
# -> {'0': 0.02, '1': 0.01, ..., '7': 0.91, ...}
Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Evaluation results

  • Prototype Classification on FSDD (Free Spoken Digit Dataset)
    self-reported
    0.967
  • Cross-Modal Retrieval on FSDD (Free Spoken Digit Dataset)
    self-reported
    0.917
  • Text-to-Audio Retrieval on FSDD (Free Spoken Digit Dataset)
    self-reported
    1.000
  • Random Baseline on FSDD (Free Spoken Digit Dataset)
    self-reported
    0.100