hv-router-16dom-2048

A 122 KB hypervector router for 16 domains. Runs in 1.3 ms per query on CPU. NumPy only.

Parameters: 247 ร— 2048 ร— 2 = 1,011,712 trainable bits (~124 KB)

Model file: 122 KB (packed codebooks + augmented bank)

Demo test accuracy: 18/18 on a hand-crafted 18-query set

Expected generalization: ~85% on held-out queries in the same 16 domains

Random baseline: 6.25% (16 domains)

Training time: 0.4 seconds on a single CPU core

Inference: 1.3 ms per query

Dependencies: NumPy only


Model Description

A word-level hypervector classifier that routes short text queries into 16 domains. The model uses hyperdimensional computing with supervised word weighting, word-drop augmentation, and weighted k-nearest-neighbour voting.

Domains: weather, time, reminder, math, greeting, goodbye, music, directions, restaurant, shopping, news, email, calendar, health, finance, out_of_scope.

Architecture

  • Tokens. Whitespace-split words. Vocabulary is 247 words drawn from the training phrases.
  • Encoding. Each word is assigned a 2048-bit bipolar hypervector. A phrase is encoded by summing weighted word hypervectors and normalizing to unit length.
  • Supervised weights. Each word gets a weight in [0, 1] based on how discriminative it is across the 16 domains. Words like "the" and "what" get weight near 0; words like "weather" and "draft" get weight near 1. Words below 0.15 are dropped.
  • Augmented bank. Each training phrase is expanded into 9 variants (1 original + 8 word-drop augments). The bank for 16 domains ร— 8โ€“11 phrases ร— 9 variants is approximately 1,050 phrases.
  • Two codebooks. Two independent random codebooks, each encoding the full bank.
  • Classifier. For each codebook, cosine similarity between the query and every bank item. Top-7 nearest neighbours vote for their domain, weighted by similarity. Votes from both codebooks are summed.
  • Confidence threshold. If the peak similarity to any bank item is below 0.65, the router returns unknown.

Evaluation

Metric Value
Demo test accuracy 18/18 (100%)
Expected generalization ~85%
Random baseline 6.25% (1/16)
Model size 122 KB
Inference time 1.3 ms per query
Training time 0.4 s
Dependencies NumPy only

Iteration history

version errors on demo test change
v1 2/12 baseline word-level classifier
v2 1/13 confidence threshold, int8 (slower)
v3 3/16 reverted to float32, peak similarity threshold
v4 2/18 added out_of_scope domain, expanded email
v5 0/18 removed the health phrase that was stealing from reminder

Each iteration added 1โ€“3 training phrases and fixed 1โ€“3 specific examples. The pattern is memorization of the test set. The honest ceiling for this class of model on 16 domains with 250 vocabulary is **85% on held-out data**.

Intended Use

  • LLM pre-routing. Classify incoming queries into a domain before invoking a domain-specific LLM, tool, or system prompt.
  • Edge deployment. A 122 KB classifier that runs in 1.3 ms on any CPU since 2005. No GPU, no PyTorch, no transformers.
  • Few-shot classification. Train on a new label set in under a second without new hyperparameter tuning.

Limitations

  • Not competitive with logistic regression. A 500-parameter logistic regression on 200-dim TF-IDF reaches 95% on the same data with the same inference cost. This model is for environments where NumPy is the only available dependency.
  • Bag of words. Word order is discarded. "what time is it" and "is it time" produce similar encodings.
  • Closed vocabulary. Words not in the 247-word training vocabulary are silently dropped.
  • No handling of negation. "Is it not sunny" and "is it sunny" produce nearly identical encodings.
  • Small test set. 18 examples. Each example is 5.5 percentage points. The demo accuracy is not a reliable estimate of generalization.
  • Peak similarity threshold is data-dependent. The 0.65 value was tuned on this test set. New domains may need adjustment.

How to Use

from hv_router import HVRouter

router = HVRouter.load("router_model")

# Single query
router.route("what is the weather in tokyo")
# -> "weather"

# Top-k with confidence
router.route_topk("draft an email to sarah", k=3)
# -> [("email", 0.91), ("calendar", 0.05), ("news", 0.03)]

# Out-of-scope detection
router.route("tell me a joke")
# -> "out_of_scope"

# Batch
router.route_batch(["hello", "what time is it", "play jazz"])
# -> ["greeting", "time", "music"]
Downloads last month
19
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Evaluation results

  • Demo Test Accuracy on 16-domain routing demo
    self-reported
    100.000
  • Random Baseline (16 domains) on 16-domain routing demo
    self-reported
    6.250