Instructions to use espnet/voxcelebs12_ebranchformer_base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ESPnet
How to use espnet/voxcelebs12_ebranchformer_base with ESPnet:
unknown model type (must be text-to-speech or automatic-speech-recognition)
- Notebooks
- Google Colab
- Kaggle
Usage
import librosa
from espnet2.bin.spk_inference import Speech2Embedding
speech2embedding = Speech2Embedding.from_pretrained(model_tag="espnet/voxcelebs12_ebranchformer_base")
# the speaker models are trained on 16 kHz mono; librosa gives that from
# whatever the file holds
speech, rate = librosa.load("audio.wav", sr=16000, mono=True)
embedding = speech2embedding(speech) # (1, embedding_dim)
RESULTS
Environments
date: 2024-11-11 13:27:22.449116
- python version: 3.10.13 (main, Sep 11 2023, 13:44:35) [GCC 11.2.0]
- espnet version: 202409
- pytorch version: 2.2.2+rocm5.6
| Mean | Std | |
|---|---|---|
| Target | 8.1720 | 3.9750 |
| Non-target | 2.0805 | 2.0805 |
| Model name | EER(%) | minDCF |
|---|---|---|
| conf/tuning/train_ebranchformer_aggregate_sys1_D384_L6_Rel_LayerDrop | 0.638 | 0.04797 |
- Downloads last month
- 10