Instructions to use danish-foundation-models/edda-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use danish-foundation-models/edda-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="danish-foundation-models/edda-v0.1")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("danish-foundation-models/edda-v0.1") model = AutoModelForSpeechSeq2Seq.from_pretrained("danish-foundation-models/edda-v0.1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Edda v0.1 — Danish speech recognition
Edda is openai/whisper-large-v3-turbo (809 M parameters, 4-layer
decoder) fully fine-tuned on 2,600 hours of transcribed Danish speech. On the
open Danish ASR leaderboard harness it scores a mean WER of
9.20 over the five test sets, the best published result at the time of release, at 18.5x real time on a single AMD MI250X GCD.
| test set | WER | CER |
|---|---|---|
| CoRal-v3 conversation | 15.99 | 9.08 |
| CoRal-v3 read-aloud | 10.52 | 4.12 |
| Common Voice Danish (leaderboard set, 2,756 clips) | 5.99 | 1.92 |
| FLEURS da_dk | 7.56 | 2.95 |
| FTSpeech | 5.95 | 3.28 |
| mean | 9.20 | 4.27 |
Scores from the leaderboard's own harness (Rye-A1/danish-asr-leaderboard.
Usage
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="danish-foundation-models/edda-v0.1", device="cuda",
chunk_length_s=30) # clips longer than 30 s are transcribed in 30 s windows
print(asr("clip.wav")["text"])
The model's generation config defaults to Danish transcription (language="da", task="transcribe"), so no prompt tokens
are needed. Recordings longer than 30 s are not truncated: with chunk_length_s=30 they are transcribed in consecutive 30 s
windows and joined, which is also how the leaderboard scores long clips. Any sample rate the pipeline can read is accepted; audio is resampled to 16 kHz. Output is cased and punctuated
in the style of the training corpora (see Transcripts below).
Training data
Six public corpora, training splits only, gold transcripts only (no pseudo-labels), 1.62 M clips / 2,600 h after filtering:
| corpus | source | clips | hours | share of training steps |
|---|---|---|---|---|
| FTSpeech (parliament) | alexandrainst/ftspeech train |
983,989 | 1,690 | 43 % |
| CoRal-v3 read-aloud | CoRal-project/coral-v3 read_aloud/train |
299,253 | 521 | 24 % |
| NST Danish | alexandrainst/nst-da train¹ |
178,923 | 234 | 16 % |
| CoRal-v3 conversation | CoRal-project/coral-v3 conversation/train |
147,184 | 144 | 12 % |
| FLEURS da_dk | google/fleurs train |
2,463 | 7.5 | 3 % |
| Common Voice 17 Danish | mozilla-foundation/common_voice_17_0 train |
3,484 | 4.1 | 2 % |
¹ NST's own test split was held out even though it is not a benchmark set.
Corpus mixing. Clips are drawn corpus-first with P(corpus) ∝ hours^0.5, then uniformly within the corpus, with replacement and no epoch structure. Over 30,000 steps × 256 clips this is ≈3 passes over FTSpeech, ≈6 over CoRal and NST, and ≈45–90 over Common Voice and FLEURS.
Audio preprocessing. Decoded, mixed to mono, resampled to 16 kHz, stored as 16-bit FLAC. Dropped: clips shorter than 0.1 s,
clips longer than 30 s were dropped from training (11,688 FTSpeech clips — Whisper's fixed window would cut the audio but
not the transcript, so keeping them would teach words that are never heard; at inference long audio is chunked instead, see
Usage), empty transcripts,
and CoRal read-aloud recordings marked rejected by the dataset's validators. No level normalisation (a peak-normalised
variant was trained and scored within 0.03 WER of this recipe).
Transcripts. Used verbatim as distributed, whitespace-stripped only; no case folding, punctuation or number normalisation at training time, so the model reproduces each corpus's conventions (FTSpeech is lower-case without punctuation; CoRal, FLEURS, NST and Common Voice are cased and mostly punctuated). These conventions do not affect the reported WER: the evaluation normalises reference and hypothesis identically before scoring (Unicode NFKC, lower-casing, punctuation removed, numerals written as words, hesitation fillers such as øh removed, whitespace collapsed), so only genuine recognition differences count.
Training recipe
Full fine-tune of encoder and decoder from openai/whisper-large-v3-turbo, unmodified architecture (learned absolute
positions, fixed 30 s windows).
| steps / batch | 30,000 × 256 clips (8 nodes × 8 MI250X GCDs × 4 clips), 7 h 23 m |
| optimizer | AdamW, β (0.9, 0.98), weight decay 0.01, gradient clipping 1.0, bf16 |
| schedule | warmup 500 → 1e-5 flat to step 24,000 → linear decay to 1e-6 at 30,000 (WSD) |
| loss | cross-entropy with label smoothing 0.05 on `< |
| EMA | 0.9998; the released weights are the EMA at step 30,000, not a validation-selected checkpoint |
| SpecAugment | time masks p 0.05 × 10 frames, feature masks p 0.05 × 10 bins |
| audio augmentation (per clip) | speed 0.9–1.1 (p 0.5); additive coloured noise at 0–20 dB SNR (p 0.6); 0.5–2 s trailing silence (p 0.25); 2.5 % of clips replaced by low-level noise with an empty transcript |
Limitations
- Transcription convention follows the corpora. FTSpeech references are the edited parliamentary record, so on parliamentary speech the model tends to omit restarts and repetitions; CoRal references are verbatim.
- Long audio is handled by fixed 30 s windows with no explicit loop guard; on very long silences or music the decoder can occasionally repeat a phrase. Split at pauses for best results.
- Domains not represented in training (children, strong non-native accents, telephone-band audio, singing) are untested.
- Evaluation references have known noise: a small fraction of CoRal read-aloud test prompts were paraphrased by the speaker, and CoRal conversation references systematically omit the copula er; both count against every model equally.
License and attribution
The model weights are released under the Apache License 2.0 by the Alexandra Institute, which is also the licensor of the
CoRal-v3 dataset. The base model openai/whisper-large-v3-turbo
is MIT-licensed; the training corpora carry their own licenses — see the respective dataset cards. Trained by the Alexandra
Institute within the CoRal project and released by Danish Foundation Models.
- Downloads last month
- 477
Model tree for danish-foundation-models/edda-v0.1
Base model
openai/whisper-large-v3Datasets used to train danish-foundation-models/edda-v0.1
CoRal-project/coral-v3
alexandrainst/ftspeech
Spaces using danish-foundation-models/edda-v0.1 2
Evaluation results
- WER on CoRal-v3 conversation (test)test set self-reported15.990
- CER on CoRal-v3 conversation (test)test set self-reported9.080
- WER on CoRal-v3 read-aloud (test)test set self-reported10.520
- CER on CoRal-v3 read-aloud (test)test set self-reported4.120
- WER on Common Voice Danish (leaderboard settest set self-reported5.990
- CER on Common Voice Danish (leaderboard settest set self-reported1.920
- WER on FLEURS da_dk (test)test set self-reported7.560
- CER on FLEURS da_dk (test)test set self-reported2.950