Edda v0.1 — Danish speech recognition

Demo

Edda is openai/whisper-large-v3-turbo (809 M parameters, 4-layer decoder) fully fine-tuned on 2,600 hours of transcribed Danish speech. On the open Danish ASR leaderboard harness it scores a mean WER of 9.20 over the five test sets, the best published result at the time of release, at 18.5x real time on a single AMD MI250X GCD.

test set WER CER
CoRal-v3 conversation 15.99 9.08
CoRal-v3 read-aloud 10.52 4.12
Common Voice Danish (leaderboard set, 2,756 clips) 5.99 1.92
FLEURS da_dk 7.56 2.95
FTSpeech 5.95 3.28
mean 9.20 4.27

Scores from the leaderboard's own harness (Rye-A1/danish-asr-leaderboard.

Usage

from transformers import pipeline

asr = pipeline("automatic-speech-recognition", model="danish-foundation-models/edda-v0.1", device="cuda",
               chunk_length_s=30)                         # clips longer than 30 s are transcribed in 30 s windows
print(asr("clip.wav")["text"])

The model's generation config defaults to Danish transcription (language="da", task="transcribe"), so no prompt tokens are needed. Recordings longer than 30 s are not truncated: with chunk_length_s=30 they are transcribed in consecutive 30 s windows and joined, which is also how the leaderboard scores long clips. Any sample rate the pipeline can read is accepted; audio is resampled to 16 kHz. Output is cased and punctuated in the style of the training corpora (see Transcripts below).

Training data

Six public corpora, training splits only, gold transcripts only (no pseudo-labels), 1.62 M clips / 2,600 h after filtering:

corpus source clips hours share of training steps
FTSpeech (parliament) alexandrainst/ftspeech train 983,989 1,690 43 %
CoRal-v3 read-aloud CoRal-project/coral-v3 read_aloud/train 299,253 521 24 %
NST Danish alexandrainst/nst-da train¹ 178,923 234 16 %
CoRal-v3 conversation CoRal-project/coral-v3 conversation/train 147,184 144 12 %
FLEURS da_dk google/fleurs train 2,463 7.5 3 %
Common Voice 17 Danish mozilla-foundation/common_voice_17_0 train 3,484 4.1 2 %

¹ NST's own test split was held out even though it is not a benchmark set.

Corpus mixing. Clips are drawn corpus-first with P(corpus) ∝ hours^0.5, then uniformly within the corpus, with replacement and no epoch structure. Over 30,000 steps × 256 clips this is ≈3 passes over FTSpeech, ≈6 over CoRal and NST, and ≈45–90 over Common Voice and FLEURS.

Audio preprocessing. Decoded, mixed to mono, resampled to 16 kHz, stored as 16-bit FLAC. Dropped: clips shorter than 0.1 s, clips longer than 30 s were dropped from training (11,688 FTSpeech clips — Whisper's fixed window would cut the audio but not the transcript, so keeping them would teach words that are never heard; at inference long audio is chunked instead, see Usage), empty transcripts, and CoRal read-aloud recordings marked rejected by the dataset's validators. No level normalisation (a peak-normalised variant was trained and scored within 0.03 WER of this recipe).

Transcripts. Used verbatim as distributed, whitespace-stripped only; no case folding, punctuation or number normalisation at training time, so the model reproduces each corpus's conventions (FTSpeech is lower-case without punctuation; CoRal, FLEURS, NST and Common Voice are cased and mostly punctuated). These conventions do not affect the reported WER: the evaluation normalises reference and hypothesis identically before scoring (Unicode NFKC, lower-casing, punctuation removed, numerals written as words, hesitation fillers such as øh removed, whitespace collapsed), so only genuine recognition differences count.

Training recipe

Full fine-tune of encoder and decoder from openai/whisper-large-v3-turbo, unmodified architecture (learned absolute positions, fixed 30 s windows).

steps / batch 30,000 × 256 clips (8 nodes × 8 MI250X GCDs × 4 clips), 7 h 23 m
optimizer AdamW, β (0.9, 0.98), weight decay 0.01, gradient clipping 1.0, bf16
schedule warmup 500 → 1e-5 flat to step 24,000 → linear decay to 1e-6 at 30,000 (WSD)
loss cross-entropy with label smoothing 0.05 on `<
EMA 0.9998; the released weights are the EMA at step 30,000, not a validation-selected checkpoint
SpecAugment time masks p 0.05 × 10 frames, feature masks p 0.05 × 10 bins
audio augmentation (per clip) speed 0.9–1.1 (p 0.5); additive coloured noise at 0–20 dB SNR (p 0.6); 0.5–2 s trailing silence (p 0.25); 2.5 % of clips replaced by low-level noise with an empty transcript

Limitations

  • Transcription convention follows the corpora. FTSpeech references are the edited parliamentary record, so on parliamentary speech the model tends to omit restarts and repetitions; CoRal references are verbatim.
  • Long audio is handled by fixed 30 s windows with no explicit loop guard; on very long silences or music the decoder can occasionally repeat a phrase. Split at pauses for best results.
  • Domains not represented in training (children, strong non-native accents, telephone-band audio, singing) are untested.
  • Evaluation references have known noise: a small fraction of CoRal read-aloud test prompts were paraphrased by the speaker, and CoRal conversation references systematically omit the copula er; both count against every model equally.

License and attribution

The model weights are released under the Apache License 2.0 by the Alexandra Institute, which is also the licensor of the CoRal-v3 dataset. The base model openai/whisper-large-v3-turbo is MIT-licensed; the training corpora carry their own licenses — see the respective dataset cards. Trained by the Alexandra Institute within the CoRal project and released by Danish Foundation Models.

Downloads last month
477
Safetensors
Model size
0.8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for danish-foundation-models/edda-v0.1

Finetuned
(638)
this model

Datasets used to train danish-foundation-models/edda-v0.1

Spaces using danish-foundation-models/edda-v0.1 2

Evaluation results