Model enters a repetitive generation loop on music-only audio until the token limit is reached

#27
by artyomboyko - opened

Using the attached audio file, which contains music only and no speech or vocals, the model enters a repetitive generation loop instead of terminating normally.
It repeatedly generates timestamp/speaker patterns such as:
[0.00][S01][1.00][1.00][S01][1.00]...
and continues until max_new_tokens is reached.
In my test, generation did not end with EOS and stopped only because the token limit was reached.
Expected behavior would be for the model to terminate without hallucinating transcript segments when no speech is present.

Sign up or log in to comment