Audio-Text-to-Text
Transformers
Safetensors
English
Chinese
moss_transcribe_diarize
text-generation
moss
audio
speech
asr
diarization
timestamp-asr
long-form-audio
multimodal
multilingual
custom_code
Eval Results
Instructions to use OpenMOSS-Team/MOSS-Transcribe-Diarize with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OpenMOSS-Team/MOSS-Transcribe-Diarize with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("OpenMOSS-Team/MOSS-Transcribe-Diarize", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Model enters a repetitive generation loop on music-only audio until the token limit is reached
#27
by artyomboyko - opened
Using the attached audio file, which contains music only and no speech or vocals, the model enters a repetitive generation loop instead of terminating normally.
It repeatedly generates timestamp/speaker patterns such as:
[0.00][S01][1.00][1.00][S01][1.00]...
and continues until max_new_tokens is reached.
In my test, generation did not end with EOS and stopped only because the token limit was reached.
Expected behavior would be for the model to terminate without hallucinating transcript segments when no speech is present.