SpeakoFlow Mini
Dictation cleanup for transcribed speech (0.8B GGUF)
None defined yet.
Dictation cleanup for transcribed speech (0.8B GGUF)
Verbatim & intended ASR with word timestamps
Zero-shot voice cloning TTS with 120M params
Recurrent feed-forward 3D Gaussian Splatting
Thai document/handwriting OCR with OpenThai 2.0 27B VLM
Hebrew text to IPA phonemes for TTS
Multilingual text-to-speech with zero-shot voice cloning
Edit one frame, ripple the change through the whole video
4-step MiniMax-H3 with sparse attention β video + audio
Convert table images to LaTeX code with CSPO
Voice cloning, voice design, speech editing and ASR
Long-term planar tracking with SAM 2 homographies
Surgical spatio-temporal grounding with RefineRank
Browser-GUI screenshot to structured observation text
Detect PII in Korean text with a 34M param model
Anime image generation & editing via distillation
GUI grounding β locate UI elements from screenshots
torch.export + AOTInductor compiler for Zing-0.5
Predict robot action chunks with GigaBrain-0.7 3.5B
Rewrite prompts into structured MiniMax-H3 video prompts
4-step text-to-video+audio with MiniMax-H3 FlashGen LoRA
Edit human-object interactions in images with OneHOI.
Generate extensible underwater 3D scenes with flow matching
Few-shot seismic facies segmentation via GP regression