Hviske v5.3 Danish ASR
Danish speech-to-text with hviske-v5.3 Conformer model
None defined yet.
Danish speech-to-text with hviske-v5.3 Conformer model
Latent Shortcut based Co-Speech Gesture Generation
Play a game world model live - steer it frame by frame
Virtual try-on with SIFT correspondence supervision
E-commerce catalog VLM for Turkish product listings
Text-to-video world model (Evoke, 14B DiT, 3-step distilled)
Reference-to-video with ByteDance Bernini-Diffusers-v2
Uncertainty-aware world model for aerial navigation
Streaming instruction-guided video editing
Drive a video world model with a camera path from one frame
Text-to-video with Light Forcing sparse attention
World-state visual QA with BAAI Orca-4B
Unified text-to-image generation and image understanding
Extract flat albedo maps from textures with FLUX.2 Klein
Multi-modal omni-task vision model demo
Controllable image-to-video with Giga-World-1
Reference-image conditioned generation with Krea 2
Text-to-3D human motion generation
NABLA-accelerated Wan2.1 text-to-video generation
Predict robot action chunks from an image + instruction
Text/speech to spoken response + 3D talking-avatar video
Multi-modal generation with diffusion transformers
Keep identity from reference, follow lineart structure
Object and Material Selection VLM