Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation Paper • 2610.05608 • Published 4 days ago • 143
From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation Paper • 2610.02179 • Published 7 days ago • 20
Scaling and Distilling Text Embeddings for Better Diffusibility Paper • 2610.01016 • Published 7 days ago • 57
Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces Paper • 2609.40362 • Published 8 days ago • 37
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 9 days ago • 101
DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence Paper • 2609.39222 • Published 8 days ago • 42
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 8 days ago • 293
AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation Paper • 2609.35530 • Published 10 days ago • 21
Chinese-Jev: Bringing System One Model to Chinese-Language Tasks Paper • 2609.36965 • Published 9 days ago • 25
HiRAE: Hierarchical Representation Autoencoding with Residual Budgets Paper • 2609.37775 • Published 9 days ago • 25
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 9 days ago • 102
SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation Paper • 2609.33399 • Published 11 days ago • 19
FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching Paper • 2609.35673 • Published 10 days ago • 34
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 10 days ago • 49
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders Paper • 2609.31620 • Published 13 days ago • 160
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs Paper • 2609.31590 • Published 13 days ago • 13
ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation Paper • 2609.28923 • Published 14 days ago • 11