SPEAR: A Simulator for Photorealistic Embodied AI Research Paper • 2607.06701 • Published 13 days ago • 6
MAOAM: Unified Object and Material Selection with Vision-Language Models Paper • 2606.04880 • Published Jun 2 • 10
MotiMotion: Motion-Controlled Video Generation with Visual Reasoning Paper • 2605.22818 • Published May 21 • 5
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization Paper • 2605.19330 • Published May 19 • 9
What matters for Representation Alignment: Global Information or Spatial Structure? Paper • 2512.10794 • Published Dec 11, 2025 • 11
TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets Paper • 2603.27520 • Published Mar 29 • 4
TrajectoryMover: Generative Movement of Object Trajectories in Videos Paper • 2603.29092 • Published Mar 31 • 3
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion Paper • 2603.15614 • Published Mar 16 • 6
Self-Evaluation Unlocks Any-Step Text-to-Image Generation Paper • 2512.22374 • Published Dec 26, 2025 • 17
Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Paper • 2512.14008 • Published Dec 16, 2025 • 10
view post Post 32721 Want to iterate on a Hugging Face Space with an LLM? Now you can easily convert any HF entire repo (Model, Dataset or Space) to a text file and feed it to a language model! multimodalart/repo2txt See translation 3 replies · 🤗 3 3 👍 3 3 🚀 2 2 🧠 1 1 + Reply
view post Post 18506 Self-Forcing - a real-time video distilled model from Wan 2.1 by @adobe is out, and they open sourced it 🐐I've built a live real time demo on Spaces 📹💨 multimodalart/self-forcing See translation 6 replies · ❤️ 12 12 🔥 6 6 + Reply
Vec2Face: Scaling Face Dataset Generation with Loosely Constrained Vectors Paper • 2409.02979 • Published Sep 4, 2024 • 1
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers Paper • 2504.10483 • Published Apr 14, 2025 • 22
Negative Token Merging: Image-based Adversarial Feature Guidance Paper • 2412.01339 • Published Dec 2, 2024 • 22
IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion Models Paper • 2311.06792 • Published Nov 12, 2023
OpenDevin: An Open Platform for AI Software Developers as Generalist Agents Paper • 2407.16741 • Published Jul 23, 2024 • 83
view post Post 35670 New feature 🔥 Image models and LoRAs now have little previews 🤏If you don't know where to start to find them, I invite you to browse cool LoRAs in the profile of some amazing fine-tuners: @artificialguybr , @alvdansen , @DoctorDiffusion , @e-n-v-y , @KappaNeuro @ostris 3 replies · ❤️ 13 13 🚀 1 1 🤗 1 1 + Reply