MCG-NJU/DMM
Text-to-Image • 0.9B • Updated • 59 • 16
Computer Vision; Video Understanding; Action Recognition
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding