MBZUAI/LLaVA-Phi-3-mini-4k-instruct-pretrain
Text Generation • Updated • 19 • 3
Natural Language Processing, Machine Learning, and Computer Vision
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Training-Free Speech-Centric Omni Understanding with Frozen VLMs