momentslab/AstroCaptions
Viewer • Updated • 44.1k • 28 • 4
video understanding, vision language models, information retrieval
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
PEEK: Picking Essential frames via Efficient Knowledge distillation