Awesome Lists

LLaMA-VID

by dvlab-research

Pythonpushed about 2 years ago

LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)

AI summary

Video image processor

An image-based language model that uses large language models to generate visual and text features from videos

stars
748
forks
45
watching
14

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.