MA-LMM

Video understanding model

This project develops an AI model for long-term video understanding

(2024CVPR) MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

GitHub

254 stars
4 watching
27 forks
Language: Python
last commit: about 2 years ago
llmvideo-understanding

Related projects:

RepositoryDescriptionStars
luogen1996/lavinAn open-source implementation of a vision-language instructed large language model513
damo-nlp-sg/m3examA benchmark for evaluating large language models in multiple languages and formats93
dcdmllm/momentorA video Large Language Model designed for fine-grained comprehension and localization in videos with a custom Temporal Perception Module for improved temporal modeling58
lyuchenyang/macaw-llmA multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation1,568
evolvinglmms-lab/longvaAn open-source project that enables the transfer of language understanding to vision capabilities through long context processing.347
huangb23/vtimellmA PyTorch-based Video LLM designed to understand and reason about video moments in terms of time boundaries.231
alpha-vllm/wemix-llmAn LLaMA-based multimodal language model with various instruction-following and multimodal variants.17
ailab-cvc/seedAn implementation of a multimodal language model with capabilities for comprehension and generation585
mlpc-ucsd/blivaA multimodal LLM designed to handle text-rich visual questions270
360cvgroup/360vlA large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.32
evolvinglmms-lab/lmms-evalTools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance2,164
dvlab-research/llama-vidAn image-based language model that uses large language models to generate visual and text features from videos748
jiyt17/ida-vlmAn open-source project that aims to improve large vision-language models by integrating identity-aware capabilities and utilizing visual instruction tuning data for movie understanding26
ldmt-muri/morpholmThis project develops language models that incorporate morphological knowledge to improve their understanding of linguistic structures and relationships.3
llyx97/tempcompassA tool to evaluate video language models' ability to understand and describe video content91