MA-LMM
by boheumd
(2024CVPR) MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
AI summary
Video understanding model
This project develops an AI model for long-term video understanding
- stars
- 254
- forks
- 27
- watching
- 4
Similar projects
Found by comparing what the projects do, not just their names.
LLM
An open-source implementation of a vision-language instructed large language model
LM Benchmark
A benchmark for evaluating large language models in multiple languages and formats
Video LLM
A video Large Language Model designed for fine-grained comprehension and localization in videos with a custom Temporal Perception Module for improved temporal modeling
Multimodal LLM
A multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation
Long context transfer
An open-source project that enables the transfer of language understanding to vision capabilities through long context processing.
Video Moment LLM
A PyTorch-based Video LLM designed to understand and reason about video moments in terms of time boundaries.
Multimodal LLM
An LLaMA-based multimodal language model with various instruction-following and multimodal variants.
Multimodal LLM
An implementation of a multimodal language model with capabilities for comprehension and generation
VQA model
A multimodal LLM designed to handle text-rich visual questions
Image understanding model
A large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.
Model evaluation toolkit
Tools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
Identity-aware video model
An open-source project that aims to improve large vision-language models by integrating identity-aware capabilities and utilizing visual instruction tuning data for movie understanding
Morphology-aware LMs
This project develops language models that incorporate morphological knowledge to improve their understanding of linguistic structures and relationships.
Video understanding tester
A tool to evaluate video language models' ability to understand and describe video content