Momentor

Video LLM

A video Large Language Model designed for fine-grained comprehension and localization in videos with a custom Temporal Perception Module for improved temporal modeling

GitHub

58 stars
6 watching
2 forks
Language: Python
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
huangb23/vtimellmA PyTorch-based Video LLM designed to understand and reason about video moments in terms of time boundaries.231
boheumd/ma-lmmThis project develops an AI model for long-term video understanding254
damo-nlp-sg/videollama2An audio-visual language model designed to advance spatial-temporal modeling and audio understanding in video processing.957
damo-nlp-mt/polylmA polyglot large language model designed to address limitations in current LLM research and provide better multilingual instruction-following capability.77
lyuchenyang/macaw-llmA multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation1,568
llyx97/tempcompassA tool to evaluate video language models' ability to understand and describe video content91
umass-foundation-model/3d-llmDeveloping a Large Language Model capable of processing 3D representations as inputs979
dvlab-research/llama-vidAn image-based language model that uses large language models to generate visual and text features from videos748
victordibia/llmxAn API that provides a unified interface to multiple large language models for chat fine-tuning79
luogen1996/lavinAn open-source implementation of a vision-language instructed large language model513
dcdmllm/cheetahA large language model designed to understand and generate instructions with accompanying visual content360
phellonchen/x-llmA framework that enables large language models to process and understand multimodal inputs from various sources such as images and speech.308
internlm/tutorialA tutorial project for exploring large language models and their applications in natural language processing tasks.1,593
mbzuai-oryx/groundinglmmAn end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations797
danieljf24/dual_encodingA deep learning project that provides a video-text retrieval model and tools for training and evaluating it on the MSR-VTT dataset154