Momentor
by DCDmllm
AI summary
Video LLM
A video Large Language Model designed for fine-grained comprehension and localization in videos with a custom Temporal Perception Module for improved temporal modeling
- stars
- 58
- forks
- 2
- watching
- 6
Similar projects
Found by comparing what the projects do, not just their names.
Video Moment LLM
A PyTorch-based Video LLM designed to understand and reason about video moments in terms of time boundaries.
Video understanding model
This project develops an AI model for long-term video understanding
Video processor
An audio-visual language model designed to advance spatial-temporal modeling and audio understanding in video processing.
Multilingual LLM
A polyglot large language model designed to address limitations in current LLM research and provide better multilingual instruction-following capability.
Multimodal LLM
A multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation
Video understanding tester
A tool to evaluate video language models' ability to understand and describe video content
3D LLM
Developing a Large Language Model capable of processing 3D representations as inputs
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
LLM API
An API that provides a unified interface to multiple large language models for chat fine-tuning
LLM
An open-source implementation of a vision-language instructed large language model
Instruction Parser
A large language model designed to understand and generate instructions with accompanying visual content
Multimodal LLM Framework
A framework that enables large language models to process and understand multimodal inputs from various sources such as images and speech.
LLM Tutorial
A tutorial project for exploring large language models and their applications in natural language processing tasks.
Multimodal conversational model
An end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations
Video text retrieval model
A deep learning project that provides a video-text retrieval model and tools for training and evaluating it on the MSR-VTT dataset