VTimeLLM
by huangb23
[CVPR'2024 Highlight] Official PyTorch implementation of the paper "VTimeLLM: Empower LLM to Grasp Video Moments".
AI summary
Video Moment LLM
A PyTorch-based Video LLM designed to understand and reason about video moments in terms of time boundaries.
- stars
- 231
- forks
- 11
- watching
- 2
Similar projects
Found by comparing what the projects do, not just their names.
Video LLM
A video Large Language Model designed for fine-grained comprehension and localization in videos with a custom Temporal Perception Module for improved temporal modeling
Video understanding model
This project develops an AI model for long-term video understanding
LLM trainer
Transfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs
Video-language model
An efficient framework for end-to-end learning on image-text and video-text tasks
Video understanding tester
A tool to evaluate video language models' ability to understand and describe video content
Text-Video Embedding Toolkit
Provides code and tools for learning joint text-video embeddings using the HowTo100M dataset
Visual search framework
PyTorch implementation of guided visual search mechanism for multimodal LLMs
ConvLSTM module
An implementation of a spatio-temporal convolutional LSTM module for video autoencoders with differentiable memory
LLM trainer
A PyTorch-based framework for training large language models in parallel on multiple devices
Video processor
An audio-visual language model designed to advance spatial-temporal modeling and audio understanding in video processing.
Machine Learning Toolkit
A collection of miscellaneous PyTorch implementations covering various machine learning concepts and techniques
Video text retrieval model
An open-source implementation of the Mixture-of-Embeddings-Experts model in Pytorch for video-text retrieval tasks.
LLM
An open-source implementation of a vision-language instructed large language model
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
VQA system
PyTorch implementation of video question answering system based on TVQA dataset