LongVU
by Vision-CAIR
AI summary
Video describer
An artificial intelligence system designed to understand and describe long-form video content
- stars
- 329
- forks
- 22
- watching
- 5
Similar projects
Found by comparing what the projects do, not just their names.
Text generator
Enables automatic generation of descriptive text from images and videos based on user input.
Visual encoder
A vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.
Video search system
A deep learning-based video search system using pre-trained models and datasets
Video annotator
Tools for efficiently scaling up video annotation using crowdsourced marketplaces.
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
gabeur/mmt259
Video retriever
Develops a cross-modal architecture for video retrieval by combining multiple types of features from videos and text
Image encoder
An implementation of a vision transformer architecture designed for high-resolution image encoding with multiple efficient attention mechanisms
Video generator library
A comprehensive toolkit for high-performance video generation and processing
Video Assistant
An offline video assistant system powered by large language models and computer vision techniques.
Video player
A video player that uses VLC as the decoder and renders QML components on OpenGL textures.
Video processor
A web-based video processing tool that uses AI to facilitate cultural and linguistic tasks such as transcription, translation, and audio synthesis.
Video understanding optimizer
Develops a method for long video understanding by optimizing memory usage
Video preprocessor
Tools for preprocessing videos for various datasets, including video cropping and annotation.
Video captioner
PyTorch implementation of video captioning, combining deep learning and computer vision techniques.
Video processor
An audio-visual language model designed to advance spatial-temporal modeling and audio understanding in video processing.