LongVU

Video describer

An artificial intelligence system designed to understand and describe long-form video content

GitHub

329 stars
5 watching
22 forks
Language: Python
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
vision-cair/chatcaptionerEnables automatic generation of descriptive text from images and videos based on user input.457
gordonhu608/mqt-llavaA vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.101
li-xirong/w2vvppA deep learning-based video search system using pre-trained models and datasets28
cvondrick/vaticTools for efficiently scaling up video annotation using crowdsourced marketplaces.609
dvlab-research/llama-vidAn image-based language model that uses large language models to generate visual and text features from videos748
gabeur/mmtDevelops a cross-modal architecture for video retrieval by combining multiple types of features from videos and text259
microsoft/vision-longformerAn implementation of a vision transformer architecture designed for high-resolution image encoding with multiple efficient attention mechanisms243
nus-hpc-ai-lab/videosysA comprehensive toolkit for high-performance video generation and processing1,819
rupertluo/valleyAn offline video assistant system powered by large language models and computer vision techniques.210
longwei/qmlvideoA video player that uses VLC as the decoder and renders QML components on OpenGL textures.33
liuzhao1225/youdub-webuiA web-based video processing tool that uses AI to facilitate cultural and linguistic tasks such as transcription, translation, and audio synthesis.1,980
rese1f/moviechatDevelops a method for long video understanding by optimizing memory usage550
aliaksandrsiarohin/video-preprocessingTools for preprocessing videos for various datasets, including video cropping and annotation.522
xiadingz/video-caption.pytorchPyTorch implementation of video captioning, combining deep learning and computer vision techniques.402
damo-nlp-sg/videollama2An audio-visual language model designed to advance spatial-temporal modeling and audio understanding in video processing.957