Video-ChatGPT

Video conversational model

A video conversation model that generates meaningful conversations about videos using large vision and language models

[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

GitHub

1k stars
15 watching
110 forks
Language: Python
last commit: about 2 years ago
chatbotclipgpt-4llamallavamulit-modalvicunavideo-chatboatvideo-conversationvision-languagevision-language-pretraining

Related projects:

RepositoryDescriptionStars
mbzuai-oryx/groundinglmmAn end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations797
abbey4799/cutegptA conversational language model developed to improve understanding of complex instructions and Chinese vocabulary.62
neukg/techgpt-2.0An advanced language model designed to generate human-like responses in various domains and applications101
79e/chatgpt-webA commercially viable web application for conversational AI built with React and OpenAI's ChatGPT technology1,366
kendryte/toucan-llmA large language model with 70 billion parameters designed for chatbot and conversational AI tasks29
renshuhuai-andy/timechatA large language model designed to understand long videos by binding visual content with timestamps and producing video token sequences of varying lengths.314
zcli-charlie/batgptA large language model designed to support long context conversations with improved efficiency and effectiveness38
open-mmlab/multimodal-gptTrains a multimodal chatbot that combines visual and language instructions to generate responses1,478
nagi-ovo/crag-ollama-chatA conversational AI demo powered by a large language model78
opengvlab/multi-modality-arenaAn evaluation platform for comparing multi-modality models on visual question-answering tasks478
m1guelpf/chatgpt-discordA Discord bot that enables interactive conversations with ChatGPT using a single command.291
360cvgroup/seechatA multimodal chatbot with computer vision capabilities integrated into a single model99
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
ailab-cvc/gpt4toolsAn intelligent system that enables automatic control and utilization of visual foundation models to interact with images in conversational settings.762
showlab/vlogTransforms video content into a long document containing visual and audio information that can be used for chat or other applications.545