ROLL-VideoQA
VideoQA model
A PyTorch-based model for answering questions about videos based on unseen scenes and storylines
PyTorch code for ROLL, a knowledge-based video story question answering model.
19 stars
3 watching
4 forks
Language: Python
last commit: almost 6 years agoknowledge-based-reasoningvideo-question-answeringvideo-understandingvisual-question-answering
Related projects:
| Repository | Description | Stars |
|---|---|---|
| A PyTorch implementation of visual question answering with multimodal representation learning | 718 | |
| PyTorch implementation of video question answering system based on TVQA dataset | 172 | |
| Implementing reading comprehension from Wikipedia questions to answer open-domain queries using PyTorch and SQuAD dataset | 401 | |
| Implementations and tools for training and fine-tuning a visual question answering model based on the 2017 CVPR workshop winner's approach. | 164 | |
| An efficient framework for end-to-end learning on image-text and video-text tasks | 709 | |
| This project explores question-answering in movies using various machine learning approaches. | 80 | |
| A software framework for training and deploying multimodal visual question answering models using compact bilinear pooling. | 222 | |
| An implementation of a VQA system using bottom-up attention, aiming to improve the efficiency and speed of visual question answering tasks. | 755 | |
| A Visual Question Answering model using a deeper LSTM and normalized CNN architecture. | 377 | |
| An implementation of a two-stage framework designed to prompt large language models with answer heuristics for knowledge-based visual question answering tasks. | 270 | |
| A video conversation model that generates meaningful conversations about videos using large vision and language models | 1,246 | |
| This project provides code for training image question answering models using stacked attention networks and convolutional neural networks. | 108 | |
| A medical visual question-answering dataset and toolkit for training models to understand medical images and instructions. | 180 | |
| An open-source implementation of a deep learning model for video deblurring and motion estimation. | 114 | |
| A PyTorch-based simulator for quantum machine learning | 45 |