vqa-mcb
by akirafukui
AI summary
VQA model framework
A software framework for training and deploying multimodal visual question answering models using compact bilinear pooling.
- stars
- 222
- forks
- 79
- watching
- 15
Similar projects
Found by comparing what the projects do, not just their names.
VQA prompter
An implementation of a two-stage framework designed to prompt large language models with answer heuristics for knowledge-based visual question answering tasks.
Visual QA framework
A framework for training Hierarchical Co-Attention models for Visual Question Answering using preprocessed data and a specific image model.
VQA model
A PyTorch implementation of visual question answering with multimodal representation learning
VQA model
A multimodal LLM designed to handle text-rich visual questions
VQA trainer
Tools and scripts for training and evaluating a visual question answering model using transfer learning from an external data source.
VQA model
A Visual Question Answering model using a deeper LSTM and normalized CNN architecture.
Medical image understanding toolkit
A medical visual question-answering dataset and toolkit for training models to understand medical images and instructions.
Multimodal model
A large multimodal model for visual question answering, trained on a dataset of 2.1B image-text pairs and 8.2M instruction sequences.
VQA Model Trainer
Implementations and tools for training and fine-tuning a visual question answering model based on the 2017 CVPR workshop winner's approach.
VQA system
An implementation of a VQA system using bottom-up attention, aiming to improve the efficiency and speed of visual question answering tasks.
VQA challenge data
A VQA dataset with unanswerable questions designed to test the limits of large models' knowledge and reasoning abilities.
Model testing framework
An evaluation framework using Chinese high school examination questions to assess large language model capabilities
tsb0601/mmvp296
Visual model evaluation
An evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.
Visual QA Model
This project presents a neural network model designed to answer visual questions by combining question and image features in a residual learning framework.
VideoQA model
A PyTorch-based model for answering questions about videos based on unseen scenes and storylines