Ask_Attend_and_Answer

Visual QA Model

Develops a deep learning model to answer questions about visual scenes based on spatial attention and question guidance

Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering

GitHub

25 stars
4 watching
11 forks
Language: C++
last commit: almost 6 years ago

Related projects:

RepositoryDescriptionStars
jiasenlu/hiecoattenvqaA framework for training Hierarchical Co-Attention models for Visual Question Answering using preprocessed data and a specific image model.349
jnhwkim/nips-mrn-vqaThis project presents a neural network model designed to answer visual questions by combining question and image features in a residual learning framework.39
zcyang/imageqa-sanThis project provides code for training image question answering models using stacked attention networks and convolutional neural networks.108
gt-vision-lab/vqa_lstm_cnnA Visual Question Answering model using a deeper LSTM and normalized CNN architecture.377
akosiorek/attend_infer_repeatAn implementation of Attend, Infer, Repeat, a method for fast scene understanding using generative models.82
huggingface/node-question-answeringProvides a simple way to perform question answering using a pre-trained model in Node.js466
yunjey/show-attend-and-tellGenerates captions for images using an attention-based neural network907
cadene/vqa.pytorchA PyTorch implementation of visual question answering with multimodal representation learning718
jazzsaxmafia/show_attend_and_tell.tensorflowA TensorFlow implementation of a neural caption generator using attention mechanisms.506
rowanz/r2cAn open-source project providing PyTorch code and data for a deep learning model that enables visual commonsense reasoning.466
davidmascharka/tbd-netsAn open-source implementation of a deep learning model designed to improve the balance between performance and interpretability in visual reasoning tasks.348
jiasenlu/adaptiveattentionAdaptive attention mechanism for image captioning using visual sentinels335
hyeonwoonoh/vqa-transfer-externaldataTools and scripts for training and evaluating a visual question answering model using transfer learning from an external data source.20
allenai/document-qaTools and codebase for training neural question answering models on multiple paragraphs of text data435
deepseek-ai/deepseek-vlA multimodal AI model that enables real-world vision-language understanding applications2,145