AdaptiveAttention
Image Captioning Model
Adaptive attention mechanism for image captioning using visual sentinels
Implementation of "Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning"
335 stars
13 watching
74 forks
Language: Jupyter Notebook
last commit: over 8 years agoLinked from 2 awesome lists
attention-mechanismimage-captioningtorch
Related projects:
| Repository | Description | Stars |
|---|---|---|
| A deep learning model for generating image captions with semantic attention | 51 | |
| Generates captions for images using an attention-based neural network | 907 | |
| Enhances language models to generate text based on visual descriptions of images | 352 | |
| A TensorFlow implementation of a neural caption generator using attention mechanisms. | 506 | |
| Trains a bottom-up attention model using Faster R-CNN and Visual Genome annotations for image captioning and VQA tasks | 1,438 | |
| This project proposes a novel method for calibrating attention distributions in multimodal models to improve contextualized representations of image-text pairs. | 30 | |
| This implementation allows users to generate captions from images using a neural network model with visual attention. | 790 | |
| A project for pre-training models to support image captioning and question answering tasks. | 416 | |
| Generates Apple Watch activity indicator images for animation | 529 | |
| Trains image paragraph captioning models to generate diverse and accurate captions | 90 | |
| An image caption generation model that uses abstract scene graphs to fine-grained control and generate captions | 200 | |
| An unsupervised image captioning framework that allows generating captions from images without paired data. | 215 | |
| A framework for training Hierarchical Co-Attention models for Visual Question Answering using preprocessed data and a specific image model. | 349 | |
| PyTorch implementation of video captioning, combining deep learning and computer vision techniques. | 402 | |
| A PyTorch implementation of image captioning models via scene graph decomposition. | 96 |