awesome-image-captioning

Image Captioning Resources

Curated list of resources on image captioning and related areas, including papers, datasets, and implementations.

A curated list of image captioning and related area resources. :-)

GitHub

1k stars
40 watching
184 forks
last commit: over 3 years ago
Linked from 1 awesome list


Awesome Image Captioning / Change Log

here56about 4 years agoMay 25 An up-to-date paper list about vision-and-language pre-training is available

Awesome Image Captioning / Papers / Survey

A Comprehensive Survey of Deep Learning for Image CaptioningHossain M et al,

Awesome Image Captioning / Papers / Before

I2t: Image parsing to text descriptionYao B Z et al,
Im2Text: Describing Images Using 1 Million Captioned PhotographsOrdonez V et al,
Deep Captioning with Multimodal Recurrent Neural NetworksMao J et al,

Awesome Image Captioning / Papers / 2015

Show and Tell: A Neural Image Caption GeneratorVinyals O et al,
Deep Visual-Semantic Alignments for Generating Image DescriptionsKarpathy A et al,
Mind’s Eye: A Recurrent Visual Representation for Image Caption GenerationChen X et al,
Long-term Recurrent Convolutional Networks for Visual Recognition and DescriptionDonahue J et al,
Guiding the Long-Short Term Memory Model for Image Caption GenerationJia X et al,
Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of ImagesMao J et al,
Expressing an Image Stream with a Sequence of Natural SentencesPark C C et al,
Show, Attend and Tell: Neural Image Caption Generation with Visual AttentionXu K et al,
Order-Embeddings of Images and LanguageVendrov I et al,
Generating Images from Captions with AttentionMansimov E et al,
Learning FRAME Models Using CNN Filters for Knowledge VisualizationLu Y, et al,
Aligning where to see and what to tell: image caption with region-based attention and scene factorizationJin J et al,

Awesome Image Captioning / Papers / 2016

Image captioning with semantic attentionYou Q et al,
DenseCap: Fully Convolutional Localization Networks for Dense CaptioningJohnson J et al,
What value do explicit high level concepts have in vision to language problems?Wu Q et al,
Deep Compositional Captioning: Describing Novel Object Categories without Paired Training DataLisa Anne Hendricks et al,
SPICE: Semantic Propositional Image Caption EvaluationAnderson P et al,
Image Captioning with Deep Bidirectional LSTMsWang C et al,
Multimodal Pivots for Image Caption TranslationHitschler J et al,
Image Caption Generation with Text-Conditional Semantic AttentionZhou L et al,
DeepDiary: Automatic Caption Generation for Lifelogging Image StreamsFan C et al,
Learning to generalize to new compositions in image understandingAtzmon Y et al,
Generating captions without looking beyond objectsHeuer H et al,
Bootstrap, Review, Decode: Using Out-of-Domain Textual Data to Improve Image CaptioningChen W et al,
Recurrent Image Captioner: Describing Images with Spatial-Invariant Transformation and Attention FilteringLiu H et al,
Recurrent Highway Networks with Language CNN for Image CaptioningGu J et al,

Awesome Image Captioning / Papers / 2017

Captioning Images with Diverse ObjectsVenugopalan S et al,
Top-down Visual Saliency Guided by CaptionsRamanishka V et al,
Self-Critical Sequence Training for Image CaptioningSteven J et al,
Dense Captioning with Joint Inference and Visual ContextYang L et al,
Skeleton Key: Image Captioning by Skeleton-Attribute DecompositionYufei W et al,
A Hierarchical Approach for Generating Descriptive Image ParagraphsKrause J et al,
Deep Reinforcement Learning-based Image Captioning with Embedding RewardRen Z et al,
Incorporating Copying Mechanism in Image Captioning for Learning Novel ObjectsTing Y et al,
Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image CaptioningLu J et al,
Attend to You: Personalized Image Captioning with Context Sequence Memory NetworksCC Park et al,
SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioningChen L et al,
Bidirectional Beam Search: Forward-Backward Inference in Neural Sequence Models for Fill-In-The-Blank Image CaptioningQing S et al,
Areas of Attention for Image CaptioningPedersoli M et al,
Boosting Image Captioning with AttributesYao T et al,
An Empirical Study of Language CNN for Image CaptioningGu J et al,
Improved Image Captioning via Policy Gradient Optimization of SPIDErLiu S et al,
Towards Diverse and Natural Image Descriptions via a Conditional GANDai B et al,
Paying Attention to Descriptions Generated by Image Captioning ModelsTavakoliy H R et al,
Show, Adapt and Tell: Adversarial Training of Cross-domain Image CaptionerChen T H et al,
Image Caption with Global-Local AttentionLi L et al,
Reference Based LSTM for Image CaptioningChen M et al,
Attention Correctness in Neural Image CaptioningLiu C et al,
Text-guided Attention Model for Image CaptioningMun J et al,
Contrastive Learning for Image CaptioningDai B et al,
Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning ChallengeVinyals O et al,
MAT: A Multimodal Attentive Translator for Image CaptioningLiu C et al,
Actor-Critic Sequence Training for Image CaptioningZhang L et al,
What is the Role of Recurrent Neural Networks (RNNs) in an Image Caption Generator?Tanti M et al,
Self-Guiding Multimodal LSTM - when we do not have a perfect training dataset for image captioningXian Y et al,
Phrase-based Image Captioning with Hierarchical LSTM ModelTan Y H et al,
Show-and-Fool: Crafting Adversarial Examples for Neural Image CaptioningChen H et al,

Awesome Image Captioning / Papers / 2018

Neural Baby TalkLu J et al,
Convolutional Image CaptioningAneja J et al,
Learning to Evaluate Image CaptioningCui Y et al,
Discriminability Objective for Training Descriptive CaptionsLuo R et al,
SemStyle: Learning to Generate Stylised Image Captions using Unaligned TextMathews A et al,
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question AnsweringAnderson P et al,
GroupCap: Group-Based Image Captioning With Structured Relevance and Diversity ConstraintsChen F et al,
Unpaired Image Captioning by Language PivotingGu J et al,
Recurrent Fusion Network for Image CaptioningJiang W et al,
Exploring Visual Relationship for Image CaptioningYao T et al,
Rethinking the Form of Latent States in Image CaptioningDai B et al,
Boosted Attention: Leveraging Human Attention for Image CaptioningChen S et al,
"Factual" or "Emotional": Stylized Image Captioning with Adaptive Learning and AttentionChen T et al,
Learning to Guide Decoding for Image CaptioningJiang W et al,
Stack-Captioning: Coarse-to-Fine Learning for Image CaptioningGu J et al,
Temporal-difference Learning with Sampling Baseline for Image CaptioningChen H et al,
Partially-Supervised Image CaptioningAnderson P et al,
A Neural Compositional Paradigm for Image CaptioningDai B et al,
Defoiling Foiled Image CaptionsWang J et al,
Punny Captions: Witty Wordplay in Image DescriptionsChandrasekaran A et al,
Object Counts! Bringing Explicit Detections Back into Image CaptioningAneja J et al,
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image CaptioningSharma P et al,
Attacking visual language grounding with adversarial examples: A case study on neural image captioningChen H et al,
simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image CaptionsLiu et al,
Improved Image Captioning with Adversarial Semantic AlignmentMelnyk I et al,
Improving Image Captioning with Conditional Generative Adversarial NetsChen C et al,
CNN+CNN: Convolutional Decoders for Image CaptioningWang Q et al,
Diverse and Controllable Image Captioning with Part-of-Speech GuidanceDeshpande A et al,

Awesome Image Captioning / Papers / 2019

Unsupervised Image CaptioningYang F et al,
Engaging Image Captioning Via PersonalityShuster K et al,
Pointing Novel Objects in Image CaptioningLi Y et al,
Auto-Encoding Scene Graphs for Image CaptioningYang X et al,
Context and Attribute Grounded Dense CaptioningYin G et al,
Look Back and Predict Forward in Image CaptioningQin Y et al,
Self-critical n-step Training for Image CaptioningGao J et al,
Intention Oriented Image Captions with Guiding ObjectsZheng Y et al,
Describing like humans: on diversity in image captioningWang Q et al,
Adversarial Semantic Alignment for Improved Image CaptionsDognin P et al,
MSCap: Multi-Style Image Captioning With Unpaired Stylized TextGao L et al,
Fast, Diverse and Accurate Image Captioning Guided By Part-of-SpeechAditya D et al,
Good News, Everyone! Context driven entity-aware captioning for news imagesBiten A F et al,
CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection50over 6 years agoZhang L et al,
Dense Relational Captioning: Triple-Stream Networks for Relationship-Based CaptioningKim D et al,
Show, Control and Tell: A Framework for Generating Controllable and Grounded CaptionsCornia M et al,
Exact Adversarial Attack to Image Captioning via Structured Output Learning With Latent VariablesXu Y et al,
Meta Learning for Image CaptioningLi N et al,
Learning Object Context for Dense CaptioningLi X et al,
Hierarchical Attention Network for Image CaptioningWang W et al,
Deliberate Residual based Attention Network for Image CaptioningGao L et al,
Improving Image Captioning with Conditional Generative Adversarial NetsChen C et al,
Connecting Language to Images: A Progressive Attention-Guided Network for Simultaneous Image Captioning and Language GroundingSong L et al,
Dense Procedure Captioning in Narrated Instructional VideosShi B et al,
Informative Image Captioning with External Sources of InformationZhao S et al,
Bridging by Word: Image Grounded Vocabulary Construction for Visual CaptioningFan Z et al,
Image Captioning with Unseen ObjectsDemirel et al,
Look and Modify: Modification Networks for Image CaptioningSammani et al,
Show, Infer and Tell: Contextual Inference for Creative CaptioningKhare et al,
SC-RANK: Improving Convolutional Image Captioning with Self-Critical Learning and Ranking Metric-based RewardYan et al,
Hierarchy Parsing for Image CaptioningYao T et al,
Entangled Transformer for Image CaptioningLi G et al,
Attention on Attention for Image CaptioningHuang L et al,
Reflective Decoding Network for Image CaptioningKe L at al,
Learning to Collocate Neural Modules for Image CaptioningYang X et al,
Image Captioning: Transforming Objects into WordsHerdade S et al,
Adaptively Aligned Image Captioning via Adaptive Attention TimeHuang L et al,
Variational Structured Semantic Inference for Diverse Image CaptioningChen F et al,
Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image RepresentationsLiu F et al,
Image Captioning with Compositional Neural Module NetworksTian J et al,
Exploring and Distilling Cross-Modal Information for Image CaptioningLiu F et al,
Swell-and-Shrink: Decomposing Image Captioning by Transformation and SummarizationWang H et al,
Hornet: a hierarchical offshoot recurrent network for improving person re-ID via image captioningYan S et al,
Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning ApproachKim D J et al,
TIGEr: Text-to-Image Grounding for Image Caption EvaluationJiang M et al,
REO-Relevance, Extraness, Omission: A Fine-grained Evaluation for Image CaptioningJiang M et al,
Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question AnsweringChangpinyo S et al,
Compositional Generalization in Image CaptioningNikolaus M et al,

Awesome Image Captioning / Papers / 2020

MemCap: Memorizing Style Knowledge for Image CaptioningZhao et al,
Unified Vision-Language Pre-Training for Image Captioning and VQAZhou L et al,
Show, Recall, and Tell: Image Captioning with Recall MechanismWang L et al,
Reinforcing an Image Caption Generator using Off-line Human FeedbackHongsuck Seo P et al,
Interactive Dual Generative Adversarial Networks for Image CaptioningLiu et al,
Feature Deformation Meta-Networks in Image Captioning of Novel ObjectsCao et al,
Joint Commonsense and Relation Reasoning for Image and Video CaptioningHou et al,
Learning Long- and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image CaptionZhang et al,
Normalized and Geometry-Aware Self-Attention Network for Image CaptioningGuo L et al,
Object Relational Graph with Teacher-Recommended Learning for Video CaptioningZhang Z et al,
Say As You Wish: Fine-grained Control of Image Caption Generation with Abstract Scene GraphsChen S et al,
X-Linear Attention Networks for Image CaptioningPan et al,
Improving Image Captioning with Better Use of CaptionShi Z et al,
Cross-modal Coherence Modeling for Caption GenerationAlikhani M et al,
Improving Image Captioning Evaluation by Considering Inter References VarianceYi Y et al,
MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningLei J et al,
Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQAKim H et al,
Length-Controllable Image CaptioningDeng C et al,
Captioning Images Taken by People Who Are BlindGurari D et al,
Towards Unique and Informative Captioning of ImagesWang Z et al,
Learning Visual Representations with Caption AnnotationsSariyildiz M et al,
Comprehensive Image Captioning via Scene Graph DecompositionZhong Y et al,
SODA: Story Oriented Dense Video Captioning Evaluation FrameworkFujita S et al,
TextCaps: a Dataset for Image Captioning with Reading ComprehensionSidorov O et al,
Compare and Reweight: Distinctive Image Captioning Using Similar Images SetsWang J et al,
Learning to Generate Grounded Visual Captions without Localization SupervisionMa C et al,
Fashion Captioning: Towards Generating Accurate Descriptions with Semantic RewardsYang X et al,
Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in VideosChen S et al,
CapWAP: Image Captioning with a PurposeFisch A et al,
X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal TransformersCho J et al,
Video2Commonsense: Generating Commonsense Descriptions to Enrich Video CaptioningFang Z et al,
Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsLi Y et al,
Diverse Image Captioning with Context-Object Split Latent SpacesMahajan S et al,
RATT: Recurrent Attention to Transient Tasks for Continual Image CaptioningChiaro R et al,

Awesome Image Captioning / Dataset

nocaps, LANG:
MS COCO, LANG:
Flickr 8k, LANG:
Flickr 30k, LANG:
AI Challenger, LANG:
Visual Genome, LANG:
SBUCaptionedPhotoDataset, LANG:
IAPR TC-12, LANG:

Awesome Image Captioning / Image Captioning Challenge

Microsoft COCO Image Captioning
Google AI Blog: Conceptual Captions
ruotianluo/self-critical.pytorch998almost 3 years ago
ruotianluo/ImageCaptioning.pytorch1,458almost 3 years ago
jiasenlu/NeuralBabyTalk525over 7 years ago
tensorflow/models/im2txt77,258almost 2 years ago
DeepRNN/image_captioning790over 4 years ago
jcjohnson/densecap1,584about 8 years ago
karpathy/neuraltalk25,515almost 9 years ago
jiasenlu/AdaptiveAttention335over 8 years ago
emansim/text2image594over 9 years ago
apple2373/chainer-caption64over 7 years ago
peteanderson80/bottom-up-attention1,438over 3 years ago

Backlinks from these awesome lists:

More related projects: