VLP

Image Captioner

A project for pre-training models to support image captioning and question answering tasks.

Vision-Language Pre-training for Image Captioning and Question Answering

GitHub

416 stars
19 watching
62 forks
Language: Python
last commit: over 4 years ago

Related projects:

RepositoryDescriptionStars
fengyang0317/unsupervised_captioningAn unsupervised image captioning framework that allows generating captions from images without paired data.215
contextualai/lensEnhances language models to generate text based on visual descriptions of images352
ruotianluo/imagecaptioning.pytorchA Python-based framework for training and testing image captioning models using PyTorch.1,458
lukemelas/image-paragraph-captioningTrains image paragraph captioning models to generate diverse and accurate captions90
libvips/lua-vipsA Lua binding for a fast image processing library with low memory needs.129
ruotianluo/self-critical.pytorchAn implementation of Self-critical Sequence Training for Image Captioning and related techniques.998
nickjiang2378/vl-interpThis project provides an official PyTorch implementation of a method to interpret and edit vision-language representations to mitigate hallucinations in image captions.46
yiwuzhong/sub-gcA PyTorch implementation of image captioning models via scene graph decomposition.96
xiadingz/video-caption.pytorchPyTorch implementation of video captioning, combining deep learning and computer vision techniques.402
chxj1992/slide_captcha_crackerA project that uses image processing techniques to locate a sliding captcha puzzle within a background image.142
rmokady/clip_prefix_captionAn approach to image captioning that leverages the CLIP model and fine-tunes a language model without requiring additional supervision or object annotation.1,326
microsoft/vision-longformerAn implementation of a vision transformer architecture designed for high-resolution image encoding with multiple efficient attention mechanisms243
chapternewscu/image-captioning-with-semantic-attentionA deep learning model for generating image captions with semantic attention51
hasinhayder/imagecaptionhoveranimationA CSS3-based solution to create hover animations for image captions354
lumingyin/quickcaptionAutomated captioning and transcription tool for video and audio files74