virtex
by kdexd
[CVPR 2021] VirTex: Learning Visual Representations from Textual Annotations
AI summary
Caption learning
A pretraining approach that uses semantically dense captions to learn visual representations and improve image understanding tasks.
- stars
- 556
- forks
- 61
- watching
- 14
Similar projects
Found by comparing what the projects do, not just their names.
Video captioner
PyTorch implementation of video captioning, combining deep learning and computer vision techniques.
Image Captioning Model
A deep learning framework providing a model architecture and training code for image captioning using semantic compositional networks
Visual reasoning engine
A framework for training multi-modal language models with a focus on visual inputs and providing interpretable thoughts.
Video captioning model
An implementation of a dense video captioning model with attention-based fusion and context gating
Image Captioner
A project for pre-training models to support image captioning and question answering tasks.
Image captioning model
A deep learning model for generating image captions with semantic attention
Image captioning method
An approach to image captioning that leverages the CLIP model and fine-tunes a language model without requiring additional supervision or object annotation.
Deep Learning Tutorials
A collection of tutorials and resources on implementing deep learning models using Python libraries such as Keras and Lasagne.
Caption generator
Trains image paragraph captioning models to generate diverse and accurate captions
Instruction generator
Creating synthetic visual reasoning instructions to improve the performance of large language models on image-related tasks
Token representation refinement
Improves pre-trained language models by encouraging an isotropic and discriminative distribution of token representations.
Image Captioner
A PyTorch implementation of image captioning models via scene graph decomposition.
Representation learner
Reimplements a popular deep learning model for unsupervised visual representation learning using TensorFlow
Image captioning system
This implementation allows users to generate captions from images using a neural network model with visual attention.
Image Captioner
An image caption generation system using a neural network architecture with pre-trained models.