CLIP_prefix_caption
by rmokady
Simple image captioning model
AI summary
Image captioning method
An approach to image captioning that leverages the CLIP model and fine-tunes a language model without requiring additional supervision or object annotation.
- stars
- 1.3K
- forks
- 220
- watching
- 7
Similar projects
Found by comparing what the projects do, not just their names.
Image captioning model
A deep learning model for generating image captions with semantic attention
Image captioner
An unsupervised image captioning framework that allows generating captions from images without paired data.
Image Captioner
A project for pre-training models to support image captioning and question answering tasks.
Image Captioner
An image caption generation system using a neural network architecture with pre-trained models.
Caption generator
Trains image paragraph captioning models to generate diverse and accurate captions
Image captioner
Enhances language models to generate text based on visual descriptions of images
Image captioning system
This implementation allows users to generate captions from images using a neural network model with visual attention.
Captioner
Automated captioning and transcription tool for video and audio files
Caption rewriting
Automates the process of generating multiple rewritten image captions by fine-tuning large vision-language models
Image captioning model
An image caption generation model that uses abstract scene graphs to fine-grained control and generate captions
kdexd/virtex556
Caption learning
A pretraining approach that uses semantically dense captions to learn visual representations and improve image understanding tasks.
Image captioning framework
A Python-based framework for training and testing image captioning models using PyTorch.
Image generator
An image caption generation system utilizing machine learning models and deep neural networks.
Caption generator
A PyTorch implementation of a framework for generating captions with styles for images and videos.
Image Captioner
A PyTorch implementation of image captioning models via scene graph decomposition.