Caption-Anything

Captioner

A tool generating descriptive captions from images with customizable controls and text styles.

Caption-Anything is a versatile tool combining image segmentation, visual captioning, and ChatGPT, generating tailored captions with diverse controls for user preferences. https://huggingface.co/spaces/TencentARC/Caption-Anything https://huggingface.co/spaces/VIPLab/Caption-Anything

GitHub

2k stars
16 watching
103 forks
Language: Python
last commit: about 3 years ago
chatgptcontrollable-generationcontrollable-image-captioningimage-captioningsegment-anything

Related projects:

RepositoryDescriptionStars
fengyang0317/unsupervised_captioningAn unsupervised image captioning framework that allows generating captions from images without paired data.215
apple2373/chainer-captionAn image caption generation system using a neural network architecture with pre-trained models.64
tpkahlon/captcha-imageA library to generate images with distorted text and background patterns for security purposes.8
eladhoffer/captiongenA PyTorch-based tool for generating captions from images128
lumingyin/quickcaptionAutomated captioning and transcription tool for video and audio files74
kacky24/stylenetA PyTorch implementation of a framework for generating captions with styles for images and videos.63
lukemelas/image-paragraph-captioningTrains image paragraph captioning models to generate diverse and accurate captions90
xiadingz/video-caption.pytorchPyTorch implementation of video captioning, combining deep learning and computer vision techniques.402
rmokady/clip_prefix_captionAn approach to image captioning that leverages the CLIP model and fine-tunes a language model without requiring additional supervision or object annotation.1,326
cshizhe/asg2capAn image caption generation model that uses abstract scene graphs to fine-grained control and generate captions200
vision-cair/chatcaptionerEnables automatic generation of descriptive text from images and videos based on user input.457
yiwuzhong/sub-gcA PyTorch implementation of image captioning models via scene graph decomposition.96
contextualai/lensEnhances language models to generate text based on visual descriptions of images352
jaywongwang/densevideocaptioningAn implementation of a dense video captioning model with attention-based fusion and context gating149
nickjiang2378/vl-interpThis project provides an official PyTorch implementation of a method to interpret and edit vision-language representations to mitigate hallucinations in image captions.46