CoLLaVO
by ByungKwanLee
[ACL 2024 Findings] Official PyTorch Implementation code for realizing the technical part of CoLLaVO: Crayon Large Language and Vision mOdel to significantly improve zero-shot vision language performances
AI summary
Vision Language Model
Develops a PyTorch implementation of an enhanced vision language model
- stars
- 93
- forks
- 13
- watching
- 5
Similar projects
Found by comparing what the projects do, not just their names.
Vision Language Integrator
Improves performance of vision language tasks by integrating computer vision capabilities into large language models
Computer Vision Toolkit
A PyTorch toolbox for supporting research and development of domain adaptation, generalization, and semi-supervised learning methods in computer vision.
Model optimization library
An implementation of Mamba-based traversal of rationale to improve performance of numerous vision language models.
Image synthesis model
An implementation of semantic image synthesis via adversarial learning using PyTorch
Computer Vision Model Builder
A Python package for building and deploying computer vision models with PyTorch
Image descriptor model
Improves large vision-language models' ability to accurately describe images by combining global and local attention mechanisms.
Vision-Language Learning Model
Develops and trains models for vision-language learning with decoupled language pre-training
Video-language model
An efficient framework for end-to-end learning on image-text and video-text tasks
Vision-Language Model
A PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities
Hallucination mitigation
This project provides an official PyTorch implementation of a method to interpret and edit vision-language representations to mitigate hallucinations in image captions.
Visual Knowledge Model
This project integrates visual knowledge into large language models to improve their capabilities and reduce hallucinations.
Vision-Language Model Framework
Implementing a unified modal learning framework for generative vision-language models
Video captioner
PyTorch implementation of video captioning, combining deep learning and computer vision techniques.
Long context transfer
An open-source project that enables the transfer of language understanding to vision capabilities through long context processing.
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.