Visual-CoT
by deepcs233
[Neurips'24 Spotlight] Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
AI summary
Visual reasoning engine
A framework for training multi-modal language models with a focus on visual inputs and providing interpretable thoughts.
- stars
- 162
- forks
- 7
- watching
- 1
Similar projects
Found by comparing what the projects do, not just their names.
rowanz/r2c466
Visual Reasoning Model
An open-source project providing PyTorch code and data for a deep learning model that enables visual commonsense reasoning.
lxtgh/omg-seg1.3K
Visual Model
Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.
Image Captioning Model
A deep learning framework providing a model architecture and training code for image captioning using semantic compositional networks
Visual Reasoning Model
An open-source implementation of a deep learning model designed to improve the balance between performance and interpretability in visual reasoning tasks.
Word Embedding Model
A library for training and evaluating a type of word embedding model that extends the original Word2Vec algorithm
Instruction generator
Creating synthetic visual reasoning instructions to improve the performance of large language models on image-related tasks
VQA model
A PyTorch implementation of visual question answering with multimodal representation learning
Word embeddings
Multi-sense word embeddings learned from visual cooccurrences
Image understanding model
A large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.
Vision-Language Model
A multimodal AI model that enables real-world vision-language understanding applications
Benchmark
An image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy
Semantic segmentation model
PyTorch implementation of DeepLab v2 for semantic segmentation on COCO-Stuff and PASCAL VOC datasets
Video-language model
An efficient framework for end-to-end learning on image-text and video-text tasks
Image-based word embeddings
Learning word embeddings from abstract images to improve language understanding
kdexd/virtex556
Caption learning
A pretraining approach that uses semantically dense captions to learn visual representations and improve image understanding tasks.