OMG-Seg
by lxtGH
OMG-LLaVA and OMG-Seg codebase [CVPR-24 and NeurIPS-24]
AI summary
Visual Model
Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.
- stars
- 1.3K
- forks
- 50
- watching
- 22
Similar projects
Found by comparing what the projects do, not just their names.
Visual decoder
A large language model designed to process and generate visual information
Language model toolkit
A guide to using pre-trained large language models in source code analysis and generation
LLM trainer
Transfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs
Language model
An implementation of DeepMind's Relational Recurrent Neural Networks (Santoro et al. 2018) in PyTorch for word language modeling
Visual reasoning engine
A framework for training multi-modal language models with a focus on visual inputs and providing interpretable thoughts.
LLM
An open-source implementation of a vision-language instructed large language model
NLP framework
A pre-trained language model designed for various NLP tasks, including dialogue generation, code completion, and retrieval.
VQA model
A Visual Question Answering model using a deeper LSTM and normalized CNN architecture.
Image understanding model
A large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.
Visual encoder
A vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.
CV library
Provides a PyTorch implementation of several computer vision tasks including object detection, segmentation and parsing.
Visual Model
Develops a multimodal Chinese language model with visual capabilities
Multimodal model developer
Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.
Model trainer
A platform for training and deploying large language and vision models that can use tools to perform tasks
Benchmark
An image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy