OMG-Seg

Visual Model

Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.

OMG-LLaVA and OMG-Seg codebase [CVPR-24 and NeurIPS-24]

GitHub

1k stars
22 watching
50 forks
Language: Python
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
opengvlab/visionllmA large language model designed to process and generate visual information956
vhellendoorn/code-lmsA guide to using pre-trained large language models in source code analysis and generation1,789
vpgtrans/vpgtransTransfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs270
l0sg/relational-rnn-pytorchAn implementation of DeepMind's Relational Recurrent Neural Networks (Santoro et al. 2018) in PyTorch for word language modeling245
deepcs233/visual-cotA framework for training multi-modal language models with a focus on visual inputs and providing interpretable thoughts.162
luogen1996/lavinAn open-source implementation of a vision-language instructed large language model513
opennlg/openbaA pre-trained language model designed for various NLP tasks, including dialogue generation, code completion, and retrieval.94
gt-vision-lab/vqa_lstm_cnnA Visual Question Answering model using a deeper LSTM and normalized CNN architecture.377
360cvgroup/360vlA large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.32
gordonhu608/mqt-llavaA vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.101
openseg-group/openseg.pytorchProvides a PyTorch implementation of several computer vision tasks including object detection, segmentation and parsing.1,191
airaria/visual-chinese-llama-alpacaDevelops a multimodal Chinese language model with visual capabilities429
yfzhang114/slimeDevelops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.143
llava-vl/llava-plus-codebaseA platform for training and deploying large language and vision models that can use tools to perform tasks717
tianyi-lab/hallusionbenchAn image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy259