Osprey

Visual guidance

This project presents a new approach to fine-grained visual understanding using pixel-wise mask regions in language instructions

[CVPR2024] The code for "Osprey: Pixel Understanding with Visual Instruction Tuning"

GitHub

781 stars
14 watching
42 forks
Language: Python
last commit: about 2 years ago
mllmpixel-understandingsamvisual-instruction-tuning

Related projects:

RepositoryDescriptionStars
rucaibox/comvintCreating synthetic visual reasoning instructions to improve the performance of large language models on image-related tasks18
salt-nlp/llavarAn open-source project that enhances visual instruction tuning for text-rich image understanding by integrating GPT-4 models with multimodal datasets.259
roboflow/maestroA tool to streamline fine-tuning of multimodal models for vision-language tasks1,415
ys-zong/vlguardImproves safety and helpfulness of large language models by fine-tuning them using safety-critical tasks47
jshilong/gpt4roiTraining and deploying large language models on computer vision tasks using region-of-interest inputs517
aidc-ai/parrotA method and toolkit for fine-tuning large language models to perform visual instruction tasks in multiple languages.34
aidc-ai/ovisAn MLLM architecture designed to align visual and textual embeddings through structural alignment575
penghao-wu/vstarPyTorch implementation of guided visual search mechanism for multimodal LLMs541
bigredt/vicoMulti-sense word embeddings learned from visual cooccurrences25
codeplant/simple-navigationA Ruby gem for creating hierarchical navigation structures in web applications886
baai-dcai/visual-instruction-tuningA dataset and model designed to scale visual instruction tuning using language-only GPT-4 models.164
byungkwanlee/moaiImproves performance of vision language tasks by integrating computer vision capabilities into large language models314
kunpengli1994/vsrnAn open-source PyTorch implementation of a visual semantic reasoning model for image-text matching294
sy-xuan/pinkThis project enables multi-modal language models to understand and generate text about visual content using referential comprehension.79
dannnylo/rtesseractA Ruby library providing an interface to the Tesseract OCR system.838