bubogpt

Visual grounding framework

An open-source framework enabling multi-modal LLMs to jointly understand text, vision, and audio and ground knowledge into visual objects.

BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

GitHub

505 stars
10 watching
35 forks
Language: Python
last commit: about 3 years ago

Related projects:

RepositoryDescriptionStars
theshadow29/zsgnet-pytorchAn implementation of a computer vision model that grounds objects in images using natural language queries.69
google-research/visu3dAn abstraction layer between various deep learning frameworks and your program.149
hxyou/idealgptA deep learning framework for iteratively decomposing vision and language reasoning via large language models.32
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336
mbzuai-oryx/groundinglmmAn end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations797
jy0205/lavitA unified framework for training large language models to understand and generate visual content544
tianyi-lab/hallusionbenchAn image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy259
jhcho99/coformerAn implementation of a deep learning model for grounding situation recognition in images45
penghao-wu/vstarPyTorch implementation of guided visual search mechanism for multimodal LLMs541
vpgtrans/vpgtransTransfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs270
davidmascharka/tbd-netsAn open-source implementation of a deep learning model designed to improve the balance between performance and interpretability in visual reasoning tasks.348
asappresearch/flambeAn ML framework for accelerating research and its integration into production workflows264
nvlabs/bongard-hoiA benchmarking tool and software framework for evaluating few-shot visual reasoning capabilities in computer vision models.64
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
s-gupta/visual-conceptsThis codebase provides a framework for detecting visual concepts in images by leveraging image captions and pre-trained models.150