AGLA

Image descriptor model

Improves large vision-language models' ability to accurately describe images by combining global and local attention mechanisms.

[Arxiv 2024] AGLA: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention

GitHub

18 stars
2 watching
0 forks
Language: Python
last commit: about 2 years ago

Related projects:

RepositoryDescriptionStars
byungkwanlee/collavoDevelops a PyTorch implementation of an enhanced vision language model93
deepseek-ai/deepseek-vlA multimodal AI model that enables real-world vision-language understanding applications2,145
dvlab-research/lisaA system that uses large language models to generate segmentation masks for images based on complex queries and world knowledge.1,923
baaivision/eveA PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities246
yiyangzhou/lureAnalyzing and mitigating object hallucination in large vision-language models to improve their accuracy and reliability.136
andy971022/auto-lamaAutomates object removal from images using computer vision techniques99
yfzhang114/llava-alignDebiasing techniques to minimize hallucinations in large visual language models75
damo-nlp-sg/vcdAn approach to reduce object hallucinations in large vision-language models by contrasting output distributions derived from original and distorted visual inputs222
ayoolaolafenwa/pixellibA deep learning library for image segmentation and object detection using PyTorch.1,054
umass-foundation-model/3d-llmDeveloping a Large Language Model capable of processing 3D representations as inputs979
opengvlab/visionllmA large language model designed to process and generate visual information956
byungkwanlee/moaiImproves performance of vision language tasks by integrating computer vision capabilities into large language models314
mshukor/evalign-iclEvaluating and improving large multimodal models through in-context learning21
algolzw/daclip-uirThis project controls vision-language models to restore degraded images in various environments and conditions.673
uclanlp/elmo-cEfficient Contextual Representation Learning Model with Continuous Outputs4