IAIS

Attention calibrator

This project proposes a novel method for calibrating attention distributions in multimodal models to improve contextualized representations of image-text pairs.

[ACL 2021] Learning Relation Alignment for Calibrated Cross-modal Retrieval

GitHub

30 stars
4 watching
4 forks
Language: Python
last commit: over 3 years ago
multimodalretrievalvision-and-language

Related projects:

RepositoryDescriptionStars
pku-yuangroup/languagebindExtending pretraining models to handle multiple modalities by aligning language and video representations751
mshukor/evalign-iclEvaluating and improving large multimodal models through in-context learning21
pku-alignment/align-anythingAligns large multimodal models with human intentions and values using various algorithms and fine-tuning methods.270
hekj/fdaThis project proposes a novel data augmentation technique to enhance visual-textual matching in vision-and-language navigation tasks.13
aidc-ai/ovisAn MLLM architecture designed to align visual and textual embeddings through structural alignment575
isekai-portal/link-context-learningAn implementation of a multimodal learning approach to improve language models' ability to recognize unseen images and understand novel concepts.91
ifl-camp/easy_handeyeAutomated calibration tool for robotic vision systems893
pkunlp-icler/pca-evalAn open-source benchmark and evaluation tool for assessing multimodal large language models' performance in embodied decision-making tasks99
szagoruyko/attention-transferImproves performance of convolutional neural networks by transferring knowledge from teacher models to student models using attention mechanisms.1,449
bryanplummer/pl-clcThis implementation provides a framework for phrase localization and visual relationship detection using comprehensive image-language cues.39
jiasenlu/adaptiveattentionAdaptive attention mechanism for image captioning using visual sentinels335
byungkwanlee/moaiImproves performance of vision language tasks by integrating computer vision capabilities into large language models314
tiger-ai-lab/uniirTrains and evaluates a universal multimodal retrieval model to perform various information retrieval tasks.114
mop/bierThis project implements a deep metric learning framework using an adversarial auxiliary loss to improve robustness.39
megvii-research/tlcImproves image restoration performance by converting global operations to local ones during inference231