Ovis

Multimodal aligner

An MLLM architecture designed to align visual and textual embeddings through structural alignment

A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.

GitHub

575 stars
7 watching
33 forks
Language: Python
last commit: almost 2 years ago
chatbotllama3multimodalmultimodal-large-language-modelsmultimodalityqwenvision-language-learningvision-language-model

Related projects:

RepositoryDescriptionStars
aidc-ai/parrotA method and toolkit for fine-tuning large language models to perform visual instruction tasks in multiple languages.34
rlhf-v/rlhf-vAligns large language models' behavior through fine-grained correctional human feedback to improve trustworthiness and accuracy.245
ailab-cvc/seedAn implementation of a multimodal language model with capabilities for comprehension and generation585
ucsc-vlaa/sight-beyond-textAn implementation of a multimodal LLM training paradigm to enhance truthfulness and ethics in language models19
pku-alignment/align-anythingAligns large multimodal models with human intentions and values using various algorithms and fine-tuning methods.270
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
pku-yuangroup/languagebindExtending pretraining models to handle multiple modalities by aligning language and video representations751
salt-nlp/llavarAn open-source project that enhances visual instruction tuning for text-rich image understanding by integrating GPT-4 models with multimodal datasets.259
deepseek-ai/deepseek-vlA multimodal AI model that enables real-world vision-language understanding applications2,145
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
tanloong/interlaced.nvimA plugin for aligning bilingual parallel texts by re-positioning text and applying highlighting.7
lancopku/iaisThis project proposes a novel method for calibrating attention distributions in multimodal models to improve contextualized representations of image-text pairs.30
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
pku-yuangroup/moe-llavaA large vision-language model using a mixture-of-experts architecture to improve performance on multi-modal learning tasks2,023
opengvlab/visionllmA large language model designed to process and generate visual information956