JiuTian-LION

Visual Knowledge Model

This project integrates visual knowledge into large language models to improve their capabilities and reduce hallucinations.

[CVPR 2024] LION: Empowering Multimodal Large Language Model with Dual-Level Visual Knowledge

GitHub

124 stars
13 watching
6 forks
Language: Jupyter Notebook
last commit: about 2 years ago

Related projects:

RepositoryDescriptionStars
liaoning97/revo-lionA comprehensive dataset and evaluation framework for Vision-Language Instruction Tuning models11
yunxinli/lingcloudEnhances language models by incorporating human-like eyes to improve visual comprehension and interaction with external world48
byungkwanlee/collavoDevelops a PyTorch implementation of an enhanced vision language model93
yfzhang114/llava-alignDebiasing techniques to minimize hallucinations in large visual language models75
yiren-jian/blitextDevelops and trains models for vision-language learning with decoupled language pre-training24
meituan-automl/mobilevlmAn implementation of a vision language model designed for mobile devices, utilizing a lightweight downsample projector and pre-trained language models.1,076
deepseek-ai/deepseek-vlA multimodal AI model that enables real-world vision-language understanding applications2,145
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
ys-zong/vl-iclA benchmarking suite for multimodal in-context learning models31
yuliang-liu/monkeyAn end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.1,849
pleisto/yuren-baichuan-7bA multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks73
brucesherwood/vpython-jupyterAn integration of VPython with Jupyter Notebook for interactive 3D visualization and simulation in scientific computing.64
sy-xuan/pinkThis project enables multi-modal language models to understand and generate text about visual content using referential comprehension.79
jiasenlu/vilbert_betaA pre-trained model and toolset for performing vision-and-language tasks using a specific neural network architecture.473
yiyangzhou/lureAnalyzing and mitigating object hallucination in large vision-language models to improve their accuracy and reliability.136