Awesome-Multimodal-Large-Language-Models

Conversational AI resources

A collection of resources and papers on multimodal large language models for understanding and building advanced conversational AI systems.

sparklessparklesLatest Advances on Multimodal Large Language Models

GitHub

13k stars
256 watching
837 forks
last commit: almost 2 years ago
Linked from 2 awesome lists

chain-of-thoughtin-context-learninginstruction-followinginstruction-tuninglarge-language-modelslarge-vision-language-modellarge-vision-language-modelsmulti-modalitymultimodal-chain-of-thoughtmultimodal-in-context-learningmultimodal-instruction-tuningmultimodal-large-language-modelsvisual-instruction-tuning

Awesome Papers / Multimodal Instruction Tuning

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Github396almost 2 years ago
Apollo: An Exploration of Video Understanding in Large Multimodal Models
Github
Demo
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Github2,616almost 2 years ago
StreamChat: Chatting with Streaming Video
CompCap: Improving Multimodal Large Language Models with Composite Captions
LinVT: Empower Your Image-level Large Language Model to Understand Videos
Github13almost 2 years ago
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Github6,394almost 2 years ago
Demo
NVILA: Efficient Frontier Visual Language Models
Github2,146almost 2 years ago
Demo
T2Vid: Translating Long Text into Multi-Image is the Catalyst for Video-LLMs
Github44almost 2 years ago
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Github67almost 2 years ago
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Github106almost 2 years ago
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Github329almost 2 years ago
Demo
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
Github89almost 2 years ago
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Github57almost 2 years ago
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models
Huggingface
Demo
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Github3,613almost 2 years ago
Demo
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via Hybrid Architecture
Github183almost 2 years ago
EAGLE: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Github549about 2 years ago
Demo
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
Github69almost 2 years ago
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Github2,365almost 2 years ago
VITA: Towards Open-Source Interactive Omni Multimodal LLM
Github1,005almost 2 years ago
LLaVA-OneVision: Easy Visual Task Transfer
Github3,099almost 2 years ago
Demo
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Github12,870almost 2 years ago
Demo
VILA^2: VILA Augmented VILA
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
EVLM: An Efficient Vision-Language Model for Visual Understanding
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
Github26almost 2 years ago
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Github2,616almost 2 years ago
Demo
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
Github1,336almost 2 years ago
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
Github9almost 2 years ago
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Github1,799almost 2 years ago
Long Context Transfer from Language to Vision
Github347almost 2 years ago
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Github1,091almost 2 years ago
TroL: Traversal of Layers for Large Language and Vision Models
Github88about 2 years ago
Unveiling Encoder-Free Vision-Language Models
Github246almost 2 years ago
VideoLLM-online: Online Video Large Language Model for Streaming Video
Github251about 2 years ago
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
Github64almost 2 years ago
Demo
Comparison Visual Instruction Tuning
Github
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Github143almost 2 years ago
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Github957almost 2 years ago
Parrot: Multilingual Visual Instruction Tuning
Github34about 2 years ago
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Github575almost 2 years ago
Matryoshka Query Transformer for Large Vision-Language Models
Github101about 2 years ago
Demo
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models
Github106about 2 years ago
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
Github102over 2 years ago
Demo
Libra: Building Decoupled Vision System on Large Language Models
Github153almost 2 years ago
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
Github136over 2 years ago
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Github6,394almost 2 years ago
Demo
Graphic Design with Large Multimodal Model
Github102over 2 years ago
BRAVE: Broadening the visual encoding of vision-language models
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Github2,616almost 2 years ago
Demo
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
Github254about 2 years ago
VITRON: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
Github406almost 2 years ago
TOMGPT: Reliable Text-Only Training Approach for Cost-Effective Multi-modal Large Language Model
LITA: Language Instructed Temporal-Localization Assistant
Github151almost 2 years ago
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Github3,229over 2 years ago
Demo
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
MoAI: Mixture of All Intelligence for Large Language and Vision Models
Github314over 2 years ago
DeepSeek-VL: Towards Real-World Vision-Language Understanding
Github2,145over 2 years ago
Demo
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Github1,849almost 2 years ago
Demo
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
Github466about 2 years ago
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Github798about 2 years ago
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Github58almost 2 years ago
ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language Model
Github249about 2 years ago
Demo
CoLLaVO: Crayon Large Language and Vision mOdel
Github93about 2 years ago
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
Github494about 2 years ago
CogCoM: Train Large Vision-Language Models Diving into Details through Chain of Manipulations
Github153about 2 years ago
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Github1,076over 2 years ago
GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning
Github43almost 2 years ago
Enhancing Multimodal Large Language Models with Vision Detection Models: An Empirical Study
Coming soon
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge
Github20,683about 2 years ago
Demo
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Github2,023almost 2 years ago
Demo
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Github2,616almost 2 years ago
Demo
Yi-VL7,743almost 2 years ago
Github7,743almost 2 years ago
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
Github108about 2 years ago
MobileVLM : A Fast, Reproducible and Strong Vision Language Assistant for Mobile Devices
Github1,076over 2 years ago
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Github6,394almost 2 years ago
Demo
Osprey: Pixel Understanding with Visual Instruction Tuning
Github781about 2 years ago
Demo
CogAgent: A Visual Language Model for GUI Agents
Github6,182over 2 years ago
Coming soon
Pixel Aligned Language Models
Coming soon
VILA: On Pre-training for Visual Language Models
Github2,146almost 2 years ago
See, Say, and Segment: Teaching LMMs to Overcome False Premises
Coming soon
Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Github1,831almost 2 years ago
Demo
Honeybee: Locality-enhanced Projector for Multimodal LLM
Github435over 2 years ago
Gemini: A Family of Highly Capable Multimodal Models
OneLLM: One Framework to Align All Modalities with Language
Github601almost 2 years ago
Demo
Lenna: Language Enhanced Reasoning Detection Assistant
Github78over 2 years ago
VaQuitA: Enhancing Alignment in LLM-Assisted Video Understanding
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Github314almost 2 years ago
Making Large Multimodal Models Understand Arbitrary Visual Prompts
Github302about 2 years ago
Demo
Dolphins: Multimodal Language Model for Driving
Github51about 2 years ago
LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning
Github255about 2 years ago
Coming soon
VTimeLLM: Empower LLM to Grasp Video Moments
Github231over 2 years ago
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
Github1,958almost 2 years ago
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Github748about 2 years ago
Coming soon
LLMGA: Multimodal Large Language Model based Generation Assistant
Github463about 2 years ago
Demo
ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Github202almost 3 years ago
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Github2,616almost 2 years ago
Demo
LION : Empowering Multimodal Large Language Model with Dual-Level Visual Knowledge
Github124about 2 years ago
An Embodied Generalist Agent in 3D World
Github379almost 2 years ago
Demo
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Github3,071almost 2 years ago
Demo
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Github895almost 2 years ago
To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Github131over 2 years ago
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Github2,732over 2 years ago
Demo
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Github1,849almost 2 years ago
Demo
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Github717over 2 years ago
Demo
NExT-Chat: An LMM for Chat, Detection and Segmentation
Github227over 2 years ago
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Github2,365almost 2 years ago
Demo
OtterHD: A High-Resolution Multi-modality Model
Github3,570over 2 years ago
CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding
Coming soon
GLaMM: Pixel Grounding Large Multimodal Model
Github797almost 2 years ago
Demo
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
Github18almost 3 years ago
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Github25,490about 2 years ago
SALMONN: Towards Generic Hearing Abilities for Large Language Models
Github1,091almost 2 years ago
Ferret: Refer and Ground Anything Anywhere at Any Granularity
Github8,509almost 2 years ago
CogVLM: Visual Expert For Large Language Models
Github6,182over 2 years ago
Demo
Improved Baselines with Visual Instruction Tuning
Github20,683about 2 years ago
Demo
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Github751over 2 years ago
Demo
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
Github79over 2 years ago
Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants
Github59over 2 years ago
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Github2,616almost 2 years ago
DreamLLM: Synergistic Multimodal Comprehension and Creation
Github402almost 2 years ago
Coming soon
An Empirical Study of Scaling Instruction-Tuned Large Multimodal Models
Coming soon
TextBind: Multi-turn Interleaved Multimodal Instruction-following
Github47about 3 years ago
Demo
NExT-GPT: Any-to-Any Multimodal LLM
Github3,344almost 2 years ago
Demo
Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics
Github19about 3 years ago
ImageBind-LLM: Multi-modality Instruction Tuning
Github5,775over 2 years ago
Demo
Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
PointLLM: Empowering Large Language Models to Understand Point Clouds
Github670almost 2 years ago
Demo
✨Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
Github43over 2 years ago
MLLM-DataEngine: An Iterative Refinement Approach for MLLM
Github39over 2 years ago
Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models
Github37about 3 years ago
Demo
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Github5,179about 2 years ago
Demo
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
Github1,098over 2 years ago
Demo
StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data
Github93over 2 years ago
BLIVA: A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions
Github270over 2 years ago
Demo
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
Github360over 2 years ago
The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
Github466about 2 years ago
Demo
LISA: Reasoning Segmentation via Large Language Model
Github1,923about 2 years ago
Demo
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Github550almost 2 years ago
3D-LLM: Injecting the 3D World into Large Language Models
Github979over 2 years ago
ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning
Demo
BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
Github505about 3 years ago
Demo
SVIT: Scaling up Visual Instruction Tuning
Github164over 2 years ago
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Github517over 2 years ago
Demo
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
Github231about 3 years ago
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Github1,958almost 2 years ago
Demo
Visual Instruction Tuning with Polite Flamingo
Github63almost 3 years ago
Demo
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Github259over 2 years ago
Demo
Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Github748about 2 years ago
Demo
MotionGPT: Human Motion as a Foreign Language
Github1,531over 2 years ago
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Github1,568over 2 years ago
Coming soon
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
Github305over 2 years ago
Demo
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Github1,246about 2 years ago
Demo
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Github3,570over 2 years ago
Demo
M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Github2,842over 2 years ago
Demo
LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
Github1,622about 2 years ago
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Github762almost 3 years ago
Demo
PandaGPT: One Model To Instruction-Follow Them All
Github772over 3 years ago
Demo
ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst
Github49about 3 years ago
Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
Github513over 2 years ago
DetGPT: Detect What You Need via Reasoning
Github761about 2 years ago
Demo
Pengi: An Audio Language Model for Audio Tasks
Github295over 2 years ago
VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Github956almost 2 years ago
Listen, Think, and Understand
Github396over 2 years ago
Demo396over 2 years ago
Github4,110about 2 years ago
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
Github180almost 2 years ago
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Github10,058almost 2 years ago
VideoChat: Chat-Centric Video Understanding
Github3,106almost 2 years ago
Demo
MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Github1,478over 3 years ago
Demo
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Github308about 3 years ago
LMEye: An Interactive Perception Network for Large Language Models
Github48about 2 years ago
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Github5,775over 2 years ago
Demo
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Github2,365almost 2 years ago
Demo
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Github25,490about 2 years ago
Visual Instruction Tuning
GitHub20,683about 2 years ago
Demo
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Github5,775over 2 years ago
Demo
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Github134over 3 years ago

Awesome Papers / Multimodal Hallucination

Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
Github28almost 2 years ago
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
Github46almost 2 years ago
FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs
Link
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
Github83almost 2 years ago
Evaluating and Analyzing Relationship Hallucinations in LVLMs
Github20almost 2 years ago
AGLA: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
Github18about 2 years ago
CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
Coming soon
Mitigating Object Hallucination via Data Augmented Contrastive Tuning
Coming soon
VDGD: Mitigating LVLM Hallucinations in Cognitive Prompts by Bridging the Visual Perception Gap
Coming soon
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
What if...?: Counterfactual Inception to Mitigate Hallucination Effects in Large Multimodal Models
Github15almost 2 years ago
Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization
Debiasing Multimodal Large Language Models
Github75over 2 years ago
HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
Github72almost 2 years ago
IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
Github39almost 2 years ago
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
Github19over 2 years ago
The Instinctive Bias: Spurious Images lead to Hallucination in MLLMs
Github8over 2 years ago
Unified Hallucination Detection for Multimodal Large Language Models
Github48over 2 years ago
A Survey on Hallucination in Large Vision-Language Models
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
Github82over 2 years ago
MOCHa: Multi-Objective Reinforcement Mitigating Caption Hallucinations
Github13almost 2 years ago
Mitigating Fine-Grained Hallucination by Fine-Tuning Large Vision-Language Models with Caption Rewrites
Github8over 2 years ago
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
Github245about 2 years ago
Demo
OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
Github293about 2 years ago
Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
Github222almost 2 years ago
Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
Github73over 2 years ago
Comins Soon
Mitigating Hallucination in Visual Language Models with Visual Supervision
HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
Github41about 2 years ago
An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
Github98over 2 years ago
FAITHSCORE: Evaluating Hallucinations in Large Vision-Language Models
Github27almost 2 years ago
Woodpecker: Hallucination Correction for Multimodal Large Language Models
Github617over 2 years ago
Demo
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
HallE-Switch: Rethinking and Controlling Object Existence Hallucinations in Large Vision Language Models for Detailed Caption
Github28over 2 years ago
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
Github136over 2 years ago
Aligning Large Multimodal Models with Factually Augmented RLHF
Github328almost 3 years ago
Demo
Evaluation and Mitigation of Agnosia in Multimodal Large Language Models
CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning
Evaluation and Analysis of Hallucination in Large Vision-Language Models
Github17almost 3 years ago
VIGC: Visual Instruction Generation and Correction
Github91over 2 years ago
Demo
Detecting and Preventing Hallucinations in Large Vision Language Models
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Github262over 2 years ago
Demo
Evaluating Object Hallucination in Large Vision-Language Models
Github187over 2 years ago

Awesome Papers / Multimodal In-Context Learning

Visual In-Context Learning for Large Vision-Language Models
RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning in Multi-Modal Large Language Model
Github76almost 2 years ago
Can MLLMs Perform Text-to-Image In-Context Learning?
Github30almost 2 years ago
Generative Multimodal Models are In-Context Learners
Github1,672almost 2 years ago
Demo
Hijacking Context in Large Multi-modal Models
Towards More Unified In-context Visual Understanding
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Github337almost 3 years ago
Demo
Link-Context Learning for Multimodal LLMs
Github91over 2 years ago
Demo
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Github3,781about 2 years ago
Demo
Med-Flamingo: a Multimodal Medical Few-shot Learner
Github396about 3 years ago
Generative Pretraining in Multimodality
Github1,672almost 2 years ago
Demo
AVIS: Autonomous Visual Information Seeking with Large Language Models
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Github3,570over 2 years ago
Demo
Exploring Diverse In-Context Configurations for Image Captioning
Github33almost 2 years ago
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
Github1,095over 2 years ago
Demo
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace
Github23,801almost 2 years ago
Demo
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Github940over 2 years ago
Demo
ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information Extraction
Github50about 3 years ago
Prompting Large Language Models with Answer Heuristics for Knowledge-based Visual Question Answering
Github270over 3 years ago
Visual Programming: Compositional visual reasoning without training
Github697about 2 years ago
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
Github85over 4 years ago
Flamingo: a Visual Language Model for Few-Shot Learning
Github3,781about 2 years ago
Demo
Multimodal Few-Shot Learning with Frozen Language Models

Awesome Papers / Multimodal Chain-of-Thought

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Github113almost 2 years ago
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
Github73over 2 years ago
Visual CoT: Unleashing Chain-of-Thought Reasoning in Multi-Modal Language Models
Github162almost 2 years ago
Compositional Chain-of-Thought Prompting for Large Multimodal Models
Github90over 2 years ago
DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models
Github35over 2 years ago
Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Github748about 2 years ago
Demo
Explainable Multimodal Emotion Reasoning
Github123over 2 years ago
EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought
Github346over 2 years ago
Let’s Think Frame by Frame: Evaluating Video Chain of Thought with Video Infilling and Prediction
T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering
Caption Anything: Interactive Image Description with Diverse Multimodal Controls
Github1,693about 3 years ago
Demo
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
Coming soon
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
Github1,095over 2 years ago
Demo
Chain of Thought Prompt Tuning in Vision Language Models
Coming soon
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Github940over 2 years ago
Demo
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Github34,555over 2 years ago
Demo
Multimodal Chain-of-Thought Reasoning in Language Models
Github3,833over 2 years ago
Visual Programming: Compositional visual reasoning without training
Github697about 2 years ago
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Github615about 2 years ago

Awesome Papers / LLM-Aided Visual Reasoning

Beyond Embeddings: The Promise of Visual Table in Multi-Modal Models
Github14almost 2 years ago
V∗: Guided Visual Search as a Core Mechanism in Multimodal LLMs
Github541over 2 years ago
LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing
Github353about 2 years ago
Demo
MM-VID: Advancing Video Understanding with GPT-4V(vision)
ControlLLM: Augment Language Models with Tools by Searching on Graphs
Github187about 2 years ago
Woodpecker: Hallucination Correction for Multimodal Large Language Models
Github617over 2 years ago
Demo
MindAgent: Emergent Gaming Interaction
Github79over 2 years ago
Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language
Github352almost 3 years ago
Demo
Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models
AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
Github66over 3 years ago
AVIS: Autonomous Visual Information Seeking with Large Language Models
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Github762almost 3 years ago
Demo
Mindstorms in Natural Language-Based Societies of Mind
LayoutGPT: Compositional Visual Planning and Generation with Large Language Models
Github306over 2 years ago
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
Github32almost 3 years ago
Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation
Github7over 3 years ago
Caption Anything: Interactive Image Description with Diverse Multimodal Controls
Github1,693about 3 years ago
Demo
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
Github1,095over 2 years ago
Demo
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace
Github23,801almost 2 years ago
Demo
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Github940over 2 years ago
Demo
ViperGPT: Visual Inference via Python Execution for Reasoning
Github1,666over 2 years ago
ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions
Github457over 3 years ago
ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information Extraction
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Github34,555over 2 years ago
Demo
Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners
Github41over 3 years ago
From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models
Github10,058almost 2 years ago
Demo
SuS-X: Training-Free Name-Only Transfer of Vision-Language Models
Github94about 3 years ago
PointCLIP V2: Adapting CLIP for Powerful 3D Open-world Learning
Github235about 3 years ago
Visual Programming: Compositional visual reasoning without training
Github697about 2 years ago
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Github34,478almost 2 years ago

Awesome Papers / Foundation Models

Emu3: Next-Token Prediction is All You Need
Github1,911almost 2 years ago
Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
Demo
Pixtral-12B
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
Github10,058almost 2 years ago
The Llama 3 Herd of Models
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Hello GPT-4o
The Claude 3 Model Family: Opus, Sonnet, Haiku
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini: A Family of Highly Capable Multimodal Models
Fuyu-8B: A Multimodal Architecture for AI Agents
Huggingface
Demo
Unified Model for Image, Video, Audio and Language Tasks
Github224over 2 years ago
Demo
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
GPT-4V(ision) System Card
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
Github544almost 2 years ago
Multimodal Foundation Models: From Specialists to General-Purpose Assistants
Bootstrapping Vision-Language Learning with Decoupled Language Pre-training
Github24almost 3 years ago
Generative Pretraining in Multimodality
Github1,672almost 2 years ago
Demo
Kosmos-2: Grounding Multimodal Large Language Models to the World
Github20,400almost 2 years ago
Demo
Transfer Visual Prompt Generator across LLMs
Github270almost 3 years ago
Demo
GPT-4 Technical Report
PaLM-E: An Embodied Multimodal Language Model
Demo
Prismer: A Vision-Language Model with An Ensemble of Experts
Github1,299over 2 years ago
Demo
Language Is Not All You Need: Aligning Perception with Language Models
Github20,400almost 2 years ago
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Github10,058almost 2 years ago
Demo
VIMA: General Robot Manipulation with Multimodal Prompts
Github781over 2 years ago
MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
Github1,843over 2 years ago
Write and Paint: Generative Vision-Language Models are Unified Modal Learners
Github43over 3 years ago
Language Models are General-Purpose Interfaces
Github20,400almost 2 years ago

Awesome Papers / Evaluation

MMGenBench: Evaluating the Limits of LMMs from the Text-to-Image Generation Perspective
Github106almost 2 years ago
OmniBench: Towards The Future of Universal Omni-Language Models
Github15almost 2 years ago
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Github86almost 2 years ago
UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models
Github3about 2 years ago
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
Github22almost 2 years ago
Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs
Github67almost 2 years ago
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
Github85almost 2 years ago
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
Github95about 2 years ago
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Github422almost 2 years ago
Benchmarking Large Multimodal Models against Common Corruptions
Github27over 2 years ago
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Github296over 2 years ago
A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Github13,117almost 2 years ago
BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models
Github84about 2 years ago
How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs
Github72almost 3 years ago
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
Github24about 2 years ago
MLLM-Bench, Evaluating Multi-modal LLMs using GPT-4V
Github56almost 2 years ago
VLM-Eval: A General Evaluation on Video Large Language Models
Coming soon
Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges
Github53over 2 years ago
On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving
Github288over 2 years ago
Towards Generic Anomaly Detection and Understanding: Large-scale Visual-linguistic Model (GPT-4V) Takes the Lead
A Comprehensive Study of GPT-4V's Multimodal Capabilities in Medical Imaging
An Early Evaluation of GPT-4V(ision)
Github11almost 3 years ago
Exploring OCR Capabilities of GPT-4V(ision) : A Quantitative and In-depth Evaluation
Github121almost 3 years ago
HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models
Github259almost 2 years ago
MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal Models
Github253almost 2 years ago
Fool Your (Vision and) Language Model With Embarrassingly Simple Permutations
Github14almost 3 years ago
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
Github21over 2 years ago
Can We Edit Multimodal Large Language Models?
Github1,981almost 2 years ago
REVO-LION: Evaluating and Refining Vision-Language Instruction Tuning Datasets
Github11almost 3 years ago
The Dawn of LMMs: Preliminary Explorations with GPT-4V(vision)
TouchStone: Evaluating Vision-Language Models by Language Models
Github79over 2 years ago
✨Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
Github43over 2 years ago
SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs
Github38almost 2 years ago
Tiny LVLM-eHub: Early Multimodal Experiments with Bard
Github478over 2 years ago
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Github274almost 2 years ago
SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Github322about 2 years ago
MMBench: Is Your Multi-modal Model an All-around Player?
Github168about 2 years ago
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Github13,117almost 2 years ago
LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
Github478over 2 years ago
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
Github305over 2 years ago
M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models
Github93over 3 years ago
On The Hidden Mystery of OCR in Large Multimodal Models
Github484almost 2 years ago

Awesome Papers / Multimodal RLHF

Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
Silkie: Preference Distillation for Large Visual Language Models
Github88over 2 years ago
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
Github245about 2 years ago
Demo
Aligning Large Multimodal Models with Factually Augmented RLHF
Github328almost 3 years ago
Demo
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
Github2almost 2 years ago

Awesome Papers / Others

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Github7almost 2 years ago
Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
Github47about 2 years ago
VCoder: Versatile Vision Encoders for Multimodal Large Language Models
Github266over 2 years ago
Prompt Highlighter: Interactive Control for Multi-Modal LLMs
Github135about 2 years ago
Planting a SEED of Vision in Large Language Model
Github585almost 2 years ago
Can Large Pre-trained Models Help Vision Models on Perception Tasks?
Github1,218almost 2 years ago
Contextual Object Detection with Multimodal Large Language Models
Github208almost 2 years ago
Demo
Generating Images with Multimodal Language Models
Github440over 2 years ago
On Evaluating Adversarial Robustness of Large Vision-Language Models
Github165almost 3 years ago
Grounding Language Models to Images for Multimodal Inputs and Outputs
Github478almost 3 years ago
Demo

Awesome Datasets / Datasets of Pre-Training for Alignment

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
COYO-700M: Image-Text Pair Dataset1,172almost 4 years ago
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Microsoft COCO: Common Objects in Context
Im2Text: Describing Images Using 1 Million Captioned Photographs
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding
Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark
Kosmos-2: Grounding Multimodal Large Language Models to the World
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline
AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Awesome Datasets / Datasets of Multimodal Instruction Tuning

E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
Link42almost 2 years ago
Multi-modal Situated Reasoning in 3D Scenes
Link
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
Link
UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models
Link3about 2 years ago
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
Link33about 2 years ago
ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language Model
Link
Visually Dehallucinative Instruction Generation: Know What You Don't Know
Link6over 2 years ago
Visually Dehallucinative Instruction Generation
Link5over 2 years ago
M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts
Link58almost 2 years ago
Making Large Multimodal Models Understand Arbitrary Visual Prompts
Link
To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Link
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
Link18almost 3 years ago
✨Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
Link43over 2 years ago
StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data
Link93over 2 years ago
Detecting and Preventing Hallucinations in Large Vision Language Models
Coming soon
ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning
BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
Link
SVIT: Scaling up Visual Instruction Tuning
Link
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Link1,958almost 2 years ago
Visual Instruction Tuning with Polite Flamingo
Link
ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Link
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Link
MotionGPT: Human Motion as a Foreign Language
Link1,531over 2 years ago
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Link262over 2 years ago
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Link1,568over 2 years ago
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
Link305over 2 years ago
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Link1,246about 2 years ago
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Link3,570over 2 years ago
M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Link
LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
Coming soon1,622about 2 years ago
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Link762almost 3 years ago
ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst
Coming soon
DetGPT: Detect What You Need via Reasoning
Link761about 2 years ago
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
Coming soon
VideoChat: Chat-Centric Video Understanding
Link1,467almost 2 years ago
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Link308about 3 years ago
LMEye: An Interactive Perception Network for Large Language Models
Link
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Link
Visual Instruction Tuning
Link
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Link134over 3 years ago

Awesome Datasets / Datasets of In-Context Learning

MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Link
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Link3,570over 2 years ago

Awesome Datasets / Datasets of Multimodal Chain-of-Thought

Explainable Multimodal Emotion Reasoning
Coming soon123over 2 years ago
EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought
Coming soon346over 2 years ago
Let’s Think Frame by Frame: Evaluating Video Chain of Thought with Video Infilling and Prediction
Coming soon
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Link615about 2 years ago

Awesome Datasets / Datasets of Multimodal RLHF

Silkie: Preference Distillation for Large Visual Language Models
Link

Awesome Datasets / Benchmarks for Evaluation

M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Link47over 2 years ago
MMGenBench: Evaluating the Limits of LMMs from the Text-to-Image Generation Perspective
Link106almost 2 years ago
MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps
Link3almost 2 years ago
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
Link
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Link
OmniBench: Towards The Future of Universal Omni-Language Models
Link
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Link
VELOCITI: Can Video-Language Models Bind Semantic Concepts through Time?
Link5about 2 years ago
Seeing Clearly, Answering Incorrectly: A Multimodal Robustness Benchmark for Evaluating MLLMs on Leading Questions
Link43almost 2 years ago
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
Link
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Link422almost 2 years ago
VL-ICL Bench: The Devil in the Details of Benchmarking Multimodal In-Context Learning
Link31over 2 years ago
TempCompass: Do Video LLMs Really Understand Videos?
Link91almost 2 years ago
GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning
Link
Can MLLMs Perform Text-to-Image In-Context Learning?
Link
Visually Dehallucinative Instruction Generation: Know What You Don't Know
Link6over 2 years ago
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Link74almost 2 years ago
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
Link22about 2 years ago
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
Link46about 2 years ago
Benchmarking Large Multimodal Models against Common Corruptions
Link27over 2 years ago
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Link296over 2 years ago
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Link
Making Large Multimodal Models Understand Arbitrary Visual Prompts
Link
M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts
Link58almost 2 years ago
Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Link121over 2 years ago
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
Link24about 2 years ago
MLLM-Bench, Evaluating Multi-modal LLMs using GPT-4V
Link56almost 2 years ago
BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models
Link
MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
Link87almost 2 years ago
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Link3,106almost 2 years ago
Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges
Link53over 2 years ago
OtterHD: A High-Resolution Multi-modality Model
Link
HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models
Link259almost 2 years ago
Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond
Link99over 2 years ago
Aligning Large Multimodal Models with Factually Augmented RLHF
Link
MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal Models
Link
✨Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
Link43over 2 years ago
Link-Context Learning for Multimodal LLMs
Link
Detecting and Preventing Hallucinations in Large Vision Language Models
Coming soon
Empowering Vision-Language Models to Follow Interleaved Vision-Language Instructions
Link360over 2 years ago
SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs
Link38almost 2 years ago
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Link274almost 2 years ago
SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Link322about 2 years ago
MMBench: Is Your Multi-modal Model an All-around Player?
Link168about 2 years ago
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
Link231about 3 years ago
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Link262over 2 years ago
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Link13,117almost 2 years ago
LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
Link478over 2 years ago
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
Link305over 2 years ago
M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models
Link93over 3 years ago
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Link2,365almost 2 years ago

Awesome Datasets / Others

IMAD: IMage-Augmented multi-modal Dialogue
Link4over 3 years ago
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Link1,246about 2 years ago
Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation
Link
Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation
Link
Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Link
Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities
Link

Backlinks from these awesome lists:

More related projects: