BLIVA
by mlpc-ucsd
(AAAI 2024) BLIVA: A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions
AI summary
VQA model
A multimodal LLM designed to handle text-rich visual questions
- stars
- 270
- forks
- 28
- watching
- 12
Similar projects
Found by comparing what the projects do, not just their names.
Visual Prompt Model
A system designed to enable large multimodal models to understand arbitrary visual prompts
Multimodal LLM
An implementation of a multimodal language model with capabilities for comprehension and generation
Image processor
An all-in-one demo for interactive image processing and generation
Multimodal model trainer
An implementation of a multimodal LLM training paradigm to enhance truthfulness and ethics in language models
Multimodal LLM test suite
A benchmark for evaluating large language models' ability to process multimodal input
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
VQA model
A Visual Question Answering model using a deeper LSTM and normalized CNN architecture.
Multimodal processor
A large multimodal language model designed to process and analyze video, image, text, and audio inputs in real-time.
VQA model framework
A software framework for training and deploying multimodal visual question answering models using compact bilinear pooling.
Model trainer
A platform for training and deploying large language and vision models that can use tools to perform tasks
Visual QA Model
This project presents a neural network model designed to answer visual questions by combining question and image features in a residual learning framework.
VQA prompter
An implementation of a two-stage framework designed to prompt large language models with answer heuristics for knowledge-based visual question answering tasks.
openbmb/viscpm1.1K
Multimodal Models
A family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages
nvlabs/eagle549
Multimodal model builder
Develops high-resolution multimodal LLMs by combining vision encoders and various input resolutions
lxtgh/omg-seg1.3K
Visual Model
Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.