MIC
by HaozheZhao
MMICL, a state-of-the-art VLM with the in context learning ability from ICL, PKU
AI summary
Multimodal learner
Develops a multimodal vision-language model to enable machines to understand complex relationships between instructions and images in various tasks.
- stars
- 337
- forks
- 15
- watching
- 10
Similar projects
Found by comparing what the projects do, not just their names.
Chart model trainer
Develops a large-scale dataset and benchmark for training multimodal chart understanding models using large language models.
Multimodal LLM
A multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks
Learning benchmark
A benchmarking suite for multimodal in-context learning models
Multimodal LLM
A multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation
Vision Language Integrator
Improves performance of vision language tasks by integrating computer vision capabilities into large language models
Multimodal alignment model
Extending pretraining models to handle multiple modalities by aligning language and video representations
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
Multimodal conversational model
An end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations
Mixture of Experts Model
A large vision-language model using a mixture-of-experts architecture to improve performance on multi-modal learning tasks
Multimodal model evaluator
Evaluating and improving large multimodal models through in-context learning
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
Region-of-Interest Training
Training and deploying large language models on computer vision tasks using region-of-interest inputs
Multimodal Chatbot
Trains a multimodal chatbot that combines visual and language instructions to generate responses