MGM
Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
AI summary
Vision-LM Framework
An open-source framework for training large language models with vision capabilities.
- stars
- 3.2K
- forks
- 281
- watching
- 28
Similar projects
Found by comparing what the projects do, not just their names.
haotian-liu/llava20.7K
Visual Instruction System
A system that uses large language and vision models to generate and process visual instructions
LLM toolkit
An open-source toolkit for pretraining and fine-tuning large language models
Instruction-following model tuner
An implementation of a method for fine-tuning language models to follow instructions with high efficiency and accuracy
Multimodal model developer
Develops large multimodal models for various computer vision tasks including image and video analysis
Image generator
Generates images from text prompts using a variant of the DALL-E model
Vision-Language Model
Enabling vision-language understanding by fine-tuning large language models on visual data.
Video generator
A deep learning framework for generating videos from text inputs and visual features.
openbmb/minicpm-v12.9K
Multimodal LLM
A multimodal language model designed to understand images, videos, and text inputs and generate high-quality text outputs.
nvlabs/eagle549
Multimodal model builder
Develops high-resolution multimodal LLMs by combining vision encoders and various input resolutions
Evaluation framework
Provides a unified framework to test generative language models on various evaluation tasks.
Multimodal reasoning model
An implementation of multimodal chain-of-thought reasoning in language models using a decoupled training framework for rationale generation and answer inference.
open-mmlab/mmcv5.9K
Computer Vision Library
Provides a foundational library for computer vision research and training deep learning models with high-quality implementation of common CPU and CUDA ops.
Model debiasing
Debiasing techniques to minimize hallucinations in large visual language models
qwenlm/qwen-vl5.2K
Large vision language model
A large vision language model with improved image reasoning and text recognition capabilities, suitable for various multimodal tasks
luodian/otter3.6K
Multi-modal AI model
A multi-modal AI model developed for improved instruction-following and in-context learning, utilizing large-scale architectures and various training datasets.