CuMo
by SHI-Labs
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
AI summary
Mixture-of-experts model
A method for scaling multimodal large language models by combining multiple experts and fine-tuning them together
- stars
- 136
- forks
- 10
- watching
- 2
Similar projects
Found by comparing what the projects do, not just their names.
Mixture of Experts Model
A large vision-language model using a mixture-of-experts architecture to improve performance on multi-modal learning tasks
Video text retrieval model
An open-source implementation of the Mixture-of-Embeddings-Experts model in Pytorch for video-text retrieval tasks.
Mixture-of-Experts Model
Developed by XVERSE Technology Inc. as a multilingual large language model with a unique mixture-of-experts architecture and fine-tuned for various tasks such as conversation, question answering, and natural language understanding.
Multimodal learner
Develops a multimodal vision-language model to enable machines to understand complex relationships between instructions and images in various tasks.
Multimodal LLM
A multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks
Efficient LLM
A large language model with improved efficiency and performance compared to similar models
Multimodal model developer
Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.
Language Model
A high-performance language model designed to excel in tasks like natural language understanding, mathematical computation, and code generation
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
Instruction model
A collection of multilingual language models trained on a dataset of instructions and responses in various languages.
Perception adapter
An adapter for improving large language models at object-level perception tasks with auxiliary perception modalities
Region-of-Interest Training
Training and deploying large language models on computer vision tasks using region-of-interest inputs
Chinese understanding benchmark
Measures the understanding of massive multitask Chinese datasets using large language models
Vision-Language Model Framework
Implementing a unified modal learning framework for generative vision-language models
LM Benchmark
A benchmark for evaluating large language models in multiple languages and formats