EAGLE
Multimodal model builder
Develops high-resolution multimodal LLMs by combining vision encoders and various input resolutions
EAGLE: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
549 stars
31 watching
45 forks
Language: Python
last commit: about 2 years agodemoeaglegpt4huggingfacelarge-language-modelsllamallama3llavallmlmmlvlmmllmnvdia
Related projects:
| Repository | Description | Stars |
|---|---|---|
| An implementation of a multimodal language model with capabilities for comprehension and generation | 585 | |
| A deep learning framework for training multi-modal models with vision and language capabilities. | 1,299 | |
| An all-in-one demo for interactive image processing and generation | 353 | |
| A PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities | 246 | |
| A framework to build versatile Multimodal Large Language Models with synergistic comprehension and creation capabilities | 402 | |
| A multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation | 1,568 | |
| This project provides a set of tools and techniques to design and improve diffusion-based generative models. | 1,447 | |
| An open-source implementation of a vision-language instructed large language model | 513 | |
| Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously. | 15 | |
| An evaluation platform for comparing multi-modality models on visual question-answering tasks | 478 | |
| A multimodal LLM designed to handle text-rich visual questions | 270 | |
| A framework for training GPT4-style language models with multimodal inputs using large datasets and pre-trained models | 231 | |
| Trains a multimodal chatbot that combines visual and language instructions to generate responses | 1,478 | |
| Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types. | 143 | |
| An open-source multilingual large language model designed to understand and generate content across diverse languages and cultural contexts | 92 |