DreamLLM
by RunpeiDong
[ICLR 2024 Spotlight] DreamLLM: Synergistic Multimodal Comprehension and Creation
AI summary
Multimodal Model Builder
A framework to build versatile Multimodal Large Language Models with synergistic comprehension and creation capabilities
- stars
- 402
- forks
- 7
- watching
- 16
Similar projects
Found by comparing what the projects do, not just their names.
Multimodal LLM Framework
A framework that enables large language models to process and understand multimodal inputs from various sources such as images and speech.
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Multimodal LLM framework
A framework for training GPT4-style language models with multimodal inputs using large datasets and pre-trained models
Multimodal LLM
An implementation of a multimodal language model with capabilities for comprehension and generation
Multimodal LLM
A multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks
openbmb/viscpm1.1K
Multimodal Models
A family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages
nvlabs/eagle549
Multimodal model builder
Develops high-resolution multimodal LLMs by combining vision encoders and various input resolutions
Language Model
An implementation of a neural network model for character-level language modeling.
Language model trainer
A framework for training and fine-tuning multimodal language models on various data types
Multimodal conversational model
An end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations
Multimodal processor
A large multimodal language model designed to process and analyze video, image, text, and audio inputs in real-time.
Multilingual LLM
A polyglot large language model designed to address limitations in current LLM research and provide better multilingual instruction-following capability.
Language Model
This repository provides an end-to-end language model capable of generating coherent text based on both spoken and written inputs.
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
LLM
An open-source implementation of a vision-language instructed large language model