SEED
by AILab-CVC
Official implementation of SEED-LLaMA (ICLR 2024).
AI summary
Multimodal LLM
An implementation of a multimodal language model with capabilities for comprehension and generation
- stars
- 585
- forks
- 32
- watching
- 15
Similar projects
Found by comparing what the projects do, not just their names.
Multimodal LLM test suite
A benchmark for evaluating large language models' ability to process multimodal input
Multimodal LLM
An LLaMA-based multimodal language model with various instruction-following and multimodal variants.
Multimodal LLM
An implementation of a multimodal language model using locality-enhanced projection techniques
VQA model
A multimodal LLM designed to handle text-rich visual questions
Multimodal LLM
A multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks
Image processor
An all-in-one demo for interactive image processing and generation
Multimodal LLM
A multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation
nvlabs/eagle549
Multimodal model builder
Develops high-resolution multimodal LLMs by combining vision encoders and various input resolutions
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Multimodal LLM framework
A framework for training GPT4-style language models with multimodal inputs using large datasets and pre-trained models
Text controller
An interactive control system for text generation in multi-modal language models
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
Multimodal LLM
An open-source multilingual large language model designed to understand and generate content across diverse languages and cultural contexts
Multimodal Model Builder
A framework to build versatile Multimodal Large Language Models with synergistic comprehension and creation capabilities