Emu
by baaivision
Emu Series: Generative Multimodal Models from BAAI
AI summary
Model framework
A multimodal generative model framework
- stars
- 1.7K
- forks
- 86
- watching
- 22
Similar projects
Found by comparing what the projects do, not just their names.
Vision-Language Model
A PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities
Generative Model
This project aims to develop a generative model for 3D multi-object scenes using a novel network architecture inspired by auto-encoding and generative adversarial networks.
Multimodal model framework
A framework for grounding language models to images and handling multimodal inputs and outputs
Language Models
A repository of pre-trained language models for various tasks and domains.
nvlabs/edm1.4K
Generative model framework
This project provides a set of tools and techniques to design and improve diffusion-based generative models.
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Language model trainer
This project provides code and model for improving language understanding through generative pre-training using a transformer-based architecture.
Multimodal conversational model
An end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations
Vision-Language Bridge
Pre-trains a multilingual model to bridge vision and language modalities for various downstream applications
Model evaluation framework
An evaluation toolkit and platform for assessing large models in various domains
Language model toolkit
Provides pre-trained language models and tools for fine-tuning and evaluation
nvlabs/eagle549
Multimodal model builder
Develops high-resolution multimodal LLMs by combining vision encoders and various input resolutions
DL hub
A unified interface to various deep learning architectures
openai/pixel-cnn1.9K
Generative model
A generative model with tractable likelihood and easy sampling, allowing for efficient data generation.
Mixture of Experts Model
A large vision-language model using a mixture-of-experts architecture to improve performance on multi-modal learning tasks