AnyGPT
by OpenMOSS
Code for "AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling"
AI summary
Multimodal converter
An open-source multimodal language model that can process and convert different data types such as speech, text, images, and music into a unified format.
- stars
- 798
- forks
- 64
- watching
- 20
Similar projects
Found by comparing what the projects do, not just their names.
openbmb/viscpm1.1K
Multimodal Models
A family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages
Multimodal Chatbot
Trains a multimodal chatbot that combines visual and language instructions to generate responses
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
Motion generator
Develops a unified model to generate high-quality motions and text descriptions from human motion data
Multimodal processor
A large multimodal language model designed to process and analyze video, image, text, and audio inputs in real-time.
Multimodal conversational model
An end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Model arena
An evaluation platform for comparing multi-modality models on visual question-answering tasks
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
Visual decoder
A large language model designed to process and generate visual information
r2d4/openlm367
LLM client
Library that provides a unified API to interact with various Large Language Models (LLMs)
Multimodal model developer
Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.
Language model trainer
This project provides code and model for improving language understanding through generative pre-training using a transformer-based architecture.
Multimodal LLM
An open-source multilingual large language model designed to understand and generate content across diverse languages and cultural contexts
File converter
A simple Matrix bot that listens to uploaded files and converts Quarto files to PDF.