Muffin

Multimodal bridge

A framework for building multimodal foundation models that can serve as bridges between different modalities and language models.

GitHub

59 stars
8 watching
3 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
yuliang-liu/monkeyAn end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.1,849
joez17/chatbridgeA unified multimodal language model capable of interpreting and reasoning about various modalities without paired data.49
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
matrix-org/matrix-bifrostA general-purpose bridge that connects multiple networks and protocols using various backends.164
mautrix/telegramEnables communication between Matrix and Telegram networks by bridging them together1,360
mwotton/hubrisA bridge between Ruby and Haskell allowing code reuse across the two languages262
sorunome/mx-puppet-bridgeA library that allows building bridges between Matrix and remote services by automating logins and interactions.95
bendudson/py4clA bridge between Common Lisp and Python, enabling interaction between the two languages through a separate process.235
mautrix/whatsappA software bridge connecting Matrix and WhatsApp1,301
subho406/omninetAn implementation of a unified architecture for multi-modal multi-task learning using PyTorch.515
yglukhov/nimpyA bridge between Nim and Python, allowing native language integration.1,482
metawilm/cl-pythonAn implementation of Python in Common Lisp, allowing mixed execution and library access between the two languages.369
openbmb/viscpmA family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages1,098
kohjingyu/fromageA framework for grounding language models to images and handling multimodal inputs and outputs478
vita-mllm/vitaA large multimodal language model designed to process and analyze video, image, text, and audio inputs in real-time.1,005