cambrian

Vision-based LLM

An open-source multimodal LLM project with a vision-centric design

Cambrian-1 is a family of multimodal LLMs with a vision-centric design.

GitHub

2k stars
23 watching
117 forks
Language: Python
last commit: almost 2 years ago
chatbotclipcomputer-visiondinoinstruction-tuninglarge-language-modelsllmsmllmmultimodal-large-language-modelsrepresentation-learning

Related projects:

RepositoryDescriptionStars
pleisto/yuren-baichuan-7bA multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks73
mbzuai-oryx/groundinglmmAn end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations797
bobazooba/xllm-demoA demo project showcasing customization possibilities of an XLLM library9
lyuchenyang/macaw-llmA multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation1,568
victordibia/llmxAn API that provides a unified interface to multiple large language models for chat fine-tuning79
bytedance/lynx-llmA framework for training GPT4-style language models with multimodal inputs using large datasets and pre-trained models231
bobazooba/xllmA tool for training and fine-tuning large language models using advanced techniques387
ailab-cvc/seedAn implementation of a multimodal language model with capabilities for comprehension and generation585
damo-nlp-mt/polylmA polyglot large language model designed to address limitations in current LLM research and provide better multilingual instruction-following capability.77
openbmb/viscpmA family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages1,098
terminaldweller/millaAn IRC bot that interacts with language models to provide answers and has customizable syntax highlighting.5
nvlabs/eagleDevelops high-resolution multimodal LLMs by combining vision encoders and various input resolutions549
open-mmlab/multimodal-gptTrains a multimodal chatbot that combines visual and language instructions to generate responses1,478
evolvinglmms-lab/longvaAn open-source project that enables the transfer of language understanding to vision capabilities through long context processing.347
phellonchen/x-llmA framework that enables large language models to process and understand multimodal inputs from various sources such as images and speech.308