Otter

Multi-modal AI model

A multi-modal AI model developed for improved instruction-following and in-context learning, utilizing large-scale architectures and various training datasets.

🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

GitHub

4k stars
100 watching
243 forks
Language: Python
last commit: over 2 years ago
artificial-inteligencechatgptdeep-learningembodied-aifoundation-modelsgpt-4instruction-tuninglarge-scale-modelsmachine-learningmulti-modalityvisual-language-learning

Related projects:

RepositoryDescriptionStars
haotian-liu/llavaA system that uses large language and vision models to generate and process visual instructions20,683
mlfoundations/open_flamingoA framework for training large multimodal models to generate text conditioned on images or other text.3,781
dvlab-research/mgmAn open-source framework for training large language models with vision capabilities.3,229
pku-yuangroup/video-llavaA deep learning framework for generating videos from text inputs and visual features.3,071
opengvlab/llama-adapterAn implementation of a method for fine-tuning language models to follow instructions with high efficiency and accuracy5,775
paddlepaddle/fastdeployA toolkit for easy and high-performance deployment of deep learning models on various hardware platforms3,034
eleutherai/lm-evaluation-harnessProvides a unified framework to test generative language models on various evaluation tasks.7,200
facico/chinese-vicunaAn instruction-following Chinese LLaMA-based model project aimed at training and fine-tuning models on specific hardware configurations for efficient deployment.4,152
modeltc/lightllmA Python-based framework for serving large language models with low latency and high scalability.2,691
alpha-vllm/llama2-accessoryAn open-source toolkit for pretraining and fine-tuning large language models2,732
open-mmlab/mmdeployA toolset for deploying deep learning models on various devices and platforms2,797
deep-floyd/ifA text-to-image synthesis model with a modular design, utilizing a frozen text encoder and cascaded pixel diffusion modules to generate photorealistic images.7,699
ludwig-ai/ludwigA low-code framework for building custom deep learning models and neural networks11,236
open-mmlab/mmcvProvides a foundational library for computer vision research and training deep learning models with high-quality implementation of common CPU and CUDA ops.5,948
openbmb/minicpm-vA multimodal language model designed to understand images, videos, and text inputs and generate high-quality text outputs.12,870