R2D2

Vision-Language Framework

A framework for large-scale cross-modal benchmarks and vision-language tasks in Chinese

GitHub

157 stars
2 watching
23 forks
Language: Python
last commit: almost 3 years ago

Related projects:

RepositoryDescriptionStars
zhourax/vegaDevelops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.33
hxyou/idealgptA deep learning framework for iteratively decomposing vision and language reasoning via large language models.32
shizhediao/davinciImplementing a unified modal learning framework for generative vision-language models43
yuliang-liu/monkeyAn end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.1,849
yiren-jian/blitextDevelops and trains models for vision-language learning with decoupled language pre-training24
wpiroboticsprojects/gripA computer vision framework for robotics applications that simplifies the creation of vision systems and generates code in multiple programming languages.380
vlf-silkie/vlfeedbackAn annotated preference dataset and training framework for improving large vision language models.88
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
baai-wudao/brivlPre-trains a multilingual model to bridge vision and language modalities for various downstream applications279
openuc2/uc2-gitAn open-source software framework for building modular electro-optical projects with interchangeable components468
byungkwanlee/moaiImproves performance of vision language tasks by integrating computer vision capabilities into large language models314
xiaoyufenfei/lednetA lightweight deep learning framework for real-time semantic segmentation514
yulingtianxia/core-ml-sampleA demo project demonstrating the integration of Core ML and Vision Framework with Swift 4 for image classification using an Inception V3 network.217
vishaal27/sus-xThis is an open-source project that proposes a novel method to train large-scale vision-language models with minimal resources and no fine-tuning required.94
kohjingyu/fromageA framework for grounding language models to images and handling multimodal inputs and outputs478