R2D2
by yuxie11
AI summary
Vision-Language Framework
A framework for large-scale cross-modal benchmarks and vision-language tasks in Chinese
- stars
- 157
- forks
- 23
- watching
- 2
Similar projects
Found by comparing what the projects do, not just their names.
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
Vision Reasoning Framework
A deep learning framework for iteratively decomposing vision and language reasoning via large language models.
Vision-Language Model Framework
Implementing a unified modal learning framework for generative vision-language models
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Vision-Language Learning Model
Develops and trains models for vision-language learning with decoupled language pre-training
Vision framework
A computer vision framework for robotics applications that simplifies the creation of vision systems and generates code in multiple programming languages.
Vision language model trainer
An annotated preference dataset and training framework for improving large vision language models.
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
Vision-Language Bridge
Pre-trains a multilingual model to bridge vision and language modalities for various downstream applications
Electro-optical toolkit
An open-source software framework for building modular electro-optical projects with interchangeable components
Vision Language Integrator
Improves performance of vision language tasks by integrating computer vision capabilities into large language models
Semantic Segmentation Framework
A lightweight deep learning framework for real-time semantic segmentation
Image classifier demo
A demo project demonstrating the integration of Core ML and Vision Framework with Swift 4 for image classification using an Inception V3 network.
Vision-Language Model Trainer
This is an open-source project that proposes a novel method to train large-scale vision-language models with minimal resources and no fine-tuning required.
Multimodal model framework
A framework for grounding language models to images and handling multimodal inputs and outputs