DaVinci
by shizhediao
Source code for the paper "Prefix Language Models are Unified Modal Learners"
AI summary
Vision-Language Model Framework
Implementing a unified modal learning framework for generative vision-language models
- stars
- 43
- forks
- 3
- watching
- 10
Similar projects
Found by comparing what the projects do, not just their names.
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
Computer Vision Toolkit
A PyTorch toolbox for supporting research and development of domain adaptation, generalization, and semi-supervised learning methods in computer vision.
Vision Reasoning Framework
A deep learning framework for iteratively decomposing vision and language reasoning via large language models.
Vision-Language Learning Model
Develops and trains models for vision-language learning with decoupled language pre-training
Vision language model trainer
An annotated preference dataset and training framework for improving large vision language models.
yuxie11/r2d2157
Vision-Language Framework
A framework for large-scale cross-modal benchmarks and vision-language tasks in Chinese
Vision Language Model
Develops a PyTorch implementation of an enhanced vision language model
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
CV framework
A PyTorch-based framework for building and training deep learning models in computer vision.
Optical flow model
An implementation of unsupervised learning for multi-frame optical flow with occlusions using PyTorch.
Image synthesis model
An implementation of semantic image synthesis via adversarial learning using PyTorch
Vision-Language Model
A PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities
Segmentation model library
Implementation of semantic segmentation models and datasets using PyTorch
Segmentation framework
Provides PyTorch implementations of various models and pipelines for semantic segmentation in deep learning.
Vision-Language Model Trainer
This is an open-source project that proposes a novel method to train large-scale vision-language models with minimal resources and no fine-tuning required.