ClipBERT

Video-language model

An efficient framework for end-to-end learning on image-text and video-text tasks

[CVPR 2021 Best Student Paper Honorable Mention, Oral] Official PyTorch code for ClipBERT, an efficient framework for end-to-end learning on image-text and video-text tasks.

GitHub

709 stars
10 watching
86 forks
Language: Python
last commit: about 3 years ago
Linked from 1 awesome list

cvpr2021pytorchvideo-question-answeringvideo-retrievalvision-and-languagevqa

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
jayleicn/tvqaPyTorch implementation of video question answering system based on TVQA dataset172
cadene/vqa.pytorchA PyTorch implementation of visual question answering with multimodal representation learning718
zhanghang1989/pytorch-encodingA Python framework for building deep learning models with optimized encoding layers and batch normalization.2,044
kaiyangzhou/dassl.pytorchA PyTorch toolbox for supporting research and development of domain adaptation, generalization, and semi-supervised learning methods in computer vision.1,236
zsef123/efficientnets-pytorchA PyTorch implementation of EfficientNet for computer vision tasks309
davidtvs/pytorch-enetA PyTorch implementation of a real-time semantic segmentation model using ENet architecture392
kacky24/stylenetA PyTorch implementation of a framework for generating captions with styles for images and videos.63
codeslake/pvdnetAn open-source implementation of a deep learning model for video deblurring and motion estimation.114
baaivision/eveA PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities246
xiadingz/video-caption.pytorchPyTorch implementation of video captioning, combining deep learning and computer vision techniques.402
byungkwanlee/collavoDevelops a PyTorch implementation of an enhanced vision language model93
jinsc37/difrintA PyTorch implementation of a deep learning-based method for video stabilization via frame interpolation.82
jwyang/graph-rcnn.pytorchA collection of PyTorch implementations of various scene graph generation models732
fartashf/vseppA PyTorch implementation of visual-semantic embedding methods for image-caption retrieval492
randl/shufflenetv2-pytorchAn implementation of a lightweight convolutional neural network architecture for mobile devices191