coot-videotext

Video transformer

An open-source implementation of a video-text representation learning framework using transformers and PyTorch.

COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning

GitHub

288 stars
8 watching
55 forks
Language: Python
last commit: about 4 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
swintransformer/video-swin-transformerAn implementation of the Video Swin Transformer architecture for video recognition tasks1,463
hanzhanggit/stackganA PyTorch implementation of a generative adversarial network for image synthesis from text descriptions1,863
chaoyuaw/pytorch-coviarA PyTorch implementation of a compressed video action recognition system502
xiadingz/video-caption.pytorchPyTorch implementation of video captioning, combining deep learning and computer vision techniques.402
clementpinard/sfmlearner-pytorchPytorch implementation of unsupervised depth and ego-motion learning from video sequences1,022
thudm/cogviewA framework for generating images from text using transformers.1,735
pixart-alpha/pixart-sigmaDevelops a PyTorch model for 4K text-to-image generation using diffusion transformer1,711
jeonsworld/vit-pytorchA PyTorch implementation of the Vision Transformer model for image recognition tasks.1,959
microsoft/megatron-deepspeedResearch tool for training large transformer language models at scale1,926
bigscience-workshop/megatron-deepspeedA collection of tools and scripts for training large transformer language models at scale1,342
pylons/colanderA library for serializing and deserializing data structures into strings, mappings, and lists while performing validation.451
leviswind/pytorch-transformerImplementation of a transformer-based translation model in PyTorch240
tongjilibo/bert4torchAn implementation of transformer models in PyTorch for natural language processing tasks1,257
mchong6/soatThis repository provides a PyTorch implementation of an image manipulation technique using a pretrained StyleGAN model.380
locuslab/pytorch_fftProvides an efficient wrapper around CUDA FFTs for PyTorch transformations315