Awesome-Visual-Transformer

CV transformer papers

Collects and curates papers on transformer-based computer vision research

Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)

GitHub

3k stars
102 watching
400 forks
last commit: over 3 years ago
Linked from 3 awesome lists

detrtransformertransformer-awesometransformer-cvtransformer-with-cvvisual-transformer

Awesome Visual-Transformer / Papers / Transformer original paper

Attention is All You Need(NIPS 2017)

Awesome Visual-Transformer / Papers / Technical blog

Link[English Blog] Transformers in Vision [ ]
Link[Chinese Blog] 3W字长文带你轻松入门视觉transformer [ ]
Link[Chinese Blog] Vision Transformer 超详细解读 (原理分析+代码解读) [ ]

Awesome Visual-Transformer / Papers / Survey

paperMultimodal learning with transformers: A survey (IEEE TPAMI) [ ] - 2023.05.11
paperA Survey of Visual Transformers [ ] - 2021.11.30
paperTransformers in Vision: A Survey [ ] - 2021.02.22
paperA Survey on Visual Transformer [ ] - 2021.1.30
paperA Survey of Transformers [ ] - 2020.6.09

Awesome Visual-Transformer / Papers / arXiv papers

paperUnderstanding Gaussian Attention Bias of Vision Transformers Using Effective Receptive [ ]
paperFocused Decoding Enables 3D Anatomical Detection by Transformers [ ] [ ]
paperTAG: Boosting Text-VQA via Text-aware Visual Question-answer Generation [ ] [ ]
paperCross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers [ ] [ ]
paperBatchFormer: Learning to Explore Sample Relationships for Robust Representation Learning [ ] [ ]
[paper]RelViT: Concept-guided Vision Transformer for Visual Relational Reasoning
paperImproved Multiscale Vision Transformers for Classification and Detection [ ] [ ]
paperDETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection [ ] [ ]
paperThree things everyone should know about Vision Transformers [ ]
paperDeiT III: Revenge of the ViT [ ]
paperDaViT: Dual Attention Vision Transformers [ ] [ ]
paperCollaborative Transformers for Grounded Situation Recognition [ ] [ ]
paperGrounded Situation Recognition with Transformers [ ] [ ]
[paper]MaxViT: Multi-Axis Vision Transformer
[paper]V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer
paperUnsupervised Anomaly Detection in Medical Images with a Memory-augmented Multi-level Cross-attentional Masked Autoencoder [ ] [ ]
paperContrastive Transformer-based Multiple Instance Learning for Weakly Supervised Polyp Frame Detection [ ] [ ]
paperVideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training [ ] [ ]
paperPeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers [ ]
paperResViT: Residual vision transformers for multi-modal medical image synthesis [ ]
paperCombining EfficientNet and Vision Transformers for Video Deepfake Detection [ ] [ ]
paperDiscrete Representations Strengthen Vision Transformer Robustness [ ]
paperStyleSwin: Transformer-based GAN for High-resolution Image Generation [ ] [ ]
paperSliced Recursive Transformer [ ] [ ]
paperDynamic Token Normalization Improves Vision Transformer [ ]
paperTokenLearner: What Can 8 Learned Tokens Do for Images and Videos? [ ] [ ]
paperImproved Robustness of Vision Transformer via PreLayerNorm in Patch Embedding [ ]
paperObject-Region Video Transformers [ ] [ ]
paperAdaptively Multi-view and Temporal Fusing Transformer for 3D Human Pose Estimation [ ] [ ]
paperNViT: Vision Transformer Compression and Parameter Redistribution [ ]
paper6D-ViT: Category-Level 6D Object Pose Estimation via Transformer-based Instance Representation Learning [ ]
paperAdversarial Token Attacks on Vision Transformers [ ]
paperContextual Transformer Networks for Visual Recognition [ ] [ ]
paperTranSalNet: Visual saliency prediction using transformers [ ]
paperMobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer [ ]
paperA free lunch from ViT: Adaptive Attention Multi-scale Fusion Transformer for Fine-grained Visual Recognition [ ]
paper3D-Transformer: Molecular Representation with Transformer in 3D Space [ ]
paperCCTrans: Simplifying and Improving Crowd Counting with Transformer [ ]
paperUFO-ViT: High Performance Linear Vision Transformer without Softmax [ ]
paperSparse Spatial Transformers for Few-Shot Learning [ ]
paperVision Transformer Hashing for Image Retrieval [ ]
paperOH-Former: Omni-Relational High-Order Transformer for Person Re-Identification [ ]
paperPix2seq: A Language Modeling Framework for Object Detection [ ]
paperCoAtNet: Marrying Convolution and Attention for All Data Sizes [ ]
paperLOTR: Face Landmark Localization Using Localization Transformer [ ]
paperTransformer-Unet: Raw Image Processing with Unet [ ]
paperGraFormer: Graph Convolution Transformer for 3D Pose Estimation [ ]
paperCDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation [ ]
paperPQ-Transformer: Jointly Parsing 3D Objects and Layouts from Point Clouds [ ] [ ]
paperAnchor DETR: Query Design for Transformer-Based Detector [ ] [ ]
paperDAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR [ ] [ ]
paperEfficient Transformer for Single Image Super-Resolution [ ]
paperMaskFormer: Per-Pixel Classification is Not All You Need for Semantic Segmentation [ ] [ ]
paperSwinIR: Image Restoration Using Swin Transformer [ ] [ ]
paperTrans4Trans: Efficient Transformer for Transparent Object and Semantic Scene Segmentation in Real-World Navigation Assistance [ ]
paperDo Vision Transformers See Like Convolutional Neural Networks? [ ]
paperBoosting Salient Object Detection with Transformer-based Asymmetric Bilateral U-Net [ ]
paperLight Field Image Super-Resolution with Transformers [ ] [ ]
paperFocal Self-attention for Local-Global Interactions in Vision Transformers [ ] [ ]
paperPolyp-PVT: Polyp Segmentation with Pyramid Vision Transformers [ ] [ ]
paperMobile-Former: Bridging MobileNet and Transformer [ ]
paperTriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network [ ]
paperPSViT: Better Vision Transformer via Token Pooling and Attention Sharing [ ]
paperBoosting Few-shot Semantic Segmentation with Transformers [ ] [ ]
paperCongested Crowd Instance Localization with Dilated Convolutional Swin Transformer [ ]
paperEvo-ViT: Slow-Fast Token Evolution for Dynamic Vision Transformer [ ]
paperStyleformer: Transformer based Generative Adversarial Networks with Style Vector [ ] [ ]
paperCMT: Convolutional Neural Networks Meet Vision Transformers [ ]
paperTransAttUnet: Multi-level Attention-guided U-Net with Transformer for Medical Image Segmentation [ ]
paperTransClaw U-Net: Claw U-Net with Transformers for Medical Image Segmentation [ ]
paperViTGAN: Training GANs with Vision Transformers [ ]
paperWhat Makes for Hierarchical Vision Transformer? [ ]
paperTrans4Trans: Efficient Transformer for Transparent Object Segmentation to Help Visually Impaired People Navigate in the Real World [ ]
paperFeature Fusion Vision Transformer for Fine-Grained Visual Categorization [ ]
paperTransformerFusion: Monocular RGB Scene Reconstruction using Transformers [ ]
paperEscaping the Big Data Paradigm with Compact Transformers [ ]
paperHow to train your ViT? Data, Augmentation,and Regularization in Vision Transformers [ ]
paperBeyond Self-attention: External Attention using Two Linear Layers for Visual Tasks [ ]
paperXCiT: Cross-Covariance Image Transformers [ ] [ ]
paperShuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer [ ] [ ]
paperVideo Swin Transformer [ ] [ ]
paperVOLO: Vision Outlooker for Visual Recognition [ ] [ ]
paperTransformer Meets Convolution: A Bilateral Awareness Net-work for Semantic Segmentation of Very Fine Resolution Ur-ban Scene Images [ ]
paperEnd-to-end Temporal Action Detection with Transformer [ ] [ ]
paperHow to train your ViT? Data, Augmentation, and Regularization in Vision Transformers [ ]
paperEfficient Self-supervised Vision Transformers for Representation Learning [ ]
paperSpace-time Mixing Attention for Video Transformer [ ]
paperTransformed CNNs: recasting pre-trained convolutional layers with self-attention [ ]
paperCAT: Cross Attention in Vision Transformer [ ]
paperScaling Vision Transformers [ ]
paperDETReg: Unsupervised Pretraining with Region Priors for Object Detection [ ] [ ]
paperChasing Sparsity in Vision Transformers:An End-to-End Exploration [ ]
paperMViT: Mask Vision Transformer for Facial Expression Recognition in the wild [ ]
paperDemystifying Local Vision Transformer: Sparse Connectivity, Weight Sharing, and Dynamic Weight [ ]
paperOn Improving Adversarial Transferability of Vision Transformers [ ]
paperFully Transformer Networks for Semantic ImageSegmentation [ ]
paperVisual Transformer for Task-aware Active Learning [ ] [ ]
paperEfficient Training of Visual Transformers with Small-Size Datasets [ ]
paperReveal of Vision Transformers Robustness against Adversarial Attacks [ ]
paperPerson Re-Identification with a Locally Aware Transformer [ ]
paperRefiner: Refining Self-attention for Vision Transformers [ ]
paperViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias [ ]
paperVideo Instance Segmentation using Inter-Frame Communication Transformers [ ]
paperTransformer in Convolutional Neural Networks [ ] [ ]
paperUformer: A General U-Shaped Transformer for Image Restoration [ ] [ ]
paperPatch Slimming for Efficient Vision Transformers [ ]
paperRegionViT: Regional-to-Local Attention for Vision Transformers [ ]
paperAssociating Objects with Transformers for Video Object Segmentation [ ] [ ]
paperFew-Shot Segmentation via Cycle-Consistent Transformer [ ]
paperGlance-and-Gaze Vision Transformer [ ] [ ]
paperUnsupervised MRI Reconstruction via Zero-Shot Learned Adversarial Transformers [ ]
paperDynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification [ ] [ ]
paperWhen Vision Transformers Outperform ResNets without Pretraining or Strong Data Augmentations [ ] [ ]
paperUnsupervised Out-of-Domain Detection via Pre-trained Transformers [ ]
paperTransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classication [ ]
paperTransVOS: Video Object Segmentation with Transformers [ ]
paperKVT: k-NN Attention for Boosting Vision Transformers [ ]
paperMSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens [ ] [ ]
paperSegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers [ ] [ ]
paperSDNet: mutil-branch for single image deraining using swin [ ] [ ]
paperNot All Images are Worth 16x16 Words: Dynamic Vision Transformers with Adaptive Sequence Length [ ]
paperGaze Estimation using Transformer [ ] [ ]
paperTransformer-Based Deep Image Matching for Generalizable Person Re-identification [ ]
paperLess is More: Pay Less Attention in Vision Transformers [ ]
paperFoveaTer: Foveated Transformer for Image Classification [ ]
paperTransformer-Based Source-Free Domain Adaptation [ ] [ ]
paperAn Attention Free Transformer [ ]
paperPTNet: A High-Resolution Infant MRI Synthesizer Based on Transformer [ ]
paperResT: An Efficient Transformer for Visual Recognition [ ] [ ]
paperCogView: Mastering Text-to-Image Generation via Transformers [ ]
paperAggregating Nested Transformers [ ]
paperTemporal Action Proposal Generation with Transformers [ ]
paperBoosting Crowd Counting with Transformers [ ]
paperCOTR: Convolution in Transformer Network for End to End Polyp Detection [ ]
paperEnd-to-End Video Object Detection with Spatial-Temporal Transformers [ ] [ ]
paperIntriguing Properties of Vision Transformers [ ] [ ]
paperCombining Transformer Generators with Convolutional Discriminators [ ]
paperRethinking the Design Principles of Robust Vision Transformer [ ]
paperVision Transformers are Robust Learners [ ] [ ]
paperManipulation Detection in Satellite Images Using Vision Transformer [ ]
paperSwin-Unet: Unet-like Pure Transformer for Medical Image Segmentation [ ] [ ]
paperSelf-Supervised Learning with Swin Transformers [ ] [ ]
paperSCTN: Sparse Convolution-Transformer Network for Scene Flow Estimation [ ]
paperRelationTrack: Relation-aware Multiple Object Tracking with Decoupled Representation [ ]
paperVisual Grounding with Transformers [ ]
paperVisual Composite Set Detection Using Part-and-Sum Transformers [ ]
paperTrTr: Visual Tracking with Transformer [ ] [ ]
paperMOTR: End-to-End Multiple-Object Tracking with TRansformer [ ] [ ]
paperAttention for Image Registration (AiR): an unsupervised Transformer approach [ ]
paperTransHash: Transformer-based Hamming Hashing for Efficient Image Retrieval [ ]
paperISTR: End-to-End Instance Segmentation with Transformers [ ] [ ]
paperCAT: Cross-Attention Transformer for One-Shot Object Detection [ ]
paperCoSformer: Detecting Co-Salient Object with Transformers [ ]
paperEnd-to-End Attention-based Image Captioning [ ]
paperPyramid Medical Transformer for Medical Image Segmentation [ ]
paperHandsFormer: Keypoint Transformer for Monocular 3D Pose Estimation ofHands and Object in Interaction [ ]
paperGasHis-Transformer: A Multi-scale Visual Transformer Approach for Gastric Histopathology Image Classification [ ]
paperEmerging Properties in Self-Supervised Vision Transformers [ ]
paperInpainting Transformer for Anomaly Detection [ ]
paperTwins: Revisiting Spatial Attention Design in Vision Transformers [ ] [ ]
paperPoint Cloud Learning with Transformer [ ]
paperMedical Transformer: Universal Brain Encoder for 3D MRI Analysis [ ]
paperConTNet: Why not use convolution and transformer at the same time? [ ] [ ]
paperDual Transformer for Point Cloud Analysis [ ]
paperImprove Vision Transformers Training by Suppressing Over-smoothing [ ] [ ]
paperTransformer Meets DCFAM: A Novel Semantic Segmentation Scheme for Fine-Resolution Remote Sensing Images [ ]
paperM3DeTR: Multi-representation, Multi-scale, Mutual-relation 3D Object Detection with Transformers [ ] [ ]
paperSkeletor: Skeletal Transformers for Robust Body-Pose Estimation [ ]
paperLearning to Cluster Faces via Transformer [ ]
paperMultiscale Vision Transformers [ ] [ ]
paperVATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text [ ]
paperSo-ViT: Mind Visual Tokens for Vision Transformer [ ] [ ]
paperToken Labeling: Training a 85.5% Top-1 Accuracy Vision Transformer with 56M Parameters on ImageNet [ ] [ ]
paperTransRPPG: Remote Photoplethysmography Transformer for 3D Mask Face Presentation Attack Detection [ ]
paperVideoGPT: Video Generation using VQ-VAE and Transformers [ ]
paperM2TR: Multi-modal Multi-scale Transformers for Deepfake Detection [ ]
paperTransformer Transforms Salient Object Detection and Camouflaged Object Detection [ ]
paperTransCrowd: Weakly-Supervised Crowd Counting with Transformer [ ] [ ]
paperVisual Transformer Pruning [ ]
paperSelf-supervised Video Retrieval Transformer Network [ ]
paperVision Transformer using Low-level Chest X-ray Feature Corpus for COVID-19 Diagnosis and Severity Quantification [ ]
paperTransGAN: Two Transformers Can Make One Strong GAN [ ] [ ]
paperGeometry-Free View Synthesis: Transformers and no 3D Priors [ ] [ ]
paperCo-Scale Conv-Attentional Image Transformers [ ] [ ]
paperLocalViT: Bringing Locality to Vision Transformers [ ] [ ]
paperCloth Interactive Transformer for Virtual Try-On [ ] [ ]
paperHandwriting Transformers [ ]
paperSiT: Self-supervised vIsion Transformer [ ] [ ]
paperOn the Robustness of Vision Transformers to Adversarial Examples [ ]
paperAn Empirical Study of Training Self-Supervised Visual Transformers [ ]
paperA Video Is Worth Three Views: Trigeminal Transformers for Video-based Person Re-identification [ ]
paperAggregated Contextual Transformations for High-Resolution Image Inpainting [ ] [ ]
paperDeepfake Detection Scheme Based on Vision Transformer and Distillation [ ]
paperAugmented Transformer with Adaptive Graph for Temporal Action Proposal Generation [ ]
paperTubeR: Tube-Transformer for Action Detection [ ]
paperAAformer: Auto-Aligned Transformer for Person Re-Identification [ ]
paperTFill: Image Completion via a Transformer-Based Architecture [ ]
paperGroup-Free 3D Object Detection via Transformers [ ] [ ]
paperSpatial-Temporal Graph Transformer for Multiple Object Tracking [ ]
paperGoing deeper with Image Transformers[ ]
paperMeta-DETR: Few-Shot Object Detection via Unified Image-Level Meta-Learning [ [ ]
paperDA-DETR: Domain Adaptive Detection Transformer by Hybrid Attention [ ]
paperRobust Facial Expression Recognition with Convolutional Visual Transformers [ ]
paperThinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers [ ]
paperSpatiotemporal Transformer for Video-based Person Re-identification[ ]
paperTransUNet: Transformers Make Strong Encoders for Medical Image Segmentation [ ] [ ]
paperCvT: Introducing Convolutions to Vision Transformers [ ] [ ]
paperTFPose: Direct Human Pose Estimation with Transformers [ ]
paperTransCenter: Transformers with Dense Queries for Multiple-Object Tracking [ ]
paperFace Transformer for Recognition [ ]
paperOn the Adversarial Robustness of Visual Transformers [ ]
paperUnderstanding Robustness of Transformers for Image Classification [ ]
paperLifting Transformer for 3D Human Pose Estimation in Video [ ]
paperGlobal Self-Attention Networks for Image Recognition[ ]
paperHigh-Fidelity Pluralistic Image Completion with Transformers [ ] [ ]
paperVision Transformers for Dense Prediction [ ] [ ]
paperTransFG: A Transformer Architecture for Fine-grained Recognition? [ ]
paperIs Space-Time Attention All You Need for Video Understanding? [ ]
paperMulti-view 3D Reconstruction with Transformer [ ]
paperCan Vision Transformers Learn without Natural Images? [ ] [ ]
paperEnd-to-End Trainable Multi-Instance Pose Estimation with Transformers [ ]
paperInstance-level Image Retrieval using Reranking Transformers [ ] [ ]
paperBossNAS: Exploring Hybrid CNN-transformers with Block-wisely Self-supervised Neural Architecture Search [ ] [ ]
paperIncorporating Convolution Designs into Visual Transformers [ ]
paperDeepViT: Towards Deeper Vision Transformer [ ]
paperEnhancing Transformer for Video Understanding Using Gated Multi-Level Attention and Temporal Adversarial Training [ ]
paper3D Human Pose Estimation with Spatial and Temporal Transformers [ ] [ ]
paperSUNETR: Transformers for 3D Medical Image Segmentation [ ]
paperScalable Visual Transformers with Hierarchical Pooling [ ]
paperConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases [ ]
paperTransMed: Transformers Advance Multi-modal Medical Image Classification [ ]
paperU-Net Transformer: Self and Cross Attention for Medical Image Segmentation [ ]
paperSpecTr: Spectral Transformer for Hyperspectral Pathology Image Segmentation [ ] [ ]
paperTransBTS: Multimodal Brain Tumor Segmentation Using Transformer [ ] [ ]
paperSSTN: Self-Supervised Domain Adaptation Thermal Object Detection for Autonomous Driving [ ]
paperTransformer is All You Need: Multimodal Multitask Learning with a Unified Transformer [ ] [ ]
paperDo We Really Need Explicit Position Encodings for Vision Transformers? [ ] [ ]
paperDeepfake Video Detection Using Convolutional Vision Transformer[ ]
paperTraining Vision Transformers for Image Retrieval[ ]
paperVideo Transformer Network[ ]
paperBottleneck Transformers for Visual Recognition [ ]
paperCPTR: Full Transformer Network for Image Captioning [ ]
paperLearn to Dance with AIST++: Music Conditioned 3D Dance Generation [ ] [ ]
paperSegmenting Transparent Object in the Wild with Transformer [ ] [ ]
paperInvestigating the Vision Transformer Model for Image Retrieval Tasks [ ]
paperTrear: Transformer-based RGB-D Egocentric Action Recognition [ ]
paperVisualSparta: Sparse Transformer Fragment-level Matching for Large-scale Text-to-Image Search [ ]
paperTrackFormer: Multi-Object Tracking with Transformers [ ]
paperTransformer Guided Geometry Model for Flow-Based Unsupervised Visual Odometry [ ]
paperTransformer for Image Quality Assessment [ ] [ ]
paperTransTrack: Multiple-Object Tracking with Transformer [ ] [ ]
paperTraining data-efficient image transformers & distillation through attention [ ] [ ]
paper3D Object Detection with Pointformer [ ]
paperToward Transformer-Based Object Detection [ ]
paperTaming Transformers for High-Resolution Image Synthesis [ ] [ ]
paperSceneFormer: Indoor Scene Generation with Transformers [ ]
paperPCT: Point Cloud Transformer [ ]
paperDETR for Pedestrian Detection[ ]
paperTransformer Guided Geometry Model for Flow-Based Unsupervised Visual Odometry[ ]
paperGeneral Multi-label Image Classification with Transformers [ ]

Awesome Visual-Transformer / Papers / 2022

paperP2T: Pyramid Pooling Transformer for Scene Understanding [ ]
paperExpanding Language-Image Pretrained Models for General Video Recognition [ ] [ ]
paperTinyViT: Fast Pretraining Distillation for Small Vision Transformers [ ] [ ]
paperCross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers [ ] [ ]
paperAiATrack: Attention in Attention for Transformer Visual Tracking [ ] [ ]
paperJoint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework [ ] [ ]
paperTowards Grand Unification of Object Tracking [ ] [ ]
paperTracking Objects as Pixel-wise Distributions [ ] [ ]
paperMasked Autoencoders Are Scalable Vision Learners [ ]
paperCSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows [ ] [ ]
paperFast Point Transformer [ ]
paperEDTER: Edge Detection With Transformer [ ] [ ]
paperBridged Transformer for Vision and Point Cloud 3D Object Detection [ ]
paperMNSRNet: Multimodal Transformer Network for 3D Surface Super-Resolution [ ]
paperHyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening [ ] [ ]
paperKeypoint Transformer: Solving Joint Identification in Challenging Hands and Object Interactions for Accurate 3D Pose Estimation [ ]
paperMPViT: Multi-Path Vision Transformer for Dense Prediction [ ]
paperA-ViT: Adaptive Tokens for Efficient Vision Transformer [ ]
paperTopFormer: Token Pyramid Transformer for Mobile Semantic Segmentation [ ] [ ]
paperContinual Learning With Lifelong Vision Transformer [ ]
paperSwin Transformer V2: Scaling Up Capacity and Resolution [ ]
paperVoxel Set Transformer: A Set-to-Set Approach to 3D Object Detection From Point Clouds [ ] [ ]
paperMulti-Class Token Transformer for Weakly Supervised Semantic Segmentation [ ]
paperHuman-Object Interaction Detection via Disentangled Transformer [ ]
paperLGT-Net: Indoor Panoramic Room Layout Estimation With Geometry-Aware Transformer Network [ ]
paperSparse Local Patch Transformer for Robust Face Alignment and Landmarks Inherent Relation Learning [ ]
paperVision Transformer With Deformable Attention [ ]
paperDearKD: Data-Efficient Early Knowledge Distillation for Vision Transformers [ ]
paperRestormer: Efficient Transformer for High-Resolution Image Restoration [ ] [ ]
paperAccelerating DETR Convergence via Semantic-Aligned Matching [ ] [ ]
paperBEVT: BERT Pretraining of Video Transformers [ ] [ ]
paperMobile-Former: Bridging MobileNet and Transformer [ ]
paperSpatio-temporal Relation Modeling for Few-shot Action Recognition [ ] [ ]
paperMiniViT: Compressing Vision Transformers with Weight Multiplexing [ ] [ ]
paperCollaborative Transformers for Grounded Situation Recognition [ ] [ ]
paperBeyond Fixation: Dynamic Window Visual Transformer [ ] [ ]
paperMultimodal Token Fusion for Vision Transformers [ ]
paperConvolutional Neural Networks Meet Vision Transformers [ ]
paperFine-tuning Image Transformers using Learnable Memory [ ]
paperAttend to Mix for Vision Transformers [ ] [ ]
paperNominate Synergistic Context in Vision Transformer for Visual Recognition [ ] [ ]
paperShunted Self-Attention via Multi-Scale Token Aggregation [ ] [ ]
paperTowards Robust Vision Transformer [ [ ]
paperLite Vision Transformer with Enhanced Self-Attention [ [ ]
paperStyTr2: Image Style Transfer with Transformers [ ] [ ]
paperImage-Adaptive Hint Generation via Vision Transformer for Outpainting [ ] [ ]

Awesome Visual-Transformer / Papers / 2021

paperProTo: Program-Guided Transformer for Program-Guided Tasks [ ] [ ]
paperAugmented Shortcuts for Vision Transformers [ ] [ ]
paperYou Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection [ ] [ ]
paperSemantic Correspondence with Transformers [ ] [ ]
paperQVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries [ ] [ ]
paperDual-stream Network for Visual Recognition [ ] [ ]
paperContainer: Context Aggregation Network [ ] [ ]
paperTransformer in Transformer [ ] [ ]
paperT6D-Direct: Transformers for Multi-Object 6D Pose Direct Regression [ ]
paperLong Short-Term Transformer for Online Action Detection [ ]
paperTransformerFusion: Monocular RGB Scene Reconstruction using Transformers [ ]
paperTransMatcher: Deep Image Matching Through Transformers for Generalizable Person Re-identification [ ]
paperTransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification [ ]
paperAssociating Objects with Transformers for Video Object Segmentation [ ]
paperTest-Time Personalization with a Transformer for Human Pose Estimation [ ]
paperRevitalizing CNN Attention via Transformers in Self-Supervised Visual Representation Learning [ ]
paperDynamic Grained Encoder for Vision Transformers [ ]
paperHRFormer: High-Resolution Vision Transformer for Dense Predict [ ]
paperSearching the Search Space of Vision Transformer [ ]
paperNot All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition [ ]
paperSegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers [ ]
paperDo Vision Transformers See Like Convolutional Neural Networks? [ ]
paperKeeping Your Eye on the Ball: Trajectory Attention in Video Transformers [ ]
paperGlance-and-Gaze Vision Transformer [ ]
paperMST: Masked Self-Supervised Transformer for Visual Representation [ ]
paperDynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification [ ]
paperTransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up [ ]
paperAugmented Shortcuts for Vision Transformers [ ]
paperImproved Transformer for High-Resolution GANs [ ]
paperAll Tokens Matter: Token Labeling for Training Better Vision Transformers [ ]
paperXCiT: Cross-Covariance Image Transformers [ ]
paperEfficient Training of Visual Transformers with Small Datasets [ ]
paperSwin Transformer: Hierarchical Vision Transformer using Shifted Windows ( ) [ ] [ ]
paperHigh-Fidelity Pluralistic Image Completion with Transformers [ ] [ ]
paperPoinTr: Diverse Point Cloud Completion with Geometry-Aware Transformers ( ) [ ] [ ]
paperRevisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with Transformers [ ] [ ]
paperRethinking Transformer-based Set Prediction for Object Detection [ ]
paperPaint Transformer: Feed Forward Neural Painting with Stroke Prediction ( ) ) [ [ ]
paper3DVG-Transformer: Relation Modeling for Visual Grounding on Point Clouds [ ]
paperTraining Vision Transformers from Scratch on ImageNet [ ] [ ]
paperTHUNDR: Transformer-Based 3D Human Reconstruction With Markers [ ]
paperMulti-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding [ ]
paperPyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions [ ] [ ]
paperSpatial-Temporal Transformer for Dynamic Scene Graph Generation [ ]
paperGLiT: Neural Architecture Search for Global and Local Image Transformer [ ]
paperTRAR: Routing the Attention Spans in Transformer for Visual Question Answering [ ]
paperUniT: Multimodal Multitask Learning With a Unified Transformer [ ] [ ]
paperStochastic Transformer Networks With Linear Competing Units: Application To End-to-End SL Translation [ ]
paperTransformer-Based Dual Relation Graph for Multi-Label Image Recognition [ ]
paperLocalTrans: A Multiscale Local Transformer Network for Cross-Resolution Homography Estimation [ ]
paperImproving 3D Object Detection With Channel-Wise Transformer [ ]
paperA Latent Transformer for Disentangled Face Editing in Images and Videos [ ] [ ]
paperGroupFormer: Group Activity Recognition With Clustered Spatial-Temporal Transformer [ ]
paperUnified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual Dialogue [ ]
paperWB-DETR: Transformer-Based Detector Without Backbone [ ]
paperThe Animation Transformer: Visual Correspondence via Segment Matching [ ]
paperThe Animation Transformer: Visual Correspondence via Segment Matching [ ]
paperRelaxed Transformer Decoders for Direct Action Proposal Generation [ ]
paperPyramid Point Cloud Transformer for Large-Scale Place Recognition [ ] [ ]
paperMultimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide Images [ ]
paperUncertainty-Guided Transformer Reasoning for Camouflaged Object Detection [ ]
paperImage Harmonization With Transformer [ ] [ ]
paperCOTR: Correspondence Transformer for Matching Across Images [ ]
paperMUSIQ: Multi-Scale Image Quality Transformer [ ]
paperEpisodic Transformer for Vision-and-Language Navigation [ ]
paperAction-Conditioned 3D Human Motion Synthesis With Transformer VAE [ ]
paperCrackFormer: Transformer Network for Fine-Grained Crack Detection [ ]
paperHiT: Hierarchical Transformer With Momentum Contrast for Video-Text Retrieval [ ]
paperEvent-Based Video Reconstruction Using Transformer [ ]
paperSTVGBert: A Visual-Linguistic Transformer Based Framework for Spatio-Temporal Video Grounding [ ]
paperHiFT: Hierarchical Feature Transformer for Aerial Tracking [ ] [ ]
paperDocFormer: End-to-End Transformer for Document Understanding [ ]
paperLeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference [ ] [ ]
paperSignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition[ ]
paperVidTr: Video Transformer Without Convolutions [ ]
paperAction-Conditioned 3D Human Motion Synthesis with Transformer VAE [ ]
paperSegmenter: Transformer for Semantic Segmentation [ ] [ ]
paperVisformer: The Vision-friendly Transformer [ ] [ ]
paperPnP-DETR: Towards Efficient Visual Analysis with Transformers ( ) [ ] [ ]
paper[ ] Voxel Transformer for 3D Object Detection [ ]
paperTransVG: End-to-End Visual Grounding with Transformers [ ]
paperAn End-to-End Transformer Model for 3D Object Detection [ ] [ ]
paperEformer: Edge Enhancement based Transformer for Medical Image Denoising [ ]
paperTransFER: Learning Relation-aware Facial Expression Representations with Transformers [ ]
paperOriented Object Detection with Transformer [ ]
paperViViT: A Video Vision Transformer [ ]
paperLearning Spatio-Temporal Transformer for Visual Tracking [ ] [ ]
paperImproving 3D Object Detection with Channel-wise Transformer [ ]
paperVisual Saliency Transformer [ ]
paperRethinking Spatial Dimensions of Vision Transformers [ ] [ ]
paperCrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification [ ] [ ]
paperPoint Transformer [ ]
paperTS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object Localization [ ] [ ]
paperVisual Transformers: Token-based Image Representation and Processing for Computer Vision [ ]
paperTransformer-Based Attention Networks for Continuous Pixel-Wise Prediction [ ] [ ]
paperConditional DETR for Fast Training Convergence [ ] [ ]
paperPIT: Position-Invariant Transform for Cross-FoV Domain Adaptation [ ] [ ]
paperSOTR: Segmenting Objects with Transformers [ ] [ ]
paperSnowflakeNet: Point Cloud Completion by Snowflake Point Deconvolution with Skip-Transformer [ ] [ ]
paperTransPose: Keypoint Localization via Transformer [ ] [ ]
paperTransReID: Transformer-based Object Re-Identification [ ] [ ]
paperSimpler is Better: Few-shot Semantic Segmentation with Classifier Weight Transformer [ ] [ ]
paperAnticipative Video Transformer [ ] [ ]
paperRethinking and Improving Relative Position Encoding for Vision Transformer [ ] [ ]
paperVision Transformer with Progressive Sampling [ ] [ ]
paperFast Convergence of DETR with Spatially Modulated Co-Attention [ ] [ ]
paperAutoFormer: Searching Transformers for Visual Recognition [ ] [ ]
paperDiverse Part Discovery: Occluded Person Re-identification with Part-Aware Transformer [ ]
paperHOTR: End-to-End Human-Object Interaction Detection with Transformers ( ) [ ]
paperEnd-to-End Human Pose and Mesh Reconstruction with Transformers [ ]
paperLine Segment Detection Using Transformers without Edges [ ]
paperMulti-Modal Fusion Transformer for End-to-End Autonomous Driving [ ] [ ]
paperPose Recognition with Cascade Transformers [ ]
paperSeeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning [ ]
paperLoFTR: Detector-Free Local Feature Matching with Transformers [ ] [ ]
paperThinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers [ ]
paperRethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers [ ] [ ]
paperTransformer Tracking [ ] [ ]
paperTransformer Meets Tracker: Exploiting Temporal Context for Robust Visual Tracking (** oral**) [ ]
paperEnd-to-End Video Instance Segmentation with Transformers [ ]
paperTransformer Interpretability Beyond Attention Visualization [ ] [ ]
paperPre-Trained Image Processing Transformer [ ]
paperUP-DETR: Unsupervised Pre-training for Object Detection with Transformers [ ]
paperPerceptual Image Quality Assessment with Transformers ( ) [ ]
paperHigh-Resolution Complex Scene Synthesis with Transformers ( ) [ ]
paperCollaborative Transformers for Grounded Situation Recognition [ ] [ ]
paperGenerative Video Transformer: Can Objects be the Words? [ ]
paperGenerative Adversarial Transformers [ ] [ ]
paperNDT-Transformer: Large-Scale 3D Point Cloud Localisation using the Normal Distribution Transform Representation [ ]
paperVTNet: Visual Transformer Network for Object Goal Navigation [ ]
paperAn Image is Worth 16x16 Words: Transformers for Image Recognition at Scale [ ] [ ]
paperDeformable DETR: Deformable Transformers for End-to-End Object Detection [ ] [ ]
paperMODELING LONG-RANGE INTERACTIONS WITHOUT ATTENTION [ ] [ ]
paperVideo Transformer for Deepfake Detection with Incremental Learning[ ]
paperHAT: Hierarchical Aggregation Transformers for Person Re-identification [ ]
paperToken Shift Transformer for Video Classification [ ] [ ]
paperDPT: Deformable Patch-based Transformer for Visual Recognition [ ] [ ]
paperUTNet: A Hybrid Transformer Architecture for Medical Image Segmentation [ ] [ ]
paperMedical Transformer: Gated Axial-Attention for Medical Image Segmentation [ ] [ ]
paperMulti-Compound Transformer for Accurate Biomedical Image Segmentation [ ] [ ]
paperProgressively Normalized Self-Attention Network for Video Polyp Segmentation [ ] [ ]
paperA Multi-Branch Hybrid Transformer Networkfor Corneal Endothelial Cell Segmentation [ ]
paperEnd-to-End Object Detection with Adaptive Clustering Transformer [ ]
paperGrounded Situation Recognition with Transformers [ ] [ ]
paperTransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation [ ] [ ]
paperVT-ADL: A Vision Transformer Network for Image Anomaly Detection and Localization ( ) [ ]
paperDETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries [ ]
paperMedical Image Segmentation using Squeeze-and-Expansion Transformers [ ]
paperYou Only Group Once: Efficient Point-Cloud Processing with Token Representation and Relation Inference Module ( ) [ ] [ ]
paperPTT: Point-Track-Transformer Module for 3D Single Object Tracking in Point Clouds [ ] [ ]
paperEnd-to-end Lane Shape Prediction with Transformers [ ] [ ]
paperVision Transformer for Fast and Efficient Scene Text Recognition [ ]

Awesome Visual-Transformer / Papers / 2020

paperEnd-to-End Object Detection with Transformers ( ) [ ] [ ]
paper[ ] Feature Pyramid Transformer ( ) [ ] [ ]

Awesome Visual-Transformer / Papers / Other resource

Awesome-Transformer-Attention4,679about 2 years ago[ ]

Backlinks from these awesome lists:

More related projects: