RelViT
by NVlabs
[ICLR 2022] RelViT: Concept-guided Vision Transformer for Visual Relational Reasoning
AI summary
Visual reasoning tool
A deep learning framework designed to improve visual reasoning capabilities by utilizing concepts and semantic relations.
- stars
- 64
- forks
- 3
- watching
- 6
Similar projects
Found by comparing what the projects do, not just their names.
Visual Reasoning Benchmark
A benchmarking tool and software framework for evaluating few-shot visual reasoning capabilities in computer vision models.
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
Visual Reasoning Model
An open-source implementation of a deep learning model designed to improve the balance between performance and interpretability in visual reasoning tasks.
Visual representation generator
A project that generates visual representations tailored for general visual reasoning, leveraging hierarchical scene descriptions and instance-level world knowledge.
jy0205/lavit544
Visual understanding and generation framework
A unified framework for training large language models to understand and generate visual content
rowanz/r2c466
Visual Reasoning Model
An open-source project providing PyTorch code and data for a deep learning model that enables visual commonsense reasoning.
Visual detection model
An implementation of reinforcement learning for visual relationship and attribute detection using PyTorch.
Instruction generator
Creating synthetic visual reasoning instructions to improve the performance of large language models on image-related tasks
Visual encoder
A vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.
Label correction tool
Develops a method to create high-quality training data from noisy labels in semantic segmentation tasks.
Visual descriptor learner
A system for learning deep representations of fine-grained visual descriptions from images
Visual QA Model
This project presents a neural network model designed to answer visual questions by combining question and image features in a residual learning framework.
Behavior alignment tool
Aligns large language models' behavior through fine-grained correctional human feedback to improve trustworthiness and accuracy.
Visual QA framework
A framework for training Hierarchical Co-Attention models for Visual Question Answering using preprocessed data and a specific image model.
productivity tool
A customizable, high-contrast Nvim color scheme designed to enhance coding focus and productivity