GroundingDINO

Open-world detector

An implementation of an object detection model designed to work in open-world scenarios with the ability to detect and recognize objects based on language descriptions.

[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"

GitHub

7k stars
43 watching
711 forks
Language: Python
last commit: about 2 years ago
Linked from 1 awesome list

object-detectionopen-worldopen-world-detectionvision-languagevision-language-transformer

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
idea-research/dinoAn implementation of a deep learning-based object detection model with improved anchor boxes for end-to-end detection tasks.2,295
facebookresearch/dinov2A PyTorch implementation of a self-supervised learning method for learning robust visual features without supervision.9,425
theshadow29/zsgnet-pytorchAn implementation of a computer vision model that grounds objects in images using natural language queries.69
jhcho99/coformerAn implementation of a deep learning model for grounding situation recognition in images45
tencentarc/gfpganAn algorithm for restoring damaged or obscured faces in images36,009
amdegroot/ssd.pytorchAn implementation of a deep learning-based object detection system in PyTorch.5,160
cszn/kairImage restoration toolbox with training and testing codes for various deep learning-based methods2,994
doubiiu/dynamicrafterThis project generates animated videos from open-domain images by leveraging pre-trained video diffusion priors.2,668
junyanz/interactive-deep-colorizationA system for automatically colorizing black and white images with user interactions.2,701
thu-mig/yolov10Real-time object detection using a neural network architecture10,116
huawei-noah/efficient-ai-backbonesA collection of efficient AI backbone architectures developed by Huawei Noah's Ark Lab.4,098
layumi/person_reid_baseline_pytorchA PyTorch implementation of an Object Re-ID baseline with various training methods and architectures4,149
roboflow/notebooksThis repository contains tutorials and examples on using state-of-the-art computer vision models and techniques5,678
devendrachaplot/deeprl-groundingTrains an RL agent to execute natural language instructions in a 3D environment using a combination of A3C and gated attention mechanisms.237
mlfoundations/open_flamingoA framework for training large multimodal models to generate text conditioned on images or other text.3,781