w2vv
Text-to-image predictor
A deep neural network architecture that predicts visual features from text to improve image and video caption retrieval
Word2VisualVec : Predicting Visual Features from Text for Image and Video Caption Retrieval
69 stars
4 watching
13 forks
Language: Python
last commit: over 6 years agoLinked from 1 awesome list
Related projects:
| Repository | Description | Stars |
|---|---|---|
| A deep learning project that provides a video-text retrieval model and tools for training and evaluating it on the MSR-VTT dataset | 154 | |
| This project proposes a solution to predict salient areas in images using convolutional neural networks. | 186 | |
| A deep learning framework for scene text recognition with rectification and attention mechanisms. | 639 | |
| An implementation of a deep learning model for predicting chemical properties from molecular structure data | 30 | |
| A deep learning-based video search system using pre-trained models and datasets | 28 | |
| Develops a deep learning framework for video retrieval using text and computer vision | 87 | |
| This project aims to predict motion and perceive scenes from bird's eye view maps for autonomous driving | 169 | |
| An implementation of Deepmind's Visual Interaction Networks using PyTorch to predict future events in physical scenes. | 166 | |
| This project enables users to generate images using convolutional neural networks (CNNs) and visualize their activations. | 499 | |
| A Python framework for deep learning-based drug-target interaction prediction using a DBN architecture. | 49 | |
| Predicts and improves visual realism in composite images using deep learning techniques | 64 | |
| This implementation allows users to generate captions from images using a neural network model with visual attention. | 790 | |
| A Python script for training and testing deep learning models to predict drug-target interactions | 73 | |
| A deep learning-based system for generating precise grasp poses for robots in real-time | 525 | |
| A deep learning framework for generating natural language descriptions of images by detecting objects and their attributes | 1,584 |