GroundedTranslation
by elliottd
Multilingual image description
AI summary
Image Description Model
Trains multilingual image description models using neural sequence models and extracts hidden features from trained models.
- stars
- 46
- forks
- 25
- watching
- 8
Similar projects
Found by comparing what the projects do, not just their names.
Image embedding
An implementation of a deep learning-based image representation learning approach using a modified fully connected layer and transfer learning from VGG16
Model distiller
Automatically trains models from large foundation models to perform specific tasks with minimal human intervention.
Image translator
A framework for unsupervised image-to-image translation using Generative Adversarial Networks (GANs) and deep learning.
Image translator
An unsupervised deep learning framework for translating images between different modalities
Vision-Language Learning Model
Develops and trains models for vision-language learning with decoupled language pre-training
Federated Image Segmentation Framework
This project presents a framework for federated domain generalization in medical image segmentation using continuous frequency space and episodic learning.
microsoft/som1.2K
Image marking tool
Enables visual grounding in large language models by overlaying spatial and speakable marks on images
Image segmentation model
An open-source implementation of an image segmentation model that combines background removal and object detection capabilities.
ELMo model variant
Efficient Contextual Representation Learning Model with Continuous Outputs
ML image recognition model
An implementation of a multimodal learning approach to improve language models' ability to recognize unseen images and understand novel concepts.
tobypde/frrn280
Image segmentation framework
A software framework for training and evaluating full-resolution residual networks for semantic image segmentation tasks
Segmentation model
A deep learning implementation of an object segmentation algorithm.
Multimodal model framework
A framework for grounding language models to images and handling multimodal inputs and outputs
Multimodal conversational model
An end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.