IMAD
[AINL 2023] IMAD: IMage Augmented multi-modal Dialogue
AI summary
Dialogue analyzer
A toolkit for analyzing and generating multi-modal dialogue with images
- stars
- 4
- forks
- 0
- watching
- 1
Similar projects
Found by comparing what the projects do, not just their names.
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
Robot learner
An implementation of a general-purpose robot learning model using multimodal prompts
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
bfelbo/deepmoji1.5K
Text analyzer
A deep learning model for analyzing sentiment and emotion in text based on emojis.
Ada editor
Provides tools and features to support development in the Ada programming language within Vim/NeoVim text editors.
Modal dialog library
An Ember addon for building modal dialogs using a consistent pattern and layout approach.
Image processor
An all-in-one demo for interactive image processing and generation
Video Analysis Toolkit
A collection of resources and tools for video analysis using deep learning and multi-modal learning techniques.
Multimodal alignment model
Extending pretraining models to handle multiple modalities by aligning language and video representations
Document conversational AI
An application that enables users to upload documents and converse with an AI-powered language model.
Image aesthetic analyzer
A deep learning-based framework for image aesthetics assessment using a convolutional neural network structure
Dialogue model
Develops multimodal instruction-following models for open-ended dialogues across multiple images
vcciv/blvd171
Autonomous driving dataset
A large-scale 5D semantics benchmark for autonomous driving
DL framework
A collection of modular deep learning components that can be easily configured and reused in various applications.
Conversational image interface
An intelligent system that enables automatic control and utilization of visual foundation models to interact with images in conversational settings.