Accountable-Textual-Visual-Chat
by matrix-alpha
The official repository for Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation.
AI summary
Instruction rejection model
Develops accountability in image generation models by learning to reject human instructions
- stars
- 7
- forks
- 2
- watching
- 1
Similar projects
Found by comparing what the projects do, not just their names.
Attention model training
Trains a bottom-up attention model using Faster R-CNN and Visual Genome annotations for image captioning and VQA tasks
Visual QA Model
This project presents a neural network model designed to answer visual questions by combining question and image features in a residual learning framework.
Instruction generator
Creating synthetic visual reasoning instructions to improve the performance of large language models on image-related tasks
Visual Instruction Toolkit
A method and toolkit for fine-tuning large language models to perform visual instruction tasks in multiple languages.
aidc-ai/ovis575
Multimodal aligner
An MLLM architecture designed to align visual and textual embeddings through structural alignment
Instruction dataset generator
Autonomously generates high-quality image-text instruction fine-tuning datasets
LLM trainer
Transfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs
Image Captioning Model
Adaptive attention mechanism for image captioning using visual sentinels
eric-xw/arel136
Story generator
This codebase provides an implementation of a novel adversarial reward learning algorithm for generating human-like visual stories from image sequences.
Model alignment
This project demonstrates the effectiveness of reinforcement learning from human feedback (RLHF) in improving small language models like GPT-2.
VQA model
A Visual Question Answering model using a deeper LSTM and normalized CNN architecture.
Visual QA Model
Develops a deep learning model to answer questions about visual scenes based on spatial attention and question guidance
Multimodal model trainer
An implementation of a multimodal LLM training paradigm to enhance truthfulness and ethics in language models
Visual Prompt Model
A system designed to enable large multimodal models to understand arbitrary visual prompts
Text generator
Enables automatic generation of descriptive text from images and videos based on user input.