LLaVA-RLHF
by llava-rlhf
Aligning LMMs with Factually Augmented RLHF
AI summary
Reward alignment system
Aligns large multimodal models with factually enhanced reward functions to improve performance and mitigate hacking in reinforcement learning
- stars
- 328
- forks
- 24
- watching
- 9
Similar projects
Found by comparing what the projects do, not just their names.
Behavior alignment tool
Aligns large language models' behavior through fine-grained correctional human feedback to improve trustworthiness and accuracy.
Model alignment
This project demonstrates the effectiveness of reinforcement learning from human feedback (RLHF) in improving small language models like GPT-2.
Reinforcement Learning framework
Replication of Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks in PyTorch for reinforcement learning tasks
Visual Prompt Model
A system designed to enable large multimodal models to understand arbitrary visual prompts
RL framework
A framework for parallel population-based reinforcement learning
Image processor
An all-in-one demo for interactive image processing and generation
Model trainer
A platform for training and deploying large language and vision models that can use tools to perform tasks
RL simulator
A framework for simulating and evaluating reinforcement learning from human feedback methods
RL framework
A Python library implementing state-of-the-art deep reinforcement learning algorithms for Keras and OpenAI Gym environments.
kaixhin/rainbow1.6K
RL framework
A Python implementation of a deep reinforcement learning algorithm combining multiple techniques for improved performance in Atari games
Visual Instruction Tuning
An open-source project that enhances visual instruction tuning for text-rich image understanding by integrating GPT-4 models with multimodal datasets.
aidc-ai/ovis575
Multimodal aligner
An MLLM architecture designed to align visual and textual embeddings through structural alignment
iffix/machin402
RL framework
An open-source reinforcement learning library for PyTorch, providing a simple and clear implementation of various algorithms.
Model debiasing
Debiasing techniques to minimize hallucinations in large visual language models
Model inversion attack
This implementation allows an attacker to directly obtain user data from federated learning gradient updates by modifying the shared model architecture.