LLaVA-RLHF

Reward alignment system

Aligns large multimodal models with factually enhanced reward functions to improve performance and mitigate hacking in reinforcement learning

Aligning LMMs with Factually Augmented RLHF

GitHub

328 stars
9 watching
24 forks
Language: Python
last commit: almost 3 years ago

Related projects:

RepositoryDescriptionStars
rlhf-v/rlhf-vAligns large language models' behavior through fine-grained correctional human feedback to improve trustworthiness and accuracy.245
ethanyanjiali/minchatgptThis project demonstrates the effectiveness of reinforcement learning from human feedback (RLHF) in improving small language models like GPT-2.214
tristandeleu/pytorch-maml-rlReplication of Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks in PyTorch for reinforcement learning tasks830
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
sjtu-marl/malibA framework for parallel population-based reinforcement learning507
llava-vl/llava-interactive-demoAn all-in-one demo for interactive image processing and generation353
llava-vl/llava-plus-codebaseA platform for training and deploying large language and vision models that can use tools to perform tasks717
tatsu-lab/alpaca_farmA framework for simulating and evaluating reinforcement learning from human feedback methods786
matthiasplappert/keras-rlA Python library implementing state-of-the-art deep reinforcement learning algorithms for Keras and OpenAI Gym environments.8
kaixhin/rainbowA Python implementation of a deep reinforcement learning algorithm combining multiple techniques for improved performance in Atari games1,591
salt-nlp/llavarAn open-source project that enhances visual instruction tuning for text-rich image understanding by integrating GPT-4 models with multimodal datasets.259
aidc-ai/ovisAn MLLM architecture designed to align visual and textual embeddings through structural alignment575
iffix/machinAn open-source reinforcement learning library for PyTorch, providing a simple and clear implementation of various algorithms.402
yfzhang114/llava-alignDebiasing techniques to minimize hallucinations in large visual language models75
lhfowl/robbing_the_fedThis implementation allows an attacker to directly obtain user data from federated learning gradient updates by modifying the shared model architecture.23