minChatGPT

Model alignment

This project demonstrates the effectiveness of reinforcement learning from human feedback (RLHF) in improving small language models like GPT-2.

A minimum example of aligning language models with RLHF similar to ChatGPT

GitHub

214 stars
5 watching
28 forks
Language: Python
last commit: about 3 years ago

Related projects:

RepositoryDescriptionStars
rlhf-v/rlhf-vAligns large language models' behavior through fine-grained correctional human feedback to improve trustworthiness and accuracy.245
pku-yuangroup/languagebindExtending pretraining models to handle multiple modalities by aligning language and video representations751
llava-rlhf/llava-rlhfAligns large multimodal models with factually enhanced reward functions to improve performance and mitigate hacking in reinforcement learning328
openai/lm-human-preferencesTraining methods and tools for fine-tuning language models using human preferences1,240
pku-alignment/align-anythingAligns large multimodal models with human intentions and values using various algorithms and fine-tuning methods.270
jshilong/gpt4roiTraining and deploying large language models on computer vision tasks using region-of-interest inputs517
tristandeleu/pytorch-maml-rlReplication of Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks in PyTorch for reinforcement learning tasks830
brightmart/xlnet_zhTrains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks230
xrsrke/instructgooseA framework for training language models using human feedback and reinforcement learning171
minqi/learning-to-communicate-pytorchThis project implements a PyTorch-based framework for learning discrete communication protocols in multi-agent reinforcement learning environments.349
ymcui/macbertImproves pre-trained Chinese language models by incorporating a correction task to alleviate inconsistency issues with downstream tasks646
sjtu-marl/malibA framework for parallel population-based reinforcement learning507
guopengf/auto-fedrlA reinforcement learning-based framework for optimizing hyperparameters in distributed machine learning environments.15
wangrongsheng/ivygptDevelops large language models to support medical diagnoses and provide helpful suggestions59
x-plug/cvaluesEvaluates and aligns the values of Chinese large language models with safety and responsibility standards481