minChatGPT
A minimum example of aligning language models with RLHF similar to ChatGPT
AI summary
Model alignment
This project demonstrates the effectiveness of reinforcement learning from human feedback (RLHF) in improving small language models like GPT-2.
- stars
- 214
- forks
- 28
- watching
- 5
Similar projects
Found by comparing what the projects do, not just their names.
Behavior alignment tool
Aligns large language models' behavior through fine-grained correctional human feedback to improve trustworthiness and accuracy.
Multimodal alignment model
Extending pretraining models to handle multiple modalities by aligning language and video representations
Reward alignment system
Aligns large multimodal models with factually enhanced reward functions to improve performance and mitigate hacking in reinforcement learning
language model tuning
Training methods and tools for fine-tuning language models using human preferences
Model aligner
Aligns large multimodal models with human intentions and values using various algorithms and fine-tuning methods.
Region-of-Interest Training
Training and deploying large language models on computer vision tasks using region-of-interest inputs
Reinforcement Learning framework
Replication of Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks in PyTorch for reinforcement learning tasks
Chinese language model
Trains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks
RLHF framework
A framework for training language models using human feedback and reinforcement learning
multi-agent learning protocol
This project implements a PyTorch-based framework for learning discrete communication protocols in multi-agent reinforcement learning environments.
Chinese language model fine-tuning tool
Improves pre-trained Chinese language models by incorporating a correction task to alleviate inconsistency issues with downstream tasks
RL framework
A framework for parallel population-based reinforcement learning
Hyperparameter optimizer
A reinforcement learning-based framework for optimizing hyperparameters in distributed machine learning environments.
Medical chatbot
Develops large language models to support medical diagnoses and provide helpful suggestions
Model value alignment
Evaluates and aligns the values of Chinese large language models with safety and responsibility standards