Training a Wikipedia Manual of Style Assistant with RLHF
How RLHF Could Train an AI Assistant to Follow Wikipedia’s Manual of Style Reinforcement Learning from Human Feedback (RLHF) is a post-training pipeline that uses human judgements to shape a pre-trained language model’s behaviour. Starting from a base large language model (LLM), RLHF can be used to train the model into a helpful assistant that…