RLHF (Reinforcement Learning from Human Feedback)
Fine-tuning models using human preference signals to make their outputs more helpful, harmless, and aligned.
Fine-tuning models using human preference signals to make their outputs more helpful, harmless, and aligned.