A training method using human feedback to make AI answers fit people’s expectations.
RLHF is like sending AI to charm school. Humans mark “nice answer” or “yikes, try again.” It slowly learns manners.
It helps an LLM act more helpful and less wild. It can also make answers too smooth.
Alignment
RLHF uses human feedback to move the model toward Alignment goals.
Fine-tuning
RLHF often happens during Fine-tuning to tune model preferences.
LLM
After RLHF, an LLM is more likely to answer in a way people expect.