AI Rookies

RLHF — Reinforcement Learning from Human Feedback

Fact

A training method using human feedback to make AI answers fit people’s expectations.

In Plain Words

RLHF is like sending AI to charm school. Humans mark “nice answer” or “yikes, try again.” It slowly learns manners.

It helps an LLM act more helpful and less wild. It can also make answers too smooth.

Related Concepts

Alignment
RLHF uses human feedback to move the model toward Alignment goals.

Fine-tuning
RLHF often happens during Fine-tuning to tune model preferences.

LLM
After RLHF, an LLM is more likely to answer in a way people expect.