1 / paper

Alignment & RLHF

What training on human preference does to a model over time, and what it quietly costs.

← All topics