4 / papers

Alignment & RLHF

What training on human preference does to a model over time, and what it quietly costs. 4 open-access papers by Pranay Mahendrakar, each with a permanent DOI.

article

Four Things Called Forgetting: Interference, Transience, Reset and Unlearning Remove Different Things, Why Forgetting Looks Easy by Accident and Hard on Purpose Only Under Different Instruments, and the Savings Measurement That Would Tell Them Apart

Machine learning uses one word for four operations. Catastrophic forgetting is damage that fine-tuning does to earlier capabilities. Transience is the fading of individual training examples during ordinary training. Resets deliberately…

article

Consensus Too Soon, or Agreement From the Start? Shared Prior, Social Coupling and Pool Coverage in Decentralised LLM Collectives, and Why Prompted Diversity and Model Heterogeneity Act on Different Terms

Groups of language-model agents that exchange answers and settle on a common one are now a standard way to build decentralised decision systems. A common worry is that they agree too soon: agents copy each other, diversity collapses, and…

About this research line

What has Pranay Mahendrakar published on Alignment & RLHF?

Pranay Mahendrakar has published 4 open-access papers on Alignment & RLHF: "Four Things Called Forgetting", "Consensus Too Soon, or Agreement From the Start? Shared Prior, Social Coupling and Pool Coverage in Decentralised LLM Collectives, and Why Prompted Diversity and Model Heterogeneity Act on Different Terms", "Three Guarantees Under One Word" and "Catastrophic Forgetting in Continual RLHF". Each is deposited on Zenodo with a permanent DOI and is free to read.

← All research topics