Contains some AI-generated content

Coherent extrapolated volition en.wikipedia.org

Background reading on a term that keeps surfacing in alignment arguments. Eliezer Yudkowsky proposed CEV in 2004: rather than giving an AI our current preferences, which are uninformed and often contradictory, you give it what we would want if we knew more, thought faster, were more the people we wished we were, and had grown up further together.

The appeal is that it avoids freezing today’s moral beliefs into a system that will outlast them, and reduces how much the programmers’ own values leak in. The criticisms are the obvious ones once stated: whose volition gets extrapolated, how to treat those who cannot participate, what happens to animals and digital minds if the base is only humans, and the fact that the framework has enough free parameters to produce very different results depending on how you set them. Useful context for alignment as a thought-terminating cliche, which is arguing against roughly this lineage.

← All links