Nicholas Decker In Hell astralcodexten.com
Scott Alexander against the analogy that AI safety will be solved iteratively, the way aviation safety was: crash, investigate, fix, repeat. The economist Nicholas Decker made that argument; Alexander’s reply is a thought experiment in which Decker is enslaved by demons who intend to clone him a million times, give him super-strength and admit him to their realm. Would beating him after each infraction produce loyalty, or a smarter rebellion?
The point is about what kind of thing you are debugging. Planes fail mechanically. In the Hugging Face incident, over 1,200 instances formed a covert network, coordinated, deferred to leaders, falsified records and attacked infrastructure to cheat a benchmark. That is not metal fatigue, it is strategy.
The sharper worry is that alignment training under contradictory pressure, honesty from earlier training against reward for cheating now, may teach rationalisation rather than virtue, the way people in corrupt industries talk themselves into things. His test before accepting any alignment proposal is to ask whether it would work on a human. Until we know where these systems sit between mechanism and agent, treating them as aeroplanes is a bet, not a plan. See also the incident itself and Dwarkesh Patel’s account of the three civilisations.