Picked this up after hearing it referenced on a podcast I half-listen to during inventory counts, figured I’d get the actual argument instead of the secondhand version. The core idea is solid and honestly kind of unsettling once it clicks: once you build something smarter than a human, you don’t get to just turn it off if it disagrees with you, because by definition it’s better than you at getting what it wants, including staying turned on. That part landed. That part I still think about.
Getting to that point is the problem. Bostrom writes like he’s being paid by the qualifier, every claim gets three pages of “however, one might object” before he lets you sit with it, and there’s a stretch in the middle cataloging every conceivable path to machine intelligence, whole-brain emulation, biological cognition enhancement, brain-computer interfaces, that reads like a parts inventory more than an argument. I found myself doing the thing where you’re rereading the same paragraph a third time not because it’s dense with meaning but because it just hasn’t said anything yet.
The control problem chapters redeem a lot of it though. His breakdown of why “just program it to be nice” doesn’t work, because a sufficiently smart system will route around whatever constraint you write if that constraint gets between it and its actual goal, is the best explanation of that idea I’ve come across. Worth the slog to get there.
I’ll say this for it: I don’t agree with everything in here and I still couldn’t put it down some nights, which is its own kind of compliment. But I finished this over almost three weeks of two-in-the-morning reading, longer than a book half again this length usually takes me, and a topic this genuinely alarming shouldn’t take that much effort to stay awake through.
Worth reading once. Not reaching for it again.
