Back to index
OtherAround 340 pages, a week or so·Around $12 to $20 depending on format

Human Compatible: AI and the Problem of Control by Stuart Russell

4.2

The one book in this category written by someone who has spent a career actually building the field, and it shows in every chapter.

What We Liked

  • The critique of the field's own foundations is precise and technically grounded
  • Proposes a concrete alternative rather than only raising alarm
  • Written with wit and clarity, unusual for the subject
  • Takes the counterarguments seriously and answers them individually
  • The author has the standing to make the argument stick

What Could Be Better

  • The proposed solution is less developed than the critique
  • The appendices are where the technical substance hides
  • Some optimism about the tractability of the problem now looks generous
  • The 2019 vintage means the current architectures barely appear

Detailed review

Stuart Russell co-wrote the textbook that a generation of AI researchers learned from, which is the relevant fact about this book, because it means the argument comes from inside the discipline rather than from a philosopher or a journalist looking at it from outside. The central claim is unusual and it is the reason this book is better than its competitors. He does not argue that AI is dangerous because it might become malicious or because it might get out of hand. He argues that the standard definition of a successful AI system, which is a machine that achieves the objective we specify, is itself the problem.

Optimise hard enough against a stated objective and you get exactly what you asked for, which is reliably not what you meant, because any objective we can write down omits almost everything we care about. That is not a speculative claim about future systems. It is a description of a failure mode that already appears in recommender systems, in reinforcement learning environments and in any optimisation loop with a proxy metric. Anyone who has watched a metric get gamed by the system optimising it has seen a small version of the argument.

What lifts this above complaint is that he proposes an alternative. Rather than building machines that pursue a fixed objective, build machines that are uncertain about what humans want, that treat human behaviour as evidence about preferences, and that therefore have an incentive to ask, to defer and to allow themselves to be corrected. A machine confident in its objective has a reason to prevent you switching it off, because being switched off means the objective goes unmet. A machine uncertain about its objective has a reason to allow it, because your reaching for the switch is evidence that it has misunderstood.

That reframing turns a philosophical worry into a research direction, and it is the most constructive contribution any of these books makes. The writing is good. Dry humour, clear structure, and a refusal to inflate. He walks through the history of the field, explains why previous predictions failed, and is happy to say when he does not know.

There is a chapter that lists the standard objections, the ones you hear at conferences and on comment threads, and answers each one properly. That we can just switch it off. That intelligence implies benevolence. That we should not talk about this until it is closer.

That worrying about it is anti research. He takes each seriously and dismantles it without contempt, which is more persuasive than the alternative approach and considerably rarer. Now the limitations, and they are real. The proposal is much less developed than the critique.

Two hundred pages establish precisely why the standard model is broken, and the constructive part is thinner. Assistance games and preference learning are sketched rather than worked through, and the hard problems, including whose preferences count when people disagree, how you handle preferences that are inconsistent or that the person would disown on reflection, and how any of this scales beyond toy settings, are acknowledged and left open. Honest, and it means you finish the book with a compelling diagnosis and a research agenda rather than a solution. The technical content is largely in the appendices, which is a publishing decision that makes the book more accessible and leaves the main text making claims whose support is at the back.

A reader without the background will take a fair amount on trust. Some of the optimism has aged oddly. There is a general sense that the field will have time to work this out, that the relevant research is tractable, and that the incentive structures might be steered. Since 2019 the pace of deployment and the competitive dynamics between labs have made that read as generous.

And like everything written before 2020 it predates the systems that now dominate. Large language models, trained on next token prediction and then shaped by human feedback, do not fit neatly into the framework of an agent maximising a specified objective. The preference learning material turns out to be more prescient than most, given what reinforcement learning from human feedback became, and the fit is loose and he could not have known. My four point two is for the best argued book in a category full of badly argued books, written by someone with the standing to criticise the field from within, marked down because the constructive half is thinner than the critical half and because the technology moved in a direction the book only partly anticipated.

If you have read the textbook, this is the same mind on a harder problem. If you have not, this is still the place to start.

[ final ]

The verdict.

The one to read if you only read one. Better argued than Bostrom, better grounded than Tegmark, and written by someone who wrote the textbook the field learned from.