Back to index
OtherAround 360 pages across 8 chapters, two to three weeks·Around $25 in hardback, cheaper in paperback and ebook

AI Snake Oil (Arvind Narayanan and Sayash Kapoor)

4.5

The most useful non technical AI book I have read, because it gives you a working test for whether a claim is plausible rather than a general attitude of scepticism. Deliberately light on how anything works.

What We Liked

  • The distinction between predictive and generative AI is the most useful framing in the whole debate
  • Every claim is backed by research, including the authors' own reproducibility work
  • Sceptical without being dismissive, which is a much harder and more useful position
  • The material on flawed benchmarks and leaked evaluations is important and largely unknown outside research
  • Genuinely accessible, so you can give it to a non technical executive who needs it

What Could Be Better

  • No technical depth at all, so it teaches you nothing about building anything
  • The generative AI chapters are less confident than the predictive ones and will date faster
  • Some critique lands more on institutions and procurement than on the technology itself
  • Occasionally repetitive in restating the central thesis
  • Offers a diagnosis more readily than a prescription for what to do instead

Detailed review

The discourse around artificial intelligence has settled into two camps that are both useless. One holds that these systems are on the verge of doing everything and the other that they are a statistical parlour trick. Both are positions rather than analyses, and neither helps you decide whether the specific tool a vendor is pitching to your organisation will work. This book is the corrective and its value comes from one central distinction, made early and applied relentlessly.

Predictive AI and generative AI are different technologies with different track records, and lumping them together is the source of most confused thinking on the subject. Generative systems have made remarkable progress. Predictive systems, the ones claiming to forecast which employee will succeed, which defendant will reoffend, which student will drop out, which applicant is worth interviewing, largely do not work, have never worked well, and are deployed at enormous scale anyway. The authors are careful and the case is devastating.

They go through the actual evaluations, and the pattern repeats. A system is claimed to predict some human outcome with impressive accuracy. On examination the accuracy is barely above a simple baseline, or the evaluation had leakage, or the impressive number came from a population that does not resemble where it is deployed. The chapter on criminal risk assessment is the one that should be most widely read, because these tools are making decisions about people's liberty on the strength of performance that would not pass review in a research setting.

Narayanan and Kapoor's own work on reproducibility failures in machine learning research runs underneath the whole argument and gives it weight. They found systematic problems, particularly data leakage, across a large number of published papers in fields applying machine learning to scientific and social questions. That is a serious finding and it explains a good deal of why so many systems perform far worse in deployment than on paper. This is not hostile outsiders complaining about a field they do not understand.

It is researchers documenting a methodological problem from inside. What makes the book good rather than merely correct is the refusal to overcorrect. It would have been easy to write a book arguing that all of this is hype. They do not.

Generative systems are treated as genuinely capable and genuinely useful, with limitations described precisely rather than dismissively. Content moderation gets a chapter that takes the difficulty of the problem seriously instead of pretending a better model would solve it. This calibration is why I trust the book. A purely negative account would be as unhelpful as the boosterism it opposes, and considerably easier to write.

The material on benchmarks deserves particular attention from technical readers, who often assume this is a solved administrative matter. Contaminated evaluation sets, benchmarks that measure something other than the capability they are named after, results that do not survive contact with a different population. Anybody making a purchasing decision on the basis of a benchmark number should read this chapter first and then ask the vendor some very specific questions. Now the limitations.

There is no technical content, by design. You will not learn how any of these systems work, and the book is explicit that this is not its purpose. It gives you the ability to evaluate claims and no ability to build or diagnose anything. That is a legitimate scope and it means this complements the technical material in this catalogue rather than substituting for any of it.

Someone who reads only this will be well defended against nonsense and unable to contribute to building an alternative. The generative AI chapters are noticeably less confident than the predictive ones, and I think the authors would agree. The evidence base on predictive systems goes back years and the conclusions are firm. The generative material is assessing a fast moving target, and some of it is already less true than when written.

The predictive critique will hold for a decade. The generative chapters will need revising, and the book is honest about that uncertainty. Some of the critique also lands more on institutions than on technology. Many of the failures described are procurement failures, accountability failures, or organisations wanting a technical answer to a political problem.

That is an important observation and it slightly undercuts the framing, since the snake oil is often being bought as eagerly as it is sold. The book notices this and could have made more of it. There is repetition, particularly in restating the central distinction, which is a common shape for books built from an argument the authors have been making publicly for years. And the prescriptive side is thinner than the diagnostic side.

You finish with a clear sense of what does not work and less guidance on what a responsible deployment looks like, which is the harder question. My 4.5 is for a book that changes how you read every AI claim you encounter afterwards. It is evidence based, carefully argued, accessible enough to hand to a decision maker, and it gives you a specific test rather than a general suspicion. In a field where the loudest voices are selling something, two researchers with no stake in the outcome are worth listening to.

[ final ]

The verdict.

Read it if you evaluate AI claims, buy AI products, or have to explain to somebody why a vendor's promise is not credible. Not a substitute for understanding how the systems work.