Back to index
OtherSelf-paced, several hours for the core essays·Free

Hamel Husain's Blog and Writing

4.6

The best available writing on evaluation, which is the part of building with language models that most teams get wrong and almost nobody teaches properly. Opinionated, specific and grounded in real consulting work.

What We Liked

  • Evaluation writing is the clearest treatment of the topic anywhere
  • Advice comes from real client work rather than theory
  • Willing to say that popular tools and practices are wrong, with reasons
  • Focuses on the unglamorous problems that actually determine success

What Could Be Better

  • Blog format means no structured path through the material
  • Assumes engineering competence and offers no on ramp for beginners
  • Some posts are tied to specific tools and age faster than the principles
  • Strong opinions occasionally read as more settled than the evidence supports

Detailed review

There is a pattern that repeats across almost every team building on language models. They ship a prototype quickly, it demos well, and then they discover they have no way to tell whether a change makes things better or worse. Someone tweaks a prompt, everyone agrees the new output looks nicer, and nobody knows if quality improved or if three other things quietly broke. Six months in, the team is making changes based on vibes and defending them with anecdotes.

Hamel Husain writes about exactly this failure, and does so better than anyone else publishing publicly. The evaluation writing is the core of it and the reason I rate this so highly. The central argument is that you need to look at your data, that you need error analysis grounded in actual failures rather than generic metrics, and that domain experts must be involved in defining what good looks like. That sounds obvious written down and almost nobody does it.

Teams reach for automated scoring frameworks before they have looked at a hundred failures themselves, which produces numbers that move without anyone understanding why. The writing on building a proper evaluation process, starting from manual review and moving towards automation only once you know what you are measuring, is the most practically valuable material I have read on building with these models. The credibility comes from real work. This is someone who consults with companies actually shipping these systems, and it shows in the specificity.

The failure modes described are recognisable, the examples have the texture of things that really happened, and the advice anticipates the objections you would raise. There is a substantial gap between writing derived from experience and writing derived from other writing, and this is firmly on the right side of it. The willingness to criticise popular practice is valuable and rarer than it should be. Much of the public conversation is promotional, because the people writing are often selling something or hoping to be hired by someone who is.

Writing that says a widely adopted framework encourages bad habits, or that a fashionable evaluation approach measures the wrong thing, provides a correction that the field genuinely needs. You do not have to agree with every position to benefit from the pushback. The focus on unglamorous problems is the right focus. Evaluation, data inspection, error analysis and iteration discipline are what separate systems that work from demos that impressed someone once.

They are also boring compared with new architectures and agent frameworks, which is precisely why they are undersupplied in the public conversation and why writing that takes them seriously is disproportionately useful. The blog format is the main structural limitation. Posts were written at different times for different reasons, and while the important ones are findable, there is no curriculum and no obvious order. A newcomer will bounce between topics without a sense of what builds on what.

The material would make an excellent structured course and it is not one, so you assemble the path yourself. The assumed background is real. This is written for people who build software, who have shipped things, who understand the engineering context. Someone learning what a language model is will find the writing sails past them, not because it is needlessly complex but because it is aimed elsewhere.

That is a correct choice for the audience and it does mean this belongs later in a learning path rather than at the start. Tool specific posts age faster than principle based ones. Writing about a particular library or workflow reflects a moment, and the tooling moves. The underlying arguments about evaluation and error analysis will still be right in five years, and some of the implementation detail around them will not.

Check dates and take the principle rather than the specific recipe. The strength of opinion is mostly a virtue and occasionally overshoots. Some positions are stated with a confidence that the available evidence does not fully support, particularly where the field genuinely has not settled. That is the cost of writing that is willing to take positions at all, and I would rather read someone who commits and is sometimes wrong than someone who hedges everything into uselessness.

Four point six. This is among the most valuable free material available for anyone building language model applications seriously, addressing the specific gap where most teams fail, from someone who has watched them fail. The lack of structure and the assumed background are real limitations, and if you build with these models professionally and have not read the evaluation writing, that is the highest value gap in your reading list.

[ final ]

The verdict.

Required reading for anyone building language model applications professionally. Start with the evaluation writing and read it twice, because most teams are getting this wrong.