What I Mean by Evaluation
When I talk about evaluation, I mean the work of figuring out whether an AI system is actually good enough to trust. A model can look impressive in a demo and still be unreliable in practice. The interesting question is not whether it can produce a strong example, but whether the evidence is strong enough to support a decision.
I care about making that gap visible: define what good means, measure behavior, understand uncertainty, and be willing to say the result is not ready yet. Evaluation is not a final checkbox after the system is built. It is part of how I decide what the system is for and whether it is doing that job.
The clearest public example is my LLM Judge Eval Harness. The first step is to calibrate a quality threshold on purpose, instead of treating a score as meaningful just because it looks high.
I separate development work from the score that counts. Tuning and exploration stay on one side of the split; the decision measurement runs blind against held-out examples so I do not reward myself for fitting the set I already saw.
A single point estimate is not enough. I use bootstrap confidence intervals and pay attention to the lower bound, because a promising mean can still sit on thin evidence.
The harness fails closed when that evidence is weak. I bias toward high precision so a false positive does not burn reviewer trust: better to withhold a pass than to ship a score that looks strong and later falls apart.
That is the kind of AI work I want to keep doing: systems where trustworthiness depends on evidence, not on how polished the demo looks. I expect to keep sharpening this habit across more evaluation problems as I go.