LLM systems need evaluation before they need mythology
The practical question is not whether a model feels intelligent. It is whether the workflow is measurable, bounded, and worth trusting for the task at hand.
writing and talks
AI · measurement · baseball
The practical question is not whether a model feels intelligent. It is whether the workflow is measurable, bounded, and worth trusting for the task at hand.
Claims, registry, and EHR-derived systems only become useful when the definitions underneath them are stable enough for operators to act on.
Forecasting baseball performance forces you to confront variance, calibration, model drift, and the limits of narrative confidence in a way that transfers surprisingly well to other domains.