nyc-weather: An Hourly Forecaster That Learned to Look Upstream
An hourly forecaster for New York City — twelve surface variables over a 500-km box, with probability of rain as the product that actually matters to anybody. Gen1 looked respectable on temperature and hopeless on rain, and an adversarial review established that the fault lay with the experiment rather than with the model: five verified defects, running from a no-op precipitation transform to a deployed checkpoint that was effectively 5%-trained. Everything afterward followed from that one lesson, which is to fix the measurement before touching the model. Gen4 was then won on geometry rather than on scale, on the observation that CDS's request quota counts fields, not area — the same price therefore buys pressure-level data over a Midwest→Gulf→Atlantic box that actually contains tomorrow's weather. The one-shot test read confirmed it: calibration verified out-of-sample, temperature error down 28% at both 24h and 48h, and twenty-four-hour rain skill doubled.