A Better Average Forecast Can Still Miss the Dangerous Tail
Ensemble forecasts can improve both ordinary weather prediction and extreme-event skill. Establishing that a particular local probability deserves trust requires a more specific test than winning an average score.
The threshold before the score
A forecast can get the temperature almost right and still put it on the wrong side of a threshold. In a hypothetical equipment test, a facilities team cares whether outdoor temperature exceeds its specified limit tomorrow afternoon. A small error near that limit matters more to this question than a larger error on a mild day. The team needs a probability for the crossing, with evidence that the probability means what it appears to mean.
That requirement differs from asking which model predicts temperature most accurately across the globe. Our argument is about the inference between those questions: improved overall performance is a reason to investigate local usefulness, while the local claim needs evidence matched to its event. This is a narrower issue than the institutional responsibilities discussed in our essay on public forecasting.
What the ensemble evidence establishes
An ensemble supplies multiple possible weather evolutions. Counting the members that cross a threshold gives a raw estimate of its probability. The average of their temperatures answers a different question: two ensembles can have the same mean while placing very different amounts of probability beyond the limit.
The WeatherBench 2 paper includes the continuous ranked probability score, or CRPS, which evaluates predictive distributions, and a spread-to-error diagnostic for ensemble calibration. Its authors explicitly present headline scores as incomplete summaries. In their formulation, CRPS is averaged over time and geographical coordinates; the spread-to-error ratio is only an initial calibration check.
Our inference is straightforward: an improvement in an average does not establish improvement in every component. Gains elsewhere can outweigh a deterioration in one region or event category. This is a limitation of the claim made from the score, rather than evidence that any particular model has that deterioration.
Indeed, the original GenCast study provides evidence against a blanket claim that AI forecasts fail at extremes. It compared 50-member ensembles on a 0.25° grid using 2019 as the test year. Beyond overall probabilistic gains, it reported better Brier skill for extreme temperature, wind and pressure thresholds, with exceptions to statistical significance at some thresholds and lead times. For wind above the 99.99th percentile, improvement was not significant beyond seven days. Verification used ERA5 for GenCast and HRES initial-condition analyses for ENS; precipitation was excluded from the main results because of concerns about ERA5 precipitation quality.
Those findings deserve their actual scope. They establish meaningful performance in the study's tests. They do not settle tomorrow's probability at an individual instrument, or establish the performance of a different model version. The research comparison is historical evidence, not an October 2026 operational ranking.
Checking a probability
The ECMWF guide to probabilistic verification explains reliability by comparing issued probabilities with the frequency of subsequent events. When these agree, the reliability diagram follows its diagonal. The guide also separates forecast accuracy, performance against a reference, and usefulness: those are distinct properties.
Apply that idea to a hypothetical archive of forecasts near a 20% exceedance probability. Roughly one fifth should verify over a sufficiently informative sample if that probability category is reliable. One crossing does not refute the forecast; one non-crossing does not vindicate it. The archive tests the repeated meaning of the number.
Our proposed reading of such a diagram also asks what was pooled. A reassuring result across many locations might conceal offsetting errors between coastal and inland sites. Conversely, dividing the archive into ever smaller categories can leave too little evidence to judge anything. Local evaluation therefore involves a trade: relevance increases as the question narrows, while the amount of available evidence usually shrinks.
This is also why adding simulated members cannot substitute for collecting verifying cases. In a hypothetical run, generating more possible futures may refine the model's estimated probability. It does not provide more independent occasions on which that probability has been tested against the world.
A fair test of extremes
There is a trap in demanding an extremes-only score. Lerch and colleagues' Forecaster's Dilemma shows how restricting ordinary evaluation to extreme observed outcomes can favor distorted forecasts. The paper discusses proper weighted scoring rules, including threshold-weighted CRPS, as ways to emphasize a region of the outcome distribution without abandoning incentives for honest probabilities. Its evidence includes theory and simulations; it is not a test of today's AI weather systems.
Consider our hypothetical facilities archive. If the reviewer retains only afternoons when the equipment threshold was crossed, a model that habitually predicts excessive heat escapes scrutiny on all the mild afternoons. The selection has removed precisely the cases that would reveal its overprediction. This example illustrates the selection problem; it is not an observed model result.
A case study can still explain a particular miss. Our recommendation is to give it that diagnostic role while retaining the full evaluation period for comparative probability scores. Investigating the worst failures and estimating overall threshold skill are compatible activities, provided the conclusions identify which question each answers.
Designing a local comparison
The following is our proposed evaluation design, not a new benchmark result. Begin by writing the event in ordinary language precisely enough that another analyst could reconstruct it. Specify the measured quantity, instrument or area, threshold, time window and forecast lead. A temperature at one instant and the afternoon maximum are different targets. A regional exceedance and a crossing at one station are different targets too.
Next, choose the comparison before inspecting the winners. Use the same forecast dates and event definition for the candidate and the reference, and document what information each had at issuance. If a local correction is fitted, reserve separate data for testing the corrected probabilities. Otherwise the apparent improvement may partly reflect reuse of the answers.
Include every eligible forecast date, recording both crossings and non-crossings. Report an event probability score alongside reliability by probability category, with the number of forecasts behind each category. Keep a simple reference based on historical event frequency as well as the competing forecasting system. These comparisons ask whether complexity buys useful information for this event.
Specify how model output becomes the verifying quantity. Interpolation to a station, a change of elevation, or taking a maximum over a region introduces a transformation. Test the resulting product that the user actually receives. Success for one transformation should not silently certify another. This parallels the distinction between atmospheric prediction and vegetation response in our satellite weather-stress essay.
Finally, report uncertainty in the comparison and explain how related weather episodes were handled. Many nearby measurements during one persistent episode should not be presented as though each were an entirely new weather experiment. If a narrow category contains little evidence, retain that limitation in the result rather than disguising it with a precise percentage.
Reporting a useful result
A useful conclusion may be modest: the candidate improved a specified threshold score, while the rarest probability categories remain uncertain. It may instead show that an overall winner offers no demonstrated advantage for this local event. Either result helps determine what evidence to gather next.
Progress in ensemble forecasting deserves evaluation that can recognize it. That means testing relevant probabilities fairly, preserving non-events, and describing the scale at which the evidence holds. A local probability earns its meaning through repeated comparison with what actually happened.
Sources
- Rasp et al., WeatherBench 2, manuscript version dated January 26, 2024.
- Price et al., Probabilistic weather forecasting with machine learning, Nature, published December 4, 2024.
- ECMWF, Forecast User Guide: probabilistic verification concepts, undated source view.
- Lerch et al., Forecaster's Dilemma: Extreme Events and Forecast Evaluation, manuscript dated December 31, 2015.
Sources accessed October 2, 2026. Historical experiments are described with their original evaluation periods.
Related reading
- The AI Weather Model Becomes the Public Forecast
- The Satellite Forecast Becomes the Weather-Stress Ledger
Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.