Evaluating real-world forecasts

Forecasting & evaluation of infectious disease dynamics

Challenges in evaluation

Challenges in evaluation

Based on what you’ve seen so far, what are the main challenges with forecasting evaluation?

Challenges in evaluation

Based on what you’ve seen so far, what are the main challenges with forecasting evaluation? (my list)

  • it can be an overwhelming data analysis project
  • forecasts (and the scores) are correlated over time
  • forecats (and the scores) can be scale-dependent
  • models don’t all make the same predictions

Data analysis

Scores are often correlated, with outliers

Flusight Eval website

Scores can be scale-dependent

Flusight Eval website

Models don’t make the same predictions

Caution

When models have made forecasts for different subsets of tasks, it can make comparison of models complicated. For example, what if one model never forecasts for the harder prediction tasks?

Relative Skill Scores are a possible solution

See e.g., (Cramer et al. 2022)

How to draw valid inference about differences between models?

Hubs facilitate forecast evaluation

Hubs accelerate scientific discovery

“Comparing the accuracy of forecasting applications is difficult because forecasting methods, forecast outcomes, and reported validation metrics varied widely.”

– Chretien et al., PLOS ONE, 2014

The US COVID-19 Forecast Hub

Launched April 2020 by the Reich Lab in collaboration with CDC. Goals were to:

  • Provide decision-makers and general public with reliable information about where the pandemic is headed in the next month.
  • Assess reliability of forecasts and gain insight into which modeling approaches do well.
  • Create a community of infectious disease modelers underpinned by an open-science ethos.

…provided CDC with weekly forecasts

…had prominent uses

Hubverse is the tool we wish we had

Hubverse

Who works with a hub?

Hubverse

What does hubverse data look like?

Some columns define a prediction task

Others are always present

Why is standardization important?

You will learn by doing in this session.

From a practical standpoint, standards facilitate

  • visualization
  • evaluation
  • ensembling

COVID Hub reboot (2024-present)

In 2024, CDC rebooted the hub, using hubverse standards.

These are the data that you will use in this session.

References

Cramer, Estee Y., Evan L. Ray, Velma K. Lopez, et al. 2022. “Evaluation of Individual and Ensemble Probabilistic Forecasts of COVID-19 Mortality in the United States.” Proceedings of the National Academy of Sciences 119 (15): e2113561119. https://doi.org/10.1073/pnas.2113561119.
Lopez, Velma K., Estee Y. Cramer, Robert Pagano, et al. 2024. “Challenges of COVID-19 Case Forecasting in the US, 2020–2021.” PLOS Computational Biology 20 (5): e1011200. https://doi.org/10.1371/journal.pcbi.1011200.

Your Turn

Use COVID-19 Forecast Hub forecasts to …

  1. create forecast visualizations.
  2. evaluate multiple forecast models.

Return to the session