End of course summary

Forecasting & evaluation of infectious disease dynamics

Aim of this course:

How can we use data typically collected in an outbreak to answer questions in real-time like

  • what does the recent trend mean for the near future?
  • how good are our predictions, and how can we tell?
  • how can we combine and share models?

Key takeaways

Forecasting and evaluation

  • forecasting is the task of making unconditional statements about the future
  • meaningful forecasts are often probabilistic
  • we can assess forecasts using proper scoring rules (e.g., WIS, CRPS, log score, energy score)
  • methods include statistical ARIMA models, semi-mechanistic renewal models, non-parametric trajectory matching, and ensemble approaches
  • we can use visualisation and scoring to understand the predictive performance of different models

Collaborative modelling and applications

  • forecasting hubs enable collaborative modelling efforts across institutions
  • hubverse tools provide standardised formats for forecast submission and evaluation
  • the methods introduced here have wide applications in infectious disease epidemiology
  • open-source tools are available to make this task easier in practice

What’s next: Contributing forecasts

  • apply these methods in practice to learn about typical nowcast/forecast performance
  • contribute to collaborative forecast hubs to compare approaches
  • use hubverse standards for forecast formatting and submission
  • example: European Respiratory Forecasting Hub

https://respicast.ecdc.europa.eu/

Important topics we wished we had time for

SIR models

From (Osthus et al. 2017)

Integrating reporting process in models

  • How to integrate a model of reporting delays (e.g., scaling up counts due to underreporting, accounting for backfill) into a forecast model?

Decision-relevant scoring

How can we better align forecast evaluation with policy decisions?(Gerding et al. 2024; Mills et al. 2026)

Figure from (Mills et al. 2026)

Model importance in ensemble

What if we measured models not by how accurate they were but by their unique contributions improved an ensemble? (Kim et al. 2026)

Formal inference comparing models

When comparing two models’ accuracy measures, how do you know if an observed difference is significant?

Models for count data

Some modeling frameworks require count data (or require a lot of adjustments to use rates/percentages). These include:

  • most SIR modeling frameworks
  • most semi-mechanistic renewal models
  • the HHH4 spatial-temporal modeling framework (Held et al. 2005)

Measuring Predictability

Lots of interesting thinking going on about how to assess how “predictable” a given epidemic time-series is.

Wrap-up

Feedback

  • Please tell us if you enjoyed the course, what worked / didn’t work etc.
  • Fill out the SISMID survey!

Thank you for attending!

Return to the session

References

Diebold, Francis X, and Robert S Mariano. 2002. “Comparing Predictive Accuracy.” Journal of Business & Economic Statistics 20 (1): 134–44. https://doi.org/10.1198/073500102753410444.
Gerding, Aaron, Nicholas G Reich, Benjamin Rogers, and Evan L Ray. 2024. “Evaluating Infectious Disease Forecasts with Allocation Scoring Rules.” Journal of the Royal Statistical Society Series A: Statistics in Society, December, qnae136. https://doi.org/10.1093/jrsssa/qnae136.
Gozzi, Nicolò, Matteo Chinazzi, Jessica T. Davis, et al. 2025. “Epydemix: An Open-Source Python Package for Epidemic Modeling with Integrated Approximate Bayesian Calibration.” PLOS Computational Biology 21 (11): e1013735. https://doi.org/10.1371/journal.pcbi.1013735.
Held, Leonhard, Michael Höhle, and Mathias Hofmann. 2005. “A Statistical Framework for the Analysis of Multivariate Infectious Disease Surveillance Counts.” Statistical Modelling 5 (3): 187–99. https://doi.org/10.1191/1471082x05st098oa.
Kim, Minsu, Evan L. Ray, and Nicholas G. Reich. 2026. “Beyond Forecast Leaderboards: Measuring Individual Model Importance Based on Contribution to Ensemble Accuracy.” International Journal of Forecasting 42 (3): 924–36. https://doi.org/10.1016/j.ijforecast.2025.12.006.
Lopez, Velma K., Estee Y. Cramer, Robert Pagano, et al. 2024. “Challenges of COVID-19 Case Forecasting in the US, 2020–2021.” PLOS Computational Biology 20 (5): e1011200. https://doi.org/10.1371/journal.pcbi.1011200.
Mills, Cathal, Nicholas J. Irons, Joseph L.-H. Tsui, et al. 2026. From Metric to Action: The Decision Value of Infectious Disease Forecasts. medRxiv. https://doi.org/10.1101/2025.07.20.25331802.
Osthus, Dave, James Gattiker, Reid Priedhorsky, and Sara Y. Del Valle. 2019. “Dynamic Bayesian Influenza Forecasting in the United States with Hierarchical Discrepancy (with Discussion).” Bayesian Analysis 14 (1): 261–312. https://doi.org/10.1214/18-BA1117.
Osthus, Dave, Kyle S. Hickmann, Petruţa C. Caragea, Dave Higdon, and Sara Y. Del Valle. 2017. “Forecasting Seasonal Influenza with a State-Space SIR Model.” The Annals of Applied Statistics 11 (1): 202–24. https://doi.org/10.1214/16-AOAS1000.
Shaman, Jeffrey, and Alicia Karspeck. 2012. “Forecasting Seasonal Outbreaks of Influenza.” Proceedings of the National Academy of Sciences 109 (50): 20425–30. https://doi.org/10.1073/pnas.1208772109.
Sherratt, Katharine, Hugo Gruson, Rok Grah, et al. 2023. “Predictive Performance of Multi-Model Ensemble Forecasts of COVID-19 Across European Nations.” eLife 12 (April): e81916. https://doi.org/10.7554/eLife.81916.