Introduction to ensembles

What is a forecast, and is it any good?

Ensembles: many forecasts into one

Figure credit: Evan Ray and Nick Reich

Why ensemble?

  1. Models are specialists and you want all the perspectives
    • different data sources
    • different philosophies, e.g. more mechanistic or more statistical approaches
    • different methodologies and parameterizations

Single model in Respicast

Many models in Respicast

Why ensemble?

  1. “Whole is greater than sum of parts”

Respicast ensemble

Ensemble methods: how to combine?

Typically… take the (equal) average!

Combination depends on your view of why models are different:

  • Noisy expressions of a single underlying truth?
  • Competing uncertainty among multiple plausible hypotheses?

Figure credit: Howerton et al. (2023)

Ensemble methods: unequal weights?

  • Can weight models by past forecast performance
    • e.g. using forecast scores
  • Rarely better than equal average
    • lots of uncertainty in weight estimation
    • past performance is not guaranteed to predict future performance … put a “strong prior” on equal weights

Ensembles from multiple modellers

Collaborative ensembles increasingly used in infectious disease modelling

Reich et al. (2022)

Collaborative modelling “hubs”

  • Projects run by research groups, public health agencies
  • Participation generally open
  • Standard format enables
    • data validation
    • ensemble-building
    • model evaluation
    • visualization

Is an ensemble any good?

Ensemble as a product

  • typically reliable performance benefit over single model
  • creates a single product for downstream forecast users

Ensemble as a process

  • principled approach to model comparison and synthesis
  • opportunity for modeller challenge, consensus, and shared responsibility for outputs (Medley 2022)

Your Turn

  1. Create unweighted and weighted ensembles from multiple models.
  2. Evaluate the ensembles against their constituent models.
  3. Consider when an ensemble is any good, and for whom.

Return to the session

References

Colón-González, Felipe J., Leonardo Soares Bastos, Barbara Hofmann, et al. 2021. “Probabilistic Seasonal Dengue Forecasting in Vietnam: A Modelling Study Using Superensembles.” PLOS Medicine 18 (3): e1003542. https://doi.org/10.1371/journal.pmed.1003542.
Cramer, Estee Y., Evan L. Ray, Velma K. Lopez, et al. 2022. “Evaluation of Individual and Ensemble Probabilistic Forecasts of COVID-19 Mortality in the United States.” Proceedings of the National Academy of Sciences 119 (15): e2113561119. https://doi.org/10.1073/pnas.2113561119.
Funk, Sebastian, Anton Camacho, Adam J. Kucharski, Rachel Lowe, Rosalind M. Eggo, and W. John Edmunds. 2019. “Assessing the Performance of Real-Time Epidemic Forecasts: A Case Study of Ebola in the Western Area Region of Sierra Leone, 2014-15.” PLOS Computational Biology 15 (2): e1006785. https://doi.org/10.1371/journal.pcbi.1006785.
Howerton, Emily, Michael C. Runge, Tiffany L. Bogich, et al. 2023. “Context-Dependent Representation of Within- and Between-Model Uncertainty: Aggregating Probabilistic Predictions in Infectious Disease Epidemiology.” Journal of The Royal Society Interface 20 (198): 20220659. https://doi.org/10.1098/rsif.2022.0659.
Medley, Graham F. 2022. “A Consensus of Evidence: The Role of SPI-M-O in the UK COVID-19 Response.” Advances in Biological Regulation 86 (December): 100918. https://doi.org/10.1016/j.jbior.2022.100918.
Reich, Nicholas G, Justin Lessler, Sebastian Funk, et al. 2022. “Collaborative Hubs: Making the Most of Predictive Epidemic Modeling.” Am. J. Public Health, April, e1–4. https://doi.org/10.2105/ajph.2022.306831.
Reich, Nicholas G., Craig J. McGowan, Teresa K. Yamana, et al. 2019. “Accuracy of Real-Time Multi-Model Ensemble Forecasts for Seasonal Influenza in the U.S. PLOS Computational Biology 15 (11): e1007486. https://doi.org/10.1371/journal.pcbi.1007486.