Multi-model ensembles

Forecasting & evaluation of infectious disease dynamics

Ensembles: many forecasts into one

Figure credit (Shandross et al. 2026)

Why ensemble?

  1. Models are specialists and you want all the perspectives
    • different data sources
    • different philosophies, e.g. more mechanistic or more statistical approaches
    • different methodologies and parameterizations
  1. A single “consensus forecast” is easier for decision-makers to digest

“Whole is greater than sum of parts”

Average of multiple predictions is often (not always) more performant than any individual model

Ensemble methods: how to average?

Figure credit: Shandross et al. (2026)

Linear opinion pool ensemble

Let \(F_m(x)\) be a cumulative density function (CDF) given by forecast model \(m\).

\[F_{LOP}(x) = \sum_{m=1}^M w_m F_m(x) \]

Figure credit: Shandross et al. (2026)

Quantile (Vincent) average ensemble

Let \(F_m^{-1}(\theta)\) be the inverse CDF, or the quantile function, for quantile level \(\theta\).

\[F_Q^{-1}(\theta) = \sum_{m=1}^M w_m F_m^{-1}(\theta) \]

Figure credit: (Shandross et al. 2026)

Ensemble methods: to weight or not?

\[F_{LOP}(x) = \sum_{m=1}^M w_m F_m(x) \]

  • How do you estimate weights?
    • by past performance: e.g., using forecast scores
    • by human judgment, using expertise and knowledge of mdoels
  • Rarely better than equal average
    • lots of uncertainty in weight estimation!
    • put a “strong prior” on equal weights, both in your mental and statistical models

Logistics of building ensembles

  • With a handful of models you have fit yourself, you can build an ensemble with a few lines of code.

  • What gets hard is the logistics at scale: collecting many forecasts in a shared format, validating them, and synthesizing them.

  • Collaborative “hubs” exist to handle that logistics problem. They are not a prerequisite for ensembling – they just make it a whole lot easier once a few models or teams are involved.

Collaborative modelling “hubs”

  • Projects run by research groups, public health agencies

  • Participation generally open

  • Standard format enables

    • data validation
    • ensemble-building
    • model evaluation
    • visualization

Hubs increasingly used in epi

Reich et al. (2022)

… e.g., the European Respicast Hub

Single model

… Multiple models

… … Multi-model ensemble

Your Turn

  1. Create unweighted and weighted ensembles using forecasts from multiple models.
  2. Evaluate the forecasts from ensembles compared to their constituent models.

Return to the session

References

Colón-González, Felipe J., Leonardo Soares Bastos, Barbara Hofmann, et al. 2021. “Probabilistic Seasonal Dengue Forecasting in Vietnam: A Modelling Study Using Superensembles.” PLOS Medicine 18 (3): e1003542. https://doi.org/10.1371/journal.pmed.1003542.
Cramer, Estee Y., Evan L. Ray, Velma K. Lopez, et al. 2022. “Evaluation of Individual and Ensemble Probabilistic Forecasts of COVID-19 Mortality in the United States.” Proceedings of the National Academy of Sciences 119 (15): e2113561119. https://doi.org/10.1073/pnas.2113561119.
Funk, Sebastian, Anton Camacho, Adam J. Kucharski, Rachel Lowe, Rosalind M. Eggo, and W. John Edmunds. 2019. “Assessing the Performance of Real-Time Epidemic Forecasts: A Case Study of Ebola in the Western Area Region of Sierra Leone, 2014-15.” PLOS Computational Biology 15 (2): e1006785. https://doi.org/10.1371/journal.pcbi.1006785.
Reich, Nicholas G, Justin Lessler, Sebastian Funk, et al. 2022. “Collaborative Hubs: Making the Most of Predictive Epidemic Modeling.” Am. J. Public Health, April, e1–4. https://doi.org/10.2105/ajph.2022.306831.
Reich, Nicholas G., Craig J. McGowan, Teresa K. Yamana, et al. 2019. “Accuracy of Real-Time Multi-Model Ensemble Forecasts for Seasonal Influenza in the U.S.” PLOS Computational Biology 15 (11): e1007486. https://doi.org/10.1371/journal.pcbi.1007486.
Shandross, Li, Emily Howerton, Lucie Contamin, et al. 2026. “Multi-Model Ensembles in Infectious Disease and Public Health: Methods, Interpretation, and Implementation in R.” Statistics in Medicine 45 (1-2): e70333. https://doi.org/10.1002/sim.70333.