Google leads CDC flu forecast evaluation for 2026 season
The CDC ranks Googleโs flu forecasting system first among 39 eligible models. The result concerns hospital admission forecasts across a defined season.

A Google forecasting system ranked first among 39 eligible models in the CDCโs newly published evaluation of influenza hospital admissions for the 2025 to 2026 season. The September 30 FluSight report names Google SAI FluEns as the leading individual team submission. The combined FluSight ensemble used for CDC forecast messaging ranked seventh.
This is an evaluation of population level forecasting. It measures how well submitted projections matched hospital admission data during a defined season. The result does not establish an ability to diagnose individual patients, choose treatments or predict every kind of disease outbreak.
How the ranking was calculated
FluSight teams forecast weekly hospital admissions for the current week and up to three weeks ahead. The main ranking uses relative weighted interval score, which evaluates how well a forecastโs uncertainty ranges match observed outcomes. Lower scores indicate better performance relative to the comparison baseline.
The CDC excluded national forecasts from this ranking because of differences in scale and excluded Puerto Rico because of data availability. Models also needed to submit at least 75 percent of the remaining forecast targets. These conditions matter when interpreting the first place result. It is a comparison within the reportโs selected forecasts and jurisdictions.
AI helped develop the forecasting software
In its September 30 announcement, Google connects the forecasting work to Empirical Research Assistance, or ERA. The system helps researchers develop and optimize scientific software.
A May research preprint explains the disease forecasting approach. A language model generates and revises executable code, while a search process explores alternatives and uses measured forecasting performance to guide further changes. Researchers supplied ideas from established methods and combined selected models into ensembles.
The paper describes forecasts submitted prospectively during the respiratory season, with public submission timestamps. That design matters because a forecast is recorded before the outcome it is trying to predict becomes available. The work also emphasizes expertise in guiding development and selecting models. Readers should distinguish that research account from the CDCโs later season evaluation.
The underlying ERA system was described in a Nature paper published in May. The fresh development is the September evaluation, which adds an external assessment of the flu forecasts rather than announcing ERA for the first time.
A useful result with a defined scope
The CDC report also notes that forecast performance declined around rapid changes in influenza trends. A season ranking should therefore be read alongside uncertainty and performance at difficult turning points. It is not a percentage accuracy claim or a guarantee about the next season.
The practical significance is evidence that AI assisted software development can contribute to a demanding forecasting workflow. The useful next questions concern repeatability across seasons, behavior during sharp changes and the reliability of the data feeding the models. Those questions are more informative than treating a leaderboard position as general clinical intelligence.
For related work on reusable scientific workflows, read ByteForwardโs report on the BootLoops research toolkit.
Featured image is an original AI generated editorial illustration.



