Safe drinking water depends not only on treatment, but also on knowing when harmful microorganisms may be present in source waters. Researchers have now developed a data-driven framework that uses routinely measured water quality indicators to predict pathogen concentrations and estimate their potential health risks.
The study, published in Biocontaminant, combines machine learning with quantitative microbial risk assessment, or QMRA, to create an ML-QMRA framework for drinking water source monitoring. The approach could complement conventional microbial testing by providing faster estimates of contamination risks from readily available water quality data.
Routine monitoring already generates a large amount of environmental information. Our goal was to determine whether these commonly measured variables could also help us anticipate microbial contamination and translate those predictions into meaningful health risk estimates."
Changzheng Cui, corresponding author of the study, East China University of Science and Technology
The researchers collected 95 surface water samples from two drinking water sources in a major city in Eastern China between May 2024 and December 2025. They monitored three commonly used indicator bacteria, fecal coliforms, Escherichia coli, and Enterococcus faecalis, together with six pathogens: Pseudomonas aeruginosa, Salmonella spp., Shigella spp., adenovirus, norovirus, and enterovirus.
The results revealed an important limitation of conventional microbial indicators. Although the three fecal indicator bacteria were significantly correlated with one another, their relationships with viral pathogens were generally weak or inconsistent. This means bacterial indicators alone may not always reflect viral contamination accurately.
To improve prediction, the team compared six machine learning approaches, including Multiple Linear Regression, Least Squares Boosting, Decision Tree, Support Vector Machine, Random Forest, and Multilayer Perceptron models.
Random Forest and Decision Tree models performed particularly well, and all optimized models achieved R² values above 0.75. The Decision Tree model showed especially strong performance for P. aeruginosa, with an R² above 0.90. Independent data collected in January and February 2026 were also used for temporal validation, supporting the ability of most models to make predictions beyond the original training period.
The researchers then linked predicted pathogen concentrations to QMRA calculations expressed as disability-adjusted life years, or DALYs. Most estimated risks remained below the World Health Organization benchmark of 10−6 DALYs per person per year, but Salmonella spp., Shigella spp., and enterovirus showed probabilities of exceeding this benchmark under unfavorable exposure conditions.
The analysis also identified disinfection efficiency as the dominant factor influencing estimated health risk, emphasizing the importance of stable and effective drinking water treatment.
To make the machine learning models more transparent, the researchers used SHapley Additive exPlanations, or SHAP. Turbidity was among the strongest predictors for fecal indicator bacteria, accounting for 41.6% to 62.1% of predictive importance in those models. Temperature, dissolved oxygen, rainfall, and other water quality variables contributed differently across individual pathogens.
The study shows that routine physicochemical measurements can provide useful predictive information about pathogen levels and their associated potential health risks. However, the authors stress that the framework still requires validation across different watersheds, seasons, treatment systems, and land-use conditions.
Future integration of ML-QMRA models with real-time monitoring systems could help water managers identify periods of elevated microbial risk earlier and support more targeted water safety interventions.
Source:
Journal reference:
Guo, B., et al. (2026). A machine learning-quantitative microbial risk assessment (ML-QMRA) framework for predicting potential health risks from pathogens in drinking water sources. Biocontaminant. DOI: 10.48130/biocontam-0026-0009. https://www.maxapress.com/data/article/biocontam/preview/pdf/biocontam-0026-0009.pdf
Facts Only
* Researchers from East China University of Science and Technology developed an ML-QMRA framework.
* The study was published in the journal Biocontaminant.
* 95 surface water samples were collected from two drinking water sources in Eastern China.
* Sampling occurred between May 2024 and December 2025.
* Monitored indicator bacteria included fecal coliforms, Escherichia coli, and Enterococcus faecalis.
* Monitored pathogens included Pseudomonas aeruginosa, Salmonella spp., Shigella spp., adenovirus, norovirus, and enterovirus.
* Six machine learning models were tested: Multiple Linear Regression, Least Squares Boosting, Decision Tree, Support Vector Machine, Random Forest, and Multilayer Perceptron.
* Temporal validation data was collected in January and February 2026.
* Health risks were calculated as disability-adjusted life years (DALYs) against a WHO benchmark of 10−6.
* SHapley Additive exPlanations (SHAP) were used to determine predictor importance.
Executive Summary
A new data-driven framework, ML-QMRA, integrates machine learning with quantitative microbial risk assessment to predict pathogen concentrations in drinking water sources. By utilizing routinely measured water quality indicators—such as turbidity, temperature, and dissolved oxygen—the system provides faster risk estimates than conventional microbial testing. Research indicates that while traditional fecal indicator bacteria correlate with each other, they are poor predictors of viral contamination, highlighting a gap in standard monitoring protocols.
Among several tested models, Random Forest and Decision Trees demonstrated the highest predictive accuracy, with the latter performing exceptionally well for P. aeruginosa. While most estimated health risks remained within World Health Organization safety benchmarks, certain pathogens like enterovirus and Salmonella spp. showed potential to exceed these limits under unfavorable conditions. The effectiveness of water disinfection remains the primary factor in mitigating these risks. While promising, the framework's reliability across different seasons, land-use conditions, and watersheds remains to be fully validated.
Full Take
This study employs a robust academic approach by combining predictive modeling with established health metrics (DALYs). The methodology is sound in its comparison of six different ML architectures and its use of SHAP for model transparency, avoiding the "black box" problem common in AI. However, a peer reviewer would likely flag the sample size (n=95) as relatively small for training complex models like Multilayer Perceptrons, which increases the risk of overfitting despite the reported R² values. Furthermore, the temporal validation window (two months) is narrow compared to the training period, leaving questions about long-term seasonal stability.
The findings challenge the long-held reliance on fecal indicator bacteria as a universal proxy for water safety, specifically revealing their failure to track viral pathogens. This suggests that current regulatory reliance on these indicators may provide a false sense of security regarding viral loads. The conclusion that disinfection efficiency is the dominant risk factor is a proportionate finding—it confirms that while prediction is useful, it does not replace the physical necessity of treatment.
For this to move from a field study to a systemic utility, the models must be validated across diverse hydrogeological contexts. A critical follow-up study would involve deploying this framework in a watershed with significantly different land-use patterns (e.g., heavy industrial vs. agricultural) to test the universality of the predictors.
Bridge Questions:
1. If viral pathogens do not correlate with bacterial indicators, what specific physicochemical signatures consistently signal viral presence across different environments?
2. How does the introduction of predictive AI change the legal liability of water managers if a model predicts "safe" levels but a contamination event occurs?
Counterstrike Scan: A coordinated campaign would use these findings to undermine public trust in current water safety standards to push a specific monitoring product. The actual content remains focused on scientific validation and framework development, showing no structural alignment with such a pattern.
