We thrive on Big Data, but hydrology has a dangerous sample size problem. We often use only 30 to 50 years of data to predict extreme 10,000-year flood events.
Training AI models to extrapolate using statistical distributions into a non-stationary future driven by climate change creates massive epistemic uncertainty. Without quantifying parameter error through Monte Carlo simulations, we are building high-tech infrastructure on low-probability sand.
We need to stop selling "certainty" in flood risk and start building Intelligent Systems that explicitly embrace the variance.
Are we ready to admit our training sets are simply too short for the risks we face?