Interpretable and Risk-Aware Machine Learning for Environmental Suitability, Exposure, and Decision Systems
View/ Open
Date
2026-08-02Type of Degree
PhD DissertationDepartment
Computer Science and Software Engineering
Metadata
Show full item recordAbstract
Environmental and biological systems increasingly require prediction and decision-making under uncertainty. This dissertation develops and applies interpretable and risk-aware machine-learning approaches for environmental suitability, environmental expo-sure, and sequential decision systems. The central objective is to translate complex ecological, public-health, and decision-process data into useful predictions while making model interpretation, validation limits, and decision risk explicit. Chapter 2 synthesizes related work on environmental machine learning, forest species distribution modeling, pollen-centered environmental health prediction, model interpretability, and risk-aware policy learning. It provides the shared conceptual frame-work for the three studies that follow: machine-learning outputs should be interpreted as decision-relevant evidence only when their data limitations, predictor correlations, validation design, and uncertainty are made explicit. Chapter 3 develops a climate-regeneration-fire informatics framework for future suit-ability mapping of loblolly pine, shortleaf pine, slash pine, and longleaf pine in the south-eastern United States. The framework integrated 79,349 adult and 22,969 seedling oc-currence records compiled from the U.S. Department of Agriculture Forest Service Forest Inventory and Analysis database (FIADB), mapped fire-history intersection, four historical climate baselines, variance-inflation screening, seven species distribution modeling algo-rithms, Random Forest interpretation, future climate projections, percentile-envelope ex-trapolation screening, and management-oriented spatial summaries. Adult and seedling records overlapped broadly but were not spatially or climatically interchangeable, indi-cating that adult persistence and early regeneration should not be treated as the same suitability signal. Wildfire stratification was interpreted as a comparison between mapped 2 fire-history-intersected and non-intersected landscape strata rather than as a causal fire-effects analysis. XGBoost and Random Forest were the most consistent algorithms, and Random Forest interpretation showed that moisture availability, thermal energy, cold limi-tation, growing-season timing, evaporative demand, heat-moisture balance, and season-ality contributed to modeled suitability. Future projections were translated into manage-ment classes such as seedling-supported suitability, adult-only suitability, future gain, fu-ture loss, non-analog uncertainty, and aggregate mill-proximity consensus. Chapter 4 applies ecological machine learning to environmental health by evaluat-ing whether airborne pollen signatures provide meaningful environmental information for county-level lung cancer incidence and mortality prediction. The analysis used county-year observations for 11 U.S. counties from 2015 through 2021 and combined five daily-derived annual active-sum pollen variables with air pollutants, Air Quality Index summaries, humidity, lag-window exposure engineering, Random Forest prediction, permutation im-portance, SHAP interpretation, and county-grouped validation. Pollen was not treated as an established lung carcinogen; instead, it was evaluated as a climate-sensitive aeroaller-gen and environmental-regime marker that may be informative alongside air pollution and humidity. Pooled Random Forest models reproduced county-year patterns more strongly for incidence than mortality, but county-grouped validation showed poor geographic trans-fer. The chapter therefore interprets pollen-associated signals as exploratory, ecological, and hypothesis-generating rather than causal or ready for external prediction. Chapter 5 develops TOPS, a transition-based volatility-controlled policy-search frame-work for risk-averse reinforcement learning. TOPS addresses a general problem exposed by the applied chapters: decisions based only on average predicted performance can be unsafe when outcomes are uncertain, variable, or difficult to validate externally. The frame-work controls reward volatility, learns from short trajectories or transition segments rather than requiring long uninterrupted rollouts, and provides global convergence guarantees under over-parameterized neural-network policy representations. This theoretical chapter 3 connects environmental prediction to risk-aware decision-making by showing how learning systems can optimize policies while accounting for variability and hazardous outcomes. Chapter 6 outlines future work that links the three studies into a broader research program. Key directions include applying risk-averse policy learning to forest adaptation and environmental-health decision support, expanding pollen and lung cancer analysis with richer individual-level and spatial-temporal data, validating forest suitability classes with independent survival and recruitment data, improving uncertainty quantification and causal interpretation, and developing reproducible tools that communicate model uncer-tainty to stakeholders.
