Projects

Selected research in atmospheric modeling, geostatistical data fusion, and scientific software. Each entry states the problem addressed, the approach taken, and the principal findings.

Global surface ozone data fusion

A 33-Year Global Reanalysis of Surface Ozone by Bayesian Maximum Entropy Data Fusion

Monthly MDA8 ozone · 0.1° global grid · 1990–2022

Long-term exposure to surface ozone is associated with respiratory mortality, cardiovascular disease, and chronic lung disease, yet the spatially resolved and temporally complete exposure fields required to quantify these effects at the global scale do not exist. Regulatory monitoring networks are concentrated in North America and Europe, leaving much of Africa, South Asia, and South America with sparse or no station coverage; chemical transport models (CTMs) and satellite retrievals achieve global coverage but carry systematic biases and spatially varying uncertainty. This study applies a Bayesian Maximum Entropy (BME) data fusion framework that integrates TOAR-II ground observations with bias-corrected CTM ensemble output (M3Fusion) and satellite-derived surface ozone retrievals (OMI-MLS, IASI+GOME2, CrIS) to produce monthly maximum daily 8-hour average (MDA8) ozone estimates at 0.1° resolution worldwide for 1990–2022. Biases in each model and satellite product are corrected using flexible RAMP (f-RAMP), a localized nonlinear method that accommodates heteroscedasticity and spatial non-stationarity and returns a corrected mean and variance at every grid cell and time step.

Under spatially independent checkerboard validation (5° boxes), fusing M3Fusion with ground observations improved R2 from 0.68 to 0.77 and reduced RMSE from 7.97 to 6.93 ppb relative to an observations-only baseline over 1990–2004; the further addition of satellite retrievals raised R2 from 0.807 to 0.812 and lowered RMSE from 5.74 to 5.60 ppb over 2017–2020. The resulting dataset is novel in three respects: it extends BME-based global ozone estimation to monthly resolution across 33 years; it introduces f-RAMP, which converts each model and satellite product into a corrected mean with a localized variance rather than treating model output as exact; and it is the first BME application for surface MDA8 ozone to fuse a CTM ensemble and multiple satellite retrievals within a single estimation step. The product is formatted for use in global burden of disease assessment and agricultural yield-loss evaluation.

Methods and tools: Bayesian Maximum Entropy (BMElib), f-RAMP bias correction, spatiotemporal covariance modeling, checkerboard cross-validation, MATLAB, Python, HPC

Gulf States air toxics exposure disparities

Exposure Disparities to Petrochemical Air Toxics in the U.S. Gulf States, 2011–2016

Daily 4 km BME estimates · SBTEX and 1,3-butadiene · Population weighted exposure ratios

The U.S. Gulf States harbor a high density of petrochemical industries that emit air toxics including styrene, benzene, toluene, ethylbenzene, and xylenes (SBTEX). Although fine-resolution concentration estimates exist for this region, little is known about the population weighted exposure (PWE) experienced by population subgroups, and no estimates existed for the carcinogen 1,3-butadiene. This study developed daily fine-resolution (4 km) Bayesian Maximum Entropy (BME) data fusion estimates of 1,3-butadiene concentrations across the Gulf States for 2011–2016, and used them, together with existing SBTEX estimates, to calculate the domain-wide population weighted exposure ratio (PWER) for non-white relative to white populations. Values greater than one quantify the extent to which a subgroup is more exposed than the reference group.

In 2011, the domain-wide PWER for non-white versus white populations exceeded one for all six pollutants, ranging from 1.035 for benzene to 1.215 for 1,3-butadiene. Among counties with populations above 100,000, the highest PWER values were found in East Baton Rouge Parish, Louisiana, and Nueces County, Texas. The PWER increased for all six pollutants between 2011 and 2016, indicating that although ambient concentrations of several pollutants declined over the period, the exposure gap between non-white and white populations widened. These results identify both the pollutants and the specific jurisdictions where the environmental justice burden is greatest, and provide the first fine-resolution 1,3-butadiene exposure surface for the region.

Methods and tools: Bayesian Maximum Entropy data fusion, bias-corrected CAMx output, emissions-informed land use regression, census-block population weighting, MATLAB, Python, R

Aircraft Dispersion Model chemistry module

Implementation and Evaluation of Chemistry Schemes in an Aircraft Dispersion Model

2021–2023 · FAA ASCENT · UNC Institute for the Environment

Airports are a significant urban source of NOx, and NO2 is a criteria pollutant governed by a 1-hour standard, yet regulatory dispersion models represent NOx to NO2 conversion through simplified schemes that assume near-instantaneous or ratio-scaled conversion. These approximations are least defensible for the hot, buoyant plumes released along runways, taxiways, and flight tracks. This work developed the chemistry module of the Aircraft Dispersion Model (ADM), a research model formulated specifically for airport plumes, and established its performance relative to established alternatives. Six NO2 chemistry treatments were implemented in ADM (GRS, MGRS, TTM/TTRM, OLM, ARM/ARM2, PVMRM2, GRSM) and evaluated alongside five AERMOD schemes at four monitoring sites in the Los Angeles International Airport domain over summer and winter 2012 episodes, with explicit chemistry (MCM v3.3.1 in F0AM) applied to the same inputs as a benchmark. A complete meteorological and emissions input chain (WRF v4.2.2 → MMIF → AERMOD v21112; AEDT → ADM) was constructed to support the evaluation, and the module was extended to secondary PM2.5 through an MGRS aerosol treatment with COBRA thermodynamic partitioning.

The principal finding is that NO2 prediction skill is governed by NOx prediction skill rather than by the sophistication of the chemistry scheme: expressed in normalized NO2 versus NOx space, results from all schemes collapse onto the 1:1 line, attributing the persistent NO2 underprediction to the dispersion treatment and to unrepresented ground support equipment emissions rather than to the chemistry. Observed NO2/NOx ratios track the diurnal ozone cycle, consistent with ozone controlling conversion while rarely limiting it at these receptors. OLM and TTRM yield nearly identical ratios, indicating that plume travel time is not a controlling factor at these source-receptor distances, whereas GRSM produces near-identical ratios across all four receptors, revealing an over-dependence on the prescribed background ozone field. Of the simple schemes, GRS performed best on the Carruthers et al. (2017) correct-triangle metric (78.95% versus 34.21% for OLM). Six substantive defects were identified and corrected in the inherited model code, two of which had silently corrupted all previous multi-day results.

Methods and tools: MATLAB, Python, Fortran, ADM, AERMOD v21112, WRF v4.2.2, MMIF, MCIP, F0AM/MCM v3.3.1, stiff ODE solvers, SLURM HPC, Git

Low-cost sensor data fusion

Integrating Low-Cost Sensor Networks into Air Quality Data Fusion

2019–2022 · U.S. Environmental Protection Agency

Regulatory monitoring is accurate but sparse, low-cost sensor (LCS) networks such as PurpleAir are dense but uncalibrated and of unknown quality, and chemical transport models and satellite retrievals provide complete coverage at the cost of systematic bias. The question posed by the sponsor was what is already established about combining these sources, and whether and how LCS observations can responsibly be incorporated. The work proceeded through a systematic literature review, a Quality Assurance Project Plan, a methods specification, and a production implementation. An ISI Web of Science search syntax was developed and refined to yield 45 high-relevance articles from an initial 169; the corpus was reviewed and scored, and summary tables pairing calibration methods with the fusion method used in each study formed the analytical basis of the delivered report. Four fusion methods were specified for the study: universal kriging, Bayesian Maximum Entropy, machine learning, and land use regression implemented by machine learning.

An acquisition pipeline was built covering the public PurpleAir network as it expanded from approximately 11,000 to 23,000 sensors, with quality classification based on dual-channel (A/B) agreement and geographic assignment performed against shapefiles, the network providing no state or county field. Verification against the provider's own published averages identified an upstream data-integrity defect: PM2.5 is published under a different column name prior to 20 October 2020, which would otherwise have silently corrupted downstream analysis. Sensors were co-located with EPA AQS monitors within a 100 m radius across California, and seven calibration approaches spanning the complexity range, from published correction factors through linear regression to random forest, k-nearest neighbors, and artificial neural networks, were evaluated on R2 and MSE. BME estimation with global-mean-trend detrending demonstrated that estimates distant from observations relax correctly to a non-zero background, which they do not without detrending. All deliverables were accepted by the sponsor under a formal review cycle.

Methods and tools: Python (xarray, Dask, scikit-learn, pymc3, holoviews), R, MATLAB (BMElib), SQL, Google Earth Engine, NetCDF/HDF, Longleaf HPC

AEDT airport emissions inventories

Construction of Airport Emissions Inventories from Radar Flight-Track Data

2022–2023 · FAA ASCENT · with Volpe National Transportation Systems Center

The Aviation Environmental Design Tool (AEDT) is the FAA's official system for computing aircraft noise, fuel burn, and emissions, and every U.S. aviation environmental assessment depends on it. Two inventories were required: one for Los Angeles International Airport, to supply the dispersion modeling described above and to independently reproduce the sponsor's reference runs, and one for Boston Logan International Airport, for which no prepared AEDT input existed. Establishing the platform required determining an undocumented version compatibility matrix and resolving a sequence of application, database, and operating-system failures in coordination with AEDT support, Volpe, and institutional IT. For LAX, approximately 140 GB of PDARS radar data were ingested and reformatted into the AEDT study database, and the sponsor's published results were reproduced; residual differences were traced to stochasticity in AEDT's internal solver rather than to configuration error, establishing the installation as validated.

The Boston Logan inventory was principally a data reconstruction problem: radar records describe where aircraft flew, whereas AEDT requires the aircraft type, engine, runway, and track for each operation. Airframe model, engine code, and equipment identifiers were recovered through SQL joins against the AEDT reference databases after the fleet detail table was found to be empty, a finding confirmed by the vendor. Operations were assigned to one of six runways by nearest point-to-line geometry, and flight tracks were synthesized from taxiway shapefiles and centroids using the AEDT Airport Designer, with the approach validated by perturbing tracks in an existing study and observing the effect on emissions and dispersion output. The completed study executed 324,000 operations in approximately 90 minutes, and a colour-coded data-gaps document was delivered to the sponsor and collaborating institution to drive the outstanding external data requests.

Methods and tools: AEDT 3d/3e, AEDT Airport Designer, SQL Server and T-SQL, R, Python, geospatial nearest-feature joins, shapefile processing, PDARS radar data

CMAQ-DDM source apportionment

Source Apportionment of Oil and Gas Emissions Using CMAQ-DDM

2021–2022 · Environmental Defense Fund · 100 simulations · ~25 TB

The Decoupled Direct Method (DDM) computes sensitivity coefficients during a CMAQ simulation rather than requiring a separate brute-force run per source group, and is the efficient means of establishing which states' oil and gas emissions contribute to PM2.5 and where. In operational terms the campaign nonetheless comprised 100 simulations, stratified by state and emission sector, generating approximately 25 TB of output on shared high-performance computing storage subject to hard quotas and a 21-day scratch expiry. The campaign was taken from initiation to completion, post-processing, compression, and archival. Job scripts were produced by Python and csh generators rather than by hand, yielding roughly 200 control files; staging scripts moved completed output to bulk storage without disturbing in-flight simulations. An attempt to write DDM output directly to bulk storage failed with MPI errors and is recorded as a negative result.

The most consequential outcome was a quality assurance finding in a parallel campaign. Verification plots of monthly-average PM2.5 returned NaN for one scenario group across all segments, with no corresponding indication in the model logs. Because the failure was scenario- and period-specific rather than random, the fault was localized to the inputs and traced to point-source emission files having been supplied to runs requiring a different sector input, the consequence of a reused build script. The scripts for every other case were audited, regenerated, and the twenty affected simulations re-run. Visualization was moved from VERDI, which renders one figure at a time and proved unworkable at this scale, to custom Python routines producing more than 80 comparable per-state figures with consistent legend ranges, together with the location of concentration extrema within the CONUS domain.

Methods and tools: CMAQ v5.2 with DDM-3D, SMOKE, Python (xarray, pandas, monetio), csh/bash, SLURM, tmux, tiered HPC storage

NASA Airathon machine learning

Machine Learning Prediction of Surface PM2.5 and NO2 from Satellite Retrievals

2022 · NASA Airathon (DrivenData) · 9th place, trace gas track

Satellites observe column-integrated quantities (aerosol optical depth for particulates and tropospheric vertical column density for NO2), whereas human exposure depends on surface concentration, and the relationship between the two is mediated by boundary layer height, humidity, aerosol composition, and vertical mixing, none of which the instrument observes directly. This project developed pipelines predicting daily surface PM2.5 from MAIAC (MCD19A2) and daily NO2 from TROPOMI across Los Angeles, Delhi, and Taipei. Each product was radiometrically calibrated from scaled integers using its valid range, fill value, scale factor, and offset; rasters were spatially aligned to the estimation grid, subset to the extents actually required, and aggregated from raster to vector per cell and time step. Cloud-obscured retrievals in the NO2 track were filled with a backward-looking moving average rather than kriging or a centered interpolation, since any method drawing on subsequent observations leaks future information into historical predictions under a temporal train-test split.

Eighteen models were implemented and compared per track under an 80/20 validation split of the 2018–2020 training period, with final models retrained on the complete training set before prediction. Gradient boosting on decision trees (CatBoost) performed best in both tracks (R2 = 0.75, RMSE = 43.36 µg/m3 for PM2.5; R2 = 0.63, RMSE = 14.00 ppb for NO2), a result consistent with its native handling of the categorical date-time features and with the threshold-like column-to-surface relationship. Random forest ranked second and third and multiple linear regression eighth in both tracks, confirming that the relationship is substantively non-linear. The entry reached ninth place on the NO2 leaderboard.

Methods and tools: Python, CatBoost, random forest, k-nearest neighbors, kriging, MAIAC MCD19A2, TROPOMI (Sentinel-5P), HDF/NetCDF/GeoTIFF

MBMEGUI Software

MBMEGUI: Software for Modern Bayesian Maximum Entropy Space/Time Analysis

Scientific software · UNC Chapel Hill, Environmental Sciences and Engineering

The Bayesian Maximum Entropy framework is mathematically demanding to apply, which limits its use by researchers and students whose questions it would otherwise serve well. MBMEGUI is a graphical software package implementing the Modern BME framework for space/time mapping and analysis, providing data integration, covariance modeling, estimation, and visualization through an interface that does not require the user to write BMElib code directly. The package is used in environmental science instruction at UNC-Chapel Hill and preserves the rigor of the underlying formalism while removing the implementation barrier to its use.

Methods and tools: MATLAB, BMElib, GUI development, spatiotemporal geostatistics, scientific software packaging and documentation