Paper · Source package · Reproduction · Saved results · Static graphic
Marco Trotta · Irrigant · m@irrigant.xyz
MLET studies when a neural residual model should modify an existing satellite evapotranspiration estimate at an unseen location. The repository contains the research manuscript, experiment code, fixed protocols, saved predictions, and verification scripts. The manuscript is a preprint draft. It has no assigned arXiv identifier or peer-reviewed acceptance.
Can a selector identify when a neural correction improves OpenET, rather than only estimate the neural model's uncertainty?
For satellite estimate
Rejecting a correction returns OpenET. Every test observation receives a prediction and contributes to the reported error. The selector's target is the realized absolute-error benefit,
Positive benefit means that the correction helps relative to the satellite estimate. Low neural error and positive correction benefit are different objectives.
This study contributes an empirical analysis of selector transfer under spatial withholding and controlled input corruption. It does not introduce a new deferral algorithm or establish a universally better ET estimator. The literature comparison relates the study to regression with deferral, selective regression, and spatial applicability domains.
The predictor is an ensemble of three neural residual predictors, each with two 32-unit ReLU hidden layers.
Training uses Adam, 120 epochs, and training-only input and target standardization.
The fixed seeds are 20260713, 20260714, and 20260715.
All selectors within an outer configuration use the same fitted neural predictions.
| Method | Rule for using the correction |
|---|---|
| OpenET | Always return the satellite estimate. |
| Full | Always apply the ensemble mean correction. |
| Spread95 | Accept below the calibrated ensemble-disagreement threshold. |
| Support95 | Accept below the calibrated five-neighbor input-distance threshold. |
| Gain | Accept when a boosted regressor predicts positive benefit over OpenET. |
| SupportGain | Require both the support test and positive predicted benefit. |
| Uniform | Apply one correction multiplier selected from 0, 0.25, 0.5, 0.75, 1. |
| Clip | Clip weather inputs to training percentiles before applying the full correction. |
Inner spatial cross-fitting supplies the selector targets and thresholds. Outer test labels do not fit the selector, scaler, predictor, or thresholds. The earlier benchmark also includes affine calibration, weather ridge, combined ridge, and matched direct and residual nonlinear models. See the selective protocol and earlier benchmark protocol for complete specifications.
The target is energy-balance-corrected actual evapotranspiration, in mm/day, from processed flux measurements. Reference evapotranspiration is an input, not the response. The archive contains sparse satellite validation dates from 2001 through 2020, rather than complete daily time series.
| Evaluation population | Observations | Stations | Proximity groups |
|---|---|---|---|
| Complete-weather cohort | 7,923 | 85 | 63 |
| Matched input-fault probes | 7,873 | 84 | 62 |
| Unseen groups and later years | 649 | 27 | 24 |
| Later-year input-fault probes | 646 | 27 | 24 |
| Cropland subset of the complete cohort | 2,670 | 32 | 18 |
Sources: OpenET validation archive and processed flux archive. The data manifest, cohort, and split assignments bind the evaluated data. The complete-weather cohort selects 85 of 152 joined stations. Cropland classification does not establish irrigation status.
The primary evaluation withholds proximity groups across ten outer folds. Groups are connected components of station pairs within 10 km; transitive connections remain together. The joint evaluation also restricts training to dates before 2019 and tests 2019 through 2020 at unseen groups. Three inner spatial folds supply selection data. The joint inner split also separates dates before 2016 from later validation dates.
The primary metric is station-macro MAE: average absolute error within each station, then average equally across stations.
The paper also reports pooled MAE, RMSE, signed bias, cropland results, and correction acceptance rates.
Paired intervals use 2,000 bootstrap draws over proximity groups, with seed 20260922.
They condition on fitted models and ranks, omit retraining uncertainty, and have no multiplicity adjustment.
A learned benefit selector can fail under input corruption even when its training objective correctly compares against the fallback. However, returning OpenET more often can explain an apparent robustness advantage. The analysis therefore compares both fixed thresholds and equal acceptance budgets.
After omitting vapor-pressure deficit (VPD), the paired wind probe multiplies the test wind input by ten. It holds labels, satellite estimates, and fitted models fixed. On 7,873 observations at 84 stations in 62 groups, Full changes from 0.829 to 1.415 mm/day station MAE. OpenET scores 0.853 mm/day on these same rows. Support95 scores 0.857 mm/day but accepts only 1.74% of the station-weighted population. These point estimates describe the fixed probe; they do not establish a general accuracy gain.
The following post hoc comparison fixes acceptance at 50%. Each difference is the named method's expected station MAE minus Support's expected station MAE, in mm/day. Positive values favor Support.
| Probe | Comparison | Difference | Paired 95% interval |
|---|---|---|---|
| Spatial groups, wind input multiplied by 10 | Gain minus Support | 0.243 | [0.110, 0.412] |
| Spatial groups, wind input multiplied by 10 | Spread minus Support | 0.023 | [-0.022, 0.066] |
| Groups and later years, wind input multiplied by 10 | Gain minus Support | 0.032 | [-0.105, 0.194] |
The first two rows use 7,873 observations and 62 groups. The third uses 646 observations and 24 groups. The support-versus-disagreement comparison is inconclusive, and later years do not confirm the benefit-selector ordering. Equal-budget ranking uses label-free test scores and a randomized boundary decision. It does not establish a deployable threshold. Read the analysis protocol, paired comparisons, and complete curves.
Agricultural transfer also limits the conclusion. On 2,670 cropland observations from 32 stations in 18 groups, no-VPD Uniform scores 0.961 versus OpenET's 0.931 mm/day. Its paired improvement interval is [-0.074, 0.017] mm/day. Every no-VPD selective alternative has a higher cropland point error than OpenET in that spatial evaluation. The complete results retain all methods, input specifications, faults, and subgroups.
- The archive was inspected before this study. The selective protocol preceded its new run; the matched-budget analysis followed the threshold results.
- Synthetic faults test predictor corruption. They do not simulate weather interventions or estimate fault prevalence.
- Three seeds measure initialization variation. They do not provide calibrated posterior uncertainty.
- The study does not implement matched leading deferral surrogates, GeoQ, or graph neural processes.
- Spatial withholding applies to the added MLET models. It does not establish independence from upstream OpenET development.
- The evidence does not establish operational forecast skill, irrigation benefits, or water savings.
See study limitations for the full scope.
Use Python 3.13.5 to match the recorded numerical environment. Run these commands from the repository root:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements-paper.lock
PYTHONPATH=src python scripts/verify_ml_paper.py
PYTHONPATH=src python scripts/verify_selective_results.py
python -m pytest tests/test_ml_transfer_audit.py tests/test_ml_selective_residual.py -qThese commands verify saved evidence without downloading raw archives or fitting new models. The audits check partitions, target alignment, source hashes, selector arithmetic, negative controls, and original figure preservation. The pandas compatibility patch copies two masks; the exact recorded runner remains available for provenance. All 90 recorded inner partitions remain identical under the patch.
For data acquisition, full refits, figure generation, and paper compilation, follow the reproduction instructions. Completed experiments refuse to overwrite recorded results. The standalone reproduction package includes code, protocols, tests, and saved predictions.
The selective experiment runs 480 neural fits across 40 outer configurations and 120 inner partitions. Its recorded end-to-end wall time is 161.98 seconds, on macOS ARM with one numerical thread, for one complete run. This timing has no repeated-run uncertainty estimate and is not a comparative speed claim. Paid training compute is USD 0; electricity cost is unmeasured. Per-model fit and prediction timings for the earlier baseline comparison appear in the paper's compute table. The receipt records fit timings, seeds, and warnings; software versions record the environment.
| Path | Content |
|---|---|
scripts/ml_selective_residual.py |
Fixed selective-correction experiment. |
scripts/ml_transfer_audit.py |
Earlier model benchmark and shared partition logic. |
docs/evaluation/ |
Scientific protocols and literature positioning. |
docs/results/ml_selective/ |
Selectors, inner predictions, outer predictions, and robustness probes. |
docs/results/ml_transfer/ |
Earlier benchmark, humidity audit, and sensitivity results. |
manuscript/arxiv/ |
Canonical LaTeX, references, tables, and figures. |
output/arxiv/ |
Submission source, metadata, and reproduction package. |
tests/ |
Software, partition, loss, provenance, and manuscript checks. |
For general package development, install the test extra and run the repository gate before a commit:
python3 -m pip install -e ".[test]"
./scripts/verify.shThe gate runs the test suite, checks serving-path isolation, and builds the recorded ETo candidate site.
See software reproducibility for the general environment contract.
The exact paper environment is pinned separately in requirements-paper.lock.
The repository also retains earlier research components:
- Idaho reference-ETo outlook and its evaluation protocol.
- Historical OpenET information-value benchmark.
- FAO-56 soil-water scaffold and separate residual-model protocol.
- NeuralHydrology provenance and vendored pyfao56 provenance.
The outlook remains a research candidate with incomplete validation. These components do not supply additional evidence for the selective-correction paper.
Cite the current manuscript as an unpublished research preprint. No arXiv identifier is assigned.
@misc{trotta2026mlet,
author = {Trotta, Marco},
title = {MLET: Selective Neural Residual Correction for Spatial Evapotranspiration},
year = {2026},
howpublished = {Research manuscript and accompanying code},
url = {https://github.com/marco-trotta1/MLET}
}Meetpal S. Kukal receives acknowledgement for research mentorship and earlier feedback. The paper cites the source data and discloses AI assistance in coding, analysis, figures, and drafting. New paper graphics adapt figures4papers, with the stated CC BY-NC 4.0 license. The original MLET visuals and Irrigant logo remain separate existing assets. The README wordmark is an original decorative SVG; its animation is not a model result. No repository-wide license is currently declared. Vendored components retain their own licenses and provenance.
This README uses Chronos, TimesFM, and NeuralHydrology as structural precedents. Their paper links, concise usage paths, and citation sections inform the organization. Their models and results are not MLET baselines.