QMED
Can Machine Learning Improve QMED?
- Version
- 1.0.0
- Status
- Current
- Published
- Evidence cut-off
Abstract
A national revalidation asked whether an ungauged machine-learning challenger improved QMED accuracy once the evaluation matched geographic transfer and donor information was kept inside the training folds. It did not. The result is useful because it separates a failed accuracy hypothesis from the access, repeatability and auditability that remain operationally valuable.
Findings at a glance
- Evaluation cohort
- 918Held-out stations in the national comparison918 held-out stations
- ML challenger R²
- 0.898Spatial-cluster cross-validation918 held-out stations
- Donor-adjusted SM2025 R²
- 0.922Train-only donor custody918 held-out stations
- Paired ΔR²
- −0.02495% CI −0.031 to −0.016918 held-out stations
Evidence distributions
Ungauged QMED under spatial-cluster cross-validation
918 held-out stations
| Method | R² | Median absolute error | FSE |
|---|---|---|---|
| ML challenger | 0.898 | 26.1% | 1.634 |
| Donor-adjusted SM2025 | 0.922 | 23.4% | 1.538 |
The challenger did not beat the benchmark. Paired ΔR² was −0.024 with 95% CI −0.031 to −0.016.
How validation design and donor custody change R²
918 stations in each validation-design comparison
| Question | More permissive design | Geographic-transfer design |
|---|---|---|
| ML validation | Random K-fold: 0.915 | Spatial-cluster cross-validation: 0.898 |
| SM2025 donor custody | All donors: 0.940 | Train-only donors: 0.922 |
Random K-fold evaluation made the ML result look stronger than spatial-cluster cross-validation. Allowing held-out neighbours to remain available as donors likewise overstated the donor-adjusted baseline; the fair comparison used train-only donors.
Frozen input and information-custody audit
918 stations in the frozen cohort
| Audit check | Frozen status | Consequence |
|---|---|---|
| SM2025 estimate | Input to the challenger | Secondary challenger |
| URBEXT | Zero throughout | Urban-adjustment chain not exercised |
The challenger was not informationally separate from SM2025, and the frozen cohort did not test an active urban-adjustment path. These are claim boundaries, not implementation footnotes.
Small-catchment performance under spatial-cluster cross-validation
26 held-out stations below 10 km² and 55 held-out stations from 10–25 km²
| Catchment area | Sample size | ML R² | ML median absolute error | Donor-adjusted SM2025 R² | Donor-adjusted SM2025 median absolute error |
|---|---|---|---|---|---|
| Below 10 km² | 26 | 0.300 | 32.6% | 0.305 | 26.9% |
| 10–25 km² | 55 | 0.743 | 30.8% | 0.818 | 29.0% |
Neither small-catchment band supports an accuracy win. The below-10 km² comparison is especially unstable because its sample is small, and both methods have weak R² there; the 10–25 km² band also favours donor-adjusted SM2025.
What survived
National open-input access
The study retained a route to national QMED screening from open-input feature families, without turning that access claim into an accuracy claim.
Repeatability
A frozen cohort, explicit fold construction, train-only donor rule and evidence revision make the negative result repeatable.
Auditability
Exact paired and size-band results expose where the challenger failed and which information each side received.
What failed
The accuracy-win hypothesis
The challenger remained behind donor-adjusted SM2025 on R², median absolute error and FSE under the geographic-transfer evaluation.
Research question
Can a machine-learning challenger improve ungauged QMED accuracy over donor-adjusted SM2025 when both methods are judged on the same held-out stations and only training-fold donors are available? The competing hypotheses were an accuracy improvement from nonlinear combination of national inputs, or no improvement once geographic transfer and donor custody were handled fairly.
Data and cohort
The frozen cohort contains 918 stations whose QMED targets were derived from at least eight non-rejected AMAX years. The challenger used broad families describing catchment scale, hydroclimate, terrain, land cover and river-network context, together with the statistical-method estimate. Exact feature construction and selection are outside the public disclosure boundary.
Results are reported for the national cohort and for two deliberately visible small-catchment strata. The below 10 km² stratum has 26 stations and the 10–25 km² stratum has 55, so the subgroup evidence must be read with its sample size rather than as a stable guarantee for an individual catchment.
Method and validation question
Random K-fold validation asks how the fitted relationship interpolates when training stations may remain geographically close to evaluation stations. Spatial-cluster cross-validation asks the harder operational question: how well does the method transfer when a centroid-based geographic group is held out? The experiment used five unbuffered centroid/KMeans clusters. No distance buffer was imposed, so the result does not assert a protected separation distance.
Performance was evaluated in log space using R² and FSE, with median absolute error reported in flow space. The paired station result is the primary verdict; random-fold and donor-custody comparisons are sensitivity checks that explain why a more permissive design can produce a stronger-looking headline.
Information asymmetry and donor custody
Donor selection is part of the evaluation. An all-donor baseline can use nearby gauges that belong to the held-out population, whereas the train-only baseline limits donor information to what would be available within the fitted side of each fold. The primary comparison therefore uses train-only donors.
SM2025 was an input to the challenger. It is therefore a secondary challenger that tries to refine an established statistical signal, not a separate estimate built without that signal. URBEXT was zero throughout the frozen cohort, so the urban-adjustment chain was not exercised. Both facts narrow the claim before any score is read.
Results and the negative verdict
The challenger did not beat the benchmark. It was behind donor-adjusted SM2025 on the national R², median absolute error and FSE results, and the paired R² interval remained on the unfavourable side of zero. The result is a failed accuracy hypothesis, not an ambiguous tie.
The validation-design table shows why the random-fold result was not enough. Geographic grouping reduced the challenger score, while enforcing train-only donor custody reduced the apparent donor result. Fairness required both corrections at once. The small-catchment table also shows no recovery of the hypothesis: SM2025 retained lower median error in both reported bands and stronger R² in the larger of the two bands.
Small-catchment instability
Small catchments are not a footnote to the national score. The below-10 km² band combines a small sample with weak R² for both methods, so small changes in station composition can materially alter the statistic. The 10–25 km² band is larger but still modest and remains adverse to the challenger. These are cohort diagnostics, not local uncertainty intervals.
Operational implications
The accuracy-win claim failed, but the exercise retained three useful properties. National open-input access supports screening without presenting access as superior prediction. A frozen evidence revision, stated fold construction and train-only donor rule make the result repeatable. Exact tables, subgroup results and explicit information custody make it auditable by a hydrologist who cannot see the private implementation.
In practice, the challenger should not displace donor-adjusted SM2025 on the evidence reported here. Its surviving value is as an auditable research and review layer: it can expose sensitivity, support national comparison and test future hypotheses under the same fair protocol.
Limitations
- The geographic clusters were centroid-based and unbuffered, so the study did not test transfer across a guaranteed separation distance.
- SM2025 was already present in the challenger inputs, which limits the comparison to a secondary challenger rather than a standalone alternative.
- URBEXT carried no variation, so the experiment provides no evidence about performance when the urban-adjustment chain is active.
- The smallest area bands contain limited station counts and cannot support a catchment-specific performance guarantee.
- Cross-validation describes this frozen cohort and evidence cut-off; it does not remove rating, target or descriptor uncertainty.
What we do not conclude
- We do not conclude that the challenger improves QMED point accuracy over donor-adjusted SM2025.
- We do not conclude that machine learning has no role in flood hydrology; this verdict applies to the tested QMED challenger and frozen validation design.
- We do not conclude that the result validates urban QMED behaviour, establishes site-specific uncertainty or replaces analyst review.
- We do not claim approval, certification or replacement of official flood-estimation methods.
Disclosure boundary
We disclose the research question, broad cohort rules, feature families, validation logic, donor custody, information asymmetry, exact national and subgroup results, failed hypothesis, limitations and evidence revision. We withhold exact feature engineering and selection, learned coefficients, model weights, full training recipes, private data-construction mechanics, code-level workflow logic and operational infrastructure. The public record is sufficient to audit the claim without providing a deployable reconstruction.
Claim record
Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.
Demonstrated
Under spatial-cluster cross-validation on 918 stations, the ungauged ML challenger achieved R² 0.898, median absolute error 26.1% and FSE 1.634; donor-adjusted SM2025 achieved R² 0.922, median absolute error 23.4% and FSE 1.538. ΔR² was −0.024 [95% CI −0.031, −0.016].
- Interpretation
- The challenger did not beat the benchmark under the validation design aligned with geographic transfer.
- Boundary
- SM2025 was an input to the challenger, so this is a secondary challenger comparison. URBEXT was zero, so the urban-adjustment chain was not exercised.
- Linked result
- Ungauged QMED under spatial-cluster cross-validation; sample: 918 held-out stations
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Demonstrated
The ML challenger scored R² 0.915 under random K-fold validation and 0.898 under spatial-cluster cross-validation; the donor-adjusted SM2025 baseline scored 0.940 with all donors and 0.922 with train-only donors.
- Interpretation
- Validation design and donor custody both change the apparent result, so both belong in the claim rather than in implementation footnotes.
- Boundary
- The geographic design used five unbuffered centroid/KMeans clusters. It did not impose a distance buffer around held-out stations.
- Linked result
- How validation design and donor custody change R²; sample: 918 stations in each validation-design comparison
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Demonstrated
Below 10 km², ML achieved R² 0.300 and median absolute error 32.6% versus SM2025 at 0.305 and 26.9% across 26 stations; from 10–25 km², ML achieved 0.743 and 30.8% versus SM2025 at 0.818 and 29.0% across 55 stations.
- Interpretation
- The reported small-catchment strata do not recover the accuracy-win hypothesis.
- Boundary
- The strata are small, especially below 10 km², and are descriptive subgroup results rather than stable site-specific performance guarantees.
- Linked result
- Small-catchment performance under spatial-cluster cross-validation; sample: 26 held-out stations below 10 km² and 55 held-out stations from 10–25 km²
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Demonstrated
SM2025 was an input to the challenger, while URBEXT was zero throughout the frozen cohort.
- Interpretation
- The model was tested as a secondary challenger to SM2025, and the urban-adjustment chain received no exercise in this cohort.
- Boundary
- The experiment cannot establish performance for a separate-from-SM2025 estimator or for urban catchments with non-zero URBEXT.
- Linked result
- Frozen input and information-custody audit; sample: 918 stations in the frozen cohort
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Retracted
The accuracy-win hypothesis for the ML challenger was withdrawn.
- Interpretation
- It did not survive spatial-cluster cross-validation against donor-adjusted SM2025 with train-only donor custody.
- Boundary
- This withdraws the tested accuracy claim; it does not show that every possible use of machine learning in flood estimation will fail.
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Inference
Open-input access, repeatability and auditability remain useful operational properties even without an accuracy advantage.
- Interpretation
- A method can improve evidence handling and national screening without improving held-out point accuracy.
- Boundary
- This is an operational interpretation, not a demonstrated reduction in project cost, delivery time or design-flow uncertainty.
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Primary sources
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
Version history
- Version 1.0.0current
Initial publication of the ML QMED adversarial revalidation.
Related reports
- National hydrologyWhat National-Scale Analysis Reveals About UK HydrologyA network-scale examination of gauge consistency, QMED validation, record-length limits and temporal sensitivity, with each result kept inside its evidential boundary.
- Statistical methodsThe 2025 Statistical Method: What Changed and Why It MattersA practitioner account of the FEH 2025 descriptor, QMED, donor, pooling, urban, distribution and uncertainty changes, with Hydrometric’s implementation evidence kept inside its local verification boundary.
- Growth curvesSkew, Scale and the Shape of UK Flood GrowthA bounded account of index-flood scale, L-CV dispersion, L-skew shape, GF100 growth, a gauged-only rating result and one frozen finite-sample model failure.
- UncertaintyHow Certain Is a Design Flood?A layered account of population validation scatter, network inconsistency, within-record sampling and the still-open model-family protocol.
- Rainfall-runoff modellingFrom ReFH Design Events to Continuous HydrologyA bounded record of the baseflow-lag refit, the public ReFH1 baseline gap, the failed daily two-store branch and the gated status of regional continuous evidence.
- Flood estimation practiceThe Modern Flood Estimation ReportA practical evidence-record model for study purpose, data review, method selection, uncertainty, provenance and approval, with an honest audit of the current Hydrometric generator.