National hydrology
What National-Scale Analysis Reveals About UK Hydrology
- Version
- 1.0.0
- Status
- Current
- Published
- Evidence cut-off
Abstract
What changes when UK flood-hydrology methods are examined across networks rather than one catchment at a time? Four bounded experiments show where gauge records disagree, where validation design changes a model comparison, where record length limits growth-curve targets, and where a temporal result must remain a case study.
Findings at a glance
- Adjacent pairs
- 85Same-river network diagnostic85 adjacent same-river pairs
- Median disagreement
- 12.2%Area-scaled QMED between paired gauges85 adjacent same-river pairs
- QMED validation cohort
- 918Stations in the spatial-cluster comparison918 stations
- Median GF100 sampling variability
- ±13.0%Within-station result near the studied record length874 gauged stations in the frozen revalidation
Evidence distributions
Area-scaled QMED disagreement between adjacent same-river gauges
85 adjacent same-river pairs
| Statistic | Disagreement |
|---|---|
| P10 | 3.4% |
| Q1 | 6.6% |
| Median | 12.2% |
| Q3 | 34.5% |
| P90 | 58.8% |
| Pairs above 50% | 15.3% |
Nearby gauges disagree before any model enters the comparison. This is a network self-consistency diagnostic, not an irreducible site-specific error floor or confidence interval. River-name pairing is heuristic.
Ungauged QMED under spatial-cluster cross-validation
918 stations
| Method | R² | Median absolute error | FSE |
|---|---|---|---|
| ML challenger | 0.898 | 26.1% | 1.634 |
| Donor-adjusted SM2025 | 0.922 | 23.4% | 1.538 |
| SM2025 with all donors available | 0.940 | — | — |
The challenger did not beat the statistical benchmark. Restricting donors for each held-out station to the training folds reduced the apparent SM2025 result from R² 0.940 to 0.922. Paired ΔR² was −0.024 with 95% CI −0.031 to −0.016.
Estimated target ceilings for an approximately 40-year AMAX record
874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
| Target | Estimated R² ceiling |
|---|---|
| L-CV | 0.857 |
| L-skew | 0.581 |
| GF100 | 0.723 |
Finite-record sampling constrains L-CV, L-skew and GF100 differently. Median within-station GF100 sampling variability was about ±13.0%; this is not universal design-flow uncertainty. The published artefact does not report an uncertainty band on the ceiling estimates.
Held-out log-KGE under named temporal partitions
Two-station methodological case study
| Station | P-A · early → late | P-B · late → early | P-C · middle → outer |
|---|---|---|---|
| 43007 | 0.584 | 0.495 | 0.815 |
| 33029 control | 0.805 | 0.721 | 0.724 |
Each partition used 60% calibration and 40% validation after a 730-day spin-up. P-A tests early → late, P-B late → early and P-C middle → outer. This two-station case shows why temporal partitioning changes the scientific question. It does not estimate national non-stationarity prevalence or attribute change to climate.
What survived
Network comparison as a review signal
Adjacent-gauge disagreement remains useful for identifying records and pairings that deserve scrutiny before a model comparison begins.
Validation matched to geographic transfer
The spatial-cluster comparison remains the relevant test for an ungauged claim because random station folds can preserve nearby information across training and evaluation sets.
Separate treatment of scale and skew
Relative spread, asymmetry and growth-factor sampling behaviour remain distinct reporting quantities; one cannot stand in for the others.
What failed
The QMED accuracy-win hypothesis
The optimistic interpretation from random validation was withdrawn when the challenger remained behind donor-adjusted SM2025 under spatial-cluster validation.
A station-level interval from paired gauges
The same-river distribution cannot be converted into an uncertainty interval for an individual site or assigned to one source of disagreement.
A universal tail allowance from record length
Within-station GF100 sampling behaviour does not establish a general uncertainty allowance for every design-flow estimate.
Research question
What becomes visible when UK flood-hydrology evidence is examined across a national network rather than one catchment at a time? The practical question is not whether a large dataset automatically produces certainty. It is whether repeated comparisons reveal inconsistencies, validation leakage, sampling limits and temporal sensitivity that a one-site workflow can leave untested.
Data and bounded cohorts
The evidence comes from four distinct experiments. The network study area-scales QMED before comparing adjacent gauges paired by river name. The QMED revalidation freezes a national station cohort and compares random station validation with five-fold unbuffered KMeans spatial-cluster cross-validation. The growth artefact selected 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage for the rating fit. The temporal study changes the held-out periods for a case station and a control.
The growth ceilings were estimated by resampling years within each station, recomputing L-moments and GF100, and using within-station sampling variance to bound observable-target R². These cohorts answer different questions and are not pooled into a single uncertainty calculation. Their exact sample sizes, statistics and evidence identifiers are recorded in the distributions and evidence record on this page.
Validation design changes the claim
Random folds test interpolation among stations that may still have nearby examples in training. Spatial-cluster folds instead hold out centroid-based geographic groups, making the experiment closer to transfer into an ungauged area. The KMeans folds were unbuffered: they do not assert a protected distance around each evaluation station. Donor custody was also enforced by restricting donor selection to the training folds.
Information symmetry matters here. The frozen challenger used the statistical-method estimate as an input, while its frozen urban-extent field carried no variation. The comparison is therefore a test of that specific challenger against donor-adjusted SM2025, not evidence that a wholly separate estimator surpassed the benchmark.
What the national view exposes
A single-site calculation can be internally tidy while neighbouring records remain mutually difficult to reconcile. Adjacent-gauge comparison brings that inconsistency into view before a predictive model is introduced. It is most useful as a review trigger: check the pairing, catchment areas, record periods, rating evidence and local hydrological explanation before deciding what the disagreement means.
The diagnostic does not identify which gauge is closer to truth and does not allocate disagreement among measurement, catchment or record effects. That is why its distribution remains separate from model validation and site-specific uncertainty statements.
Scale, skew and growth are different quantities
QMED is the index-flood magnitude. L-CV is the second L-moment divided by the first, a dimensionless measure of relative spread. L-skew is the third L-moment divided by the second and describes asymmetry, which gives the sparsely observed upper tail greater leverage. GF100 is the fitted one-hundred-year quantile divided by QMED: a tail growth ratio, not a substitute for diagnosing either relative spread or skew.
Finite-record sampling constrains L-CV, L-skew and GF100 differently. The ceiling estimates bound observable-target R² using the measured within-station sampling variance. The published artefact does not report an uncertainty band on those ceiling estimates, and the within-record GF100 spread must not be promoted into a general allowance for design-flow uncertainty.
The temporal example is a case, not a national trend
Each partition uses 60% calibration and 40% validation after a 730-day spin-up. P-A asks early → late, P-B late → early and P-C middle → outer. Holding out these different periods asks whether a model transfers across different parts of a record, not merely whether it reproduces a random sample from the same mixture of years. The case and control illustrate why the chosen periods should be disclosed and why a single aggregate score can hide temporal sensitivity.
The experiment was designed as a methodological demonstration. It has neither the cohort nor the attribution design needed to estimate how widespread non-stationarity is or to assign a cause.
Withdrawn interpretation and ruled-out conclusions
The random-validation accuracy-win interpretation is the only retraction recorded here. It was withdrawn because it did not survive the geographic-transfer test. The other two failed outcomes are ruled-out conclusions, not retractions or tested hypotheses: paired-gauge disagreement does not justify a station-level interval, and within-record GF100 variability does not justify a universal design-flow allowance.
Operational implications
- Review neighbouring gauges and their evidence before treating one at-site estimate as an unquestioned reference.
- Match validation to the deployment claim, disclose the fold construction, and keep held-out information out of donor choices.
- Report index-flood magnitude, relative spread, skew and tail growth separately so that sensitivity is visible rather than compressed into one number.
- Keep network consistency, population cross-validation, sampling variability and temporal sensitivity as distinct uncertainty evidence unless a calibrated combination has been demonstrated.
Limitations and disclosure boundary
River-name pairing is heuristic. The spatial clusters are unbuffered centroid groups. The growth result applies near the studied record length, and the published artefact gives no uncertainty band for the estimated ceilings. The temporal result is deliberately a bounded case study. Each limitation constrains interpretation even though the underlying cohort is broader than a conventional single-site analysis.
We disclose cohorts, validation logic, exact result distributions, failed hypotheses, evidence identifiers and operational meaning. We withhold feature engineering, learned parameters, training recipes, deployable artefacts, private data mechanics, operational infrastructure, customer material and licensed outputs. This is enough to audit the claim without reconstructing the engine.
What we do not conclude
- The same-river diagnostic is not a site-specific uncertainty interval and does not identify a single cause of disagreement.
- The QMED experiment does not show that the challenger is separate from SM2025 or that random validation is adequate for ungauged transfer.
- The growth experiment does not provide a universal uncertainty percentage for a design flood.
- The temporal case does not estimate national prevalence and does not attribute the observed sensitivity to climate.
- None of these experiments establishes regulatory approval or the replacement of established flood-estimation methods.
Claim record
Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.
Demonstrated
Across 85 adjacent same-river gauge pairs, area-scaled QMED disagreement had a median of 12.2%, Q1–Q3 of 6.6–34.5%, P90 of 58.8%, and 15.3% of pairs differed by more than 50%.
- Interpretation
- National-scale pairing exposes inconsistency that a one-site workflow cannot show.
- Boundary
- This is a network self-consistency diagnostic, not an irreducible site-specific error floor or confidence interval. River-name pairing is heuristic.
- Linked result
- Area-scaled QMED disagreement between adjacent same-river gauges; sample: 85 adjacent same-river pairs
- Evidence
- Hydrometric same-river network consistency study (2026).Evidence ID:
db10a70d
- Hydrometric same-river network consistency study (2026).Evidence ID:
Demonstrated
Under spatial-cluster cross-validation on 918 stations, the ungauged ML challenger achieved R² 0.898, median absolute error 26.1% and FSE 1.634; donor-adjusted SM2025 achieved R² 0.922, median absolute error 23.4% and FSE 1.538. ΔR² was −0.024 [95% CI −0.031, −0.016].
- Interpretation
- The challenger did not beat the statistical benchmark under the validation design aligned with geographic transfer.
- Boundary
- The comparison used five-fold unbuffered KMeans spatial-cluster cross-validation. The frozen challenger included the statistical-method estimate among its inputs, and the frozen urban-extent field contained zeros throughout.
- Linked result
- Ungauged QMED under spatial-cluster cross-validation; sample: 918 stations
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Demonstrated
The apparent all-donor result fell from R² 0.940 to 0.922 when donors for each held-out station were restricted to the training folds.
- Interpretation
- Donor availability is part of the validation design, not a neutral implementation detail.
- Boundary
- The restricted result tests geographic transfer without retaining held-out neighbours as donors.
- Linked result
- Ungauged QMED under spatial-cluster cross-validation; sample: 918 stations
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Retracted
The random-validation accuracy-win interpretation for the QMED challenger was withdrawn.
- Interpretation
- Five-fold unbuffered KMeans spatial-cluster cross-validation placed the challenger behind donor-adjusted SM2025.
- Boundary
- Only the accuracy-win interpretation is retracted; this does not establish that machine learning is unusable for QMED.
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Demonstrated
The growth artefact selected 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage available for the rating fit.
- Interpretation
- Target ceilings were estimated by resampling years within each station, recomputing L-moments and GF100, and using within-station sampling variance to bound observable-target R².
- Boundary
- The published artefact does not report an uncertainty band on the ceiling estimates.
- Linked result
- Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Demonstrated
For approximately 40-year AMAX records, estimated target ceilings were R² 0.857 for L-CV, 0.581 for L-skew and 0.723 for GF100; median within-station GF100 uncertainty was about ±13.0%.
- Interpretation
- Finite-record sampling constrains L-CV, L-skew and GF100 differently.
- Boundary
- The GF100 result is median within-station sampling variability near the studied record length, not universal design-flow uncertainty. The published artefact does not report an uncertainty band on the ceiling estimates.
- Linked result
- Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Demonstrated
Station 43007 produced held-out log-KGE scores of 0.584, 0.495 and 0.815 under different temporal partitions; control station 33029 produced 0.805, 0.721 and 0.724.
- Interpretation
- Each partition used 60% calibration and 40% validation after a 730-day spin-up: P-A evaluated early → late, P-B late → early and P-C middle → outer.
- Boundary
- This two-station methodological case does not establish national non-stationarity prevalence or climate causation.
- Linked result
- Held-out log-KGE under named temporal partitions; sample: Two-station methodological case study
- Evidence
- Hydrometric temporal-partition case study (2026).Evidence ID:
e5da4413
- Hydrometric temporal-partition case study (2026).Evidence ID:
Inference
A defensible review should keep network disagreement, population validation scatter, record-length sampling variability and temporal sensitivity as separate evidence layers.
- Interpretation
- Separation preserves the meaning and scope of each experiment.
- Boundary
- This is a reporting inference; the components have not been calibrated as a combined site-specific probability model.
- Evidence
- Hydrometric same-river network consistency study (2026).Evidence ID:
db10a70d - Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b - Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898 - Hydrometric temporal-partition case study (2026).Evidence ID:
e5da4413
- Hydrometric same-river network consistency study (2026).Evidence ID:
Primary sources
- Hydrometric same-river network consistency study (2026).Evidence ID:
db10a70d - Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b - Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898 - Hydrometric temporal-partition case study (2026).Evidence ID:
e5da4413
Version history
- Version 1.0.0current
Initial publication of the national hydrology evidence synthesis.
Related reports
- QMEDCan Machine Learning Improve QMED?A fair national comparison showing why an ML accuracy claim failed under geographic validation, while open-input access, repeatability and auditability survived.
- Statistical methodsThe 2025 Statistical Method: What Changed and Why It MattersA practitioner account of the FEH 2025 descriptor, QMED, donor, pooling, urban, distribution and uncertainty changes, with Hydrometric’s implementation evidence kept inside its local verification boundary.
- Growth curvesSkew, Scale and the Shape of UK Flood GrowthA bounded account of index-flood scale, L-CV dispersion, L-skew shape, GF100 growth, a gauged-only rating result and one frozen finite-sample model failure.
- UncertaintyHow Certain Is a Design Flood?A layered account of population validation scatter, network inconsistency, within-record sampling and the still-open model-family protocol.
- Rainfall-runoff modellingFrom ReFH Design Events to Continuous HydrologyA bounded record of the baseflow-lag refit, the public ReFH1 baseline gap, the failed daily two-store branch and the gated status of regional continuous evidence.
- Flood estimation practiceThe Modern Flood Estimation ReportA practical evidence-record model for study purpose, data review, method selection, uncertainty, provenance and approval, with an honest audit of the current Hydrometric generator.