Skip to main content

National hydrology

What National-Scale Analysis Reveals About UK Hydrology

Version
1.0.0
Status
Current
Published
Evidence cut-off

Abstract

What changes when UK flood-hydrology methods are examined across networks rather than one catchment at a time? Four bounded experiments show where gauge records disagree, where validation design changes a model comparison, where record length limits growth-curve targets, and where a temporal result must remain a case study.

Findings at a glance

Adjacent pairs
85Same-river network diagnostic85 adjacent same-river pairs
Median disagreement
12.2%Area-scaled QMED between paired gauges85 adjacent same-river pairs
QMED validation cohort
918Stations in the spatial-cluster comparison918 stations
Median GF100 sampling variability
±13.0%Within-station result near the studied record length874 gauged stations in the frozen revalidation

Evidence distributions

Area-scaled QMED disagreement between adjacent same-river gauges

85 adjacent same-river pairs

StatisticDisagreement
P103.4%
Q16.6%
Median12.2%
Q334.5%
P9058.8%
Pairs above 50%15.3%

Nearby gauges disagree before any model enters the comparison. This is a network self-consistency diagnostic, not an irreducible site-specific error floor or confidence interval. River-name pairing is heuristic.

Ungauged QMED under spatial-cluster cross-validation

918 stations

MethodMedian absolute errorFSE
ML challenger0.89826.1%1.634
Donor-adjusted SM20250.92223.4%1.538
SM2025 with all donors available0.940

The challenger did not beat the statistical benchmark. Restricting donors for each held-out station to the training folds reduced the apparent SM2025 result from R² 0.940 to 0.922. Paired ΔR² was −0.024 with 95% CI −0.031 to −0.016.

Estimated target ceilings for an approximately 40-year AMAX record

874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit

TargetEstimated R² ceiling
L-CV0.857
L-skew0.581
GF1000.723

Finite-record sampling constrains L-CV, L-skew and GF100 differently. Median within-station GF100 sampling variability was about ±13.0%; this is not universal design-flow uncertainty. The published artefact does not report an uncertainty band on the ceiling estimates.

Held-out log-KGE under named temporal partitions

Two-station methodological case study

StationP-A · early → lateP-B · late → earlyP-C · middle → outer
430070.5840.4950.815
33029 control0.8050.7210.724

Each partition used 60% calibration and 40% validation after a 730-day spin-up. P-A tests early → late, P-B late → early and P-C middle → outer. This two-station case shows why temporal partitioning changes the scientific question. It does not estimate national non-stationarity prevalence or attribute change to climate.

What survived

Network comparison as a review signal

Adjacent-gauge disagreement remains useful for identifying records and pairings that deserve scrutiny before a model comparison begins.

Validation matched to geographic transfer

The spatial-cluster comparison remains the relevant test for an ungauged claim because random station folds can preserve nearby information across training and evaluation sets.

Separate treatment of scale and skew

Relative spread, asymmetry and growth-factor sampling behaviour remain distinct reporting quantities; one cannot stand in for the others.

What failed

The QMED accuracy-win hypothesis

The optimistic interpretation from random validation was withdrawn when the challenger remained behind donor-adjusted SM2025 under spatial-cluster validation.

A station-level interval from paired gauges

The same-river distribution cannot be converted into an uncertainty interval for an individual site or assigned to one source of disagreement.

A universal tail allowance from record length

Within-station GF100 sampling behaviour does not establish a general uncertainty allowance for every design-flow estimate.

Research question

What becomes visible when UK flood-hydrology evidence is examined across a national network rather than one catchment at a time? The practical question is not whether a large dataset automatically produces certainty. It is whether repeated comparisons reveal inconsistencies, validation leakage, sampling limits and temporal sensitivity that a one-site workflow can leave untested.

Data and bounded cohorts

The evidence comes from four distinct experiments. The network study area-scales QMED before comparing adjacent gauges paired by river name. The QMED revalidation freezes a national station cohort and compares random station validation with five-fold unbuffered KMeans spatial-cluster cross-validation. The growth artefact selected 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage for the rating fit. The temporal study changes the held-out periods for a case station and a control.

The growth ceilings were estimated by resampling years within each station, recomputing L-moments and GF100, and using within-station sampling variance to bound observable-target R². These cohorts answer different questions and are not pooled into a single uncertainty calculation. Their exact sample sizes, statistics and evidence identifiers are recorded in the distributions and evidence record on this page.

Validation design changes the claim

Random folds test interpolation among stations that may still have nearby examples in training. Spatial-cluster folds instead hold out centroid-based geographic groups, making the experiment closer to transfer into an ungauged area. The KMeans folds were unbuffered: they do not assert a protected distance around each evaluation station. Donor custody was also enforced by restricting donor selection to the training folds.

Information symmetry matters here. The frozen challenger used the statistical-method estimate as an input, while its frozen urban-extent field carried no variation. The comparison is therefore a test of that specific challenger against donor-adjusted SM2025, not evidence that a wholly separate estimator surpassed the benchmark.

What the national view exposes

A single-site calculation can be internally tidy while neighbouring records remain mutually difficult to reconcile. Adjacent-gauge comparison brings that inconsistency into view before a predictive model is introduced. It is most useful as a review trigger: check the pairing, catchment areas, record periods, rating evidence and local hydrological explanation before deciding what the disagreement means.

The diagnostic does not identify which gauge is closer to truth and does not allocate disagreement among measurement, catchment or record effects. That is why its distribution remains separate from model validation and site-specific uncertainty statements.

Scale, skew and growth are different quantities

QMED is the index-flood magnitude. L-CV is the second L-moment divided by the first, a dimensionless measure of relative spread. L-skew is the third L-moment divided by the second and describes asymmetry, which gives the sparsely observed upper tail greater leverage. GF100 is the fitted one-hundred-year quantile divided by QMED: a tail growth ratio, not a substitute for diagnosing either relative spread or skew.

Finite-record sampling constrains L-CV, L-skew and GF100 differently. The ceiling estimates bound observable-target R² using the measured within-station sampling variance. The published artefact does not report an uncertainty band on those ceiling estimates, and the within-record GF100 spread must not be promoted into a general allowance for design-flow uncertainty.

The temporal example is a case, not a national trend

Each partition uses 60% calibration and 40% validation after a 730-day spin-up. P-A asks early → late, P-B late → early and P-C middle → outer. Holding out these different periods asks whether a model transfers across different parts of a record, not merely whether it reproduces a random sample from the same mixture of years. The case and control illustrate why the chosen periods should be disclosed and why a single aggregate score can hide temporal sensitivity.

The experiment was designed as a methodological demonstration. It has neither the cohort nor the attribution design needed to estimate how widespread non-stationarity is or to assign a cause.

Withdrawn interpretation and ruled-out conclusions

The random-validation accuracy-win interpretation is the only retraction recorded here. It was withdrawn because it did not survive the geographic-transfer test. The other two failed outcomes are ruled-out conclusions, not retractions or tested hypotheses: paired-gauge disagreement does not justify a station-level interval, and within-record GF100 variability does not justify a universal design-flow allowance.

Operational implications

  • Review neighbouring gauges and their evidence before treating one at-site estimate as an unquestioned reference.
  • Match validation to the deployment claim, disclose the fold construction, and keep held-out information out of donor choices.
  • Report index-flood magnitude, relative spread, skew and tail growth separately so that sensitivity is visible rather than compressed into one number.
  • Keep network consistency, population cross-validation, sampling variability and temporal sensitivity as distinct uncertainty evidence unless a calibrated combination has been demonstrated.

Limitations and disclosure boundary

River-name pairing is heuristic. The spatial clusters are unbuffered centroid groups. The growth result applies near the studied record length, and the published artefact gives no uncertainty band for the estimated ceilings. The temporal result is deliberately a bounded case study. Each limitation constrains interpretation even though the underlying cohort is broader than a conventional single-site analysis.

We disclose cohorts, validation logic, exact result distributions, failed hypotheses, evidence identifiers and operational meaning. We withhold feature engineering, learned parameters, training recipes, deployable artefacts, private data mechanics, operational infrastructure, customer material and licensed outputs. This is enough to audit the claim without reconstructing the engine.

What we do not conclude

  • The same-river diagnostic is not a site-specific uncertainty interval and does not identify a single cause of disagreement.
  • The QMED experiment does not show that the challenger is separate from SM2025 or that random validation is adequate for ungauged transfer.
  • The growth experiment does not provide a universal uncertainty percentage for a design flood.
  • The temporal case does not estimate national prevalence and does not attribute the observed sensitivity to climate.
  • None of these experiments establishes regulatory approval or the replacement of established flood-estimation methods.

Claim record

Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.

  1. Demonstrated

    Across 85 adjacent same-river gauge pairs, area-scaled QMED disagreement had a median of 12.2%, Q1–Q3 of 6.6–34.5%, P90 of 58.8%, and 15.3% of pairs differed by more than 50%.

    Interpretation
    National-scale pairing exposes inconsistency that a one-site workflow cannot show.
    Boundary
    This is a network self-consistency diagnostic, not an irreducible site-specific error floor or confidence interval. River-name pairing is heuristic.
    Linked result
    Area-scaled QMED disagreement between adjacent same-river gauges; sample: 85 adjacent same-river pairs
    Evidence
    • Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
  2. Demonstrated

    Under spatial-cluster cross-validation on 918 stations, the ungauged ML challenger achieved R² 0.898, median absolute error 26.1% and FSE 1.634; donor-adjusted SM2025 achieved R² 0.922, median absolute error 23.4% and FSE 1.538. ΔR² was −0.024 [95% CI −0.031, −0.016].

    Interpretation
    The challenger did not beat the statistical benchmark under the validation design aligned with geographic transfer.
    Boundary
    The comparison used five-fold unbuffered KMeans spatial-cluster cross-validation. The frozen challenger included the statistical-method estimate among its inputs, and the frozen urban-extent field contained zeros throughout.
    Linked result
    Ungauged QMED under spatial-cluster cross-validation; sample: 918 stations
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  3. Demonstrated

    The apparent all-donor result fell from R² 0.940 to 0.922 when donors for each held-out station were restricted to the training folds.

    Interpretation
    Donor availability is part of the validation design, not a neutral implementation detail.
    Boundary
    The restricted result tests geographic transfer without retaining held-out neighbours as donors.
    Linked result
    Ungauged QMED under spatial-cluster cross-validation; sample: 918 stations
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  4. Retracted

    The random-validation accuracy-win interpretation for the QMED challenger was withdrawn.

    Interpretation
    Five-fold unbuffered KMeans spatial-cluster cross-validation placed the challenger behind donor-adjusted SM2025.
    Boundary
    Only the accuracy-win interpretation is retracted; this does not establish that machine learning is unusable for QMED.
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  5. Demonstrated

    The growth artefact selected 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage available for the rating fit.

    Interpretation
    Target ceilings were estimated by resampling years within each station, recomputing L-moments and GF100, and using within-station sampling variance to bound observable-target R².
    Boundary
    The published artefact does not report an uncertainty band on the ceiling estimates.
    Linked result
    Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  6. Demonstrated

    For approximately 40-year AMAX records, estimated target ceilings were R² 0.857 for L-CV, 0.581 for L-skew and 0.723 for GF100; median within-station GF100 uncertainty was about ±13.0%.

    Interpretation
    Finite-record sampling constrains L-CV, L-skew and GF100 differently.
    Boundary
    The GF100 result is median within-station sampling variability near the studied record length, not universal design-flow uncertainty. The published artefact does not report an uncertainty band on the ceiling estimates.
    Linked result
    Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  7. Demonstrated

    Station 43007 produced held-out log-KGE scores of 0.584, 0.495 and 0.815 under different temporal partitions; control station 33029 produced 0.805, 0.721 and 0.724.

    Interpretation
    Each partition used 60% calibration and 40% validation after a 730-day spin-up: P-A evaluated early → late, P-B late → early and P-C middle → outer.
    Boundary
    This two-station methodological case does not establish national non-stationarity prevalence or climate causation.
    Linked result
    Held-out log-KGE under named temporal partitions; sample: Two-station methodological case study
    Evidence
    • Hydrometric temporal-partition case study (2026).Evidence ID: e5da4413
  8. Inference

    A defensible review should keep network disagreement, population validation scatter, record-length sampling variability and temporal sensitivity as separate evidence layers.

    Interpretation
    Separation preserves the meaning and scope of each experiment.
    Boundary
    This is a reporting inference; the components have not been calibrated as a combined site-specific probability model.
    Evidence
    • Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
    • Hydrometric temporal-partition case study (2026).Evidence ID: e5da4413

Primary sources

  1. Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
  2. Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  3. Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  4. Hydrometric temporal-partition case study (2026).Evidence ID: e5da4413

Version history

  1. Version 1.0.0current

    Initial publication of the national hydrology evidence synthesis.