Skip to main content

National hydrology

What National-Scale Analysis Reveals About UK Hydrology

Version
1.0.0
Status
Current
Published
Evidence cut-off

In short

We asked what you can see about UK flood estimation from every gauged river at once that you cannot see from one catchment at a time. Four things: neighbouring gauges disagree about the same flood, an easy test flatters a new method, forty years of records only pin down so much, and which years you hold back changes the answer.

Findings at a glance

Neighbouring gauges compared
85Pairs of gauges sitting on the same river85 adjacent same-river pairs
Typical disagreement between them
12.2%How far apart two neighbours put the same yearly flood85 adjacent same-river pairs
River gauges in the method test
918Every one held back from training before it was scored918 stations
Wobble from record length alone
±13.0%Reshuffling the years inside a 40-year record moves the answer this much874 gauged stations in the frozen revalidation

What we tested

A flood estimate is usually built one catchment at a time. You pick the site, pull its descriptors, choose donor gauges nearby, and the answer looks tidy. We wanted to know what that tidiness hides, so we ran the same checks across every gauged river in the country at once.

Four questions, four separate experiments. First: when two gauges sit on the same river, one above the other, do they agree about the size of the typical yearly flood once you allow for the extra catchment area between them? We compared 85 such pairs.

Second: does it matter how you mark a new method? We scored a machine-learning model and the FEH 2025 statistical method (FEH is the Flood Estimation Handbook, the UK flood-frequency reference) on 918 gauges, twice — once with nearby rivers allowed into training, once with whole regions held back.

Third: how much can roughly forty years of annual flood records actually tell you? We took 874 gauges and reshuffled the years inside each record, then recalculated the flood statistics to see how much they moved on their own.

Fourth: does the choice of which years you test on change the answer? We took one station and one control station and moved the held-back period around.

What we saw

Two gauges, one river, two different floods

85 pairs of neighbouring gauges on the same river, compared after adjusting for catchment area.

0%20%40%60%Difference in the typical yearly flood (%) — shaded band holds the middle half of pairsDisagreement between the two gaugesDisagreement between the two gauges — Lowest tenth of pairs: 3.4%3.4%Disagreement between the two gauges — Middle pair: 12.2%12.2%Disagreement between the two gauges — Highest tenth of pairs: 58.8%58.8%
Half of neighbouring pairs disagree by more than 12.2% about the same flood, and the worst tenth by more than 58.8% — before any model is involved.

What forty years of records can and cannot pin down

874 gauged rivers, each with at least ten usable years. The bar is the best score any method could hope for, given how much the number moves when you reshuffle the years.

0.00.20.40.60.81.0Best score a 40-year record could support (R², where 1.0 is perfect)How spread out the floods are (L-CV)How spread out the floods are (L-CV) — Best achievable score: 0.8570.857How fast the big floods grow (GF100)How fast the big floods grow (GF100) — Best achievable score: 0.7230.723How lopsided the record is (L-skew)How lopsided the record is (L-skew) — Best achievable score: 0.5810.581
Forty years of floods pin down how spread out they are far better than they pin down the shape of the tail.

Move the test years and the score moves with them

Two stations, one of interest and one control, each scored on years held back from fitting.

  • Station 43007
  • Control station 33029
0.00.20.40.60.81.0Score on the held-back years (log-KGE, where 1.0 is perfect)Fitted on early years, tested on lateFitted on early years, tested on late — Station 43007: 0.5840.584Fitted on early years, tested on late — Control station 33029: 0.8050.805Fitted on late years, tested on earlyFitted on late years, tested on early — Station 43007: 0.4950.495Fitted on late years, tested on early — Control station 33029: 0.7210.721Fitted on middle years, tested on the outer onesFitted on middle years, tested on the outer ones — Station 43007: 0.8150.815Fitted on middle years, tested on the outer ones — Control station 33029: 0.7240.724
The same station scored 0.495 on one held-back period and 0.815 on another; the control barely moved. One average score would have hidden that.

Outcome

  • Neighbouring gauges disagree, and sometimes badly. Across 85 pairs the middle disagreement was 12.2%, the middle half ran from 6.6% to 34.5%, and 15.3% of pairs differed by more than half. This says the network is not perfectly self-consistent. It does not say which gauge is right, and it is not an uncertainty range for any one site.
  • The easy test flattered the new method. Letting nearby rivers stay in training put the FEH 2025 method at 0.922 and lifted the machine-learning model to 0.915 — close enough to argue about. Holding whole regions back left the model at 0.898 against 0.922. We withdrew the earlier reading that the model won.
  • Donor choice is part of the test, not a detail. Allowing the statistical method to borrow from gauges in the held-back region moved its score from 0.922 to 0.940. That is a bigger swing than the gap being argued over.
  • Forty years buys different amounts of certainty for different things. The achievable score was 0.857 for spread, 0.723 for tail growth and 0.581 for lopsidedness. Reshuffling years inside a record moved the 100-year growth factor by about ±13.0%. That is record-length wobble, not a general allowance to add to a design flow.
  • Which years you test on changes the verdict. One station ran from 0.495 to 0.815 depending on the split. Two stations is a demonstration, not a national trend, and it says nothing about the cause.

What we decided next

Decided 14 July 2026. These four numbers never get added together, and Hydrometric reports them in four separate places. Network disagreement, method performance across the country, record-length wobble and time-period sensitivity answer different questions on different sets of rivers. Stacking them into one uncertainty band would be arithmetic without a meaning.

Two things changed in how we run and present work. Every method comparison is now scored with whole regions held back and donors confined to the training side; an easier score is only ever shown beside the fair one. And adjacent-gauge disagreement became a review prompt: where two neighbours disagree sharply, the pairing, catchment areas, record periods and rating evidence get looked at before the estimate is trusted.

Left open: we have not run the same-evidence comparison of candidate distribution families, so we cannot yet show how much of a design flow depends on that choice alone.

Read the detailExact tables, method names and the full claim record

Research question

What becomes visible when UK flood-hydrology evidence is examined across a national network rather than one catchment at a time? The practical question is not whether a large dataset automatically produces certainty. It is whether repeated comparisons reveal inconsistencies, validation leakage, sampling limits and temporal sensitivity that a one-site workflow can leave untested.

Data and bounded cohorts

The evidence comes from four distinct experiments. The network study area-scales QMED before comparing adjacent gauges paired by river name. The QMED revalidation freezes a national station cohort and compares random station validation with five-fold unbuffered KMeans spatial-cluster cross-validation. The growth artefact selected 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage for the rating fit. The temporal study changes the held-out periods for a case station and a control.

The growth ceilings were estimated by resampling years within each station, recomputing L-moments and GF100, and using within-station sampling variance to bound observable-target R². These cohorts answer different questions and are not pooled into a single uncertainty calculation. Their exact sample sizes, statistics and evidence identifiers are recorded in the distributions and evidence record on this page.

Validation design changes the claim

Random folds test interpolation among stations that may still have nearby examples in training. Spatial-cluster folds instead hold out centroid-based geographic groups, making the experiment closer to transfer into an ungauged area. The KMeans folds were unbuffered: they do not assert a protected distance around each evaluation station. Donor custody was also enforced by restricting donor selection to the training folds.

Information symmetry matters here. The frozen challenger used the statistical-method estimate as an input, while its frozen urban-extent field carried no variation. The comparison is therefore a test of that specific challenger against donor-adjusted SM2025, not evidence that a wholly separate estimator surpassed the benchmark.

What the national view exposes

A single-site calculation can be internally tidy while neighbouring records remain mutually difficult to reconcile. Adjacent-gauge comparison brings that inconsistency into view before a predictive model is introduced. It is most useful as a review trigger: check the pairing, catchment areas, record periods, rating evidence and local hydrological explanation before deciding what the disagreement means.

The diagnostic does not identify which gauge is closer to truth and does not allocate disagreement among measurement, catchment or record effects. That is why its distribution remains separate from model validation and site-specific uncertainty statements.

Scale, skew and growth are different quantities

QMED is the index-flood magnitude. L-CV is the second L-moment divided by the first, a dimensionless measure of relative spread. L-skew is the third L-moment divided by the second and describes asymmetry, which gives the sparsely observed upper tail greater leverage. GF100 is the fitted one-hundred-year quantile divided by QMED: a tail growth ratio, not a substitute for diagnosing either relative spread or skew.

Finite-record sampling constrains L-CV, L-skew and GF100 differently. The ceiling estimates bound observable-target R² using the measured within-station sampling variance. The published artefact does not report an uncertainty band on those ceiling estimates, and the within-record GF100 spread must not be promoted into a general allowance for design-flow uncertainty.

The temporal example is a case, not a national trend

Each partition uses 60% calibration and 40% validation after a 730-day spin-up. P-A asks early → late, P-B late → early and P-C middle → outer. Holding out these different periods asks whether a model transfers across different parts of a record, not merely whether it reproduces a random sample from the same mixture of years. The case and control illustrate why the chosen periods should be disclosed and why a single aggregate score can hide temporal sensitivity.

The experiment was designed as a methodological demonstration. It has neither the cohort nor the attribution design needed to estimate how widespread non-stationarity is or to assign a cause.

Withdrawn interpretation and ruled-out conclusions

The random-validation accuracy-win interpretation is the only retraction recorded here. It was withdrawn because it did not survive the geographic-transfer test. The other two failed outcomes are ruled-out conclusions, not retractions or tested hypotheses: paired-gauge disagreement does not justify a station-level interval, and within-record GF100 variability does not justify a universal design-flow allowance.

Operational implications

  • Review neighbouring gauges and their evidence before treating one at-site estimate as an unquestioned reference.
  • Match validation to the deployment claim, disclose the fold construction, and keep held-out information out of donor choices.
  • Report index-flood magnitude, relative spread, skew and tail growth separately so that sensitivity is visible rather than compressed into one number.
  • Keep network consistency, population cross-validation, sampling variability and temporal sensitivity as distinct uncertainty evidence unless a calibrated combination has been demonstrated.

Limitations and disclosure boundary

River-name pairing is heuristic. The spatial clusters are unbuffered centroid groups. The growth result applies near the studied record length, and the published artefact gives no uncertainty band for the estimated ceilings. The temporal result is deliberately a bounded case study. Each limitation constrains interpretation even though the underlying cohort is broader than a conventional single-site analysis.

We disclose cohorts, validation logic, exact result distributions, failed hypotheses, evidence identifiers and operational meaning. We withhold feature engineering, learned parameters, training recipes, deployable artefacts, private data mechanics, operational infrastructure, customer material and licensed outputs. This is enough to audit the claim without reconstructing the engine.

What we do not conclude

  • The same-river diagnostic is not a site-specific uncertainty interval and does not identify a single cause of disagreement.
  • The QMED experiment does not show that the challenger is separate from SM2025 or that random validation is adequate for ungauged transfer.
  • The growth experiment does not provide a universal uncertainty percentage for a design flood.
  • The temporal case does not estimate national prevalence and does not attribute the observed sensitivity to climate.
  • None of these experiments establishes regulatory approval or the replacement of established flood-estimation methods.

Evidence distributions

Area-scaled QMED disagreement between adjacent same-river gauges

85 adjacent same-river pairs

StatisticDisagreement
P103.4%
Q16.6%
Median12.2%
Q334.5%
P9058.8%
Pairs above 50%15.3%

Nearby gauges disagree before any model enters the comparison. This is a network self-consistency diagnostic, not an irreducible site-specific error floor or confidence interval. River-name pairing is heuristic.

Ungauged QMED under spatial-cluster cross-validation

918 stations

MethodMedian absolute errorFSE
ML challenger0.89826.1%1.634
Donor-adjusted SM20250.92223.4%1.538
SM2025 with all donors available0.940

The challenger did not beat the statistical benchmark. Restricting donors for each held-out station to the training folds reduced the apparent SM2025 result from R² 0.940 to 0.922. Paired ΔR² was −0.024 with 95% CI −0.031 to −0.016.

Estimated target ceilings for an approximately 40-year AMAX record

874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit

TargetEstimated R² ceiling
L-CV0.857
L-skew0.581
GF1000.723

Finite-record sampling constrains L-CV, L-skew and GF100 differently. Median within-station GF100 sampling variability was about ±13.0%; this is not universal design-flow uncertainty. The published artefact does not report an uncertainty band on the ceiling estimates.

Held-out log-KGE under named temporal partitions

Two-station methodological case study

StationP-A · early → lateP-B · late → earlyP-C · middle → outer
430070.5840.4950.815
33029 control0.8050.7210.724

Each partition used 60% calibration and 40% validation after a 730-day spin-up. P-A tests early → late, P-B late → early and P-C middle → outer. This two-station case shows why temporal partitioning changes the scientific question. It does not estimate national non-stationarity prevalence or attribute change to climate.

What survived

Network comparison as a review signal

Adjacent-gauge disagreement remains useful for identifying records and pairings that deserve scrutiny before a model comparison begins.

Validation matched to geographic transfer

The spatial-cluster comparison remains the relevant test for an ungauged claim because random station folds can preserve nearby information across training and evaluation sets.

Separate treatment of scale and skew

Relative spread, asymmetry and growth-factor sampling behaviour remain distinct reporting quantities; one cannot stand in for the others.

What failed

The QMED accuracy-win hypothesis

The optimistic interpretation from random validation was withdrawn when the challenger remained behind donor-adjusted SM2025 under spatial-cluster validation.

A station-level interval from paired gauges

The same-river distribution cannot be converted into an uncertainty interval for an individual site or assigned to one source of disagreement.

A universal tail allowance from record length

Within-station GF100 sampling behaviour does not establish a general uncertainty allowance for every design-flow estimate.

Claim record

Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.

  1. Demonstrated

    Across 85 adjacent same-river gauge pairs, area-scaled QMED disagreement had a median of 12.2%, Q1–Q3 of 6.6–34.5%, P90 of 58.8%, and 15.3% of pairs differed by more than 50%.

    Interpretation
    National-scale pairing exposes inconsistency that a one-site workflow cannot show.
    Boundary
    This is a network self-consistency diagnostic, not an irreducible site-specific error floor or confidence interval. River-name pairing is heuristic.
    Linked result
    Area-scaled QMED disagreement between adjacent same-river gauges; sample: 85 adjacent same-river pairs
    Evidence
    • Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
  2. Demonstrated

    Under spatial-cluster cross-validation on 918 stations, the ungauged ML challenger achieved R² 0.898, median absolute error 26.1% and FSE 1.634; donor-adjusted SM2025 achieved R² 0.922, median absolute error 23.4% and FSE 1.538. ΔR² was −0.024 [95% CI −0.031, −0.016].

    Interpretation
    The challenger did not beat the statistical benchmark under the validation design aligned with geographic transfer.
    Boundary
    The comparison used five-fold unbuffered KMeans spatial-cluster cross-validation. The frozen challenger included the statistical-method estimate among its inputs, and the frozen urban-extent field contained zeros throughout.
    Linked result
    Ungauged QMED under spatial-cluster cross-validation; sample: 918 stations
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  3. Demonstrated

    The apparent all-donor result fell from R² 0.940 to 0.922 when donors for each held-out station were restricted to the training folds.

    Interpretation
    Donor availability is part of the validation design, not a neutral implementation detail.
    Boundary
    The restricted result tests geographic transfer without retaining held-out neighbours as donors.
    Linked result
    Ungauged QMED under spatial-cluster cross-validation; sample: 918 stations
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  4. Retracted

    The random-validation accuracy-win interpretation for the QMED challenger was withdrawn.

    Interpretation
    Five-fold unbuffered KMeans spatial-cluster cross-validation placed the challenger behind donor-adjusted SM2025.
    Boundary
    Only the accuracy-win interpretation is retracted; this does not establish that machine learning is unusable for QMED.
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  5. Demonstrated

    The growth artefact selected 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage available for the rating fit.

    Interpretation
    Target ceilings were estimated by resampling years within each station, recomputing L-moments and GF100, and using within-station sampling variance to bound observable-target R².
    Boundary
    The published artefact does not report an uncertainty band on the ceiling estimates.
    Linked result
    Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  6. Demonstrated

    For approximately 40-year AMAX records, estimated target ceilings were R² 0.857 for L-CV, 0.581 for L-skew and 0.723 for GF100; median within-station GF100 uncertainty was about ±13.0%.

    Interpretation
    Finite-record sampling constrains L-CV, L-skew and GF100 differently.
    Boundary
    The GF100 result is median within-station sampling variability near the studied record length, not universal design-flow uncertainty. The published artefact does not report an uncertainty band on the ceiling estimates.
    Linked result
    Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  7. Demonstrated

    Station 43007 produced held-out log-KGE scores of 0.584, 0.495 and 0.815 under different temporal partitions; control station 33029 produced 0.805, 0.721 and 0.724.

    Interpretation
    Each partition used 60% calibration and 40% validation after a 730-day spin-up: P-A evaluated early → late, P-B late → early and P-C middle → outer.
    Boundary
    This two-station methodological case does not establish national non-stationarity prevalence or climate causation.
    Linked result
    Held-out log-KGE under named temporal partitions; sample: Two-station methodological case study
    Evidence
    • Hydrometric temporal-partition case study (2026).Evidence ID: e5da4413
  8. Inference

    A defensible review should keep network disagreement, population validation scatter, record-length sampling variability and temporal sensitivity as separate evidence layers.

    Interpretation
    Separation preserves the meaning and scope of each experiment.
    Boundary
    This is a reporting inference; the components have not been calibrated as a combined site-specific probability model.
    Evidence
    • Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
    • Hydrometric temporal-partition case study (2026).Evidence ID: e5da4413

Primary sources

  1. Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
  2. Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  3. Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  4. Hydrometric temporal-partition case study (2026).Evidence ID: e5da4413

Version history

  1. Version 1.0.0current

    Initial publication of the national hydrology evidence synthesis.