National hydrology
What National-Scale Analysis Reveals About UK Hydrology
- Version
- 1.0.0
- Status
- Current
- Published
- Evidence cut-off
In short
We asked what you can see about UK flood estimation from every gauged river at once that you cannot see from one catchment at a time. Four things: neighbouring gauges disagree about the same flood, an easy test flatters a new method, forty years of records only pin down so much, and which years you hold back changes the answer.
Findings at a glance
- Neighbouring gauges compared
- 85Pairs of gauges sitting on the same river85 adjacent same-river pairs
- Typical disagreement between them
- 12.2%How far apart two neighbours put the same yearly flood85 adjacent same-river pairs
- River gauges in the method test
- 918Every one held back from training before it was scored918 stations
- Wobble from record length alone
- ±13.0%Reshuffling the years inside a 40-year record moves the answer this much874 gauged stations in the frozen revalidation
What we tested
A flood estimate is usually built one catchment at a time. You pick the site, pull its descriptors, choose donor gauges nearby, and the answer looks tidy. We wanted to know what that tidiness hides, so we ran the same checks across every gauged river in the country at once.
Four questions, four separate experiments. First: when two gauges sit on the same river, one above the other, do they agree about the size of the typical yearly flood once you allow for the extra catchment area between them? We compared 85 such pairs.
Second: does it matter how you mark a new method? We scored a machine-learning model and the FEH 2025 statistical method (FEH is the Flood Estimation Handbook, the UK flood-frequency reference) on 918 gauges, twice — once with nearby rivers allowed into training, once with whole regions held back.
Third: how much can roughly forty years of annual flood records actually tell you? We took 874 gauges and reshuffled the years inside each record, then recalculated the flood statistics to see how much they moved on their own.
Fourth: does the choice of which years you test on change the answer? We took one station and one control station and moved the held-back period around.
What we saw
Two gauges, one river, two different floods
85 pairs of neighbouring gauges on the same river, compared after adjusting for catchment area.
What forty years of records can and cannot pin down
874 gauged rivers, each with at least ten usable years. The bar is the best score any method could hope for, given how much the number moves when you reshuffle the years.
Move the test years and the score moves with them
Two stations, one of interest and one control, each scored on years held back from fitting.
- Station 43007
- Control station 33029
Outcome
- Neighbouring gauges disagree, and sometimes badly. Across 85 pairs the middle disagreement was 12.2%, the middle half ran from 6.6% to 34.5%, and 15.3% of pairs differed by more than half. This says the network is not perfectly self-consistent. It does not say which gauge is right, and it is not an uncertainty range for any one site.
- The easy test flattered the new method. Letting nearby rivers stay in training put the FEH 2025 method at 0.922 and lifted the machine-learning model to 0.915 — close enough to argue about. Holding whole regions back left the model at 0.898 against 0.922. We withdrew the earlier reading that the model won.
- Donor choice is part of the test, not a detail. Allowing the statistical method to borrow from gauges in the held-back region moved its score from 0.922 to 0.940. That is a bigger swing than the gap being argued over.
- Forty years buys different amounts of certainty for different things. The achievable score was 0.857 for spread, 0.723 for tail growth and 0.581 for lopsidedness. Reshuffling years inside a record moved the 100-year growth factor by about ±13.0%. That is record-length wobble, not a general allowance to add to a design flow.
- Which years you test on changes the verdict. One station ran from 0.495 to 0.815 depending on the split. Two stations is a demonstration, not a national trend, and it says nothing about the cause.
What we decided next
Decided 14 July 2026. These four numbers never get added together, and Hydrometric reports them in four separate places. Network disagreement, method performance across the country, record-length wobble and time-period sensitivity answer different questions on different sets of rivers. Stacking them into one uncertainty band would be arithmetic without a meaning.
Two things changed in how we run and present work. Every method comparison is now scored with whole regions held back and donors confined to the training side; an easier score is only ever shown beside the fair one. And adjacent-gauge disagreement became a review prompt: where two neighbours disagree sharply, the pairing, catchment areas, record periods and rating evidence get looked at before the estimate is trusted.
Left open: we have not run the same-evidence comparison of candidate distribution families, so we cannot yet show how much of a design flow depends on that choice alone.
Read the detailExact tables, method names and the full claim record
Research question
What becomes visible when UK flood-hydrology evidence is examined across a national network rather than one catchment at a time? The practical question is not whether a large dataset automatically produces certainty. It is whether repeated comparisons reveal inconsistencies, validation leakage, sampling limits and temporal sensitivity that a one-site workflow can leave untested.
Data and bounded cohorts
The evidence comes from four distinct experiments. The network study area-scales QMED before comparing adjacent gauges paired by river name. The QMED revalidation freezes a national station cohort and compares random station validation with five-fold unbuffered KMeans spatial-cluster cross-validation. The growth artefact selected 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage for the rating fit. The temporal study changes the held-out periods for a case station and a control.
The growth ceilings were estimated by resampling years within each station, recomputing L-moments and GF100, and using within-station sampling variance to bound observable-target R². These cohorts answer different questions and are not pooled into a single uncertainty calculation. Their exact sample sizes, statistics and evidence identifiers are recorded in the distributions and evidence record on this page.
Validation design changes the claim
Random folds test interpolation among stations that may still have nearby examples in training. Spatial-cluster folds instead hold out centroid-based geographic groups, making the experiment closer to transfer into an ungauged area. The KMeans folds were unbuffered: they do not assert a protected distance around each evaluation station. Donor custody was also enforced by restricting donor selection to the training folds.
Information symmetry matters here. The frozen challenger used the statistical-method estimate as an input, while its frozen urban-extent field carried no variation. The comparison is therefore a test of that specific challenger against donor-adjusted SM2025, not evidence that a wholly separate estimator surpassed the benchmark.
What the national view exposes
A single-site calculation can be internally tidy while neighbouring records remain mutually difficult to reconcile. Adjacent-gauge comparison brings that inconsistency into view before a predictive model is introduced. It is most useful as a review trigger: check the pairing, catchment areas, record periods, rating evidence and local hydrological explanation before deciding what the disagreement means.
The diagnostic does not identify which gauge is closer to truth and does not allocate disagreement among measurement, catchment or record effects. That is why its distribution remains separate from model validation and site-specific uncertainty statements.
Scale, skew and growth are different quantities
QMED is the index-flood magnitude. L-CV is the second L-moment divided by the first, a dimensionless measure of relative spread. L-skew is the third L-moment divided by the second and describes asymmetry, which gives the sparsely observed upper tail greater leverage. GF100 is the fitted one-hundred-year quantile divided by QMED: a tail growth ratio, not a substitute for diagnosing either relative spread or skew.
Finite-record sampling constrains L-CV, L-skew and GF100 differently. The ceiling estimates bound observable-target R² using the measured within-station sampling variance. The published artefact does not report an uncertainty band on those ceiling estimates, and the within-record GF100 spread must not be promoted into a general allowance for design-flow uncertainty.
The temporal example is a case, not a national trend
Each partition uses 60% calibration and 40% validation after a 730-day spin-up. P-A asks early → late, P-B late → early and P-C middle → outer. Holding out these different periods asks whether a model transfers across different parts of a record, not merely whether it reproduces a random sample from the same mixture of years. The case and control illustrate why the chosen periods should be disclosed and why a single aggregate score can hide temporal sensitivity.
The experiment was designed as a methodological demonstration. It has neither the cohort nor the attribution design needed to estimate how widespread non-stationarity is or to assign a cause.
Withdrawn interpretation and ruled-out conclusions
The random-validation accuracy-win interpretation is the only retraction recorded here. It was withdrawn because it did not survive the geographic-transfer test. The other two failed outcomes are ruled-out conclusions, not retractions or tested hypotheses: paired-gauge disagreement does not justify a station-level interval, and within-record GF100 variability does not justify a universal design-flow allowance.
Operational implications
- Review neighbouring gauges and their evidence before treating one at-site estimate as an unquestioned reference.
- Match validation to the deployment claim, disclose the fold construction, and keep held-out information out of donor choices.
- Report index-flood magnitude, relative spread, skew and tail growth separately so that sensitivity is visible rather than compressed into one number.
- Keep network consistency, population cross-validation, sampling variability and temporal sensitivity as distinct uncertainty evidence unless a calibrated combination has been demonstrated.
Limitations and disclosure boundary
River-name pairing is heuristic. The spatial clusters are unbuffered centroid groups. The growth result applies near the studied record length, and the published artefact gives no uncertainty band for the estimated ceilings. The temporal result is deliberately a bounded case study. Each limitation constrains interpretation even though the underlying cohort is broader than a conventional single-site analysis.
We disclose cohorts, validation logic, exact result distributions, failed hypotheses, evidence identifiers and operational meaning. We withhold feature engineering, learned parameters, training recipes, deployable artefacts, private data mechanics, operational infrastructure, customer material and licensed outputs. This is enough to audit the claim without reconstructing the engine.
What we do not conclude
- The same-river diagnostic is not a site-specific uncertainty interval and does not identify a single cause of disagreement.
- The QMED experiment does not show that the challenger is separate from SM2025 or that random validation is adequate for ungauged transfer.
- The growth experiment does not provide a universal uncertainty percentage for a design flood.
- The temporal case does not estimate national prevalence and does not attribute the observed sensitivity to climate.
- None of these experiments establishes regulatory approval or the replacement of established flood-estimation methods.
Evidence distributions
Area-scaled QMED disagreement between adjacent same-river gauges
85 adjacent same-river pairs
| Statistic | Disagreement |
|---|---|
| P10 | 3.4% |
| Q1 | 6.6% |
| Median | 12.2% |
| Q3 | 34.5% |
| P90 | 58.8% |
| Pairs above 50% | 15.3% |
Nearby gauges disagree before any model enters the comparison. This is a network self-consistency diagnostic, not an irreducible site-specific error floor or confidence interval. River-name pairing is heuristic.
Ungauged QMED under spatial-cluster cross-validation
918 stations
| Method | R² | Median absolute error | FSE |
|---|---|---|---|
| ML challenger | 0.898 | 26.1% | 1.634 |
| Donor-adjusted SM2025 | 0.922 | 23.4% | 1.538 |
| SM2025 with all donors available | 0.940 | — | — |
The challenger did not beat the statistical benchmark. Restricting donors for each held-out station to the training folds reduced the apparent SM2025 result from R² 0.940 to 0.922. Paired ΔR² was −0.024 with 95% CI −0.031 to −0.016.
Estimated target ceilings for an approximately 40-year AMAX record
874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
| Target | Estimated R² ceiling |
|---|---|
| L-CV | 0.857 |
| L-skew | 0.581 |
| GF100 | 0.723 |
Finite-record sampling constrains L-CV, L-skew and GF100 differently. Median within-station GF100 sampling variability was about ±13.0%; this is not universal design-flow uncertainty. The published artefact does not report an uncertainty band on the ceiling estimates.
Held-out log-KGE under named temporal partitions
Two-station methodological case study
| Station | P-A · early → late | P-B · late → early | P-C · middle → outer |
|---|---|---|---|
| 43007 | 0.584 | 0.495 | 0.815 |
| 33029 control | 0.805 | 0.721 | 0.724 |
Each partition used 60% calibration and 40% validation after a 730-day spin-up. P-A tests early → late, P-B late → early and P-C middle → outer. This two-station case shows why temporal partitioning changes the scientific question. It does not estimate national non-stationarity prevalence or attribute change to climate.
What survived
Network comparison as a review signal
Adjacent-gauge disagreement remains useful for identifying records and pairings that deserve scrutiny before a model comparison begins.
Validation matched to geographic transfer
The spatial-cluster comparison remains the relevant test for an ungauged claim because random station folds can preserve nearby information across training and evaluation sets.
Separate treatment of scale and skew
Relative spread, asymmetry and growth-factor sampling behaviour remain distinct reporting quantities; one cannot stand in for the others.
What failed
The QMED accuracy-win hypothesis
The optimistic interpretation from random validation was withdrawn when the challenger remained behind donor-adjusted SM2025 under spatial-cluster validation.
A station-level interval from paired gauges
The same-river distribution cannot be converted into an uncertainty interval for an individual site or assigned to one source of disagreement.
A universal tail allowance from record length
Within-station GF100 sampling behaviour does not establish a general uncertainty allowance for every design-flow estimate.
Claim record
Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.
Demonstrated
Across 85 adjacent same-river gauge pairs, area-scaled QMED disagreement had a median of 12.2%, Q1–Q3 of 6.6–34.5%, P90 of 58.8%, and 15.3% of pairs differed by more than 50%.
- Interpretation
- National-scale pairing exposes inconsistency that a one-site workflow cannot show.
- Boundary
- This is a network self-consistency diagnostic, not an irreducible site-specific error floor or confidence interval. River-name pairing is heuristic.
- Linked result
- Area-scaled QMED disagreement between adjacent same-river gauges; sample: 85 adjacent same-river pairs
- Evidence
- Hydrometric same-river network consistency study (2026).Evidence ID:
db10a70d
- Hydrometric same-river network consistency study (2026).Evidence ID:
Demonstrated
Under spatial-cluster cross-validation on 918 stations, the ungauged ML challenger achieved R² 0.898, median absolute error 26.1% and FSE 1.634; donor-adjusted SM2025 achieved R² 0.922, median absolute error 23.4% and FSE 1.538. ΔR² was −0.024 [95% CI −0.031, −0.016].
- Interpretation
- The challenger did not beat the statistical benchmark under the validation design aligned with geographic transfer.
- Boundary
- The comparison used five-fold unbuffered KMeans spatial-cluster cross-validation. The frozen challenger included the statistical-method estimate among its inputs, and the frozen urban-extent field contained zeros throughout.
- Linked result
- Ungauged QMED under spatial-cluster cross-validation; sample: 918 stations
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Demonstrated
The apparent all-donor result fell from R² 0.940 to 0.922 when donors for each held-out station were restricted to the training folds.
- Interpretation
- Donor availability is part of the validation design, not a neutral implementation detail.
- Boundary
- The restricted result tests geographic transfer without retaining held-out neighbours as donors.
- Linked result
- Ungauged QMED under spatial-cluster cross-validation; sample: 918 stations
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Retracted
The random-validation accuracy-win interpretation for the QMED challenger was withdrawn.
- Interpretation
- Five-fold unbuffered KMeans spatial-cluster cross-validation placed the challenger behind donor-adjusted SM2025.
- Boundary
- Only the accuracy-win interpretation is retracted; this does not establish that machine learning is unusable for QMED.
- Evidence
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b
- Hydrometric QMED adversarial revalidation (2026).Evidence ID:
Demonstrated
The growth artefact selected 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage available for the rating fit.
- Interpretation
- Target ceilings were estimated by resampling years within each station, recomputing L-moments and GF100, and using within-station sampling variance to bound observable-target R².
- Boundary
- The published artefact does not report an uncertainty band on the ceiling estimates.
- Linked result
- Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Demonstrated
For approximately 40-year AMAX records, estimated target ceilings were R² 0.857 for L-CV, 0.581 for L-skew and 0.723 for GF100; median within-station GF100 uncertainty was about ±13.0%.
- Interpretation
- Finite-record sampling constrains L-CV, L-skew and GF100 differently.
- Boundary
- The GF100 result is median within-station sampling variability near the studied record length, not universal design-flow uncertainty. The published artefact does not report an uncertainty band on the ceiling estimates.
- Linked result
- Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations; each had at least 10 non-rejected AMAX years and per-peak stage for the rating fit
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Demonstrated
Station 43007 produced held-out log-KGE scores of 0.584, 0.495 and 0.815 under different temporal partitions; control station 33029 produced 0.805, 0.721 and 0.724.
- Interpretation
- Each partition used 60% calibration and 40% validation after a 730-day spin-up: P-A evaluated early → late, P-B late → early and P-C middle → outer.
- Boundary
- This two-station methodological case does not establish national non-stationarity prevalence or climate causation.
- Linked result
- Held-out log-KGE under named temporal partitions; sample: Two-station methodological case study
- Evidence
- Hydrometric temporal-partition case study (2026).Evidence ID:
e5da4413
- Hydrometric temporal-partition case study (2026).Evidence ID:
Inference
A defensible review should keep network disagreement, population validation scatter, record-length sampling variability and temporal sensitivity as separate evidence layers.
- Interpretation
- Separation preserves the meaning and scope of each experiment.
- Boundary
- This is a reporting inference; the components have not been calibrated as a combined site-specific probability model.
- Evidence
- Hydrometric same-river network consistency study (2026).Evidence ID:
db10a70d - Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b - Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898 - Hydrometric temporal-partition case study (2026).Evidence ID:
e5da4413
- Hydrometric same-river network consistency study (2026).Evidence ID:
Primary sources
- Hydrometric same-river network consistency study (2026).Evidence ID:
db10a70d - Hydrometric QMED adversarial revalidation (2026).Evidence ID:
2438954b - Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898 - Hydrometric temporal-partition case study (2026).Evidence ID:
e5da4413
Version history
- Version 1.0.0current
Initial publication of the national hydrology evidence synthesis.
Related reports
- QMEDCan Machine Learning Improve QMED?We gave a machine-learning model every advantage and it still lost to the FEH 2025 statistical method on 918 river gauges. Here is the losing scoreline, and why the losing test was the right one.
- Statistical methodsThe 2025 Statistical Method: What Changed and Why It MattersFive linked stages of the UK flood-estimation standard changed at once in 2025. What each one does, which three numbers our software is held to, and which decisions still belong to the hydrologist.
- Growth curvesSkew, Scale and the Shape of UK Flood GrowthA flood growth curve has four moving parts, and two promising shortcuts to it did not work. One needed the very gauge it was predicting; the other was not saved by the shortness of real records.
- UncertaintyHow Certain Is a Design Flood?We have three measurements that all sound like the uncertainty in a design flood. They answer three different questions on three different sets of rivers, so we publish them separately and never add them up.
- Rainfall-runoff modellingFrom ReFH Design Events to Continuous HydrologyA stocktake of four rainfall-to-flood modelling threads: one regional refit held up, one gap measurement landed near 50%, one daily model failed its own test, and one result was withheld for missing evidence.
- Flood estimation practiceThe Modern Flood Estimation ReportWe marked our own report generator against the eleven things a Flood Estimation Report must record. One is complete, three are partial and seven are missing — so we call the output a calculation report.