Rainfall-runoff modelling
From ReFH Design Events to Continuous Hydrology
- Version
- 1.0.0
- Status
- Current
- Published
- Evidence cut-off
In short
We asked how much of our rainfall-to-flood modelling work is solid enough to publish. One regional refit held up across 640 catchments, one measurement showed the open version of the model sitting about half a world away from the licensed one, one daily model failed its own test and was closed, and one result was left out because the evidence for it was not where it should be.
Findings at a glance
- Catchments in the regional refit
- 640Scored with whole hydrometric areas withheld one at a time640 catchments
- What the refit explained
- 0.528Share of the variation, range 0.451 to 0.583, against a published figure of 0.35 scored on its own fitting data640 catchments
- Gap to the licensed model at 1 in 100 years
- 47.6%Our open ReFH1 implementation against genuine ReFH2 pilot outputs18 complete cases
- Catchments the daily model passed
- 2 of 9All three targets had to be met at the same catchmentNine stations
What we tested
The statistical method estimates how big a flood gets. It does not tell you the shape of the flood — how fast the river rises, how long it stays up, how much water passes in total. For that you need a rainfall-runoff model, and in the UK that usually means ReFH: the Revitalised Flood Hydrograph model. ReFH2 is the current licensed version; ReFH1, its predecessor, is published openly.
We had four threads of work in this area and wanted to know which of them were solid enough to publish. This report is that stocktake.
The first thread refitted one of the model's parameters — how quickly a catchment's slow, sustained flow responds — across 640 catchments. To check it would work somewhere new, we withheld whole hydrometric areas at a time rather than scattered catchments.
The second measured a gap. Our open implementation runs on the published ReFH1 coefficients. We compared it with genuine ReFH2 outputs on 18 pilot cases, under controlled access to the licensed results, purely to find out how far apart the two sit.
The third tested a daily rainfall-runoff model on nine gauged catchments, against a test declared before we ran it: the model had to get the typical flood size, the year-to-year variation and the overall fit right, all three, at the same station.
The fourth is an hourly regional continuous model. It has results. They are not in this report, and the reason is below.
What we saw
The regional refit survived having whole regions taken away
640 catchments, withholding one hydrometric area at a time. The published figure beside it was scored on the same data it was fitted to.
How far the open model sits from the licensed one
18 complete pilot cases, comparing our open ReFH1 implementation with genuine ReFH2 outputs.
The daily model failed the test we set it in advance
Nine gauged catchments for the declared test; eight complete catchments for the loosened version.
Outcome
- The regional refit held up. Across 640 catchments with whole hydrometric areas withheld it explained 0.528 of the variation, range 0.451 to 0.583, against a published figure of 0.35. The published figure was scored on its own fitting data, so it is a reference point rather than an opponent — and our targets were derived from flow records and bias-corrected, which is a real caveat, not a footnote.
- The open model is a long way from the licensed one. Across 18 pilot cases the gap on the 1 in 100 year peak was 47.6%, with an average absolute gap of 53.5%. That is the size of the starting gap. It is not a claim to have beaten ReFH2, and no licensed case output crosses into this publication.
- The daily model failed, on its own terms. It cleared the combined test at two of nine catchments. Loosening the test reached at most four of eight, and the year-to-year variation it produced stayed too narrow throughout.
- One result was withheld for lack of evidence, not lack of performance. The regional hourly model's scores were not backed by the fixed, addressable evidence files our publication rule requires, so no number for it appears here. That is a gate we set ourselves and chose not to step around.
What we decided next
Decided 14 July 2026. We stopped work on the daily lumped two-store model as a route to flood growth curves. It failed the test we set before running it, and the loosened version did not rescue it. Effort moved to model classes that carry state between catchments and through time.
Decided 14 July 2026. The baseflow-lag refit stays, with its target construction and the withheld-region test reported alongside it every time the number is quoted.
Decided 14 July 2026. No regional continuous score gets published until its evidence files exist, are fixed and match the record they claim to come from. When that gate is met, the strongest thing such a result could say is that the model reproduces past flows from measured rainfall — not that it forecasts weather, not that it estimates rare floods, and not that it replaces ReFH2. Stating that boundary now stops a future result quietly growing into a bigger claim.
The pilot comparison also settled a procurement question: closing a gap of that size needs controlled access to licensed reference outputs, so validation work is planned around obtaining it rather than around further open-coefficient tuning.
Read the detailExact tables, method names and the full claim record
Research question
Which parts of the ReFH and continuous-hydrology research record are strong enough to publish, which tested branch failed, and which regional-continuous claims remain outside the evidence gate? The report separates parameter transfer, a public-baseline comparison, daily lumped simulation and future hourly continuous evidence because they do not share a target or validation design.
Evidence and method
Three completed experiments enter the numerical ledger. A 640-catchment refit withheld one hydrometric area at a time. A commercially controlled pilot compared aggregate outputs from the public ReFH1 coefficient and equation stack with genuine ReFH2 outputs across 18 complete cases. A daily two-store study applied a combined QMED, L-CV and KGE decision gate to nine stations and retained an eight-station bounded sensitivity check.
Regional continuous Stage A is handled differently. Publication required immutable score artefacts in this implementation worktree and a match to the cited evidence ledger. Those artefacts are absent, so the regional section carries a structured not-yet-demonstrated status and no score distribution.
Baseflow-lag refit
The flow-derived baseflow-lag refit reached leave-one-hydrometric-area-out R² 0.528 [0.451, 0.583] across 640 catchments. The published WHS in-sample R² 0.35 is retained as a reference, giving a difference of 0.178 [0.101, 0.233]. Withholding whole hydrometric areas makes the refit result relevant to regional transfer rather than within-sample fit alone.
The target custody is material: targets are flow-derived and bias-mapped. The refit retained held-area signal, while the published WHS figure is a contextual reference only because it uses a different validation design. This is not a like-for-like contest between identical experiments. A disattenuated statistic is omitted because it is not measured predictive performance.
Commercial evidence custody
Across 18 complete cases, the public ReFH1 stack differed from genuine ReFH2 pilot outputs by 47.6% on Q100 peak; its mean absolute gap was 53.5%. This is a baseline gap, not a ReFH2 beat. It establishes that the open baseline and the licensed reference were materially different in the pilot.
Only aggregate statistics cross the publication boundary. No licensed case output, private row, coefficient set or reconstruction detail is reproduced. The comparison therefore documents baseline adequacy and evidence custody without exposing the commercial reference surface.
Failed daily two-store branch
The tested daily lumped branch cleared its combined QMED/L-CV/KGE gate at two of nine stations. Under the pre-bounded relaxation, performance reached at most four of eight complete stations. Across the objective variants, model L-CV of 0.18–0.21 remained too narrow to establish the required between-basin growth behaviour.
Only that per-basin daily lumped branch is closed. The result does not rule out regional models, persistent hourly state, alternative model classes or continuous simulation generally. Nor are daily held-period maxima presented as instantaneous design-flood evidence.
Stage A publication gate
Regional continuous Stage A is classified as not-yet-demonstrated in this version. The required immutable score artefacts are absent from the implementation worktree, so the evidence gate fails and no Stage A performance metric, percentile or station count is published.
If a future version satisfies the immutable gate, its admissible claim would still be an observed-forcing hourly hindcast, not weather forecasting, Q100 skill or ReFH2 replacement. That future boundary is stated now so later evidence cannot silently broaden the scientific target.
What survived and what failed
The held-area baseflow-lag refit survived as bounded regional parameter evidence. The aggregate commercial-custody comparison survived as a diagnostic of the public ReFH1 starting point. The per-basin daily two-store branch failed its combined gate and remains closed within that exact implementation boundary.
Regional continuous publication remains not-yet-demonstrated because its evidence gate was not met. That is not a negative score for the model class: the necessary immutable proof is missing from the publication workspace and the public claim remains open.
Operational implications
- Keep parameter-transfer results tied to the target construction and geographic holdout used to obtain them.
- Use the ReFH pilot only to establish the open baseline gap and the need for controlled licensed validation.
- Do not spend further effort tuning the closed per-basin daily lumped branch as a standalone growth-curve solution.
- Require immutable, report-addressable artefacts before publishing a regional continuous score or distribution.
- Keep hydrograph response, event replay and flood-frequency targets in separate claim ledgers.
Limitations
- Baseflow-lag targets are flow-derived and bias-mapped rather than direct observations of a unique physical store parameter.
- The WHS reference is in-sample and therefore does not share the held-area validation design of the refit.
- The genuine ReFH2 comparison contains 18 complete pilot cases and publishes aggregate gaps only.
- The daily evidence concerns a small frozen cohort, daily resolution and one per-basin lumped architecture.
- Stage A contributes no public performance result because its immutable evidence gate was not satisfied in this worktree.
What we do not conclude
- We do not present a disattenuated association as measured held-area predictive performance.
- We do not treat the public-stack gap as superiority to genuine ReFH2 or as evidence of product equivalence.
- We do not generalise the daily two-store failure beyond the tested per-basin daily lumped branch.
- We do not convert daily maxima or hydrograph response into Q100 evidence.
- We do not publish regional continuous performance without the immutable artefact and evidence-ledger match.
Disclosure boundary
This report discloses the research questions, broad cohorts, validation designs, aggregate distributions, failed gate, claim statuses, source revisions and scientific boundaries. Its only sources are the frozen baseflow-lag refit, public-stack gap, daily Arm A, flood-objective and regional Stage A evidence records.
It withholds licensed case outputs, private rows, learned coefficients, fitted parameter artefacts, exact feature construction, calibration sequences, runnable recipes and deployable model internals. The public evidence is sufficient to audit each claim without reconstructing the proprietary method or importing unverified Stage A results.
Evidence distributions
Baseflow-lag refit and published reference
640 catchments in the Hydrometric refit. The WHS figure is a published in-sample reference, not a second result from the same cohort or validation folds.
| Evidence row | Validation design | R² | Interval | Cohort and custody |
|---|---|---|---|---|
| Hydrometric baseflow-lag refit | Leave-one-hydrometric-area-out | 0.528 | [0.451, 0.583] | 640 catchments; flow-derived and bias-mapped targets |
| Published WHS reference | In-sample | 0.35 | Not reported in the cited comparison | Published model-development statistic; not the same validation design |
| Difference in R² | Hydrometric held-area result minus the published in-sample reference | 0.178 | [0.101, 0.233] | Arithmetic comparison across differently validated results |
The refit reached leave-one-hydrometric-area-out R² 0.528 [0.451, 0.583], compared with published WHS in-sample R² 0.35; the difference was 0.178 [0.101, 0.233]. The validation designs differ, and the refit targets are flow-derived and bias-mapped.
Public ReFH1 stack against genuine ReFH2 pilot outputs
18 complete pilot cases. Only aggregate statistics are published; case rows and licensed outputs remain withheld.
| Comparison statistic | Result | Cases | Interpretation |
|---|---|---|---|
| Q100 peak gap | 47.6% | 18 complete cases | Aggregate gap between the public ReFH1 coefficient/equation stack and genuine ReFH2 pilot outputs |
| Mean absolute gap | 53.5% | 18 complete cases | Aggregate absolute difference across the complete pilot cases |
This is a baseline gap, not a ReFH2 beat. It measures how far the public ReFH1 stack sat from genuine ReFH2 pilot outputs under controlled commercial custody; it does not expose or reconstruct those outputs.
Daily two-store branch adjudication
Nine stations for the frozen combined gate; eight complete stations for the bounded relaxation.
| Gate or diagnostic | Result | Cohort | Boundary |
|---|---|---|---|
| Frozen combined QMED/L-CV/KGE gate | Two of nine stations | Nine-station daily Arm A | The pre-declared combined gate was the decision rule |
| Bounded relaxation | At most four of eight | Eight complete stations in the relaxed comparison | A sensitivity bound, not a replacement success criterion |
| Between-basin model L-CV | 0.18–0.21 | Tested daily branch | Narrow dispersion persisted across the objective variants |
The daily lumped model passed the combined gate at two of nine stations. Under the bounded relaxation it reached at most four of eight, while model L-CV remained 0.18–0.21. Only that per-basin daily lumped branch is closed.
What survived
Baseflow-lag regional refit
Held-area validation retained predictive signal across 640 catchments, with the target-construction and comparator differences kept visible.
Commercial-custody baseline check
The pilot established the size of the public ReFH1 baseline gap without publishing licensed case outputs or treating that gap as superiority.
What failed
Per-basin daily lumped two-store branch
The combined gate passed at only two of nine stations, and bounded relaxation did not resolve the narrow between-basin L-CV.
Claim record
Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.
Demonstrated
On 640 catchments, the flow-derived baseflow-lag refit reached leave-one-hydrometric-area-out R² 0.528 [0.451, 0.583], compared with published WHS in-sample R² 0.35.
- Interpretation
- The observed difference in R² was 0.178 [0.101, 0.233], providing evidence that the regional refit retained signal when whole hydrometric areas were withheld.
- Boundary
- The targets are flow-derived and bias-mapped, and the WHS reference is in-sample; no disattenuated headline is presented as measured predictive performance.
- Linked result
- Baseflow-lag refit and published reference; sample: 640 catchments in the Hydrometric refit. The WHS figure is a published in-sample reference, not a second result from the same cohort or validation folds.
- Evidence
- Hydrometric ReFH parameter re-fit study (2026).Evidence ID:
bcb1200f
- Hydrometric ReFH parameter re-fit study (2026).Evidence ID:
Demonstrated
Across 18 complete cases, the public ReFH1 coefficient/equation stack differed from genuine ReFH2 pilot outputs by 47.6% on Q100 peak, with a 53.5% mean absolute gap.
- Interpretation
- The comparison quantifies the starting distance between an open baseline and licensed pilot outputs under controlled custody.
- Boundary
- This is a baseline gap, not a ReFH2 beat; it does not establish product equivalence, disclose case-level commercial outputs or transfer proprietary method details.
- Linked result
- Public ReFH1 stack against genuine ReFH2 pilot outputs; sample: 18 complete pilot cases. Only aggregate statistics are published; case rows and licensed outputs remain withheld.
- Evidence
- Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID:
a7f921fb
- Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID:
Demonstrated
The tested daily two-store branch cleared its combined QMED/L-CV/KGE gate at two of nine stations.
- Interpretation
- The branch failed its own multi-objective decision rule across most of the frozen station cohort.
- Boundary
- Only that per-basin daily lumped branch is closed; the result does not generalise to regional continuous simulation or every physical rainfall-runoff model.
- Linked result
- Daily two-store branch adjudication; sample: Nine stations for the frozen combined gate; eight complete stations for the bounded relaxation.
- Evidence
- Hydrometric daily two-store Arm A adjudication (2026).Evidence ID:
b153a973 - Hydrometric flood-weighted objective adjudication (2026).Evidence ID:
1510b52c
- Hydrometric daily two-store Arm A adjudication (2026).Evidence ID:
Demonstrated
Under a bounded relaxation, the daily branch reached at most four of eight complete stations, while model L-CV remained 0.18–0.21.
- Interpretation
- Objective changes moved some station-level outcomes without demonstrating adequate combined behaviour across the cohort.
- Boundary
- The relaxed count is a sensitivity bound, not a post-hoc success criterion; daily maxima are not described as instantaneous design-flood performance.
- Linked result
- Daily two-store branch adjudication; sample: Nine stations for the frozen combined gate; eight complete stations for the bounded relaxation.
- Evidence
- Hydrometric daily two-store Arm A adjudication (2026).Evidence ID:
b153a973 - Hydrometric flood-weighted objective adjudication (2026).Evidence ID:
1510b52c
- Hydrometric daily two-store Arm A adjudication (2026).Evidence ID:
Not yet demonstrated
Regional continuous Stage A remains not-yet-demonstrated in this public report because the immutable score artefacts required by its evidence gate are absent from the implementation worktree.
- Interpretation
- The evidence identifier is retained so a future version can be checked against admissible immutable artefacts without importing unverified performance claims now.
- Boundary
- If the publication gate is later satisfied, an admissible Stage A claim would remain an observed-forcing hourly hindcast, not weather forecasting, Q100 skill or ReFH2 replacement.
- Evidence
- Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID:
1a8d59c9
- Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID:
Inference
Hydrograph-response evidence and flood-frequency evidence should remain separate until each target is validated directly.
- Interpretation
- Parameter transfer, event-shape behaviour, continuous hindcast response and upper-tail design quantiles answer different scientific questions.
- Boundary
- This reporting distinction does not itself demonstrate a continuous model, a design quantile or superiority to an established method.
- Evidence
- Hydrometric ReFH parameter re-fit study (2026).Evidence ID:
bcb1200f - Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID:
a7f921fb - Hydrometric daily two-store Arm A adjudication (2026).Evidence ID:
b153a973 - Hydrometric flood-weighted objective adjudication (2026).Evidence ID:
1510b52c - Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID:
1a8d59c9
- Hydrometric ReFH parameter re-fit study (2026).Evidence ID:
Primary sources
- Hydrometric ReFH parameter re-fit study (2026).Evidence ID:
bcb1200f - Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID:
a7f921fb - Hydrometric daily two-store Arm A adjudication (2026).Evidence ID:
b153a973 - Hydrometric flood-weighted objective adjudication (2026).Evidence ID:
1510b52c - Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID:
1a8d59c9
Version history
- Version 1.0.0current
Initial publication of bounded ReFH parameter, baseline-gap, failed daily-branch and gated regional-continuous evidence.
Related reports
- National hydrologyWhat National-Scale Analysis Reveals About UK HydrologyLook at every gauged river at once and four things show up that one catchment never does: neighbours disagree, easy tests flatter new methods, forty years only buys so much, and the years you pick change the answer.
- QMEDCan Machine Learning Improve QMED?We gave a machine-learning model every advantage and it still lost to the FEH 2025 statistical method on 918 river gauges. Here is the losing scoreline, and why the losing test was the right one.
- Statistical methodsThe 2025 Statistical Method: What Changed and Why It MattersFive linked stages of the UK flood-estimation standard changed at once in 2025. What each one does, which three numbers our software is held to, and which decisions still belong to the hydrologist.
- Growth curvesSkew, Scale and the Shape of UK Flood GrowthA flood growth curve has four moving parts, and two promising shortcuts to it did not work. One needed the very gauge it was predicting; the other was not saved by the shortness of real records.
- UncertaintyHow Certain Is a Design Flood?We have three measurements that all sound like the uncertainty in a design flood. They answer three different questions on three different sets of rivers, so we publish them separately and never add them up.
- Flood estimation practiceThe Modern Flood Estimation ReportWe marked our own report generator against the eleven things a Flood Estimation Report must record. One is complete, three are partial and seven are missing — so we call the output a calculation report.