Skip to main content

Rainfall-runoff modelling

From ReFH Design Events to Continuous Hydrology

Version
1.0.0
Status
Current
Published
Evidence cut-off

In short

We asked how much of our rainfall-to-flood modelling work is solid enough to publish. One regional refit held up across 640 catchments, one measurement showed the open version of the model sitting about half a world away from the licensed one, one daily model failed its own test and was closed, and one result was left out because the evidence for it was not where it should be.

Findings at a glance

Catchments in the regional refit
640Scored with whole hydrometric areas withheld one at a time640 catchments
What the refit explained
0.528Share of the variation, range 0.451 to 0.583, against a published figure of 0.35 scored on its own fitting data640 catchments
Gap to the licensed model at 1 in 100 years
47.6%Our open ReFH1 implementation against genuine ReFH2 pilot outputs18 complete cases
Catchments the daily model passed
2 of 9All three targets had to be met at the same catchmentNine stations

What we tested

The statistical method estimates how big a flood gets. It does not tell you the shape of the flood — how fast the river rises, how long it stays up, how much water passes in total. For that you need a rainfall-runoff model, and in the UK that usually means ReFH: the Revitalised Flood Hydrograph model. ReFH2 is the current licensed version; ReFH1, its predecessor, is published openly.

We had four threads of work in this area and wanted to know which of them were solid enough to publish. This report is that stocktake.

The first thread refitted one of the model's parameters — how quickly a catchment's slow, sustained flow responds — across 640 catchments. To check it would work somewhere new, we withheld whole hydrometric areas at a time rather than scattered catchments.

The second measured a gap. Our open implementation runs on the published ReFH1 coefficients. We compared it with genuine ReFH2 outputs on 18 pilot cases, under controlled access to the licensed results, purely to find out how far apart the two sit.

The third tested a daily rainfall-runoff model on nine gauged catchments, against a test declared before we ran it: the model had to get the typical flood size, the year-to-year variation and the overall fit right, all three, at the same station.

The fourth is an hourly regional continuous model. It has results. They are not in this report, and the reason is below.

What we saw

The regional refit survived having whole regions taken away

640 catchments, withholding one hydrometric area at a time. The published figure beside it was scored on the same data it was fitted to.

0.300.400.500.60Share of the variation explained (R²) — axis starts at 0.30Our refit, whole regions withheld (640 catchments)Our refit, whole regions withheld (640 catchments) — Central value: 0.5280.528Published figure, scored on its own fitting dataPublished figure, scored on its own fitting data — Published value: 0.350.35
The refit explained more of the variation than the published figure even with whole regions withheld — but the two were scored under different rules, so this is context rather than a contest.

How far the open model sits from the licensed one

18 complete pilot cases, comparing our open ReFH1 implementation with genuine ReFH2 outputs.

0%20%40%60%Difference from genuine ReFH2 pilot outputs (%)1 in 100 year peak flow1 in 100 year peak flow — Difference from genuine ReFH2 outputs: 47.6%47.6%Average across all comparisonsAverage across all comparisons — Difference from genuine ReFH2 outputs: 53.5%53.5%
Roughly half a world apart on the same 18 cases — this measures the size of the starting gap, and says nothing about which model is better.

The daily model failed the test we set it in advance

Nine gauged catchments for the declared test; eight complete catchments for the loosened version.

Declared test: all three targets at once (9 catchments)Declared test: all three targets at once (9 catchments) — Cleared the test: 2 cleared2 clearedDeclared test: all three targets at once (9 catchments) — Did not clear the test: 7 did not7 did notLoosened check (8 complete catchments)Loosened check (8 complete catchments) — Cleared the test: up to 4 clearedup to 4 clearedLoosened check (8 complete catchments) — Did not clear the test: 4 did not4 did not
Two stations out of nine cleared all three targets at once — and loosening the test reached at most four of eight, which is not a pass.

Outcome

  • The regional refit held up. Across 640 catchments with whole hydrometric areas withheld it explained 0.528 of the variation, range 0.451 to 0.583, against a published figure of 0.35. The published figure was scored on its own fitting data, so it is a reference point rather than an opponent — and our targets were derived from flow records and bias-corrected, which is a real caveat, not a footnote.
  • The open model is a long way from the licensed one. Across 18 pilot cases the gap on the 1 in 100 year peak was 47.6%, with an average absolute gap of 53.5%. That is the size of the starting gap. It is not a claim to have beaten ReFH2, and no licensed case output crosses into this publication.
  • The daily model failed, on its own terms. It cleared the combined test at two of nine catchments. Loosening the test reached at most four of eight, and the year-to-year variation it produced stayed too narrow throughout.
  • One result was withheld for lack of evidence, not lack of performance. The regional hourly model's scores were not backed by the fixed, addressable evidence files our publication rule requires, so no number for it appears here. That is a gate we set ourselves and chose not to step around.

What we decided next

Decided 14 July 2026. We stopped work on the daily lumped two-store model as a route to flood growth curves. It failed the test we set before running it, and the loosened version did not rescue it. Effort moved to model classes that carry state between catchments and through time.

Decided 14 July 2026. The baseflow-lag refit stays, with its target construction and the withheld-region test reported alongside it every time the number is quoted.

Decided 14 July 2026. No regional continuous score gets published until its evidence files exist, are fixed and match the record they claim to come from. When that gate is met, the strongest thing such a result could say is that the model reproduces past flows from measured rainfall — not that it forecasts weather, not that it estimates rare floods, and not that it replaces ReFH2. Stating that boundary now stops a future result quietly growing into a bigger claim.

The pilot comparison also settled a procurement question: closing a gap of that size needs controlled access to licensed reference outputs, so validation work is planned around obtaining it rather than around further open-coefficient tuning.

Read the detailExact tables, method names and the full claim record

Research question

Which parts of the ReFH and continuous-hydrology research record are strong enough to publish, which tested branch failed, and which regional-continuous claims remain outside the evidence gate? The report separates parameter transfer, a public-baseline comparison, daily lumped simulation and future hourly continuous evidence because they do not share a target or validation design.

Evidence and method

Three completed experiments enter the numerical ledger. A 640-catchment refit withheld one hydrometric area at a time. A commercially controlled pilot compared aggregate outputs from the public ReFH1 coefficient and equation stack with genuine ReFH2 outputs across 18 complete cases. A daily two-store study applied a combined QMED, L-CV and KGE decision gate to nine stations and retained an eight-station bounded sensitivity check.

Regional continuous Stage A is handled differently. Publication required immutable score artefacts in this implementation worktree and a match to the cited evidence ledger. Those artefacts are absent, so the regional section carries a structured not-yet-demonstrated status and no score distribution.

Baseflow-lag refit

The flow-derived baseflow-lag refit reached leave-one-hydrometric-area-out R² 0.528 [0.451, 0.583] across 640 catchments. The published WHS in-sample R² 0.35 is retained as a reference, giving a difference of 0.178 [0.101, 0.233]. Withholding whole hydrometric areas makes the refit result relevant to regional transfer rather than within-sample fit alone.

The target custody is material: targets are flow-derived and bias-mapped. The refit retained held-area signal, while the published WHS figure is a contextual reference only because it uses a different validation design. This is not a like-for-like contest between identical experiments. A disattenuated statistic is omitted because it is not measured predictive performance.

Commercial evidence custody

Across 18 complete cases, the public ReFH1 stack differed from genuine ReFH2 pilot outputs by 47.6% on Q100 peak; its mean absolute gap was 53.5%. This is a baseline gap, not a ReFH2 beat. It establishes that the open baseline and the licensed reference were materially different in the pilot.

Only aggregate statistics cross the publication boundary. No licensed case output, private row, coefficient set or reconstruction detail is reproduced. The comparison therefore documents baseline adequacy and evidence custody without exposing the commercial reference surface.

Failed daily two-store branch

The tested daily lumped branch cleared its combined QMED/L-CV/KGE gate at two of nine stations. Under the pre-bounded relaxation, performance reached at most four of eight complete stations. Across the objective variants, model L-CV of 0.18–0.21 remained too narrow to establish the required between-basin growth behaviour.

Only that per-basin daily lumped branch is closed. The result does not rule out regional models, persistent hourly state, alternative model classes or continuous simulation generally. Nor are daily held-period maxima presented as instantaneous design-flood evidence.

Stage A publication gate

Regional continuous Stage A is classified as not-yet-demonstrated in this version. The required immutable score artefacts are absent from the implementation worktree, so the evidence gate fails and no Stage A performance metric, percentile or station count is published.

If a future version satisfies the immutable gate, its admissible claim would still be an observed-forcing hourly hindcast, not weather forecasting, Q100 skill or ReFH2 replacement. That future boundary is stated now so later evidence cannot silently broaden the scientific target.

What survived and what failed

The held-area baseflow-lag refit survived as bounded regional parameter evidence. The aggregate commercial-custody comparison survived as a diagnostic of the public ReFH1 starting point. The per-basin daily two-store branch failed its combined gate and remains closed within that exact implementation boundary.

Regional continuous publication remains not-yet-demonstrated because its evidence gate was not met. That is not a negative score for the model class: the necessary immutable proof is missing from the publication workspace and the public claim remains open.

Operational implications

  • Keep parameter-transfer results tied to the target construction and geographic holdout used to obtain them.
  • Use the ReFH pilot only to establish the open baseline gap and the need for controlled licensed validation.
  • Do not spend further effort tuning the closed per-basin daily lumped branch as a standalone growth-curve solution.
  • Require immutable, report-addressable artefacts before publishing a regional continuous score or distribution.
  • Keep hydrograph response, event replay and flood-frequency targets in separate claim ledgers.

Limitations

  • Baseflow-lag targets are flow-derived and bias-mapped rather than direct observations of a unique physical store parameter.
  • The WHS reference is in-sample and therefore does not share the held-area validation design of the refit.
  • The genuine ReFH2 comparison contains 18 complete pilot cases and publishes aggregate gaps only.
  • The daily evidence concerns a small frozen cohort, daily resolution and one per-basin lumped architecture.
  • Stage A contributes no public performance result because its immutable evidence gate was not satisfied in this worktree.

What we do not conclude

  • We do not present a disattenuated association as measured held-area predictive performance.
  • We do not treat the public-stack gap as superiority to genuine ReFH2 or as evidence of product equivalence.
  • We do not generalise the daily two-store failure beyond the tested per-basin daily lumped branch.
  • We do not convert daily maxima or hydrograph response into Q100 evidence.
  • We do not publish regional continuous performance without the immutable artefact and evidence-ledger match.

Disclosure boundary

This report discloses the research questions, broad cohorts, validation designs, aggregate distributions, failed gate, claim statuses, source revisions and scientific boundaries. Its only sources are the frozen baseflow-lag refit, public-stack gap, daily Arm A, flood-objective and regional Stage A evidence records.

It withholds licensed case outputs, private rows, learned coefficients, fitted parameter artefacts, exact feature construction, calibration sequences, runnable recipes and deployable model internals. The public evidence is sufficient to audit each claim without reconstructing the proprietary method or importing unverified Stage A results.

Evidence distributions

Baseflow-lag refit and published reference

640 catchments in the Hydrometric refit. The WHS figure is a published in-sample reference, not a second result from the same cohort or validation folds.

Evidence rowValidation designIntervalCohort and custody
Hydrometric baseflow-lag refitLeave-one-hydrometric-area-out0.528[0.451, 0.583]640 catchments; flow-derived and bias-mapped targets
Published WHS referenceIn-sample0.35Not reported in the cited comparisonPublished model-development statistic; not the same validation design
Difference in R²Hydrometric held-area result minus the published in-sample reference0.178[0.101, 0.233]Arithmetic comparison across differently validated results

The refit reached leave-one-hydrometric-area-out R² 0.528 [0.451, 0.583], compared with published WHS in-sample R² 0.35; the difference was 0.178 [0.101, 0.233]. The validation designs differ, and the refit targets are flow-derived and bias-mapped.

Public ReFH1 stack against genuine ReFH2 pilot outputs

18 complete pilot cases. Only aggregate statistics are published; case rows and licensed outputs remain withheld.

Comparison statisticResultCasesInterpretation
Q100 peak gap47.6%18 complete casesAggregate gap between the public ReFH1 coefficient/equation stack and genuine ReFH2 pilot outputs
Mean absolute gap53.5%18 complete casesAggregate absolute difference across the complete pilot cases

This is a baseline gap, not a ReFH2 beat. It measures how far the public ReFH1 stack sat from genuine ReFH2 pilot outputs under controlled commercial custody; it does not expose or reconstruct those outputs.

Daily two-store branch adjudication

Nine stations for the frozen combined gate; eight complete stations for the bounded relaxation.

Gate or diagnosticResultCohortBoundary
Frozen combined QMED/L-CV/KGE gateTwo of nine stationsNine-station daily Arm AThe pre-declared combined gate was the decision rule
Bounded relaxationAt most four of eightEight complete stations in the relaxed comparisonA sensitivity bound, not a replacement success criterion
Between-basin model L-CV0.18–0.21Tested daily branchNarrow dispersion persisted across the objective variants

The daily lumped model passed the combined gate at two of nine stations. Under the bounded relaxation it reached at most four of eight, while model L-CV remained 0.18–0.21. Only that per-basin daily lumped branch is closed.

What survived

Baseflow-lag regional refit

Held-area validation retained predictive signal across 640 catchments, with the target-construction and comparator differences kept visible.

Commercial-custody baseline check

The pilot established the size of the public ReFH1 baseline gap without publishing licensed case outputs or treating that gap as superiority.

What failed

Per-basin daily lumped two-store branch

The combined gate passed at only two of nine stations, and bounded relaxation did not resolve the narrow between-basin L-CV.

Claim record

Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.

  1. Demonstrated

    On 640 catchments, the flow-derived baseflow-lag refit reached leave-one-hydrometric-area-out R² 0.528 [0.451, 0.583], compared with published WHS in-sample R² 0.35.

    Interpretation
    The observed difference in R² was 0.178 [0.101, 0.233], providing evidence that the regional refit retained signal when whole hydrometric areas were withheld.
    Boundary
    The targets are flow-derived and bias-mapped, and the WHS reference is in-sample; no disattenuated headline is presented as measured predictive performance.
    Linked result
    Baseflow-lag refit and published reference; sample: 640 catchments in the Hydrometric refit. The WHS figure is a published in-sample reference, not a second result from the same cohort or validation folds.
    Evidence
    • Hydrometric ReFH parameter re-fit study (2026).Evidence ID: bcb1200f
  2. Demonstrated

    Across 18 complete cases, the public ReFH1 coefficient/equation stack differed from genuine ReFH2 pilot outputs by 47.6% on Q100 peak, with a 53.5% mean absolute gap.

    Interpretation
    The comparison quantifies the starting distance between an open baseline and licensed pilot outputs under controlled custody.
    Boundary
    This is a baseline gap, not a ReFH2 beat; it does not establish product equivalence, disclose case-level commercial outputs or transfer proprietary method details.
    Linked result
    Public ReFH1 stack against genuine ReFH2 pilot outputs; sample: 18 complete pilot cases. Only aggregate statistics are published; case rows and licensed outputs remain withheld.
    Evidence
    • Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID: a7f921fb
  3. Demonstrated

    The tested daily two-store branch cleared its combined QMED/L-CV/KGE gate at two of nine stations.

    Interpretation
    The branch failed its own multi-objective decision rule across most of the frozen station cohort.
    Boundary
    Only that per-basin daily lumped branch is closed; the result does not generalise to regional continuous simulation or every physical rainfall-runoff model.
    Linked result
    Daily two-store branch adjudication; sample: Nine stations for the frozen combined gate; eight complete stations for the bounded relaxation.
    Evidence
    • Hydrometric daily two-store Arm A adjudication (2026).Evidence ID: b153a973
    • Hydrometric flood-weighted objective adjudication (2026).Evidence ID: 1510b52c
  4. Demonstrated

    Under a bounded relaxation, the daily branch reached at most four of eight complete stations, while model L-CV remained 0.18–0.21.

    Interpretation
    Objective changes moved some station-level outcomes without demonstrating adequate combined behaviour across the cohort.
    Boundary
    The relaxed count is a sensitivity bound, not a post-hoc success criterion; daily maxima are not described as instantaneous design-flood performance.
    Linked result
    Daily two-store branch adjudication; sample: Nine stations for the frozen combined gate; eight complete stations for the bounded relaxation.
    Evidence
    • Hydrometric daily two-store Arm A adjudication (2026).Evidence ID: b153a973
    • Hydrometric flood-weighted objective adjudication (2026).Evidence ID: 1510b52c
  5. Not yet demonstrated

    Regional continuous Stage A remains not-yet-demonstrated in this public report because the immutable score artefacts required by its evidence gate are absent from the implementation worktree.

    Interpretation
    The evidence identifier is retained so a future version can be checked against admissible immutable artefacts without importing unverified performance claims now.
    Boundary
    If the publication gate is later satisfied, an admissible Stage A claim would remain an observed-forcing hourly hindcast, not weather forecasting, Q100 skill or ReFH2 replacement.
    Evidence
    • Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID: 1a8d59c9
  6. Inference

    Hydrograph-response evidence and flood-frequency evidence should remain separate until each target is validated directly.

    Interpretation
    Parameter transfer, event-shape behaviour, continuous hindcast response and upper-tail design quantiles answer different scientific questions.
    Boundary
    This reporting distinction does not itself demonstrate a continuous model, a design quantile or superiority to an established method.
    Evidence
    • Hydrometric ReFH parameter re-fit study (2026).Evidence ID: bcb1200f
    • Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID: a7f921fb
    • Hydrometric daily two-store Arm A adjudication (2026).Evidence ID: b153a973
    • Hydrometric flood-weighted objective adjudication (2026).Evidence ID: 1510b52c
    • Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID: 1a8d59c9

Primary sources

  1. Hydrometric ReFH parameter re-fit study (2026).Evidence ID: bcb1200f
  2. Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID: a7f921fb
  3. Hydrometric daily two-store Arm A adjudication (2026).Evidence ID: b153a973
  4. Hydrometric flood-weighted objective adjudication (2026).Evidence ID: 1510b52c
  5. Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID: 1a8d59c9

Version history

  1. Version 1.0.0current

    Initial publication of bounded ReFH parameter, baseline-gap, failed daily-branch and gated regional-continuous evidence.