Skip to main content

Rainfall-runoff modelling

From ReFH Design Events to Continuous Hydrology

Version
1.0.0
Status
Current
Published
Evidence cut-off

Abstract

A bounded account of what survived in ReFH parameter research, how far a public ReFH1 baseline sat from genuine ReFH2 pilot outputs, why one daily two-store branch was closed, and why regional continuous hydrology remains outside the public demonstrated ledger until its immutable score artefacts are available.

Findings at a glance

Baseflow-lag held-area R²
0.528Leave-one-hydrometric-area-out, with interval [0.451, 0.583]640 catchments
Q100 peak baseline gap
47.6%Public ReFH1 stack versus genuine ReFH2 pilot outputs18 complete cases
Mean absolute baseline gap
53.5%Aggregate pilot comparison under commercial custody18 complete cases
Daily combined gate
2 of 9QMED, L-CV and KGE all required in the tested daily two-store branchNine stations

Evidence distributions

Baseflow-lag refit and published reference

640 catchments in the Hydrometric refit. The WHS figure is a published in-sample reference, not a second result from the same cohort or validation folds.

Evidence rowValidation designIntervalCohort and custody
Hydrometric baseflow-lag refitLeave-one-hydrometric-area-out0.528[0.451, 0.583]640 catchments; flow-derived and bias-mapped targets
Published WHS referenceIn-sample0.35Not reported in the cited comparisonPublished model-development statistic; not the same validation design
Difference in R²Hydrometric held-area result minus the published in-sample reference0.178[0.101, 0.233]Arithmetic comparison across differently validated results

The refit reached leave-one-hydrometric-area-out R² 0.528 [0.451, 0.583], compared with published WHS in-sample R² 0.35; the difference was 0.178 [0.101, 0.233]. The validation designs differ, and the refit targets are flow-derived and bias-mapped.

Public ReFH1 stack against genuine ReFH2 pilot outputs

18 complete pilot cases. Only aggregate statistics are published; case rows and licensed outputs remain withheld.

Comparison statisticResultCasesInterpretation
Q100 peak gap47.6%18 complete casesAggregate gap between the public ReFH1 coefficient/equation stack and genuine ReFH2 pilot outputs
Mean absolute gap53.5%18 complete casesAggregate absolute difference across the complete pilot cases

This is a baseline gap, not a ReFH2 beat. It measures how far the public ReFH1 stack sat from genuine ReFH2 pilot outputs under controlled commercial custody; it does not expose or reconstruct those outputs.

Daily two-store branch adjudication

Nine stations for the frozen combined gate; eight complete stations for the bounded relaxation.

Gate or diagnosticResultCohortBoundary
Frozen combined QMED/L-CV/KGE gateTwo of nine stationsNine-station daily Arm AThe pre-declared combined gate was the decision rule
Bounded relaxationAt most four of eightEight complete stations in the relaxed comparisonA sensitivity bound, not a replacement success criterion
Between-basin model L-CV0.18–0.21Tested daily branchNarrow dispersion persisted across the objective variants

The daily lumped model passed the combined gate at two of nine stations. Under the bounded relaxation it reached at most four of eight, while model L-CV remained 0.18–0.21. Only that per-basin daily lumped branch is closed.

What survived

Baseflow-lag regional refit

Held-area validation retained predictive signal across 640 catchments, with the target-construction and comparator differences kept visible.

Commercial-custody baseline check

The pilot established the size of the public ReFH1 baseline gap without publishing licensed case outputs or treating that gap as superiority.

What failed

Per-basin daily lumped two-store branch

The combined gate passed at only two of nine stations, and bounded relaxation did not resolve the narrow between-basin L-CV.

Research question

Which parts of the ReFH and continuous-hydrology research record are strong enough to publish, which tested branch failed, and which regional-continuous claims remain outside the evidence gate? The report separates parameter transfer, a public-baseline comparison, daily lumped simulation and future hourly continuous evidence because they do not share a target or validation design.

Evidence and method

Three completed experiments enter the numerical ledger. A 640-catchment refit withheld one hydrometric area at a time. A commercially controlled pilot compared aggregate outputs from the public ReFH1 coefficient and equation stack with genuine ReFH2 outputs across 18 complete cases. A daily two-store study applied a combined QMED, L-CV and KGE decision gate to nine stations and retained an eight-station bounded sensitivity check.

Regional continuous Stage A is handled differently. Publication required immutable score artefacts in this implementation worktree and a match to the cited evidence ledger. Those artefacts are absent, so the regional section carries a structured not-yet-demonstrated status and no score distribution.

Baseflow-lag refit

The flow-derived baseflow-lag refit reached leave-one-hydrometric-area-out R² 0.528 [0.451, 0.583] across 640 catchments. The published WHS in-sample R² 0.35 is retained as a reference, giving a difference of 0.178 [0.101, 0.233]. Withholding whole hydrometric areas makes the refit result relevant to regional transfer rather than within-sample fit alone.

The target custody is material: targets are flow-derived and bias-mapped. The refit retained held-area signal, while the published WHS figure is a contextual reference only because it uses a different validation design. This is not a like-for-like contest between identical experiments. A disattenuated statistic is omitted because it is not measured predictive performance.

Commercial evidence custody

Across 18 complete cases, the public ReFH1 stack differed from genuine ReFH2 pilot outputs by 47.6% on Q100 peak; its mean absolute gap was 53.5%. This is a baseline gap, not a ReFH2 beat. It establishes that the open baseline and the licensed reference were materially different in the pilot.

Only aggregate statistics cross the publication boundary. No licensed case output, private row, coefficient set or reconstruction detail is reproduced. The comparison therefore documents baseline adequacy and evidence custody without exposing the commercial reference surface.

Failed daily two-store branch

The tested daily lumped branch cleared its combined QMED/L-CV/KGE gate at two of nine stations. Under the pre-bounded relaxation, performance reached at most four of eight complete stations. Across the objective variants, model L-CV of 0.18–0.21 remained too narrow to establish the required between-basin growth behaviour.

Only that per-basin daily lumped branch is closed. The result does not rule out regional models, persistent hourly state, alternative model classes or continuous simulation generally. Nor are daily held-period maxima presented as instantaneous design-flood evidence.

Stage A publication gate

Regional continuous Stage A is classified as not-yet-demonstrated in this version. The required immutable score artefacts are absent from the implementation worktree, so the evidence gate fails and no Stage A performance metric, percentile or station count is published.

If a future version satisfies the immutable gate, its admissible claim would still be an observed-forcing hourly hindcast, not weather forecasting, Q100 skill or ReFH2 replacement. That future boundary is stated now so later evidence cannot silently broaden the scientific target.

What survived and what failed

The held-area baseflow-lag refit survived as bounded regional parameter evidence. The aggregate commercial-custody comparison survived as a diagnostic of the public ReFH1 starting point. The per-basin daily two-store branch failed its combined gate and remains closed within that exact implementation boundary.

Regional continuous publication remains not-yet-demonstrated because its evidence gate was not met. That is not a negative score for the model class: the necessary immutable proof is missing from the publication workspace and the public claim remains open.

Operational implications

  • Keep parameter-transfer results tied to the target construction and geographic holdout used to obtain them.
  • Use the ReFH pilot only to establish the open baseline gap and the need for controlled licensed validation.
  • Do not spend further effort tuning the closed per-basin daily lumped branch as a standalone growth-curve solution.
  • Require immutable, report-addressable artefacts before publishing a regional continuous score or distribution.
  • Keep hydrograph response, event replay and flood-frequency targets in separate claim ledgers.

Limitations

  • Baseflow-lag targets are flow-derived and bias-mapped rather than direct observations of a unique physical store parameter.
  • The WHS reference is in-sample and therefore does not share the held-area validation design of the refit.
  • The genuine ReFH2 comparison contains 18 complete pilot cases and publishes aggregate gaps only.
  • The daily evidence concerns a small frozen cohort, daily resolution and one per-basin lumped architecture.
  • Stage A contributes no public performance result because its immutable evidence gate was not satisfied in this worktree.

What we do not conclude

  • We do not present a disattenuated association as measured held-area predictive performance.
  • We do not treat the public-stack gap as superiority to genuine ReFH2 or as evidence of product equivalence.
  • We do not generalise the daily two-store failure beyond the tested per-basin daily lumped branch.
  • We do not convert daily maxima or hydrograph response into Q100 evidence.
  • We do not publish regional continuous performance without the immutable artefact and evidence-ledger match.

Disclosure boundary

This report discloses the research questions, broad cohorts, validation designs, aggregate distributions, failed gate, claim statuses, source revisions and scientific boundaries. Its only sources are the frozen baseflow-lag refit, public-stack gap, daily Arm A, flood-objective and regional Stage A evidence records.

It withholds licensed case outputs, private rows, learned coefficients, fitted parameter artefacts, exact feature construction, calibration sequences, runnable recipes and deployable model internals. The public evidence is sufficient to audit each claim without reconstructing the proprietary method or importing unverified Stage A results.

Claim record

Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.

  1. Demonstrated

    On 640 catchments, the flow-derived baseflow-lag refit reached leave-one-hydrometric-area-out R² 0.528 [0.451, 0.583], compared with published WHS in-sample R² 0.35.

    Interpretation
    The observed difference in R² was 0.178 [0.101, 0.233], providing evidence that the regional refit retained signal when whole hydrometric areas were withheld.
    Boundary
    The targets are flow-derived and bias-mapped, and the WHS reference is in-sample; no disattenuated headline is presented as measured predictive performance.
    Linked result
    Baseflow-lag refit and published reference; sample: 640 catchments in the Hydrometric refit. The WHS figure is a published in-sample reference, not a second result from the same cohort or validation folds.
    Evidence
    • Hydrometric ReFH parameter re-fit study (2026).Evidence ID: bcb1200f
  2. Demonstrated

    Across 18 complete cases, the public ReFH1 coefficient/equation stack differed from genuine ReFH2 pilot outputs by 47.6% on Q100 peak, with a 53.5% mean absolute gap.

    Interpretation
    The comparison quantifies the starting distance between an open baseline and licensed pilot outputs under controlled custody.
    Boundary
    This is a baseline gap, not a ReFH2 beat; it does not establish product equivalence, disclose case-level commercial outputs or transfer proprietary method details.
    Linked result
    Public ReFH1 stack against genuine ReFH2 pilot outputs; sample: 18 complete pilot cases. Only aggregate statistics are published; case rows and licensed outputs remain withheld.
    Evidence
    • Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID: a7f921fb
  3. Demonstrated

    The tested daily two-store branch cleared its combined QMED/L-CV/KGE gate at two of nine stations.

    Interpretation
    The branch failed its own multi-objective decision rule across most of the frozen station cohort.
    Boundary
    Only that per-basin daily lumped branch is closed; the result does not generalise to regional continuous simulation or every physical rainfall-runoff model.
    Linked result
    Daily two-store branch adjudication; sample: Nine stations for the frozen combined gate; eight complete stations for the bounded relaxation.
    Evidence
    • Hydrometric daily two-store Arm A adjudication (2026).Evidence ID: b153a973
    • Hydrometric flood-weighted objective adjudication (2026).Evidence ID: 1510b52c
  4. Demonstrated

    Under a bounded relaxation, the daily branch reached at most four of eight complete stations, while model L-CV remained 0.18–0.21.

    Interpretation
    Objective changes moved some station-level outcomes without demonstrating adequate combined behaviour across the cohort.
    Boundary
    The relaxed count is a sensitivity bound, not a post-hoc success criterion; daily maxima are not described as instantaneous design-flood performance.
    Linked result
    Daily two-store branch adjudication; sample: Nine stations for the frozen combined gate; eight complete stations for the bounded relaxation.
    Evidence
    • Hydrometric daily two-store Arm A adjudication (2026).Evidence ID: b153a973
    • Hydrometric flood-weighted objective adjudication (2026).Evidence ID: 1510b52c
  5. Not yet demonstrated

    Regional continuous Stage A remains not-yet-demonstrated in this public report because the immutable score artefacts required by its evidence gate are absent from the implementation worktree.

    Interpretation
    The evidence identifier is retained so a future version can be checked against admissible immutable artefacts without importing unverified performance claims now.
    Boundary
    If the publication gate is later satisfied, an admissible Stage A claim would remain an observed-forcing hourly hindcast, not weather forecasting, Q100 skill or ReFH2 replacement.
    Evidence
    • Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID: 1a8d59c9
  6. Inference

    Hydrograph-response evidence and flood-frequency evidence should remain separate until each target is validated directly.

    Interpretation
    Parameter transfer, event-shape behaviour, continuous hindcast response and upper-tail design quantiles answer different scientific questions.
    Boundary
    This reporting distinction does not itself demonstrate a continuous model, a design quantile or superiority to an established method.
    Evidence
    • Hydrometric ReFH parameter re-fit study (2026).Evidence ID: bcb1200f
    • Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID: a7f921fb
    • Hydrometric daily two-store Arm A adjudication (2026).Evidence ID: b153a973
    • Hydrometric flood-weighted objective adjudication (2026).Evidence ID: 1510b52c
    • Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID: 1a8d59c9

Primary sources

  1. Hydrometric ReFH parameter re-fit study (2026).Evidence ID: bcb1200f
  2. Hydrometric public-stack to commercial-ReFH2 gap study (2026).Evidence ID: a7f921fb
  3. Hydrometric daily two-store Arm A adjudication (2026).Evidence ID: b153a973
  4. Hydrometric flood-weighted objective adjudication (2026).Evidence ID: 1510b52c
  5. Hydrometric regional continuous Stage A evidence ledger (2026).Evidence ID: 1a8d59c9

Version history

  1. Version 1.0.0current

    Initial publication of bounded ReFH parameter, baseline-gap, failed daily-branch and gated regional-continuous evidence.