Skip to main content

Growth curves

Skew, Scale and the Shape of UK Flood Growth

Version
1.0.0
Status
Current
Published
Evidence cut-off

In short

We asked what a flood growth curve is actually made of, and whether two promising shortcuts to it hold up. Neither did: a strong-looking result from gauge rating data turned out to need the gauge it was predicting, and a simulation model that produced too little year-to-year variation was not saved by the shortness of real records.

Findings at a glance

Gauged rivers examined
874Each with at least ten usable years of annual floods874 gauged stations with at least 10 non-rejected AMAX years and per-peak stage
Best score 40 years can support for spread
0.857The same ceiling for lopsidedness is only 0.581874-station within-station year-resampling experiment
Wobble from record length alone
±13.0%How much reshuffling the years moves the 100-year growth factor — not an allowance to add to a design flow874-station within-station year-resampling experiment
Basins the simulation could not reach
21 of 27Real rivers varied more than the model’s 95th percentile27 basins in the frozen finite-sample experiment

What we tested

A flood growth curve turns a typical flood into a rare one. It has four moving parts and they are easy to confuse. The first is size: how big the typical yearly flood is at this site. The second is spread: how much the yearly floods vary from one another. The third is lopsidedness: whether the record has a few very large years pulling away from the rest. The fourth is tail growth: how much bigger a 1 in 100 year flood is than the typical one, once size, spread and lopsidedness have been fitted.

Hydrologists call these the index flood, L-CV, L-skew and GF100. Getting one of them plausible does not make the others right, and most arguments about growth curves are really arguments about which of the four is wrong.

We ran two experiments. The first took 874 gauged rivers and asked how much roughly forty years of records can pin down at all: we reshuffled the years inside each record and watched how far the answers moved on their own. Inside that experiment we also tested a promising shortcut — using information from how the gauge converts water level to flow to predict the spread.

The second experiment took a rainfall-runoff simulation model that was producing too little year-to-year variation, and asked whether that was really a modelling fault or just an artefact of comparing it with short real records. We ran it for 10,000 simulated years across 27 river basins, then chopped the output into 40-year windows of the same length as the real records and compared like with like.

What we saw

Forty years of floods buys more certainty about spread than about shape

874 gauged rivers, each with at least ten usable years. The bar is the best score any method could achieve given how much the number moves when the years are reshuffled.

0.00.20.40.60.81.0Best score a 40-year record could support (R², where 1.0 is perfect)Spread of the yearly floods (L-CV)Spread of the yearly floods (L-CV) — Best achievable score: 0.8570.857Tail growth at 100 years (GF100)Tail growth at 100 years (GF100) — Best achievable score: 0.7230.723Lopsidedness of the record (L-skew)Lopsidedness of the record (L-skew) — Best achievable score: 0.5810.581
No method can score above these lines on a 40-year record — and the ceiling for lopsidedness is far lower than for spread.

A shortcut that only worked where it was not needed

Predicting the spread of yearly floods. On the left, measured at the gauge itself; on the right, transferred to gauges held back from fitting.

  • Rating-derived inputs
  • The standard comparison
0.00.20.40.60.8Share of the variation in spread explained (R²)Measured at the gauge itself (874 rivers)Measured at the gauge itself (874 rivers) — Rating-derived inputs: 0.7100.710Measured at the gauge itself (874 rivers) — The standard comparison: 0.4200.420Transferred to held-back gauges (40 rivers)Transferred to held-back gauges (40 rivers) — Rating-derived inputs: 0.2940.294Transferred to held-back gauges (40 rivers) — The standard comparison: 0.2480.248
Rating-derived inputs looked far better than catchment descriptors at the gauge — but they need that gauge's own record, and transferred to held-back gauges the advantage almost vanished.

Short records were not the excuse

27 river basins. The simulation was run for 10,000 years, then cut into windows the same length as each basin's real record.

0.100.150.20Spread of the yearly floods (L-CV)Model, over 10,000 simulated yearsModel, over 10,000 simulated years — Long-run value: 0.1310.131Model, sampled at real record lengthsModel, sampled at real record lengths — Middle window: 0.1290.129Model, sampled at real record lengths — 95th percentile: 0.1590.159The rivers themselvesThe rivers themselves — Observed middle: 0.2050.205
Real rivers vary more than the model does, and sampling the model at real record lengths never reaches them — in 21 of 27 basins the observed value sat above the model's 95th percentile.

Outcome

  • Forty years pins down spread, not shape. The best achievable score was 0.857 for spread, 0.723 for tail growth and 0.581 for lopsidedness. Reshuffling years inside a record moved the 100-year growth factor by about ±13.0% on its own — record-length wobble, not an allowance to add to a design flow.
  • The rating shortcut was circular. It scored 0.710 against 0.420 for catchment descriptors — but both the prediction and the thing being predicted came from the same gauge's record. Transferred to gauges held back from fitting it scored 0.294 against 0.248 for the standard approach. We withdrew the reading that it works without a gauge on site.
  • Short records did not explain the model's missing variation. Real rivers showed 0.205 against 0.131 for the simulation. Sampling the model over real 40-year windows closed only 38% of that gap, and in 21 of 27 basins the real value still sat above the model's 95th percentile.
  • Nor was it a fluke of short samples pushing values up. The bias from record length was −0.002 — effectively nothing, and in the wrong direction to help.
  • Matching lopsidedness cannot supply missing spread. A model with too little spread produces a flat growth curve even when the lopsidedness looks right. The two are not interchangeable.

What we decided next

Decided 14 July 2026. Rating-derived inputs are no longer pursued as a route for sites without a gauge, and the earlier claim that they were has been withdrawn in public. They remain a bounded research interest at gauged sites, where the gauge record is available anyway.

Decided 14 July 2026. We stopped tuning the c6_t3zero simulation configuration as a route to flood growth curves. The one defence available to it — that real records are too short to judge it fairly — has been tested and closed. Effort moved to model classes that carry state between years rather than treating each basin on its own.

Hydrometric now reports flood size, spread, lopsidedness and tail growth as four separate quantities rather than one growth-curve score, so a reviewer can see which of the four a disagreement comes from.

Left open: we have not yet run the same evidence through the alternative distribution families side by side, so we cannot say how much of a rare-flood estimate rests on that choice alone.

Read the detailExact tables, method names and the full claim record

Research question

Which parts of a flood growth curve are scale, dispersion, asymmetry and upper-tail growth; how much does a finite AMAX record constrain each observed target; and can finite-record sampling explain the deficient L-CV of one frozen continuous-model configuration? A parallel hypothesis asked whether rating-derived information could support transfer beyond gauged stations.

Definitions: scale, dispersion, shape and growth

The index-flood scale is the reference flow magnitude that multiplies a dimensionless growth curve. It answers “how large is the characteristic flood?” before the curve answers how that flow grows with return period. A scale error moves flows in magnitude even when the dimensionless curve is unchanged.

L-CV describes dispersion: the relative spread in the annual-maximum series. L-skew describes asymmetry: how strongly the distribution departs from a symmetric shape. GF100 is the fitted 100-year growth factor relative to the index flood. It is an output of the adopted evidence and family, not a synonym for L-skew. These roles are related, but they are not interchangeable.

Cohort and method

The growth revalidation selected 874 gauged stations. Each station had at least 10 non-rejected AMAX years and per-peak stage available for the rating analysis. The record-length experiment estimated within-station sampling variance by resampling years within each station, then recomputing L-moments and GF100. It did not substitute differences between stations for finite-record variability within a station.

The separate finite-sample experiment fixed CK1g configuration c6_t3zero across 27 basins. It used deterministic replay, 10,000 synthetic years and non-overlapping record-length windows matched to each observed series. The full synthetic record tests the model’s long-run L-CV; the matched windows test the distribution that would be seen at the observed record lengths.

Target ceilings from finite records

The estimated target ceilings were 0.857 for L-CV, 0.581 for L-skew and 0.723 for GF100 in the approximately 40-year experiment. These are estimates of how reproducible each observed target is under the frozen resampling design. They are not achieved model scores and the source does not report an uncertainty band for the ceilings.

Median within-station GF100 sampling variability was about ±13.0%. This is not universal design-flow uncertainty: it does not include index-flood error, rating error, family choice, model structure or local data-quality decisions. Finite-record sampling constrains L-CV, L-skew and GF100 differently, so the three statistics should remain separate rather than being collapsed into a monotonic rule.

Rating hypothesis and retraction

The original gauged association was strong: rating-derived features explained L-CV with R² 0.710 versus 0.420 for catchment descriptors. That comparison was self-referential for transfer because the rating evidence came from the target station’s own record—the same gauged setting in which the L-CV target was observed.

Spatial-cluster cross-validation is appropriate for neighbour and geographic transfer, but it cannot remove within-station self-reference when an evaluation station’s rating-derived feature and target L-moments come from the same observed record. The frozen held-out, train-only transfer set is therefore the relevant capability test, not the small change under spatial clustering: rating L-CV was 0.294 versus 0.248 for FEH pooling (n=40). The rating features were still extracted at each held-out target. They cannot run at an ungauged site. Any earlier reading of this work as a capability without a gauged target record is therefore withdrawn; the remaining result is a bounded gauged association.

Frozen finite-sample experiment

Observed median L-CV was 0.205, compared with 0.131 in the frozen 10,000-year model series. The model’s matched record-length windows had median 0.129 and P95 0.159. Observed L-CV exceeded that P95 in 21 of 27 basins, so the discrepancy was not confined to a small number of sites.

The finite-sample spread closed only 38% of the aggregate gap and the median record-length bias was −0.002. Short records did not inflate L-CV through upward bias. The experiment therefore closes one defence for c6_t3zero: sampling records at the observed lengths did not rescue its missing dispersion.

Implications for growth curves

A growth-curve review should begin by separating the index-flood magnitude from dimensionless growth. Within the growth evidence, low L-CV can flatten a growth curve even when L-skew is adequate. Matching asymmetry does not manufacture missing dispersion, and a plausible L-CV does not by itself validate upper-tail growth.

GF100 then records the combined consequence of the fitted L-moments and adopted distribution family at one return period. It is useful as an auditable tail summary, but it does not identify which part of a discrepancy came from scale, dispersion, asymmetry, family choice or finite-record target noise. Those evidence layers should be reported beside it.

Proposed family protocol

The proposed family protocol would compare GLO, KAP3 and GEV using the same frozen evidence record. It would report fitting validity, growth factors at common return periods and the sensitivity of the upper tail to family choice, while keeping index-flood scale and evidence custody visible.

That protocol is not yet demonstrated. No completed GLO/KAP3/GEV comparison is reported in these source artefacts, so their range is not a demonstrated uncertainty envelope and must not be presented as achieved coverage or a probability interval.

Limitations

  • The growth cohort required both usable AMAX and per-peak stage, so its selection is narrower than all gauged AMAX stations.
  • The target ceilings belong to one within-station resampling design and approximately 40-year experiment; no uncertainty band for those ceilings was reported.
  • Rating-derived inputs require the target gauge record and remain self-referential for any use that assumes no target observations.
  • The finite-sample verdict covers L-CV in one frozen configuration and 27 basins. It does not test every growth statistic, architecture or physical model.
  • The proposed candidate-family comparison has not been executed as a demonstrated experiment.

What we do not conclude

  • We do not turn ±13.0% GF100 variability or the target ceilings into a universal allowance, a site-specific interval or a design-flow uncertainty band.
  • We do not infer a universal ordering in which one of L-CV, L-skew or GF100 must always be the noisiest finite-record target.
  • We do not attribute positive L-CV bias to finite record length; the frozen experiment measured a median bias of −0.002.
  • We do not claim that target-record rating features provide a route for sites without gauged observations.
  • We do not generalise one frozen model failure to all continuous or physical growth-curve approaches.
  • We do not present the proposed GLO/KAP3/GEV protocol as a completed family range or uncertainty envelope.

Disclosure boundary

This report discloses the research questions, broad cohort rules, within-station resampling design, frozen replay design, exact aggregate tables, failed hypotheses, claim statuses, source revisions and the scope of each conclusion. Its only evidence sources are the frozen Hydrometric growth revalidation and finite-sample experiment.

It withholds raw licensed AMAX, stage and flow records, exact feature engineering and selection, learned coefficients, per-station private evidence, model-engine internals, calibration recipes and code-level workflow logic. The public result can be audited without exposing the data or implementation needed to reconstruct the proprietary method.

Evidence distributions

Estimated target ceilings for an approximately 40-year AMAX record

874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage; ceilings represent the approximately 40-year record-length experiment.

TargetEstimated R² ceiling
L-CV0.857
L-skew0.581
GF1000.723

These are statistic-specific estimates of target recoverability under within-station year resampling. Finite-record sampling constrains L-CV, L-skew and GF100 differently. Median within-station GF100 sampling variability was about ±13.0%; this is not universal design-flow uncertainty. The experiment does not report an uncertainty band for the ceilings.

Gauged rating association and frozen transfer audit for L-CV

874 gauged stations in the original association; 40 gauged stations in the frozen held-out, train-only transfer.

Evidence stageRating inputRating L-CV R²ComparatorComparator L-CV R²Cohort and scope
Original gauged associationRating-derived features0.710Catchment descriptors0.420874 gauged stations
Frozen held-out transferRating-derived features0.294FEH pooling0.24840 gauged stations

The original 0.710 versus 0.420 result was a gauged association. Under frozen held-out, train-only transfer, the rating result was 0.294 versus 0.248 for FEH pooling across 40 gauged stations. Rating-derived features still came from each target station’s own record, so neither comparison establishes operation without a gauged target record.

Observed and frozen-model L-CV

27 basins under deterministic replay of frozen CK1g configuration c6_t3zero, using 10,000 synthetic years and non-overlapping windows matched to each observed record length.

MeasureL-CV
Observed median0.205
10,000-year model median0.131
Model record-length median0.129
Model record-length P950.159
Model record-length median bias−0.002

Observed L-CV exceeded the model record-length P95 in 21 of 27 basins. The finite-sample spread closed only 38% of the aggregate gap, while median record-length bias was −0.002. Short records did not inflate L-CV through upward bias, so sample size did not rescue this frozen model configuration.

What survived

Scale, dispersion and shape remain separate evidence

Index-flood magnitude, L-CV, L-skew and GF100 answer different questions and must remain separately visible in a growth-curve assessment.

Statistic-specific record-length testing

Within-station year resampling exposed different finite-record constraints for L-CV, L-skew and GF100 without converting them into one universal allowance.

What failed

Rating evidence as a route without a gauged target record

The rating inputs were extracted from the target station’s own measurements, so the apparent predictive signal remained gauged and self-referential.

Finite record length as a rescue for the frozen model

Matched record-length windows closed only part of the L-CV gap and did not create the missing dispersion through upward bias.

Claim record

Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.

  1. Inference

    Index-flood scale, L-CV dispersion, L-skew asymmetry and GF100 upper-tail growth are distinct parts of a growth-curve evidence record.

    Interpretation
    A plausible value for one part cannot substitute for evidence about the others.
    Boundary
    This is an interpretive framework for reading the two frozen experiments, not a new fitted family or universal causal decomposition.
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
    • Hydrometric finite-sample experiment (2026).Evidence ID: 01ced6e2
  2. Demonstrated

    For the approximately 40-year record experiment, estimated R² ceilings were 0.857 for L-CV, 0.581 for L-skew and 0.723 for GF100; median within-station GF100 sampling variability was about ±13.0%.

    Interpretation
    Finite-record sampling constrains the three targets differently, so each target needs its own sampling evidence.
    Boundary
    The ceiling estimates have no reported uncertainty band, and ±13.0% is median within-record GF100 sampling variability rather than universal design-flow uncertainty.
    Linked result
    Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage; ceilings represent the approximately 40-year record-length experiment.
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  3. Demonstrated

    The original gauged L-CV association was R² 0.710 for rating-derived features versus 0.420 for catchment descriptors; frozen held-out, train-only transfer across 40 gauged stations returned 0.294 for rating-derived features versus 0.248 for FEH pooling.

    Interpretation
    The frozen transfer result is a modest gauged comparison, not evidence that the rating inputs are available without the target record.
    Boundary
    Rating-derived inputs were extracted from the target station’s own record and cannot run at an ungauged site.
    Linked result
    Gauged rating association and frozen transfer audit for L-CV; sample: 874 gauged stations in the original association; 40 gauged stations in the frozen held-out, train-only transfer.
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  4. Retracted

    The earlier interpretation that rating-derived evidence supported transfer to sites without a gauged target record was withdrawn.

    Interpretation
    Both the original association and the frozen transfer used information extracted from the target station’s own measurements.
    Boundary
    The withdrawal applies to use without a gauged target record; it does not invalidate bounded rating-informed research at gauged stations.
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  5. Demonstrated

    Across 27 basins, observed median L-CV was 0.205 versus 0.131 for the frozen 10,000-year model series; matched record-length windows had median 0.129 and P95 0.159, 21 of 27 observed values exceeded that P95, aggregate spread closure was 38%, and median bias was −0.002.

    Interpretation
    Record-length sampling did not rescue the L-CV deficit in frozen configuration c6_t3zero and did not act through systematic upward bias.
    Boundary
    This is a failure of one frozen continuous-model configuration on one L-CV experiment, not a verdict on every continuous, physical or growth-curve model.
    Linked result
    Observed and frozen-model L-CV; sample: 27 basins under deterministic replay of frozen CK1g configuration c6_t3zero, using 10,000 synthetic years and non-overlapping windows matched to each observed record length.
    Evidence
    • Hydrometric finite-sample experiment (2026).Evidence ID: 01ced6e2
  6. Inference

    Low L-CV can flatten a growth curve even when L-skew is adequate.

    Interpretation
    Matching asymmetry alone does not supply missing dispersion, while GF100 also depends on the adopted family and fitted evidence.
    Boundary
    The exact effect is family- and parameter-dependent; this is not a universal numerical mapping from an L-moment to GF100.
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
    • Hydrometric finite-sample experiment (2026).Evidence ID: 01ced6e2
  7. Not yet demonstrated

    A same-evidence GLO, KAP3 and GEV comparison is a proposed family protocol.

    Interpretation
    It would show how candidate-family choice changes fitted growth while holding the evidence record fixed.
    Boundary
    The comparison has not been run as a completed experiment and is not a demonstrated uncertainty envelope.
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898

Primary sources

  1. Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  2. Hydrometric finite-sample experiment (2026).Evidence ID: 01ced6e2

Version history

  1. Version 1.0.0current

    Initial publication separating scale, L-moment shape, growth, gauged rating evidence and finite-sample model failure.