Growth curves
Skew, Scale and the Shape of UK Flood Growth
- Version
- 1.0.0
- Status
- Current
- Published
- Evidence cut-off
Abstract
Flood growth is not one parameter. Index-flood scale sets the reference magnitude; L-CV describes dispersion, L-skew describes asymmetry, and GF100 records the fitted upper-tail growth relative to that scale. Two frozen experiments show both what finite records constrain and what they do not excuse: a promising rating association remained gauged and self-referential, while record-length sampling did not rescue one continuous-model configuration with deficient L-CV.
Findings at a glance
- Growth revalidation cohort
- 874Gauged stations passing the frozen AMAX and stage selection rules874 gauged stations with at least 10 non-rejected AMAX years and per-peak stage
- L-CV target ceiling
- 0.857Estimated R² ceiling for an approximately 40-year AMAX record874-station within-station year-resampling experiment
- GF100 sampling variability
- ±13.0%Median within-station finite-record variability, not a design-flow allowance874-station within-station year-resampling experiment
- Above model record-length P95
- 21 of 27Basins whose observed L-CV exceeded the frozen model threshold27 basins in the frozen finite-sample experiment
Evidence distributions
Estimated target ceilings for an approximately 40-year AMAX record
874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage; ceilings represent the approximately 40-year record-length experiment.
| Target | Estimated R² ceiling |
|---|---|
| L-CV | 0.857 |
| L-skew | 0.581 |
| GF100 | 0.723 |
These are statistic-specific estimates of target recoverability under within-station year resampling. Finite-record sampling constrains L-CV, L-skew and GF100 differently. Median within-station GF100 sampling variability was about ±13.0%; this is not universal design-flow uncertainty. The experiment does not report an uncertainty band for the ceilings.
Gauged rating association and frozen transfer audit for L-CV
874 gauged stations in the original association; 40 gauged stations in the frozen held-out, train-only transfer.
| Evidence stage | Rating input | Rating L-CV R² | Comparator | Comparator L-CV R² | Cohort and scope |
|---|---|---|---|---|---|
| Original gauged association | Rating-derived features | 0.710 | Catchment descriptors | 0.420 | 874 gauged stations |
| Frozen held-out transfer | Rating-derived features | 0.294 | FEH pooling | 0.248 | 40 gauged stations |
The original 0.710 versus 0.420 result was a gauged association. Under frozen held-out, train-only transfer, the rating result was 0.294 versus 0.248 for FEH pooling across 40 gauged stations. Rating-derived features still came from each target station’s own record, so neither comparison establishes operation without a gauged target record.
Observed and frozen-model L-CV
27 basins under deterministic replay of frozen CK1g configuration c6_t3zero, using 10,000 synthetic years and non-overlapping windows matched to each observed record length.
| Measure | L-CV |
|---|---|
| Observed median | 0.205 |
| 10,000-year model median | 0.131 |
| Model record-length median | 0.129 |
| Model record-length P95 | 0.159 |
| Model record-length median bias | −0.002 |
Observed L-CV exceeded the model record-length P95 in 21 of 27 basins. The finite-sample spread closed only 38% of the aggregate gap, while median record-length bias was −0.002. Short records did not inflate L-CV through upward bias, so sample size did not rescue this frozen model configuration.
What survived
Scale, dispersion and shape remain separate evidence
Index-flood magnitude, L-CV, L-skew and GF100 answer different questions and must remain separately visible in a growth-curve assessment.
Statistic-specific record-length testing
Within-station year resampling exposed different finite-record constraints for L-CV, L-skew and GF100 without converting them into one universal allowance.
What failed
Rating evidence as a route without a gauged target record
The rating inputs were extracted from the target station’s own measurements, so the apparent predictive signal remained gauged and self-referential.
Finite record length as a rescue for the frozen model
Matched record-length windows closed only part of the L-CV gap and did not create the missing dispersion through upward bias.
Research question
Which parts of a flood growth curve are scale, dispersion, asymmetry and upper-tail growth; how much does a finite AMAX record constrain each observed target; and can finite-record sampling explain the deficient L-CV of one frozen continuous-model configuration? A parallel hypothesis asked whether rating-derived information could support transfer beyond gauged stations.
Definitions: scale, dispersion, shape and growth
The index-flood scale is the reference flow magnitude that multiplies a dimensionless growth curve. It answers “how large is the characteristic flood?” before the curve answers how that flow grows with return period. A scale error moves flows in magnitude even when the dimensionless curve is unchanged.
L-CV describes dispersion: the relative spread in the annual-maximum series. L-skew describes asymmetry: how strongly the distribution departs from a symmetric shape. GF100 is the fitted 100-year growth factor relative to the index flood. It is an output of the adopted evidence and family, not a synonym for L-skew. These roles are related, but they are not interchangeable.
Cohort and method
The growth revalidation selected 874 gauged stations. Each station had at least 10 non-rejected AMAX years and per-peak stage available for the rating analysis. The record-length experiment estimated within-station sampling variance by resampling years within each station, then recomputing L-moments and GF100. It did not substitute differences between stations for finite-record variability within a station.
The separate finite-sample experiment fixed CK1g configurationc6_t3zero across 27 basins. It used deterministic replay, 10,000 synthetic years and non-overlapping record-length windows matched to each observed series. The full synthetic record tests the model’s long-run L-CV; the matched windows test the distribution that would be seen at the observed record lengths.
Target ceilings from finite records
The estimated target ceilings were 0.857 for L-CV, 0.581 for L-skew and 0.723 for GF100 in the approximately 40-year experiment. These are estimates of how reproducible each observed target is under the frozen resampling design. They are not achieved model scores and the source does not report an uncertainty band for the ceilings.
Median within-station GF100 sampling variability was about ±13.0%. This is not universal design-flow uncertainty: it does not include index-flood error, rating error, family choice, model structure or local data-quality decisions. Finite-record sampling constrains L-CV, L-skew and GF100 differently, so the three statistics should remain separate rather than being collapsed into a monotonic rule.
Rating hypothesis and retraction
The original gauged association was strong: rating-derived features explained L-CV with R² 0.710 versus 0.420 for catchment descriptors. That comparison was self-referential for transfer because the rating evidence came from the target station’s own record—the same gauged setting in which the L-CV target was observed.
Spatial-cluster cross-validation is appropriate for neighbour and geographic transfer, but it cannot remove within-station self-reference when an evaluation station’s rating-derived feature and target L-moments come from the same observed record. The frozen held-out, train-only transfer set is therefore the relevant capability test, not the small change under spatial clustering: rating L-CV was 0.294 versus 0.248 for FEH pooling (n=40). The rating features were still extracted at each held-out target. They cannot run at an ungauged site. Any earlier reading of this work as a capability without a gauged target record is therefore withdrawn; the remaining result is a bounded gauged association.
Frozen finite-sample experiment
Observed median L-CV was 0.205, compared with 0.131 in the frozen 10,000-year model series. The model’s matched record-length windows had median 0.129 and P95 0.159. Observed L-CV exceeded that P95 in 21 of 27 basins, so the discrepancy was not confined to a small number of sites.
The finite-sample spread closed only 38% of the aggregate gap and the median record-length bias was −0.002. Short records did not inflate L-CV through upward bias. The experiment therefore closes one defence for c6_t3zero: sampling records at the observed lengths did not rescue its missing dispersion.
Implications for growth curves
A growth-curve review should begin by separating the index-flood magnitude from dimensionless growth. Within the growth evidence, low L-CV can flatten a growth curve even when L-skew is adequate. Matching asymmetry does not manufacture missing dispersion, and a plausible L-CV does not by itself validate upper-tail growth.
GF100 then records the combined consequence of the fitted L-moments and adopted distribution family at one return period. It is useful as an auditable tail summary, but it does not identify which part of a discrepancy came from scale, dispersion, asymmetry, family choice or finite-record target noise. Those evidence layers should be reported beside it.
Proposed family protocol
The proposed family protocol would compare GLO, KAP3 and GEV using the same frozen evidence record. It would report fitting validity, growth factors at common return periods and the sensitivity of the upper tail to family choice, while keeping index-flood scale and evidence custody visible.
That protocol is not yet demonstrated. No completed GLO/KAP3/GEV comparison is reported in these source artefacts, so their range is not a demonstrated uncertainty envelope and must not be presented as achieved coverage or a probability interval.
Limitations
- The growth cohort required both usable AMAX and per-peak stage, so its selection is narrower than all gauged AMAX stations.
- The target ceilings belong to one within-station resampling design and approximately 40-year experiment; no uncertainty band for those ceilings was reported.
- Rating-derived inputs require the target gauge record and remain self-referential for any use that assumes no target observations.
- The finite-sample verdict covers L-CV in one frozen configuration and 27 basins. It does not test every growth statistic, architecture or physical model.
- The proposed candidate-family comparison has not been executed as a demonstrated experiment.
What we do not conclude
- We do not turn ±13.0% GF100 variability or the target ceilings into a universal allowance, a site-specific interval or a design-flow uncertainty band.
- We do not infer a universal ordering in which one of L-CV, L-skew or GF100 must always be the noisiest finite-record target.
- We do not attribute positive L-CV bias to finite record length; the frozen experiment measured a median bias of −0.002.
- We do not claim that target-record rating features provide a route for sites without gauged observations.
- We do not generalise one frozen model failure to all continuous or physical growth-curve approaches.
- We do not present the proposed GLO/KAP3/GEV protocol as a completed family range or uncertainty envelope.
Disclosure boundary
This report discloses the research questions, broad cohort rules, within-station resampling design, frozen replay design, exact aggregate tables, failed hypotheses, claim statuses, source revisions and the scope of each conclusion. Its only evidence sources are the frozen Hydrometric growth revalidation and finite-sample experiment.
It withholds raw licensed AMAX, stage and flow records, exact feature engineering and selection, learned coefficients, per-station private evidence, model-engine internals, calibration recipes and code-level workflow logic. The public result can be audited without exposing the data or implementation needed to reconstruct the proprietary method.
Claim record
Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.
Inference
Index-flood scale, L-CV dispersion, L-skew asymmetry and GF100 upper-tail growth are distinct parts of a growth-curve evidence record.
- Interpretation
- A plausible value for one part cannot substitute for evidence about the others.
- Boundary
- This is an interpretive framework for reading the two frozen experiments, not a new fitted family or universal causal decomposition.
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898 - Hydrometric finite-sample experiment (2026).Evidence ID:
01ced6e2
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Demonstrated
For the approximately 40-year record experiment, estimated R² ceilings were 0.857 for L-CV, 0.581 for L-skew and 0.723 for GF100; median within-station GF100 sampling variability was about ±13.0%.
- Interpretation
- Finite-record sampling constrains the three targets differently, so each target needs its own sampling evidence.
- Boundary
- The ceiling estimates have no reported uncertainty band, and ±13.0% is median within-record GF100 sampling variability rather than universal design-flow uncertainty.
- Linked result
- Estimated target ceilings for an approximately 40-year AMAX record; sample: 874 gauged stations, each with at least 10 non-rejected AMAX years and per-peak stage; ceilings represent the approximately 40-year record-length experiment.
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Demonstrated
The original gauged L-CV association was R² 0.710 for rating-derived features versus 0.420 for catchment descriptors; frozen held-out, train-only transfer across 40 gauged stations returned 0.294 for rating-derived features versus 0.248 for FEH pooling.
- Interpretation
- The frozen transfer result is a modest gauged comparison, not evidence that the rating inputs are available without the target record.
- Boundary
- Rating-derived inputs were extracted from the target station’s own record and cannot run at an ungauged site.
- Linked result
- Gauged rating association and frozen transfer audit for L-CV; sample: 874 gauged stations in the original association; 40 gauged stations in the frozen held-out, train-only transfer.
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Retracted
The earlier interpretation that rating-derived evidence supported transfer to sites without a gauged target record was withdrawn.
- Interpretation
- Both the original association and the frozen transfer used information extracted from the target station’s own measurements.
- Boundary
- The withdrawal applies to use without a gauged target record; it does not invalidate bounded rating-informed research at gauged stations.
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Demonstrated
Across 27 basins, observed median L-CV was 0.205 versus 0.131 for the frozen 10,000-year model series; matched record-length windows had median 0.129 and P95 0.159, 21 of 27 observed values exceeded that P95, aggregate spread closure was 38%, and median bias was −0.002.
- Interpretation
- Record-length sampling did not rescue the L-CV deficit in frozen configuration c6_t3zero and did not act through systematic upward bias.
- Boundary
- This is a failure of one frozen continuous-model configuration on one L-CV experiment, not a verdict on every continuous, physical or growth-curve model.
- Linked result
- Observed and frozen-model L-CV; sample: 27 basins under deterministic replay of frozen CK1g configuration c6_t3zero, using 10,000 synthetic years and non-overlapping windows matched to each observed record length.
- Evidence
- Hydrometric finite-sample experiment (2026).Evidence ID:
01ced6e2
- Hydrometric finite-sample experiment (2026).Evidence ID:
Inference
Low L-CV can flatten a growth curve even when L-skew is adequate.
- Interpretation
- Matching asymmetry alone does not supply missing dispersion, while GF100 also depends on the adopted family and fitted evidence.
- Boundary
- The exact effect is family- and parameter-dependent; this is not a universal numerical mapping from an L-moment to GF100.
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898 - Hydrometric finite-sample experiment (2026).Evidence ID:
01ced6e2
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Not yet demonstrated
A same-evidence GLO, KAP3 and GEV comparison is a proposed family protocol.
- Interpretation
- It would show how candidate-family choice changes fitted growth while holding the evidence record fixed.
- Boundary
- The comparison has not been run as a completed experiment and is not a demonstrated uncertainty envelope.
- Evidence
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898
- Hydrometric growth-curve revalidation (2026).Evidence ID:
Primary sources
- Hydrometric growth-curve revalidation (2026).Evidence ID:
76b42898 - Hydrometric finite-sample experiment (2026).Evidence ID:
01ced6e2
Version history
- Version 1.0.0current
Initial publication separating scale, L-moment shape, growth, gauged rating evidence and finite-sample model failure.
Related reports
- National hydrologyWhat National-Scale Analysis Reveals About UK HydrologyA network-scale examination of gauge consistency, QMED validation, record-length limits and temporal sensitivity, with each result kept inside its evidential boundary.
- QMEDCan Machine Learning Improve QMED?A fair national comparison showing why an ML accuracy claim failed under geographic validation, while open-input access, repeatability and auditability survived.
- Statistical methodsThe 2025 Statistical Method: What Changed and Why It MattersA practitioner account of the FEH 2025 descriptor, QMED, donor, pooling, urban, distribution and uncertainty changes, with Hydrometric’s implementation evidence kept inside its local verification boundary.
- UncertaintyHow Certain Is a Design Flood?A layered account of population validation scatter, network inconsistency, within-record sampling and the still-open model-family protocol.
- Rainfall-runoff modellingFrom ReFH Design Events to Continuous HydrologyA bounded record of the baseflow-lag refit, the public ReFH1 baseline gap, the failed daily two-store branch and the gated status of regional continuous evidence.
- Flood estimation practiceThe Modern Flood Estimation ReportA practical evidence-record model for study purpose, data review, method selection, uncertainty, provenance and approval, with an honest audit of the current Hydrometric generator.