Skip to main content

Uncertainty

How Certain Is a Design Flood?

Version
1.0.0
Status
Current
Published
Evidence cut-off

Abstract

A design flood does not have one universal error bar. Population validation scatter, network self-consistency and within-record sampling answer different questions; natural variability, data, fitted parameters, method structure and model-family choice add distinct but potentially overlapping concerns. This report keeps the frozen evidence in separate ledgers and shows why it cannot be multiplied into a site interval.

Findings at a glance

SM2025 population FSE
1.538Donor-adjusted QMED under spatial-cluster held-out validation918 held-out stations
SM2025 median absolute error
23.4%Population QMED validation performance, not local coverage918 held-out stations
Same-river median disagreement
12.2%Area-scaled network self-consistency diagnostic85 same-river pairs
GF100 sampling variability
±13.0%Median within-record variability near the studied record length874 gauged stations in the within-station resampling experiment

Evidence distributions

Uncertainty is layered, not one bar

Five conceptual layers. Numerical evidence comes from 918 spatial-cluster held-out stations, 85 same-river pairs and 874 gauged stations in a within-station resampling experiment; natural variability and the proposed family protocol have no standalone numerical estimate here.

LayerEvidenceWhat it meansWhat it does not mean
Natural variabilityNot isolated as a standalone numerical quantity in the cited experiments.Flood-producing weather, antecedent state and catchment response vary between events and through time.It is not measured by relabelling validation residuals, network disagreement or finite-record sampling as natural variability.
Data uncertaintyn=85 same-river pairs: median area-scaled QMED disagreement 12.2%; Q1–Q3 6.6–34.5%; P90 58.8%.A network self-consistency diagnostic showing how paired records disagree after area scaling.It is not an individual-site interval, a lower bound on error or an attribution to rating error alone.
Parameter and finite-record samplingn=874 gauged stations: median within-record GF100 sampling variability about ±13.0% near the approximately 40-year record length studied.Finite AMAX samples make fitted growth statistics vary when years within a station are resampled.It is not a universal design-flow allowance and does not include index-flood, data, structural or family effects.
Structural and transfer performancen=918 spatial-cluster held-out stations: donor-adjusted SM2025 median absolute QMED error 23.4% and FSE 1.538.Population cross-validation performance for the complete frozen transfer method across held-out stations.It does not isolate structure from every other residual source and is population cross-validation scatter, not a site-specific confidence interval.
Model-family choiceA same-evidence GLO/KAP3/GEV comparison is proposed; no completed envelope is reported.A future sensitivity protocol could keep evidence fixed while showing how fitted upper-tail results change by family.It is not achieved coverage, a probability interval or a demonstrated family envelope.

The rows answer different questions. The natural, data, parameter, structural and model-family uncertainty layers are conceptually separate, but mechanisms can overlap and are not assumed independent. The three numerical studies therefore remain labelled evidence layers and are not combined mechanically.

Worked population-factor illustration for a 10.00 m³/s estimate

The factor was estimated from 918 held-out stations. The 10.00 m³/s value is a single illustrative estimate introduced only to show the arithmetic.

DirectionFormulaResultInterpretationn and scope
Divide by one FSE factor10.00 m³/s ÷ 1.5386.50 m³/s (approximately 6.5 m³/s)Lower multiplicative population-scatter reference from the same factor.FSE estimated across n=918 spatial-cluster held-out stations; one illustrative arithmetic input, not a local sample.
Multiply by one FSE factor10.00 m³/s × 1.53815.38 m³/s (approximately 15.4 m³/s)Upper multiplicative population-scatter reference from the same factor.FSE estimated across n=918 spatial-cluster held-out stations; one illustrative arithmetic input, not a local sample.

Dividing and multiplying 10.00 m³/s by the population FSE gives approximately 6.5–15.4 m³/s. This demonstrates factor arithmetic only: it is population cross-validation scatter, not a site-specific confidence interval, and no coverage level is assigned to the example.

What survived

Keep evidence in labelled layers

Population validation, network consistency and within-record sampling remain useful when each keeps its cohort, validation design, interpretation and boundary.

Use factor arithmetic as an illustration

Dividing and multiplying an illustrative estimate by one population FSE makes the multiplicative scale legible without turning it into a local interval.

What failed

One mechanically combined uncertainty band

The cited quantities have different targets, units, cohorts and dependence structures, so adding percentages or multiplying factors has no justified interpretation.

An achieved family envelope

The same-evidence GLO/KAP3/GEV exercise remains a proposed protocol, not a completed result.

Research question

What can the available national experiments say about uncertainty in a design flood, and where must their interpretation stop? The central question is not how to manufacture one wider bar. It is how to keep unlike evidence legible: method performance across held-out stations, disagreement within the gauging network, sampling within finite AMAX records, and sensitivity to assumptions that have not yet been tested together.

Terminology and taxonomy

Natural uncertainty concerns varying weather, antecedent state and catchment response. Data uncertaintyconcerns measurement, rating, record quality and descriptor evidence.Parameter uncertainty concerns quantities fitted from finite evidence. Structural uncertainty concerns the assumptions in the transfer, pooling or response model. Model-family uncertainty concerns the adopted probability family when more than one family remains plausible.

The natural, data, parameter, structural and model-family uncertainty layers are conceptually separate, but they can overlap and are not assumed independent. A difficult flood measurement can affect a fitted parameter and a validation residual at the same time. “Layer” is therefore an evidence label, not a claim of statistical orthogonality.

Evidence layer 1: population validation scatter

Donor-adjusted SM2025 was evaluated on 918 spatial-cluster held-out stations, with donor availability restricted to the training folds. Its median absolute QMED error was 23.4% and its FSE was 1.538. Those values describe out-of-sample performance across the frozen validation population. They mix whatever residual effects remain in that design; they do not identify a structural component in isolation.

The factor is population cross-validation scatter, not a site-specific confidence interval. Applying it to one example is a useful way to show multiplicative scale, but it does not attach a local coverage level or account for evidence that was not part of that validation statistic.

Evidence layer 2: network self-consistency

Across 85 adjacent same-river pairs, area-scaled QMED disagreement had a median of 12.2%, Q1–Q3 6.6–34.5% and P90 58.8%. This is a network self-consistency diagnostic only. It asks whether paired gauged records agree after a simple area scaling, not how uncertain either member is at a named project site.

River-name pairing is heuristic, and disagreement can reflect several overlapping causes: measurement, non-linear area scaling, tributaries, regulation, differing record periods or genuinely different catchment response. The distribution cannot identify one cause, establish an irreducible floor or be transferred as an individual-site interval.

Evidence layer 3: within-record sampling

The growth revalidation resampled years within each of 874 gauged stations and recomputed the fitted targets. Near the approximately 40-year record length studied, median within-record GF100 sampling variability was about ±13.0%. The experiment isolates one finite-record question more directly than a comparison between different stations.

The result belongs to GF100, the selected cohort and the frozen resampling design. It is not a universal design-flow allowance. It does not include QMED error, network data disagreement, structural assumptions, model-family choice or future non-stationarity.

Worked population factor

For an illustrative estimate of 10.00 m³/s, dividing by 1.538 gives 6.50 m³/s and multiplying by 1.538 gives 15.38 m³/s: approximately 6.5–15.4 m³/s. The separate worked table keeps the formula, rounded result, interpretation and 918-station scope together.

This is factor arithmetic, not an interval calibrated for the example site. The lower and upper values must retain the population-validation label wherever they are reproduced.

Why mechanical combination is invalid

The three numerical results do not share a target. The FSE concerns QMED prediction residuals across held-out stations; the same-river distribution concerns paired network disagreement; and ±13.0% concerns within-record GF100 sampling. Their cohorts, estimands and dependence structures differ.

They are therefore not combined mechanically. Adding 23.4%, 12.2% and 13.0% would treat unlike summaries as commensurate components. Multiplying 1.538 by new factors inferred from the other percentages would assume conversions and independence that have not been established. A joint result would need an explicit probabilistic model, stated covariance assumptions and validation against the quantity it claims to cover; none is supplied by these studies.

Model-family protocol status

A proposed protocol would fit GLO, KAP3 and GEV in parallel to the same frozen evidence, retain valid fits, and report how the resulting growth factors differ while holding the index flood and evidence custody fixed. This would be a sensitivity display, separate from validation scatter and finite-record sampling.

The protocol is not yet demonstrated. The cited artefacts contain no completed parallel-family envelope, so no family range, coverage or performance achievement is claimed here.

Practical reporting pattern

  1. State the adopted design-flow estimate, method and evidence date.
  2. Report population validation performance with its target, cohort, validation design and sample size.
  3. Report site and network data diagnostics as diagnostics, including record quality and reasons they may not transfer.
  4. Report finite-record sampling evidence for the statistic actually studied, with its record-length and cohort scope.
  5. Report structural and family sensitivities separately, marking unexecuted protocols as open research.
  6. Do not publish a combined total unless a joint model, dependence assumptions, coverage target and validation evidence are explicit.

Limitations

  • The QMED experiment reports population performance; it does not calibrate the worked factor for the illustrative estimate.
  • Same-river pairing is heuristic, area scaling is simplified and the diagnostic does not attribute disagreement to individual causes.
  • The GF100 result belongs to the selected 874-station cohort, the approximately 40-year record-length experiment and its within-station resampling design.
  • Natural variability and non-stationarity are named in the taxonomy but are not quantified by the three cited experiments.
  • Dependence and overlap between evidence layers have not been fitted, so the report supplies no joint uncertainty distribution.
  • The proposed GLO/KAP3/GEV comparison has not been executed as a completed family-envelope study.

What we do not conclude

  • We do not convert the 918-station FSE into local coverage for the illustrative estimate.
  • We do not treat the 85-pair network distribution as an error floor or transfer it to an individual site.
  • We do not treat ±13.0% GF100 sampling variability as a universal design-flow allowance.
  • We do not add the reported percentages or multiply them into one uncertainty factor.
  • We do not assume that natural, data, parameter, structural and family effects are independent.
  • We do not present GLO/KAP3/GEV spread as a completed or achieved uncertainty envelope.

Disclosure boundary

This report discloses the three research questions, broad cohort and validation designs, exact aggregate tables, worked-factor arithmetic, interpretations, claim statuses, source revisions and the point at which each conclusion stops. Its evidence sources are the frozen Hydrometric QMED revalidation, same-river network study and growth revalidation only.

It withholds raw station records, exact donor and feature engineering, learned coefficients, per-station residuals, private data-construction mechanics, customer information and deployable workflow logic. That boundary permits audit of the public claims without exposing the implementation needed to reconstruct the proprietary method.

Claim record

Each claim is classified by the evidence that supports it. The boundary states what the claim does not establish.

  1. Inference

    The natural, data, parameter, structural and model-family uncertainty categories are conceptually separate layers in a design-flood evidence record.

    Interpretation
    Each layer needs its own target, evidence basis and reporting boundary before it can inform a project decision.
    Boundary
    The layers can overlap and are not assumed independent; this taxonomy is not a fitted probabilistic decomposition.
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
    • Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  2. Demonstrated

    Across 918 spatial-cluster held-out stations, donor-adjusted SM2025 had median absolute QMED error 23.4% and FSE 1.538.

    Interpretation
    The result measures frozen-method generalisation scatter across the validation population.
    Boundary
    This is population cross-validation scatter, not a site-specific confidence interval; it does not isolate a single uncertainty mechanism.
    Linked result
    Uncertainty is layered, not one bar; sample: Five conceptual layers. Numerical evidence comes from 918 spatial-cluster held-out stations, 85 same-river pairs and 874 gauged stations in a within-station resampling experiment; natural variability and the proposed family protocol have no standalone numerical estimate here.
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  3. Inference

    For illustration, 10.00 m³/s divided and multiplied by FSE 1.538 gives approximately 6.5–15.4 m³/s.

    Interpretation
    The calculation exposes the multiplicative scale of the population factor without assigning local coverage.
    Boundary
    It is arithmetic using one population factor, not a project uncertainty interval or a combination with other evidence layers.
    Linked result
    Worked population-factor illustration for a 10.00 m³/s estimate; sample: The factor was estimated from 918 held-out stations. The 10.00 m³/s value is a single illustrative estimate introduced only to show the arithmetic.
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  4. Demonstrated

    Across 85 same-river pairs, area-scaled QMED disagreement had median 12.2%, Q1–Q3 6.6–34.5% and P90 58.8%.

    Interpretation
    The distribution diagnoses consistency across paired records in the gauging network.
    Boundary
    It is a network self-consistency diagnostic only; river-name pairing is heuristic, and the result is not transferable to an individual-site interval or lower bound on error.
    Linked result
    Uncertainty is layered, not one bar; sample: Five conceptual layers. Numerical evidence comes from 918 spatial-cluster held-out stations, 85 same-river pairs and 874 gauged stations in a within-station resampling experiment; natural variability and the proposed family protocol have no standalone numerical estimate here.
    Evidence
    • Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
  5. Demonstrated

    Across 874 gauged stations, median within-record GF100 sampling variability was about ±13.0% near the approximately 40-year record length studied.

    Interpretation
    Resampling years within each station measures how finite-record sampling moves the fitted GF100 target in that experiment.
    Boundary
    It is not a universal design-flow allowance and does not quantify index-flood, data, structural or family uncertainty.
    Linked result
    Uncertainty is layered, not one bar; sample: Five conceptual layers. Numerical evidence comes from 918 spatial-cluster held-out stations, 85 same-river pairs and 874 gauged stations in a within-station resampling experiment; natural variability and the proposed family protocol have no standalone numerical estimate here.
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  6. Inference

    The population FSE, same-river disagreement and within-record GF100 sampling result are not combined mechanically.

    Interpretation
    They concern different targets and cohorts, and their overlap and dependence have not been estimated.
    Boundary
    No additive or multiplicative total follows from these three studies; a combined result would require an explicit joint model and validated assumptions.
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
    • Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  7. Not yet demonstrated

    A same-evidence GLO/KAP3/GEV envelope is a proposed reporting protocol.

    Interpretation
    The protocol would report candidate-family sensitivity while keeping the index flood and evidence record fixed.
    Boundary
    It is not yet demonstrated: no completed parallel-family result, coverage assessment or achieved uncertainty envelope is reported.
    Evidence
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898
  8. Inference

    A practical design-flood record should report the adopted estimate, population validation, local data diagnostics, finite-record sensitivity and structural or family sensitivities in separate labelled fields.

    Interpretation
    Readers can then see which evidence is measured, which is local, and which remains open without treating unlike quantities as interchangeable.
    Boundary
    This is a reporting pattern inferred from the frozen studies, not a demonstrated reduction in error or a prescribed project allowance.
    Evidence
    • Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
    • Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
    • Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898

Primary sources

  1. Hydrometric QMED adversarial revalidation (2026).Evidence ID: 2438954b
  2. Hydrometric same-river network consistency study (2026).Evidence ID: db10a70d
  3. Hydrometric growth-curve revalidation (2026).Evidence ID: 76b42898

Version history

  1. Version 1.0.0current

    Initial publication separating population validation, network consistency, finite-record sampling and proposed model-family sensitivity.