Staged skill against the gauge

The correction is built in three stages - Linear Scaling (LS), then LS + Empirical Quantile Mapping + a GPD tail (LSEQM), then a CNN refinement (LSEQM+DL). It is scored two ways: out-of-sample against the 172 independent BMKG stations (the strongest claim), and in-sample against the CPC-UNI target the correction was trained on. Together they show the marginal distribution moving to the gauge while the day-by-day timing does not.

What each stage adds

Each stage moves the metrics matched to its design dimension and leaves the others alone - which is why the pipeline is built in three steps, not one.

Skill by metric

Pick a reference, then a metric: the three stages side by side against the perfect-agreement target. The CPC-UNI reference adds the temporal-skill metrics that only exist in-sample.

· · perfect agreement at . Best stage shaded in the tables below.

Out-of-sample: vs 172 BMKG stations

The strongest claim: the BMKG network is independent of the correction reference. Three pillars move to target; event detection is a designed trade-off (unpacked in detection by threshold).

In-sample: vs CPC-UNI (the calibration target)

The full two-tier picture, including the Temporal Skill rows that live only in-sample. LS has zero bias by construction here; LSEQM/LSEQM+DL slightly overshoot CPC-UNI at the upper tail - yet match BMKG almost exactly, because CPC-UNI itself under-catches heavy rain.

The paradox in one place: across the four distributional pillars the corrected product reaches the gauge to within a few percent, but the Temporal Skill rows barely move - Pearson r holds near , RMSE and NSE do not improve. That flat timing track is the subject of [the timing ceiling](./ceiling), and most of it is a fixable [calendar-window artefact](./window).