Staged skill against the gauge
The correction is built in three stages - Linear Scaling (LS), then LS + Empirical Quantile Mapping + a GPD tail (LSEQM), then a CNN refinement (LSEQM+DL). It is scored two ways: out-of-sample against the 172 independent BMKG stations (the strongest claim), and in-sample against the CPC-UNI target the correction was trained on. Together they show the marginal distribution moving to the gauge while the day-by-day timing does not.
What each stage adds
Each stage moves the metrics matched to its design dimension and leaves the others alone - which is why the pipeline is built in three steps, not one.
Skill by metric
Pick a reference, then a metric: the three stages side by side against the perfect-agreement target. The CPC-UNI reference adds the temporal-skill metrics that only exist in-sample.
Out-of-sample: vs 172 BMKG stations
The strongest claim: the BMKG network is independent of the correction reference. Three pillars move to target; event detection is a designed trade-off (unpacked in detection by threshold).
In-sample: vs CPC-UNI (the calibration target)
The full two-tier picture, including the Temporal Skill rows that live only in-sample. LS has zero bias by construction here; LSEQM/LSEQM+DL slightly overshoot CPC-UNI at the upper tail - yet match BMKG almost exactly, because CPC-UNI itself under-catches heavy rain.