Detection skill across rainfall thresholds

The pooled event-detection numbers on the scorecard show POD falling from 0.78 (LS) to 0.65 after correction - which looks like a loss. It is not: it is a crossover. Raw LS wins at drizzle by over-forecasting wet days; the corrected product gives that back in exchange for a calibrated wet-day frequency and better skill at the heavy rain that actually matters. Move across the threshold axis and watch it flip.

Skill is pooled as the cross-station median per dekad, then averaged over the 36 dekads (172 BMKG stations). Above 50 mm/day events become rare - roughly 7 per dekad at 50 mm, ~1 at 100 mm - so the curves thin out and the 150 mm "extreme" class is omitted as unstable. Toggle the IQR band to see the across-station spread.

The pooled version of this trade-off is the event-detection pillar on the [staged scorecard](./staged-skill); the flat timing track it cannot fix is [the timing ceiling](./ceiling).