Four published failure modes. Three Workbench regimes where EMA (or trackers) clearly beat balanced GSRF Practical, plus one Axiom industrial composite boundary. Same tone: we lost under that protocol — do not sell as that job.
1. Feedback delay as command path
On delayed measured-plant feedback packages, MAD to measured was worse for GSRF by roughly +59% (3 min), +97% (6 min), and +106% (10 min) versus EMA. The spring fights delayed truth. We lost. Do not deploy here.
2. Fixed false-alarm residual detection
At matched FA on synthetic excursions, residual recall was about 0.44 (GSRF) vs 0.83 (EMA). Peak compression + larger normal residual raises thresholds. We lost as a detector. Use a detector; use GSRF to calm.
3. Tracking MAD vs raw
Excellence hunt multi-family package: 0 / 52 MAD wins vs best EMA. Not noise—identity. The spring is the product, not a bug. If your KPI is hug-the-raw, you want EMA/Kalman/k_return≈0.
4. Axiom industrial composite (PID-style pure evaluate)
Axiom pure evaluate on real PID-style loops (frozen GSRF Practical, industrial metrics, causal baselines):
no clear composite wins vs strong causal smoothers on the sealed eligible set.
Wear / command-churn-style terms often favour baselines. Boundary mapped; not a hero claim.
This is a different protocol from Workbench mid-band/spectral packs — do not conflate Axiom industrial
oscillation_power with mid-band ring-kill characterization.
We lost under that claim frame. Do not sell ZO as a universal industrial-PID composite champion.
Why publish this?
Because the green lights (mid-band kill, peaks, adaptive x*) only mean something next to red lights. Full cage: Evidence → Boundaries.
Sources: characterization v4 delay · probes v5 alarm · excellence hunt v2 · Axiom industrial pure evaluate (2026-08, sealed eligible set).