Technical prove-it

Locked evidence. Documented limits.

Every number below comes from the GSRF Audit Workbench characterization packs— not slideware. Competitors sell confidence. We sell receipts.

Line chart: grey raw mid-band ring and spike, yellow EMA tracks closer to raw, cyan GSRF Practical damps ring and pulls the peak toward normal
Visual identity check (synthetic). Locked step/gain packs are in sections 1–2 below.

All four sections are on this page—scroll or jump. Nothing is hidden behind tabs.

Safety cage — at a glance

StatusTopicLocked finding
Green Mid-band oscillation ~+35% osc31 vs best EMA; ~−70% sweet-band spectral power (phase6.2 full series)
Green Peak compression ~38% mean peak-dev advantage; 49/52 dataset wins — download pack
Green Adaptive x* Regime jumps 50→70→40: frozen ~9.7 error → adaptive ~1.5
Red Feedback delay MAD to measured vs EMA: +59% / +97% / +106% at τ=3 / 6 / 10 min
Red Fixed-FA detection Residual recall ~0.44 GSRF vs ~0.83 EMA at matched FA
Red Tracking MAD 0/52 wins vs best EMA — spring costs tracking by design

Full narrative, charts, and sources: sections 1–5 below. Industrial actuator battery → · Green/red detail →

Download receipts (not a curtain)

Open the ledgers behind headline numbers. Partial public packs from frozen Workbench runs. Still better: gsrf-bench on your CSV (~10 min).

1 — The identity

GSRF is not “EMA with better marketing.” It is a soft-servo to a declared normal.

update ≈ baseline + k_return·(−(x − x*)) + mem·tanh(Δx) + w_obs·tanh(x_obs − x)
state lives in log-space; output = exp(x)

Step response (the intentional trade-off)

Step from 50 → 65 with x* frozen at the pre-step normal. Plot from locked workbench CSV:

Step response: raw to 65, EMA tracks to 65, GSRF Practical settles near 55, k_return=0 tracks to 65
Source: characterization_v4 responses/step_up.csv
MethodSettles nearMeaning
EMA65Tracks the new level
GSRF k_return = 065Spring off → tracker-like
GSRF Practical (balanced)~55Halfway home — spring vs observation pull
GSRF Reference50Open attractor twin (not a live tracker)

That halfway settle is peak compression. It is also why MAD-to-raw loses. If you need full settle after a permanent jump, move x* (see section 3)—don’t pretend balanced GSRF is EMA.

Source: workbench characterization v4 · responses/step_up

2 — Frequency sweep & the sweet spot

Headline: GSRF is tuned to clamp amplitude toward normal and kill mid-band ring—not to track slow waves.

Gain plateau (pure sine, workbench)

RMS AC gain vs period. EMA → 1.0 on slow waves. GSRF plateaus ~0.37. Points from locked CSV:

AC gain vs period: EMA approaches 1.0 at long periods; GSRF plateaus near 0.37
Source: gsrf_characterization_v4 sine_frequency_response.csv

Oscillation metric + spectral power

ResultNumberNotes
osc31 vs best EMA ~+35% phase6.2 · 3 seeds · full series
Sweet-band spectral kill ~−70% Same packs · not a rolling-31-only artifact
Frequency-sweep sweet spot P ≈ 50–120 Peak ~+51% osc adv at P=80; loses at very fast (P=20) and very slow (P=400–600) on osc31

Sources: characterization v4 · probes v5 FULL PSD · excellence hunt v2

3 — Adaptive x* (the fix for permanent jumps)

Headline: We fixed the “stuck after step” failure mode for multi-regime plants—by moving the thermostat, not by lying about the spring.

Synthetic permanent shifts 50 → 70 → 40 with mid-band oscillation. Mean |error| vs true regime base on regimes B+C:

MethodMean abs mean-error (B+C)
GSRF frozen prefix x*~9.7 (stuck on old normal)
GSRF fixed full-series x*~9.7
GSRF adaptive (window 200)~1.5
EMA_low~0.5 (best pure tracker)

In the original regime, adaptive GSRF still beat EMA on oscillation (osc31 ~0.50 vs ~0.78). On real data: adaptive helps multi-regime industrial/MBP-style traces; do not blindly adapt on flat SpO2.

Sources: probes v5 adaptive · realdata v6 VitalDB/phase5 matrix

4 — The safety cage

Red light

Feedback delay

MAD to delayed measured plant vs EMA:

  • τ = 3 min: GSRF ~+59% worse
  • τ = 6 min: ~+97% worse
  • τ = 10 min: ~+106% worse

Do not use balanced GSRF as a delayed closed-loop command smoother.

Red light

Fixed-FA alarm detection

At matched false-alarm rate on synthetic excursions (residual mode):

  • GSRF recall ~0.44
  • EMA recall ~0.83

Peak compression and higher normal residual force a higher threshold when FA is matched. Use a detector for alarms; use GSRF to calm.

Green light — complete

Use freely (within band)

These are the jobs GSRF is built for. Numbers are from locked workbench packs, not demos.

  • Oscillation kill / mid-band spectral calming — osc31 ~+35% vs best EMA; sweet-band power ~−70% (phase6.2 full series, 3 seeds)
  • Peak compression toward defined normal — ~38% mean peak-dev advantage; 49 / 52 wins across characterization families
  • Adaptive x* when the plant’s normal truly moves — permanent jumps tracked without abandoning the spring (frozen ~9.7 → adaptive ~1.5 mean abs error)
  • Pre-AI / pre-model signal hygiene — calm ring and spikes before expensive inference (not sold as a detector)

Frequency sweet spot: period ≈ 50–120 samples (peak osc adv ~+51% at P=80). Outside that band, re-check with your metric—very fast (P≈20) and very slow (P≈400–600) osc31 can favor EMA.

Tracking MAD: 0 / 52 wins vs best EMA in excellence hunt v2. That is not a bug—that is the spring doing its job. Full red-light detail is in the two columns above (delay + fixed-FA alarms). Narrative of specific losses: When GSRF lost.

Sources: characterization v4 delay · probes v5 alarm & PSD · excellence hunt v2 · adaptive probe v5

5 — Industrial actuator batteries (synthetic, bounded)

Not “GSRF wins industrial actuators.” These are synthetic SCADA-style traces run through a frozen Phase 5 MAIN workbench evaluator (balanced preset) vs best EMA / Clamped_EMA. Real partner plant data is a separate, future ladder.

Two packs (19 files total). Metrics from the Audit Workbench only — shaping score = wear + churn + rate-stress + osc31 vs best EMA (≥3/4 clear, 2/4 partial, 1/4 mixed, 0/4 loss).

Aggregate (honest)

PackTypeFilesClearPartialMixedPure losses
Seven industrial actuators* Synthetic SCADA-style 7 5 1 1 0
Adversarial Tier-1 Synthetic (anti-complimentary) 12 2 6 4 0
Real / public first look DAMADICS plant + TE simulated 6 0 2 0 4

*Turbine/RO remaps documented. †Early DAMADICS first look (3 actuators × 3 days) superseded by multi-day pack below. TE Mode-1 XMVs are simulated plant, not field actuators. Same frozen evaluator — real plant is harder.

What we claim / refuse

Allowed
  • On synthetic packs: regime-bounded shaping wins (7 clear / 7 partial / 5 mixed / 0 loss on 19 files)
  • Adversarial pack included (only 2/12 clear) — not flattering-only
  • On synthetic Phase 5 MAIN, wear+osc often both favor Practical vs best EMA (pack-dependent)
  • Independent stiction stratum: osc win rate ~76% [57–89%], n=25 (labeled synthetic)
  • Independent slow-tracking stratum: wear ~68% [48–83%], n=25 (provisional)
  • On real DAMADICS multi-day: full day table — mostly composite loss; per-metric rates published
  • Tracking MAD vs raw worse for Practical when it shapes — spring identity
Refused
  • “GSRF wins industrial actuators” without qualifiers
  • Treating synthetic 0-loss as real-plant proof
  • Averaging synthetic + DAMADICS + TE into one fake win rate
  • “88% wear wins on Lublin” or other inverted metrics not in DECISION_SUMMARY
  • Universal “~95% wear on any adversarial synthetic” (generator-family only; independent overall ~18%)
  • Kalman / commercial toolbox bake-off (not run here)

Real plant multi-day (DAMADICS Lublin)

Official dumps part1–4 · 25 days × 3 valves = 75 full-rate frozen Phase 5 MAIN runs (+ 3 stride-10 all-days continuity checks). Workbench only — not hand-calc.

CV → demand, X → measured. Early 3-day first look (2 partial / 1 loss) is superseded by multi-day.

Composite verdicts (per-day, 75 runs)

ClearPartialMixedLoss
16860

Per-metric vs best EMA (same 75 days) — inversion check

When the composite looks bad, open the ledger. Lower is better for wear / osc / rate / MAD.

MetricGSRF better than best EMA
Wear (total travel)1/75 (1.3%) — GSRF usually higher travel (~1.8× mean)
Rate-stress14/75 (18.7%)
Oscillation8/75 (10.7%)
Churn0/75
MAD tracking0/75 (spring cost)

Domain pattern (honest): synthetic Phase 5 SCADA-style packs often show shaping (incl. wear) edges; multi-day Lublin is hard — composite mostly loss. Same frozen rule. No wear-champion claim on real plant.

Independent adversarial map (stratified, n=100)

Positives first: on a second generator family (100 synthetic industrial files, frozen Phase 5 MAIN), stiction-heavy files show a robust oscillation niche and slow-tracking files a provisional wear niche — with Wilson 95% CIs. Overall “any industrial synthetic” is not a free win (limits below). Same evaluator as all other industrial packs.
Domain (independent synthetic)What GSRF does vs best EMAEvidence
Stiction-heavy Reduces oscillation 76% [57–89%] Wilson 95% CI · n=25
Slow-tracking Reduces wear (provisional) 68% [48–83%] · n=25 · CI low edge near 50%
Multi-regime (this pack) No metric niche at 50% bar 0% wear · 0% osc · n=25
Mixed pathologies (this pack) No metric niche at 50% bar 0% wear · 0% osc · n=25
Overall independent (100 files) Does not win “general industrial synthetic” Wear 18% [12–27%] · osc 24% [17–33%] · loss 62/100
Honest limits (independent data): On independent adversarial data (100 files), GSRF does not win on general industrial problems. It wins on stiction-heavy oscillation (76%) and slow-tracking wear (68%, provisional). Multi-regime and mixed problems in that pack: 0%. We publish both the wins and the failures. The map is the product.

Generator-family note: Earlier tier‑1 synthetic adversarial fleets can show high wear win rates within that family. Independent generation does not reproduce a universal ~95% wear claim — label family when quoting those packs.

Domain map: independent stiction osc and slow-tracking wear niches, overall limits
Figure — Independent pack domain map (claim-safe summary visual)

True command+measured: the test that mattered

The EMA-proxy robot data was always marked as provisional. True command+measured is the real test — native controller (or teleop) demand vs independent measured, frozen Phase 5 MAIN vs best EMA.

Wear is not a proxy artifact. 298 continuous true-command samples across three platforms — 0 losses. We’re ready to run your data. Ledger: download fleet ZIP · TRUECMD_FLEET_ALL_SERIES.csv

Platform n Wear Osc Rate Loss
ABB IRB 4400 (RMPD TCP cmd vs Leica laser) 30 30/30 9/30 30/30 0
NIST UR5 joints (target vs actual, pos+vel) 72 72/72 60/72 0/72 0
Franka DROID joints (abs action vs state) 196 196/196 134/196 0/196 0
Continuous total 298 298/298 203/298 30/298* 0
  • ABB RMPD — wear 30/30, rate 30/30, osc 9/30 (position osc weak; velocity clear)
  • NIST UR5 — wear 72/72, osc 60/72, 0 losses (vel osc 36/36; some pos joints mixed)
  • Franka DROID — wear 196/196, osc 134/196, 0 losses (~15 Hz teleop abs-joint)
  • The loss: one discrete pose cloud (UR5 GTPN y-axis) where EMA won on total travel/osc — documented, not merged with continuous fleet

*Rate is surface-dependent: almost all rate wins are RMPD TCP. NIST and Franka joint packs are 0 on rate. Osc is surface-dependent too. Do not publish one robot-wide osc or rate %. Competitors sell confidence. We sell receipts — Workbench CSVs, overlays, and gate decisions on disk.

Allowed Wear/travel better vs best EMA on continuous true cmd+meas (ABB TCP + UR joints + Franka joints), with osc/rate only under surface + source labels. Ask partners for native command + joint_states at plant rate.

Banned One robot-wide osc/rate % · EMA-proxy as true cmd · controller replacement · plant wear miracles.

EMA-proxy robot packs (still labeled proxy)

Provisional public robot packs where demand = EMA proxy of measured (native command missing): Proxy osc looked strong (~91% fleet) — true-cmd surface is different (wear holds; osc depends on path vs joint tracking). Supports proprioception pre-stage under proxy protocol only. Not a joint controller; not multi-joint fleet ROI; not merged with valve DAMADICS rates.

Also locked (Workbench): ISDB public loops osc candidate (~64%, amber) · SWaT / partner plant under NDA still open for green-light plant shaping.

Sources: experiments/phase5_seven_actuator_datasets/ · phase5_adversarial_tier1_pack/ · phase5_damadics_multiday/DAMADICS_MULTIDAY_FINAL_REPORT.md · METRIC_INVERSION_AUDIT.md · experiments/massive_independent_adversarial_validation/ (n=100, stratified CIs) · robot truecmd fleet: phase5_robot_cmd_measured · phase5_robot_nist_ur5_truecmd · phase5_robot_franka_cmd_measured · frozen run_phase5_adversarial_pack / run_phase5_main · DAMADICS: iair.mchtr.pw.edu.pl/Damadics · Bartyś et al. 2006 CEP · citations

Actuators + robot true cmd → Research note → When GSRF lost → Try on your CSV →

Questions people actually ask

Google-snippet style answers that match the claim ladder. Full FAQ: /faq.

What locked evidence does GSRF publish?

Identity (gain ~0.37, step ~55), mid-band osc ~+35% / spectral ~−70%, peak ~38% (49/52), adaptive x*, industrial Phase 5 synthetic batteries, DAMADICS partial/loss, and red lights for delay + fixed-FA residual.

Does GSRF win industrial actuators?

Not without labels. Independent n=100: overall wear ~18% / osc ~24%; stiction osc ~76% and slow-tracking wear ~68% (provisional) when labeled. DAMADICS multi-day wear 1/75. Robot true-cmd continuous wear 298/298 (osc/rate surface-labeled). Flagship: actuators guide. Not a controller.

Where does GSRF fail?

MAD tracking 0/52, multi-minute delayed command paths, fixed-FA residual detection vs EMA, and extreme frequency bands. When GSRF lost →

Where do the numbers come from?

GSRF Audit Workbench locked packs. Workbench · Methodology.