Reserved English test set · real recorded noise

Real noise.
Paired truth.
Listen.

Forty held-out English clips with naturally noisy inputs and aggressive speech-isolated, bandwidth-extended references—compared across Super Sonique and single-speaker Sidon.

Checkpoint Frozen KVAE 64-dim continuous Super Sonique Euler-8 · CFG 1.5 · temperature 0.80 Corruption None added
Super Sonique MOS recovery
Sidon MOS recovery
IPTV · Sonique / Sidon
Podcast · Sonique / Sidon

Episode-disjoint test

Two restorers. One panel.

The noisy recordings are real aligned inputs; no synthetic corruption or target leakage is added. Super Sonique is trained from scratch on 500 hours of Studio clean-voice targets with Gaussian-source conditional flow matching. Inference follows cfm_preferred8_cfg_1.5 exactly: raw weights, checkpoint normalization, Euler-8, CFG 1.5, noise temperature 0.80, jitter 0.625 with pattern seed 42, BF16 denoiser calls and FP32 integration. Eight batched calls represent 16 conditional/unconditional branch evaluations. The frozen posterior-mean KVAE decoder uses 250-frame cores with 11-frame halos. The comparison remains official single-speaker Sidon v0.1—not DialogueSidon. DNSMOS P.835 OVRL scores the exact MP3s; recovery is unclamped.

Ground truthAligned clean reference
Real noisyRecording → frozen codec
Super SoniqueStudio CFM · Euler-8 · CFG 1.5
Sidon v0.1Single-speaker restoration model

Loading the real-noise evaluation panel…