Self-cascade latent Schrödinger bridge for CT field-of-view extension.
Wu X, Liu J, Yu H, Maier A, Huang Y.
Abstract
Computed tomography (CT) field-of-view (FOV) truncation leaves peripheral anatomy absent, causing errors in radiotherapy dose calculation and body-composition analysis. Existing methods rely on unavailable projection data or suffer over-smoothing and hallucination under severe truncation. Approach. We propose the self-cascade latent Schrödinger bridge (SCL-SB), combining a latent image-to-image Schrödinger bridge (I2SB) and a self-cascade mechanism in one shared-weight U-Net. I2SB constructs a diffusion bridge from the truncated-image distribution to the full-FOV distribution, enabling high-quality reconstruction in ten denoising steps. Two sequential rounds share identical weights: the first extrapolates, the second refines it using that output. A frequency-aware residual mechanism (FARM) amplifies gradients at high-frequency inter-round residuals to stabilise training and recover faithful detail. The truncation radius is sampled from a stratified distribution spanning mild to severe truncation during training, enabling a single model to cover the full severity spectrum without radius-specific fine-tuning. Main results. On a head-to-chest CT dataset spanning three truncation severities, SCL-SB, trained as a single model without radius-specific tuning, achieves the highest structural similarity index measure (SSIM) and a high-frequency ratio (HF-Ratio) closest to ideal among compared methods at every severity, running nearly 9× faster than a vision transformer-based baseline. Cascade training alone improves the first round's output without a second inference pass, termed cascade implicit enhancement (CIE). Relative to the same architecture without cascade (I2SB), the model recovers more high-frequency detail (HF-Ratio +45.3%) and a higher SSIM by 0.076 (9.0%) at severe truncation, at no extra cost. Significance. SCL-SB addresses the posterior-mean regression that blurs high-frequency detail in ill-posed FOV extrapolation, produces continuous Hounsfield unit reconstructions free of Vision-transformer patch-boundary discontinuities, and avoids spectral hallucination characteristic of GAN-based methods. Although clinical validation on larger, multi-centre cohorts is necessary, these properties may improve image guidance in applications such as adaptive radiotherapy, body-composition assessment, and spine surgery.