Enhance the training pipeline with stateful checkpointing and improve
the resilience of the data loading process against filesystem latency
and transient I/O errors.
- Implement auto-resuming in `train_ddp` by loading model, optimizer,
and scheduler states from `state.pt`.
- Add atomic state saving using temporary files to prevent corruption.
- Introduce `_read_nii` with exponential backoff retries to handle
transient NFS/filesystem failures during NIfTI reading.
- Add explicit error handling for missing or unreadable label files in
`PatchDataset`.
- Update `sliding_window_probs` to conditionally apply Test-Time
Augmentation (TTA) based on the `tta` parameter.
- Add `scripts/test_dataloader.py` for verifying dataset integrity.
Refactor the data loading and preprocessing pipeline to handle edge cases in
medical imaging data, including NaN/Inf values, shape mismatches, and
numerical instability during training.
- Update `01_build_ntuh_manifest.py` with improved regex for T1c detection,
spine exclusion, and deduplication logic based on acquisition timestamps.
- Enhance `PatchDataset` in `src/dataset.py` to handle NaN/Inf values,
clip intensity ranges, and ensure label/image shape alignment via
padding/trimming.
- Add a zero-gradient fallback in `src/training.py` to prevent DDP
synchronization failures when encountering NaN/Inf losses.
- Add `scripts/scan_procs.py` for process monitoring.
- Increase DataLoader timeout to prevent hangs during heavy I/O.