No description
Find a file
Furen Xiao 8fe1818c7d Enhance T1c series selection and reconstruction processes
- Updated regex patterns for improved matching of series names and tags.
- Added functionality to reject non-head series based on study and series descriptions.
- Implemented max voxel spacing check to filter out series with excessive spacing.
- Enhanced the reconstruction script to handle dynamic-frame series exclusions and artifact pruning.
- Modified output paths for reconstructed NIfTI files and added QA screenshot generation.
- Improved argument parsing in benchmark and pseudo-labeling scripts for better flexibility.
- Introduced a new script for generating QA screenshots from reconstructed volumes.
2026-09-27 07:14:17 +08:00
scripts Enhance T1c series selection and reconstruction processes 2026-09-27 07:14:17 +08:00
scripts_monai feat(monai): add MONAI pipeline C implementation 2026-09-26 11:04:20 +08:00
scripts_nnu refactor(nnu): transition to native volume processing for Pipeline B 2026-09-26 12:44:40 +08:00
src Enhance T1c series selection and reconstruction processes 2026-09-27 07:14:17 +08:00
.gitignore Enhance T1c series selection and reconstruction processes 2026-09-27 07:14:17 +08:00
AGENTS.md feat(monai): add MONAI pipeline C implementation 2026-09-26 11:04:20 +08:00
BENCHMARKING.md Enhance T1c series selection and reconstruction processes 2026-09-27 07:14:17 +08:00
README.md refactor(nnu): transition to native volume processing for Pipeline B 2026-09-26 12:44:40 +08:00

Longitudinal T1c Brain-Tumor Segmentation

Longitudinal (repeated-measures) analysis of brain T1 post-contrast (T1c) MRI with tumor segmentation, plus an iterative pseudo-labeling study that leverages a large unlabeled pool from the same subjects/series to improve a held-out, patient-level test Dice.

All volumes are preprocessed to a uniform 1 mm isotropic, head-cropped, percentile-normalized form (Pipelines A and C) so that tumor counts are directly comparable in mm³ across sources and timepoints. Pipeline B instead feeds the native volumes straight into nnU-Net, which performs its own preprocessing (see below).

Environment

Conda env longitudinal (Python 3.14, torch 2.14 +cu126). Activate before any command:

source /opt/conda/etc/profile.d/conda.sh && conda activate longitudinal

GPU (CUDA 12.6) is available; multi-GPU jobs use torchrun (in-house) or nnU-Net's own DDP (-num_gpus). MONAI is pip-installed in the env (Pipeline C). Run scripts from the repo root.

Data sources

Source What Notes
ntuh Labeled T1c + tumor segmentation Native MR, deduped per acquisition
m6 Labeled (GTV warped from CT) + large unlabeled pool GTV registered CT→T1c
lee Longitudinal T1c scanned as JPG + DICOM txt Volumes reconstructed from slices

Directory layout

data/
  manifests/            jsonl row tables (see Pipeline A)
  proc/                 <key>.nii.gz (1mm, cropped, normalized)
  proc/<key>_label.nii.gz
  procmeta/<key>.json   geometry (origin/direction/crop_vox) + normalization
  vols.json             labeled tumor volume stats (mm3 percentiles)
  pseudo/roundK/        in-house pseudo-labels (rows.jsonl, masks, summary.json)
  pseudo_nnu/roundK/    nnU-Net pseudo-labels (rows.jsonl, masks, summary.json)
  pseudo_monai/roundK/  MONAI pseudo-labels (rows.jsonl, masks, summary.json)
src/                    U-Net, dataset, losses, training/eval helpers
scripts/                Pipeline A (in-house 3D U-Net)
scripts_nnu/            Pipeline B (nnU-Net)
scripts_monai/          Pipeline C (MONAI)
runs/roundK/            in-house checkpoints (best.pt, final.pt, state.pt)
runs_nnu/roundK/        nnU-Net checkpoint snapshots (best_nnu.pth)
runs_monai/roundK/      MONAI checkpoints (best.pt, final.pt, state.pt)
nnu/                    nnU-Net raw / preprocessed / results trees
results/                evaluation JSON tables + plots
logs/                   run logs

Pipeline A — In-house 3D U-Net (scripts/)

3-class-capable but used as binary (background / tumor) Unet3D, 96³ patches, DDP, per-sample weighted loss.

# Script Purpose
01 01_build_ntuh_manifest.py Build NTUH2022G4 labeled T1c + seg manifest
02 02_build_m6_dataset.py Build M6-2025 manifests (GTV CT→T1c registration/warp)
03 03_scan_lee_t1c.py Scan lee for brain T1c series → raw + selected manifests
04 04_reconstruct_lee.py Reconstruct lee T1c niftis from JPG slices + txt metadata
05 05_build_splits.py Patient-level train/val/test splits + unlabeled pool + volume stats
06 06_pseudo_label.py Pseudo-label the pool with a round model (sliding-window + TTA)
07 07_train.py DDP training entrypoint (labeled + weighted pseudo rows)
08 08_eval.py Held-out test evaluation (Dice from probability map)
09 09_run_iterative.py Orchestrates rounds 0…K and produces the summary table/plot
— preprocess.py Crop/normalize source volumes → data/proc/
— scan_procs.py Flag corrupt/non-finite processed volumes
— test_dataloader.py Smoke-test the training dataloader

Pseudo-label gates (per volume): pos when p_tumor ≥ tau_pos, largest-CC fraction ≥ min_cc_frac, and volume within the labeled p2–p98 range; neg when ≥ neg_frac of interior voxels have p_bg ≥ tau_neg; then a per-subject longitudinal consistency filter.

Run the full study:

python scripts/09_run_iterative.py --rounds 4 --gpus 3

Individual stages are standalone, e.g.:

torchrun --standalone --nproc_per_node 3 scripts/07_train.py \
    --rows data/manifests/split_train.jsonl --val data/manifests/split_val.jsonl \
    --epochs 40 --lr 3e-4 --batch 3 --ckpt-dir runs/round0
python scripts/08_eval.py --rows data/manifests/split_test.jsonl --ckpt runs/round0/best.pt

Known issue in Pipeline A (longitudinal consistency)

06_pseudo_label.py's grid_info()/dice_a_on_b() resample timepoints into a shared physical space using procmeta origin/direction/crop_vox. This is not reliable across separate acquisitions:

  • Each scan has its own scanner/patient coordinate frame (table offset + head pose), so two visits of the same subject do not overlap in absolute space. Benchmarking against labeled multi-timepoint patients gives median cross-visit true-label Dice ≈ 0, meaning the consistency gate tends to over-reject.
  • crop_vox is stored as [z, y, x] array-axis starts, while direction is ordered (x_dir, y_dir, z_dir); the pairings in grid_info mix these up.

This is noted for the record; Pipeline B below replaces the gate with a frame-invariant one. Pipeline A was left unchanged.


Pipeline B — nnU-Net iterative pseudo-labeling (scripts_nnu/)

The same study driven by nnU-Net v2 (installed build nnunetv2 2.8.1) as the segmentation backbone, keeping the identical patient-level splits, unlabeled pool, and evaluation protocol so results are directly comparable to Pipeline A.

This installed nnU-Net is a modern fork: preprocessed data is .b2nd (blosc2), checkpoints store network_weights (not model), and --save_probabilities writes a per-case <case>.npz (channel-first (C, z, y, x), resampled back to the original native input grid — so are the exported argmax segs). The code below targets that build, not the upstream nnU-Net docs.

Dataset: Dataset210_NTUH_T1C_PL, single channel t1c, 2 classes (background=0, tumor=1), 3d_fullres only, nnUNetPlans, fold 0.

# Script Purpose
01 01_nnu_prepare_dataset.py Build/rebuild the raw dataset (symlinked native imagesTr, labelsTr 0/1 + dataset.json) from a rows jsonl
02 02_nnu_plan_preprocess.py Plan + preprocess (--clean), then write subject-level splits_final.json
03 03_nnu_train.py Train one round via NTUHLPLTrainer (DDP -num_gpus), optional warm start
04 04_nnu_pseudo_label.py Multi-GPU nnUNetv2_predict on native pool volumes + selection gates (physical mm)
05 05_nnu_eval_test.py Held-out test evaluation on native volumes (probability Dice@0.5 + hard-seg Dice)
06 06_nnu_run_iterative.py Orchestrates rounds 0…K, table + plot
— trainers/ntuh_pl_trainer.py NTUHLPLTrainer: env-driven epochs/LR + full-weight warm start
— nnu_common.py Shared paths, env, dataset/split helpers, native-grid selection + consistency

Preprocessing (native inputs)

Pipeline B does no resampling, cropping, or intensity normalization before nnU-Net. The raw dataset (01) symlinks the native T1c NIfTIs from the source manifests (img field) into imagesTr; labelsTr entries are symlinks when the label is already 0/1 on the image grid, otherwise nearest-warped + binarized copies (label fix-up only). All geometric/intensity preprocessing is then nnU-Net's own plan_and_preprocess: nonzero-bbox crop, resample to the plan's median-based target spacing, per-channel normalization, .b2nd storage. The split/pool row jsonls therefore carry both the processed (pimg/plabel, Pipelines A/C) and native (img/label, Pipeline B) paths — rebuild them with scripts/05_build_splits.py.

Downstream of the model (selection gates, consistency filter, evaluation) work in each case's native grid with physical units: nnUNetv2_predict itself already returns the probability map (and the argmax seg) resampled back to the native input grid, so no extra remapping is needed; tumor volumes are voxel counts × native voxel volume (mm³); centroid offsets are scaled to mm by the native spacing.

Round semantics

  • Round 0: train on the labeled patient-level train split (nnU-Net internal validation = the held-out val split, patient-disjoint via splits_final.json). Pseudo-label the remaining pool → data/pseudo_nnu/round1.
  • Round k (k ≥ 1): dataset grows with all accepted pseudo-labels from rounds 1…k (positives + zero-mask negatives, de-duplicated by key); re-plan/preprocess; warm-start training from round k−1's best checkpoint at a lower LR; predict the remaining pool; evaluate on test.

Because this build has no incremental preprocessing, each round re-plans and re-preprocesses the whole (growing) dataset; the pool shrinks each round since accepted cases are excluded from re-prediction.

Custom trainer

NTUHLPLTrainer (resolved through the nnUNet_extTrainer env var) adds:

  • NNU_PL_EPOCHS / NNU_PL_LR — epoch count and initial PolyLR (set per round).
  • Full-weight warm start (NNU_PL_WARMSTART): loads all weights including the segmentation head in on_train_start. The CLI -pretrained_weights flag deliberately skips .seg_layers. keys, which would silently re-initialize the head and break round-to-round fine-tuning — hence this hook.

Longitudinal consistency (frame-invariant)

Instead of resampling into absolute patient space (unreliable, see Pipeline A note), two timepoints are compatible if they agree with at least one accepted neighbor on both:

  • tumor centroid offset relative to the head centroid ≤ --max-rel-dist (default 40 mm), and
  • tumor volume ratio ≤ --vol-ratio (default 10×)

same greedy-removal scheme as Pipeline A, but robust to scanner/pose differences across visits.

Running

Full study from the repo root:

python scripts_nnu/06_nnu_run_iterative.py --rounds 4 --gpus 3

Defaults: baseline 250 epochs @ 1e-2; warm-started rounds 75 epochs @ 1e-3. Pseudo-label gates match Pipeline A (--tau-pos 0.95, --min-cc-frac 0.2, --neg-frac 0.9, volume p2–p98). Optional flags: --no-tta, --no-neg-pseudo, and the gate overrides. Outputs to results/round{k}_test_nnu.json, results/iterative_table_nnu.jsonl, results/iterative_dice_nnu.png; per-round checkpoints snapshotted to runs_nnu/round{k}/best_nnu.pth.

Individual stages:

python scripts_nnu/01_nnu_prepare_dataset.py --rows <rows.jsonl>
python scripts_nnu/02_nnu_plan_preprocess.py --val data/manifests/split_val.jsonl
python scripts_nnu/03_nnu_train.py --gpus 3 --epochs 250 --lr 1e-2
python scripts_nnu/04_nnu_pseudo_label.py --pool data/manifests/unlabeled_pool.jsonl \
    --out data/pseudo_nnu/round1 --gpus 3
python scripts_nnu/05_nnu_eval_test.py --rows data/manifests/split_test.jsonl \
    --out results/round0_test_nnu.json --gpus 3

Verified

End-to-end smoke-tested on a scratch 4-case native dataset (mixed anisotropic spacing, multi-class + off-grid labels): dataset build (native symlinks + label fixup) → planning → splits_final.json → training 1 epoch → growing re-plan/re-preprocess → sharded native prediction → selection gates in mm → test eval. The consistency filter's keep / reject / single-timepoint paths are unit-tested with synthetic volumes.


Pipeline C — MONAI iterative pseudo-labeling (scripts_monai/)

The same study driven by MONAI (1.6.x, pip) as the segmentation backbone, keeping the identical patient-level splits, unlabeled pool, pseudo-label gates, frame-invariant consistency filter, and evaluation protocol so results are directly comparable to Pipelines A and B.

  • Network: monai.networks.nets.UNet (1 in / 2 out, channels 16→128, ~1.2 M params) — the MONAI analogue of Pipeline A's Unet3D(base=16, depth=4).
  • Data: MONAI Dataset + transform chain (load → min-pad → per-axis flips / rot90 / brightness / noise → RandCropByPosNegLabeld 96³). Patch sampling is foreground-aware (pos=neg=1) rather than Pipeline A's uniform random crop.
  • Loss: per-sample weighted Dice+CE on MONAI DiceLoss (background+foreground averaged, dice weight 0.5) — the same convention as Pipeline A, so pseudo rows are down-weighted sample-by-sample (--pseudo-weight 0.3).
  • Schedule: AdamW + linear warmup + cosine annealing (Pipeline A's schedule).
  • Inference: MONAI sliding_window_inference (gaussian blend, 50% overlap, 96³ windows) + the same 4-view flip TTA as Pipeline A; multi-GPU = torchrun rank shards (like Pipeline A), DDP for training.
  • Shared with the other pipelines: round semantics, checkpoint layout (runs_monai/roundK/{best,final,state}.pt), pseudo outputs (data/pseudo_monai/roundK/), and the gate/consistency code itself is imported from scripts_nnu/nnu_common.py so it cannot drift.
# Script Purpose
01 01_monai_build_rows.py Assemble a round's training manifest (labeled + accepted pseudo rows, weighted, de-duped by key)
02 02_monai_train.py DDP training of the MONAI UNet for one round (torchrun), full-weight warm start, auto-resume
03 03_monai_pseudo_label.py Sliding-window + TTA prediction of the remaining pool, selection gates + consistency filter
04 04_monai_eval_test.py Held-out test evaluation (Dice from probability map @0.5)
05 05_monai_run_iterative.py Orchestrates rounds 0…K, table + plot
— monai_common.py Network / weighted loss / transform chain / inference + distributed helpers

Running

Full study from the repo root:

python scripts_monai/05_monai_run_iterative.py --rounds 4 --gpus 3

Defaults: baseline 100 epochs @ 3e-4; warm-started rounds 30 epochs @ 1e-3. Pseudo-label gates match Pipelines A/B (--tau-pos 0.95, --min-cc-frac 0.2, --neg-frac 0.9, volume p2–p98). Optional flags: --no-tta, --no-neg-pseudo, and the gate overrides. Outputs to results/round{k}_test_monai.json, results/iterative_table_monai.jsonl, results/iterative_dice_monai.png; per-round checkpoints in runs_monai/round{k}/best.pt.

Individual stages:

python scripts_monai/01_monai_build_rows.py --train data/manifests/split_train.jsonl \
    --accepted data/pseudo_monai/round1/accepted.jsonl \
    --out data/pseudo_monai/round1_rows.jsonl
torchrun --standalone --nproc_per_node 3 scripts_monai/02_monai_train.py \
    --rows data/pseudo_monai/round1_rows.jsonl --val data/manifests/split_val.jsonl \
    --epochs 30 --lr 1e-3 --batch 3 --ckpt-dir runs_monai/round1 \
    --pretrained runs_monai/round0/best.pt
torchrun --standalone --nproc_per_node 3 scripts_monai/03_monai_pseudo_label.py \
    --ckpt runs_monai/round1/best.pt --pool data/manifests/unlabeled_pool.jsonl \
    --out data/pseudo_monai/round2 --already data/pseudo_monai/added_keys.jsonl
torchrun --standalone --nproc_per_node 3 scripts_monai/04_monai_eval_test.py \
    --rows data/manifests/split_test.jsonl --ckpt runs_monai/round1/best.pt \
    --out results/round1_test_monai.json