SEER: Skill-Evolving Grounded Reasoning for Free-Text Promptable 3D Medical Image Segmentation

1Fudan University, 2University of Science and Technology of China
Overview of the SEER framework

Different words. Same clinical intent. Stable 3D masks.
SEER makes free-text promptable 3D medical segmentation robust by grounding clinical free-text prompts in patient anatomy and evolving reusable reasoning skills.

Abstract

Free-text promptable 3D medical image segmentation offers an intuitive and clinically flexible interaction paradigm. However, current methods are highly sensitive to linguistic variability: minor changes in phrasing can cause substantial performance degradation despite identical clinical intent. Existing approaches attempt to improve robustness through stronger vision-language fusion or larger vocabularies, yet they lack mechanisms to consistently align ambiguous free-form expressions with anatomically grounded representations.

We propose Skill-Evolving groundEd Reasoning (SEER), a novel framework for free-text promptable 3D medical image segmentation that explicitly bridges linguistic variability and anatomical precision through a reasoning-driven design. First, we curate the SEER-Trace dataset, which pairs raw clinical requests with image-grounded, skill-tagged reasoning traces, establishing a reproducible benchmark. Second, SEER constructs an evidence-aligned target representation via a vision-language reasoning chain that verifies clinical intent against image-derived anatomical evidence, thereby enforcing semantic consistency before voxel-level decoding. Third, we introduce SEER-Loop, a dynamic skill-evolving strategy that distills high-reward reasoning trajectories into reusable skill artifacts and progressively integrates them into subsequent inference, enabling structured self-refinement and improved robustness to diverse linguistic expressions.

Extensive experiments demonstrate superior performance of SEER over state-of-the-art baselines. Under linguistic perturbations, SEER reduces performance variance by 81.94% and improves worst-case Dice by 18.60%.

Why SEER?

Same clinical intent ≠ Same segmentation

Prior Methods under Free-text Prompts

  • Label-like prompts required
  • Fragile to noisy clinical language
  • Limited visual grounding and intent reasoning
Different descriptions of the right ilium cause prior predictions to drift or collapse, while SEER produces consistent, on-target masks

How SEER Works

01 SEER-Trace

SEER-Trace pairs clinical prompts with 22,330 image-grounded, skill-tagged reasoning traces to train the reasoning policy.

02 Grounded Reasoning

SEER uses patient-specific anatomy to interpret clinical intent.

Image evidence resolves LV as the left ventricle through evidence, rationale, and answer

03 Evolving Reusable Skills

SEER-Loop distills high-reward reasoning into reusable skills, refines SEER-Bank, and retrieves relevant skills for later reasoning rounds.

SEER-Loop distills high-reward reasoning, refines SEER-Bank, and retrieves skills for later rounds

Experiments and Findings

We evaluate SEER on two out-of-distribution benchmarks and compare it with strong 3D medical segmentation baselines under both native label prompting and realistic free-text clinical prompting.

BrainMetShare MRI

Partial OOD: unseen sources and target labels within the brain anatomical domain.

BrainMetShare: standard deviation decreases by 50.1 percent and worst-case Dice increases by 10.1 percent versus VoxTell

PENGWIN CT

Strict OOD: pelvic anatomy and target labels absent from SEER-Trace coverage.

PENGWIN: standard deviation decreases by 86.9 percent and worst-case Dice increases by 20.3 percent versus VoxTell
Full Quantitative Results

SEER improves free-text Dice and robustness across both benchmarks.

Dataset Method Label Prompting Free-text Prompting
Dice ↑ Dice ↑ Worst Dice ↑ Std. ↓
BrainMetShare SAT 22.16 0.69 0.00 2.53
BiomedParseV2 18.66 2.53 0.00 7.27
Text3DSAM 0.10 0.41 0.00 0.93
MedSAM3 11.33 16.62 10.56 5.17
VoxTell 48.19 52.15 46.71 3.35
SEER (Ours) 51.70 53.83 51.44 1.67
PENGWIN SAT 96.05 0.01 0.00 0.13
BiomedParseV2 1.35 8.53 0.00 7.50
Text3DSAM 24.75 0.01 0.00 0.16
MedSAM3 18.26 5.75 3.67 6.40
VoxTell 97.59 92.26 79.34 7.49
SEER (Ours) 97.56 97.39 95.47 0.98

Ablation Study

ConfigurationDice ↑Worst Dice ↑Std. ↓
Baseline92.2679.347.49
+ Vanilla VLM84.8461.9014.15
+ Fine-tuned VLM w/ Grounded Reasoning95.9288.273.84
+ SEER-Loop (Ours)97.3995.470.98

Grounded reasoning and SEER-Loop improve accuracy and robustness, while naive VLM-based text parsing degrades performance.

Qualitative Comparisons

Across typo noise, spatial specifiers, and clinical orders, SEER produces on-target masks grounded in the image.

Qualitative comparison across typo noise, spatial specifiers, and clinical orders

Next Steps

Stay tuned for SEER V2Coming soon

Towards Reliable Medical AI for Real-World Clinical Workflows!

BibTeX

@article{zhang2026skill,
  title={Skill-Evolving Grounded Reasoning for Free-Text Promptable 3D Medical Image Segmentation},
  author={Zhang, Tongrui and Wang, Chenhui and Li, Yongming and Chen, Zhihao and Zhan, Xufeng and Shan, Hongming},
  journal={arXiv preprint arXiv:2603.08215},
  year={2026}
}