Semantic Feature Modulation for Mammographic Lesion Classification

Shahar Mahpod and Gil Ben Artzi
Ariel University, Israel

Oral - International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), 2026

Presentation

Poster

Paper

Github


Dedicated MLP modulation networks map each BI-RADS descriptor (breast density, mass shape, mass margins, ...) to per-sample (γ, β) parameters, computed from that sample's descriptors, that replace the static affine parameters of selected Layer Normalization layers in a ConvNeXt backbone. Descriptors are injected either in block mode (each descriptor modulates consecutive layers) or in interleaved mode (descriptors alternate across layers within a stage).


Highlights


Abstract

Most mammography AI systems predict benign versus malignant from the image alone, ignoring the structured reasoning radiologists perform through the BI-RADS descriptor lexicon. We introduce a conditional modulation framework that dynamically shifts intermediate feature statistics based on BI-RADS descriptors. Rather than considering descriptors as static labels, our method produces sample-specific modulation parameters from each sample's descriptors; because these parameters act on image-dependent intermediate features, the same attribute affects the representation differently depending on the image, capturing context-dependent diagnostic meaning. We show that our approach is highly effective. Across CBIS-DDSM and CMMD, our model improves AUC by up to +8.3% and F1 by up to +5.0% over the image-only baseline. By representing descriptors as conditional continuous embeddings, our framework enables exact interpolation between attributes within the same BI-RADS category; interpolation experiments confirm monotonic risk paths aligned with clinical severity, despite training only on hard categorical labels. Our model enables a post-detection decision-support mode using probabilistic outputs from emerging BI-RADS prediction models. Perturbation analysis confirms approximately 25% of joint-setting predictions change when all descriptors are removed.

Method

We adopt ConvNeXt-Base as backbone. ConvNeXt uses Layer Normalization (LN) throughout, and since LN operates per-sample, its affine parameters are natural control points for sample-specific conditioning. We replace the static (γ, β) of selected LN layers with parameters predicted from clinical descriptors: for each descriptor type t, a dedicated two-layer MLP g(t) maps the descriptor embedding z(t) to modulation parameters,

CLN(h, z(t)) = γ(t)(z(t)) ⊙ h − μ(h)σ(h) + β(t)(z(t))

Global descriptors (e.g., breast density) modulate earlier layers and lesion-level descriptors (e.g., margins, morphology) modulate later layers. Training uses hard one-hot descriptor labels, yet at inference a descriptor may be any convex combination z(t) = (1−α) zA(t) + α zB(t), allowing exact interpolation between adjacent BI-RADS attributes without retraining.


Results

All models share the same ConvNeXt-Base backbone and training protocol. Base: image only. Concat: descriptor embeddings concatenated with pooled features. SE-Gate: SE-style channel gating conditioned on descriptors. MTL: multi-task learning with auxiliary descriptor prediction heads. Cond-LN Block / Inter: ours, with block and interleaved conditioning. Values are mean ± std (%) over 5 runs; best per setting in bold.

CBIS-DDSM

ModelAccAUCF1Rec
MassBase80.3±1.188.7±0.878.5±1.479.1±2.2
Concat79.5±2.188.3±0.777.5±2.677.4±3.5
SE-Gate80.7±1.487.8±0.578.4±1.777.1±2.2
MTL81.2±1.987.8±1.279.4±2.279.6±2.4
Cond-LN Block83.4±1.589.9±0.482.1±1.684.1±3.3
Cond-LN Inter84.7±0.990.9±1.383.0±0.981.8±2.7
CalcBase80.0±1.286.4±0.869.1±2.272.9±3.9
Concat79.1±1.886.5±1.167.7±2.571.4±2.6
SE-Gate79.5±0.886.6±0.768.9±1.074.0±1.6
MTL79.4±1.486.7±0.968.5±2.273.1±2.8
Cond-LN Block80.6±1.688.8±0.770.9±2.377.1±2.5
Cond-LN Inter82.8±1.190.9±0.474.1±0.780.2±4.0
BothBase78.8±0.686.9±0.172.3±1.173.9±2.0
Concat78.9±1.187.1±0.772.8±1.175.2±1.1
SE-Gate79.4±1.487.0±0.773.2±1.675.2±1.0
MTL80.2±0.888.2±0.474.7±0.978.2±1.5
Cond-LN Block81.2±1.590.3±0.374.2±2.472.2±3.8
Cond-LN Inter81.2±0.790.3±0.475.0±1.775.4±4.3

CMMD

ModelAccAUCF1Rec
MassBase85.9±0.786.9±0.391.8±0.594.5±1.7
Concat86.8±0.885.4±3.792.3±0.594.8±1.1
SE-Gate86.8±0.886.8±1.492.3±0.595.1±0.6
MTL86.9±1.285.4±1.692.3±0.794.4±1.4
Cond-LN Block86.3±1.485.9±1.592.0±0.894.8±1.1
Cond-LN Inter86.6±0.786.3±0.692.1±0.494.3±0.7
CalcBase86.5±2.083.3±1.792.5±1.195.3±1.0
Concat86.5±1.583.2±2.092.5±0.894.6±1.2
SE-Gate87.6±1.085.6±1.193.1±0.695.7±0.7
MTL88.0±0.982.8±1.593.4±0.596.2±1.0
Cond-LN Block87.3±2.486.4±2.792.8±1.494.2±3.0
Cond-LN Inter89.7±1.791.6±1.194.1±1.093.4±3.2
BothBase85.8±0.887.3±1.391.8±0.594.2±1.0
Concat85.5±0.685.8±1.291.7±0.393.9±0.9
SE-Gate86.5±0.787.3±0.792.2±0.494.0±0.8
MTL86.5±0.387.6±0.892.2±0.293.7±0.3
Cond-LN Block88.7±0.590.2±0.493.6±0.397.2±0.6
Cond-LN Inter88.1±0.591.7±0.893.2±0.396.6±0.4

Cond-LN Inter achieves the highest (or tied-highest) AUC in 5 of 6 dataset × setting combinations. CMMD-mass illustrates a regime where descriptor headroom is limited; the subset's imbalanced malignant prior may further constrain how much descriptor signal can add.


Descriptor Interpolation and Monotonicity

We sweep descriptor inputs as convex combinations z = (1−α) zA + α zB, α ∈ {0, 0.1, ..., 1}, across all CBIS-DDSM test images, regardless of their original descriptor, pathology, or lesion type. All adjacent-pair interpolation curves increase monotonically, and the ordering aligns with clinical knowledge: irregular shapes and spiculated margins are the strongest malignancy indicators. The model was trained only on hard one-hot labels and never saw soft inputs.

Mass shape

Mass margins

Mean malignancy probability rises from 0.29 (ROUND) to 0.55 (IRREGULAR) for mass shape, and from 0.22 (CIRCUMSCRIBED) to 0.69 (SPICULATED) for mass margins.


Descriptor Sensitivity

Zeroing individual descriptors on the CBIS-DDSM joint setting shows that the lesion-specific attributes, mass margins and shape, produce the largest probability shifts (left) and prediction flip rates (right).

When all descriptors are removed, accuracy drops from 81.0% to 68.0% and about 25% of predictions change.


Qualitative Analysis

Grad-CAM++ at multiple ConvNeXt depths for CBIS-DDSM mass cases. The benign case (top) exhibits diffuse patterns; the malignant case (bottom) shows increasingly localized activations, with deep-stage activations aligning with the lesion boundary.





Paper


Semantic Feature Modulation for Mammographic Lesion Classification, Shahar Mahpod, Gil Ben Artzi
MICCAI 2026
PDF  |  Code

@inproceedings{mahpod2026semantic,
  title     = {Semantic Feature Modulation for Mammographic Lesion Classification},
  author    = {Mahpod, Shahar and Ben Artzi, Gil},
  booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
  year      = {2026}
}


Acknowledgements

We acknowledge the Ariel HPC Center at Ariel University for providing computing resources. The website template is from here.