Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Learning to Merge Superpixels

Sponsored · Your Podcast. Everywhere. Effortlessly. Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.

Learning to Merge Superpixels

EANN 2026

Avatar for Olivier Lézoray

Olivier Lézoray

July 23, 2026

More Decks by Olivier Lézoray

Other Decks in Research

Transcript

  1. LEARNING TO MERGE SUPERPIXELS Olivier LÉZORAY Université Caen Normandie, ENSICAEN,

    Normandie Univ, GREYC UMR 6072, Caen, FRANCE [email protected] https://lezoray.users.greyc.fr
  2. Outline 1. Motivation 2. Proposed Method (SNSM) 3. Experiments &

    Results 4. Conclusion O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels 2 / 15
  3. Motivation The superpixel principle ▶ Superpixels → compact image representation

    ▶ Produces strong over-segmentation ▶ Objects are split into many superpixels Original image (pixels) Key question SLIC Can we learn to intelligently simplify this representation? ∼768 SP SNSM (ours) Our answer Learn a data-driven criterion to decide whether two adjacent superpixels should be merged. O. Lézoray et al.– EAAAI 2026 SLIC superpixels (over-segmentation) Learning to merge superpixels Merged regions (simplified) ∼60 SP 3 / 15
  4. Limits of Existing Approaches Method Limitation Heuristics [Barcelos, 2024] Manual

    thresholds, poor generalization across textures Square patches [Huang, 2024] Pixels outside both superpixels are included Average pooling [Liu, 2018] Discards internal feature diversity of superpixels SNSM (ours) NetVLAD + siamese, fully learned ▶ Existing methods rely on hand-crafted rules or lose structural information ▶ Our approach replaces heuristics with learned pairwise affinities Region Adjacency Graph: green = mergeable, red = not mergeable O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels 4 / 15
  5. Outline 1. Motivation 2. Proposed Method (SNSM) 3. Experiments &

    Results 4. Conclusion O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels 5 / 15
  6. SNSM — Architecture Overview Image I ∈ R3×H×W VGG16 (first

    layers, no pooling) X = ΦVGG (I) ∈ RC×H×W C = 64 Gather pixels of SA FA ∈ RC×NA Siamese towers with shared weights NetVLAD (K=16) vA ∈ RK×C MLP [K × C → d → d] e(A) ∈ Rd , d=256 Gather pixels of SB FB ∈ RC×NB NetVLAD (K=16) vB ∈ RK×C MLP [K × C → d → d] e(B) ∈ Rd , d=256 |−| ⊙ Concatenate zAB = [e(A); e(B); |e(A)−e(B)|; e(A) ⊙ e(B)] zAB ∈ R4d Decision MLP [ 4d → 256 → 64 → 1 ] logit lAB = g(zAB ) ∈ R σ 1. Feature extraction VGG16 (frozen) before first max-pool 2. Aggregation NetVLAD per superpixel variable size → fixed SP feature vector 3. Enriched pairwise vector Add contrast and agreement between SP features 4. Decision Pairwise MLP merge probability ∈ [0, 1] Merge probability p(A ↔ B) = σ(lAB ) O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels 6 / 15
  7. VGG Feature Extraction + NetVLAD Aggregation VGG16 — dense feature

    extraction Image I (3 × H × W ) ▶ Truncated before first max-pooling VGG16 (frozen) ΦVGG ▶ Preserves high spatial resolution ▶ C = 64 channels per pixel, weights frozen Feature map X C = 64, H ×W ▶ Low-level cues: edges, textures Superpixel mask FS ∈ R64×NS NetVLAD — fixed-size aggregation ▶ Learns K = 16 visual cluster centroids {ck } P ▶ Per-cluster residual aggregation: rk = αkp (xp − ck ) p∈PS ▶ Variable size superpixel → fixed K ×C vector ▶ Captures intra-superpixel feature distribution O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels NetVLAD (K = 16) vS ∈ RK×C Embedding e(S) ∈ R256 ø Key: variable-size input → fixed d = 256 embedding 7 / 15
  8. Pairwise Representation & Decision Head Enriched pairwise vector e(A) ∈

    R256 Given embeddings e(A) and e(B):   zAB = e(A) ; e(B) ; |e(A)−e(B)| ; e(A) ⊙ e(B) e(B) ∈ R256 |e(A)−e(B)| ▶ e(A), e(B): individual appearance ▶ |e(A)−e(B)|: contrast between superpixels ▶ e(A) ⊙ e(B): feature agreement zAB ∈ R1024 Decision MLP : estimates merge probability FC+LN+ReLU FC+ReLU e(A) ⊙ e(B) FC 4d −−−−−−−→ 256 −−−−−→ 64 −→ 1 P (yAB = 1) = σ(ℓAB ) ∈ [0, 1] FC 256 + LN + ReLU FC 64 + ReLU P (merge) ∈ [0, 1] Layer Normalization + Dropout (0.3) for regularization O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels 8 / 15
  9. Outline 1. Motivation 2. Proposed Method (SNSM) 3. Experiments &

    Results 4. Conclusion O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels 9 / 15
  10. Dataset & Training Setup ISEG Dataset Training details ▶ 151

    natural images with binary ground-truth segmentations ▶ Challenging: wide background appearance variability ▶ Superpixel pairs labeled via majority voting from GT ▶ Loss: Binary Cross-Entropy (with logits) ▶ Epochs: 50, early stopping on balanced val. set ▶ Learning rate: 10−4 ▶ Split: 80% train / 20% val. Class imbalance & rebalancing ▶ 299 566 pairs total: 94.5% mergeable (positive) ▶ Strong imbalance → random undersampling of positives ▶ Balanced set: 16 375 pos. / 16 375 neg. ▶ No class-weighted loss O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels Why uniform sampling? Class-weighted losses would shift the predicted probabilities, making threshold δ unreliable for hierarchical merging. Balanced sampling → preserves probability calibration for threshold-based merging 10 / 15
  11. Ablation Study t-SNE of MLP hidden layer (pair-level) Variant F1

    (bal.) F1 (raw) WN — Full repr. NN — Full repr. (ours) NN — w/o product (⊙) NN — w/o difference (| · |) NN — concat only Average Pooling 0.8243 0.8503 0.8389 0.8291 0.8344 0.8301 0.8959 0.9161 0.9066 0.8795 0.8866 0.9053 Negative pairs Positive pairs 40 20 0 20 40 60 75 50 25 0 25 50 75 t-SNE of MLP embedding u1 • Positive (mergeable) • Negative ▶ Classical NetVLAD normalization (WN) is damaging (−2.6 pts F1) ▶ Full representation > all reduced variants Distribution of predicted probabilities positives negatives 1750 1500 1250 1000 750 ▶ NetVLAD > average pooling (+1.1 pts F1) 500 250 0 0.0 0.2 0.4 0.6 0.8 1.0 Predicted probability distribution O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels 11 / 15
  12. Hierarchical Greedy Merging & Results Image SLIC Merged (SNSM) Merging

    strategy 1. Predict merge prob. for all adjacent pairs 2. Sort pairs by decreasing confidence 3. Greedy merges via Union-Find 4. Threshold δ = 0.4649 (F1-optimized) 5. Stop when confidence < δ 768 → 60 superpixels (−92%) Dice IoU # SP SLIC Merged 0.949 0.905 ≈768 0.913 0.850 ≈60 Moderate quality loss for a drastic simplification O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels 12 / 15
  13. Outline 1. Motivation 2. Proposed Method (SNSM) 3. Experiments &

    Results 4. Conclusion O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels 13 / 15
  14. Conclusion & Perspectives Summary — SNSM Perspectives ▶ VGG +

    NetVLAD: fixed-size embeddings from irregular superpixels ▶ Siamese + enriched pairwise representation : robust merge prediction ▶ Hierarchical greedy merging: 768 → 60 regions, Dice 0.95 → 0.91 Takeaway Learning to merge superpixels is an effective way to simplify region-based representations for downstream vision tasks. O. Lézoray et al.– EAAAI 2026 Learning to merge superpixels ▶ Generalization to other datasets ▶ More powerful backbone (ViT, DINOv2) ▶ Evaluation on downstream tasks (semantic segmentation, object recognition) ▶ Use the merge probabilities to weight edges of a GNN and perform the merging 14 / 15