Article
HEADHUNTER: Training-Free Annotated Dataset Synthesis via Self-Guided Diffusion Transformer Attention Head Selection
2026-06-29
Abstract excerpt
Pixel-level annotation remains a major bottleneck for semantic segmentation, motivating methods that synthesize image-label pairs directly from generative models. Prior synthetic dataset generators typically obtain pseudo-labels from cross-attention maps or learned decoders over generative features; however, recent text-to-image (T2I) models increasingly use multimodal diffusion transformers (MM-DiTs), where conce...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 7156debd-18e0-58b2-b2b9-f3d0ae31fad2
- DOI
- 10.20944/preprints202606.2132.v1
