Article
HEADHUNTER: Training-Free Annotated Dataset Synthesis via Self-Guided Diffusion Transformer Attention Head Selection
2026-07-01
Abstract excerpt
Pixel-level annotation remains a major bottleneck for semantic segmentation, motivating methods that synthesize image-label pairs directly from generative models. Prior synthetic dataset generators typically obtain pseudo-labels from cross-attention maps or learned decoders over generative features; however, recent text-to-image (T2I) models increasingly use multimodal diffusion transformers (MM-DiTs), where conce...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 27463918-b9f2-5f66-9c8d-d7f5b4d2011b
- DOI
- 10.20944/preprints202606.2132.v2
