Article
CAT: Content-Adaptive Image Tokenization
2025-01-17
Abstract excerpt
Most existing image tokenizers encode images into a fixed number of tokens or patches, overlooking the inherent variability in image complexity. To address this, we introduce Content-Adaptive Tokenizer (CAT), which dynamically adjusts representation capacity based on the image content and encodes simpler images into fewer tokens. We design a caption-based evaluation system that leverages large language models (LLM...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 804f7323-86a7-5915-bd05-610ef82867fd
- DOI
- 10.32388/wcbnq2
