Article
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Descriptions to Comprehensive Video Understanding
2025-02-04
Abstract excerpt
We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior general video understanding capabilities. Tarsier2 achieves significant advancements through three key upgrades: (1) Scaling pre-training data from 11M to 40M video-text pairs, enriching both volume and diversity; (2) Performing fine-grained t...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 5d395ca1-cdba-589f-a65e-6f31fcce0994
- DOI
- 10.32388/x26ilu
