Back to search

Article

Tarsier2: Advancing Large Vision-Language Models from Detailed Video Descriptions to Comprehensive Video Understanding

2025-02-04

Abstract excerpt

We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior general video understanding capabilities. Tarsier2 achieves significant advancements through three key upgrades: (1) Scaling pre-training data from 11M to 40M video-text pairs, enriching both volume and diversity; (2) Performing fine-grained t...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
5d395ca1-cdba-589f-a65e-6f31fcce0994
DOI
10.32388/x26ilu
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Descriptions to Comprehensive Video UnderstandingDOI 10.32388/x26ilu
Select a neighboring publication to make it the new centre.