Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

TransVLM: Detecting Any Shot Transition with Vision-Language Models

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Ce Chen
Word Count
1,358
Company Posts That Month
66
Language
-
Hacker News Points
-
Post removed?
No
Summary

TransVLM is a vision-language approach to Shot Transition Detection that identifies complete temporal segments for cuts, dissolves, fades, wipes, and effects rather than treating transitions as isolated frame-level boundary points. Developed by HeyGen Research and the University of Melbourne and accepted at ECCV 2026, it combines color frames with optical-flow information at the vision encoder level, enabling stronger detection of fine-grained temporal changes without increasing language-model token counts. The system also uses synthetic training data generated from 59 FFmpeg transition effects, re-annotated public data, and sliding-window inference with temporal merging to process arbitrarily long videos. Its accompanying benchmark contains 5,215 videos, more than 100 hours of footage, and 45,239 segment-labeled transitions. TransVLM reported segment-level F1 scores of 78.3% on public data and 89.5% on synthetic data, with a 0.11-second boundary error and practical real-time performance, outperforming conventional shot-boundary detectors and general-purpose vision-language models, particularly on gradual and complex transitions.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.