TransVLM: Detecting Any Shot Transition with Vision-Language Models
Blog post from Hugging Face
TransVLM is a vision-language approach to Shot Transition Detection that identifies complete temporal segments for cuts, dissolves, fades, wipes, and effects rather than treating transitions as isolated frame-level boundary points. Developed by HeyGen Research and the University of Melbourne and accepted at ECCV 2026, it combines color frames with optical-flow information at the vision encoder level, enabling stronger detection of fine-grained temporal changes without increasing language-model token counts. The system also uses synthetic training data generated from 59 FFmpeg transition effects, re-annotated public data, and sliding-window inference with temporal merging to process arbitrarily long videos. Its accompanying benchmark contains 5,215 videos, more than 100 hours of footage, and 45,239 segment-labeled transitions. TransVLM reported segment-level F1 scores of 78.3% on public data and 89.5% on synthetic data, with a 0.11-second boundary error and practical real-time performance, outperforming conventional shot-boundary detectors and general-purpose vision-language models, particularly on gradual and complex transitions.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.