MTOR study combines captions and temporal patterns to detect AI video
Researchers test synthetic video detection across 46 generator variants. Code and model releases remain pending.

An October 5 preprint introduces MTOR, a research detector that combines video captions with patterns in how visual features change between frames.
The system generates its own description of each video, then combines that text with visual representations. A second component looks for unusually persistent features and limited variation over time.
The authors evaluated five benchmarks covering 46 generator variants against 16 comparison methods. Reported accuracy ranged from 85.18% to 95.48% across the benchmarks. Adding temporal signals improved average performance, although some results declined, including on VidProM.
These are benchmark results, not a guarantee about any individual clip. The official repository still lists inference code and trained models as forthcoming. ByteForward has not independently reproduced the evaluation.
Archival photograph Desk Setup by Luke Chesser, dated October 30, 2014, via Wikimedia Commons under CC0 1.0. The image illustrates a computer workspace and does not show MTOR, a detected video or the reported experiment. Existing WebP rendition.



