Introducing DETECT-World: The First World Model for Deepfake Detection
Blog post from Resemble AI
Resemble AI introduces DETECT-World, a third-generation multimodal deepfake detector designed to identify manipulated audio, images, and video not only through signatures of known generators but also through learned physical-consistency checks involving lighting, geometry, motion, and audio-visual synchronization. Built on Meta AI’s V-JEPA 2 video foundation model and fine-tuned on hundreds of manipulation sources, it produces an overall manipulation probability alongside per-frame anomaly scores and spatial heatmaps. The company says adaptive training emphasizes sources the model currently misclassifies and applies real-world compression and platform-degradation augmentations equally to authentic and fake media. In internal testing, DETECT-World reportedly reached 95.8% image accuracy and 98.2% video accuracy, while its audio system achieved 99.47% accuracy on the externally validated Podonos benchmark, covering 54 languages and reducing latency to 399 milliseconds. Resemble AI reports that the model identified an unseen real-time face-swap tool with roughly 95% accuracy, though it notes that reported scores are probabilities rather than proof, long videos receive reduced-resolution analysis and no heatmaps beyond 60 seconds, and high-stakes decisions should include human review. DETECT-World is available through Resemble Detect’s API, including streaming, batch, on-premises, and air-gapped deployment options.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 8 | 4,432 | 1,050 | 222 | -31% |
| AI Model Fine-tuning | 2 | 554 | 154 | 60 | -43% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.