
Researchers at UCLA have developed an optical-neural processor that uses light to detect deepfake videos quickly and efficiently. Unlike conventional digital systems that typically process videos sequentially, the new architecture can examine 15 or more video streams simultaneously during a single optical pass, tells Science Daily.
The hybrid system begins with a lightweight digital encoder that extracts spatial, spectral, and temporal features from each video. These features are converted into phase patterns and displayed on a programmable spatial light modulator. The resulting wavefront passes through a passive optical decoder, where diffraction performs part of the computation. Optical detectors then generate an authenticity score for each video.
In experiments using 15 Celeb-DF videos at once, the system achieved average detection accuracy of 97.79%, with 99.86% sensitivity and 95.72% specificity. Its false-negative rate was about 0.14%, an important result for a screening system intended to prevent manipulated videos from escaping detection. When researchers increased capacity to 18 simultaneous videos, accuracy remained at 96.13%.
Additional passive diffractive layers improved performance on more difficult manipulations by about 6.8% without substantially increasing inference latency or electrical energy requirements. Because these optical layers perform computation through light diffraction, they require no additional electrical power during inference.
The researchers also tested the system on previously unseen videos generated by Google Veo 3. With minimal fine-tuning, it achieved 94.80% accuracy and 97.61% sensitivity, suggesting that the approach can adapt to newer generative AI models.
Another advantage is resistance to adversarial attacks. Key model parameters are physically embedded in the optical hardware, making them harder to reproduce or reverse engineer.
Rather than replacing digital deepfake detectors, the researchers envision optical AI as a high-throughput first screening stage. Suspicious content could then be forwarded to more computationally intensive digital systems for detailed analysis.
