Detecting a convincing deepfake has become a race against scale. Generative AI can now produce realistic video faster than any human review team could ever watch it, and the digital detectors built to keep up are themselves expensive, power-hungry, and surprisingly easy to sabotage. A team at the University of California, Los Angeles, has proposed a different approach: stop doing all the work in silicon and let light do part of the thinking.
Their new processor uses laser light to screen more than a dozen videos at once. In tests, it identified AI-generated footage with close to 98% accuracy while consuming a small fraction of the energy of a comparable digital model. The research, led by Professor Aydogan Ozcan and published in the journal eLight, points toward a future in which deepfake detection is faster, cheaper, and much harder to fool.
What UCLA Built: A Deepfake Detector That Runs on Light
Most deepfake detectors are purely digital. They pass a video through a large neural network, and a single analysis can require hundreds of billions of floating point operations. Because videos are usually processed one after another, the time and energy needed climb with every clip you add.
There is a second problem. Digital detectors are vulnerable to adversarial attacks, meaning that carefully crafted, barely visible alterations can make a fake sail through as authentic. Attackers who can obtain or reverse-engineer the model can design perturbations tailored to defeat it.
The UCLA system takes a hybrid path. A lightweight digital encoder does the initial work, compressing each video into a compact signature, and then a passive optical decoder performs the final, decisive computation using light itself. The team describes the design as a high-throughput, attack-resilient first layer of defense for screening enormous volumes of video.
How an Optical-Neural Processor Actually Works
The idea sounds exotic, but the pipeline breaks down into a few clear steps.
From Video Frames to a Phase Pattern
For each video, the encoder samples 12 frames and runs them through a face-detection stage and two parallel branches, one working in the spatial domain and one in the Fourier domain. Their features are fused and mapped into a two-dimensional phase pattern. Several videos are then aggregated into a single combined pattern, which is displayed on a programmable spatial light modulator.
Light Through a Passive Decoder
A coherent laser beam illuminates that pattern. The resulting wavefront travels through free space and, optionally, through additional phase-only diffractive layers before reaching paired detectors. Because light intensity cannot be negative, the system uses a differential scheme: it compares a positive and a negative detector for each video, and the difference between them produces the authenticity score. A stronger fake signal flags manipulated footage, while a stronger real signal clears it.
One elegant detail is that adding diffractive layers makes the decoder deeper and more capable without adding electrical power, because the layers are passive, static structures that shape light through diffraction. In the team’s tests, moving from zero to two diffractive layers on a harder dataset improved accuracy by roughly 6.8%.

The Numbers: 97.79% Accuracy Across 15 Videos at Once
On the Celeb-DF benchmark, a widely used set of face-swap deepfakes, the processor examined 15 videos in a single optical pass. Its average experimental results were striking.
| Metric | Result |
|---|---|
| Detection accuracy | 97.79% |
| Sensitivity (fake videos caught) | 99.86% |
| Specificity (real videos cleared) | 95.72% |
| False-negative rate | about 0.14% |
| Videos processed per optical pass | 15 (up to 18 demonstrated) |
Sensitivity measures how often the system correctly flags a fake, which is the number that matters most for a screening tool, since it barely lets a manipulated clip slip through. For context, human sensitivity on the same set of fake videos can be as low as about 21%.
Pushing the hardware to process 18 videos per pass reduced accuracy to 96.13%, a dip the researchers attribute to more optical cross-talk between channels. They conclude that the 15-video configuration offers the best balance.
Why Light Beats Silicon: Speed, Energy, and Parallelism
The real appeal of optical computing is not only accuracy. It is that throughput, energy efficiency, and adversarial robustness, three properties that are normally difficult to achieve at the same time, can improve together.
The Energy Math
The optical decoder sits on what engineers call the Pareto-optimal frontier of the accuracy-energy trade-off. It matches or beats far heavier digital models while using a fraction of the power.
| Decoder | Energy per video |
|---|---|
| Optical, 15 videos per pass (2 diffractive layers) | 1.38 to 4.11 millijoules |
| Optical, 18 videos per pass | about 1.15 to 3.43 millijoules |
| Optical, 10-megapixel modulator (about 120 videos) | about 0.18 to 0.66 millijoules |
| Digital Deep CNN baseline | 30.44 millijoules |
Put in system terms, the optical decoder drew roughly 3.5 to 10.5 watts, compared with about 77 watts for a performance-matched digital decoder running at the same throughput, a reduction of about 86% to 96% at the decoder level. End-to-end savings are smaller, around 13%, because the digital encoder remains the dominant energy cost. With a lighter encoder, the full pipeline’s energy use fell by roughly 38% to 42%.

Throughput and Latency
Because a single optical frame classifies all 15 videos at once, the decoder is no longer the bottleneck. The digital encoder processes about 2,542 videos per second, while the optical decoder can handle roughly 2,700 videos per second on a standard spatial light modulator, rising to about 21,600 videos per second with a faster modulator. Inference latency ranged from about 5.56 milliseconds per batch down to about 0.69 milliseconds with higher-speed hardware.
Critically, the approach scales with aperture size: a larger spatial light modulator allows more videos to be processed per pass, which also drives the energy cost per video down.
Hard to Fool: Security Built Into the Hardware
Because part of the computation happens through the physical diffraction of light, the model’s key parameters are effectively baked into the hardware itself. An attacker cannot easily read or reproduce them, which makes reverse-engineering the detector or crafting targeted adversarial changes to evade it considerably harder.
In testing, the processor showed several security advantages at once:
- It resists black-box adversarial attacks and carries inherent protection against white-box attacks.
- It keeps working when videos are degraded by image noise, blur, JPEG compression, and physical misalignment.
- Under heavy additive noise, it holds above 95% sensitivity, while lightweight digital decoders collapse to near zero.
Testing Against VEO-3 and Newer AI Video
Older detectors can be blind-sided by the newest generative models, which produce footage that lacks the telltale artifacts of earlier fakes. So the team tested two tougher challenges.
On DeepSpeak, a dataset in which identities are matched by face and voice before swapping, a basic version of the processor scored about 89%. Adding two diffractive layers lifted accuracy to about 96% and pushed specificity from 86% to nearly 95%.
The researchers then fed the system videos made by Google’s VEO-3 model. With only minimal fine-tuning, it reached 94.80% accuracy and 97.61% sensitivity on footage it had never seen, suggesting the approach can adapt as generators improve.
Why Deepfake Detection Is So Hard
It helps to understand why this problem resists easy answers. Detectors tend to perform well on manipulations they have seen before and stumble on new ones, a weakness known as the generalization gap. Watermarks and provenance labels help, but they can be stripped by screenshots, compression, or a simple re-export. We explored those limits and how chatbots often misjudge images in our analysis of why deepfake detectors and provenance signals fail to verify synthetic media.
No single tool solves the problem. The UCLA team’s own framing is telling: their processor is designed as a highly sensitive first stage, not a final verdict. Massive volumes of video pass through the optical screen, and anything flagged as suspicious is handed to heavier digital models for closer inspection.
What This Means for Moderation, Media Trust, and You
The researchers see the technology fitting into settings where platforms must triage millions of clips and can least afford to miss a well-made fake:
- Large-scale content moderation, screening huge volumes of uploaded video.
- Media authentication for newsrooms and fact-checkers.
- Security-critical screening where a missed fake carries real consequences.
- Any pipeline that needs a fast, low-power first pass before deeper digital analysis.

It is a lab prototype rather than a product you can download today, and it still needs a digital companion to confirm its suspicions. But the underlying idea, that embedding computation in physical processes can unlock new trade-offs between performance, efficiency, and security, extends well beyond deepfakes.
Frequently Asked Questions
How does a light-powered AI detect deepfakes?
A digital encoder converts each video into a phase pattern displayed on a spatial light modulator. Laser light then propagates through a passive optical decoder, and paired detectors read the resulting intensity to produce an authenticity score.
How accurate is the UCLA optical deepfake detector?
On the Celeb-DF benchmark, it reached 97.79% accuracy, 99.86% sensitivity, and 95.72% specificity, with a false-negative rate of about 0.14%, while processing 15 videos per optical pass.
Is optical computing more energy-efficient than digital AI?
For this task, yes. The optical decoder used roughly 1.4 to 4.1 millijoules per video versus about 30 millijoules for a comparable digital decoder, an 86% to 96% reduction at the decoder level.
Can this detector be fooled by adversarial attacks?
It resisted black-box attacks and carries inherent protection against white-box attacks, because its optical parameters are physically embedded in the hardware and are difficult to measure or reproduce.
Can it detect videos made by Google VEO-3?
Yes. After minimal fine-tuning, it achieved 94.80% accuracy and 97.61% sensitivity on previously unseen VEO-3 videos.
Can I use this deepfake detector, and when will it be available?
Not yet. It is a research prototype, and no consumer release has been announced. The researchers position it as a first-stage screening layer that would work alongside existing digital detectors.
The Bottom Line
UCLA’s optical processor shows that moving part of an AI’s work into light can deliver accuracy and speed that digital-only systems struggle to match while using far less energy and resisting sabotage. Deepfake detection will always be an arms race, but this is a rare case where the defender’s hardware has a natural advantage. The full study, “Scalable, energy-efficient optical-neural architecture for multiplexed deepfake video detection,” is published in eLight and was summarized in a ScienceDaily release.
