\
Insight
\
New Gartner® report — Reality Defender is named a Market Shaper in deepfake detection, as of June 2026.
Get the report\
Insight
\
Alex Lisle
CTO
A single deepfake attack now moves across three channels at once. A cloned voice opens the contact center call. A swapped face joins the follow-up video meeting. A forged image lands in the paperwork that closes the deal. Same operation, three modalities, one goal.
Gartner reported that 62% of organizations experienced at least one deepfake attack in the past 12 months. The sophisticated attacks rarely confine themselves to a single channel.
Consider the Arup case. In early 2024, a finance worker joined a video call with what looked like his CFO and several colleagues. He approved transfers worth about $25 million. He later learned that everyone on the call except him was synthetic. That's not one deepfake. That's multiple synthetic faces and voices, coordinated in a single live meeting.
The attacker's job is to move through whatever you're watching least.
A single-modality tool watches one channel well. Voice detection listens to audio. It doesn't see the video call that follows, and it doesn't read the image in the file. The moment an attack crosses from one channel to another, it crosses a seam.
In practice, nothing covers that crossing. The voice tool clears the call, the video goes to a different product or no product at all, and the coordinated attack lives in the space between them.
A seam is any point where one tool assumes another is covering the handoff, and attackers aim straight for it.
Plenty of tools now advertise more than one channel. Look closer and many are a voice-first product with a video model bolted on later, or two acquired engines wired together after the fact.
That architecture still has seams. The models were trained separately, they score separately, and they don’t correlate what they see across a single incident. Bolting a second model onto a voice-first product isn’t the same as building for all three from the start.
Adding a modality stitches channels together. Building multimodal removes the stitches.
There's a deeper reason the seams persist. For most vendors, detection is one feature inside a larger platform. It gets the attention a feature gets, and the modalities that don't drive the roadmap fall behind.
When detection is the entire company, the incentive flips. Every modality is core, because there's nothing else to sell. The coverage stays even because the business depends on it.
When detection is one feature among many, the seams are somebody else's problem.
Single-platform multimodal detection means one architecture analyzing voice, video, and image together, not three tools reporting separately.
Analyze voice, video, and image on one architecture, not three disconnected tools
Score every input in real time, before the decision or the transfer lands
Correlate signals across channels so a coordinated attack has nowhere to hide
Update as new generators ship, across every modality at the same time
Run detection-only, with nothing generative grading its own output
Single-platform multimodal detection closes the seams because there aren't any.
Reality Defender runs detection-only, multimodal analysis across voice, video, and image in real time, on a single platform. It's the difference between watching one door and watching the building.
We believe, that single architecture is also what the market is now recognizing. Reality Defender is positioned highest for Potential to Execute in the Gartner Emerging Market Quadrant for Deepfake Detection — Startup Vendors, as of June 2026.
We believe, that placement comes down to production. One platform covering voice, video, and image scales inside regulated banks and contact centers without the handoffs and integration debt that separate tools carry. It runs in real time, updates across every modality at once, and stays independent of whatever generated the media. Coverage with no seams is easier to operate than coverage stitched together after the fact.
No. A voice model analyzes audio and has no visibility into a video stream or an image. A video deepfake in the same attack passes untouched unless a system covers that modality too. Coordinated attacks exploit exactly that gap.
Multimodal detection analyzes more than one type of media — voice, video, and image — for signs of AI generation. Done well, it uses a single architecture that scores every channel and correlates the results, rather than separate tools that each see only their own channel.
Separate tools per channel create seams between them, and coordinated attacks target those seams. A single-platform approach that covers voice, video, and image together removes the handoffs an attacker would otherwise use.
Source: Gartner, Emerging Market Quadrant for Deepfake Detection — Startup Vendors, By Apeksha Kaushik, Alfredo Ramirez IV, Akif Khan, David Senf, 25 June 2026
Gartner Press Release, Gartner Survey Reveals GenAI Attacks Are on the Rise, September 22, 2025. https://www.gartner.com/en/newsroom/press-releases/2025-09-22-gartner-survey-reveals-generative-artificial-intelligence-attacks-are-on-the-rise
Gartner is a trademark of Gartner, Inc. and/or its affiliates.
Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose
\
Insights