Skip to main content

Video Detection

Learn how Reality Defender detects manipulated or AI-generated videos.

Written by Emily Essig

This article provides a high-level overview of video detection: what kinds of manipulation Reality Defender looks for, how full-frame and face-region analysis work together, what can affect confidence, and where to learn more about results.


How Reality Defender Detects Video Deepfakes

Video manipulation can take many forms. Common examples include:

  • Face swaps, where one person’s face is replaced with another person’s face

  • Facial reenactment, where one person’s expressions or movements are transferred onto another person

  • Broader frame-level generation or editing, where manipulation signals may appear in the body, background, scene, motion, or surrounding visual context

Reality Defender analyzes videos using a combination of full-frame and face-region detection. These approaches help cover a wider range of video content, including both face-forward videos and videos where a usable face may not be present.

Full-frame video analysis

Full-frame analysis evaluates the complete video frame, rather than only detected faces.

This can help identify manipulation signals that appear outside the face region, including:

  • Body or movement inconsistencies

  • Background or scene artifacts

  • Temporal or motion irregularities

  • Visual context that may indicate AI generation or manipulation

Full-frame analysis is especially useful when faces are absent, small, blurry, heavily compressed, partially occluded, or not consistently detectable.

Face-region video analysis

Face-region analysis evaluates detected faces across frames when usable face regions are available.

This can help identify manipulation artifacts in areas such as:

  • Facial texture

  • Lighting or edge consistency

  • Motion across frames

  • Facial expression or lip movement consistency

When a clear face is available, Reality Defender can use both face-region and full-frame signals as part of the overall analysis.


What Affects Video Detection Confidence

Detection confidence can vary based on the quality and characteristics of the video. Factors that may affect results include:

  • Video resolution and compression

  • Motion blur

  • Lighting conditions

  • Face size and visibility

  • Camera angle or pose

  • Whether a face is present and consistently detectable

  • The type and location of manipulation signals in the video

Because different videos contain different signals, results may vary by file and detection path.


Video Result Localization

Reality Defender may provide localization details for video results when localization data is available.

Localization can help show which portions of a video contributed to the result. Depending on the file and detection path, this may include segment-level indicators, detected face regions, or other visual markers in the product or API response.

Not every video will include the same localization details. For example, a video analyzed primarily through full-frame detection may return a result even when a usable detected face is not available.


YouTube Link Scans: Resolution Notes

When you upload a YouTube link in the Reality Defender web interface, the system downloads and analyzes the video at 360p by default to support faster, more consistent processing. Scanning at higher resolutions may not necessarily improve results.

Support for full 1080p+ scanning is planned in upcoming platform updates for enterprise customers that require maximum fidelity.


Related FAQs

Question

Answer

How are video deepfakes made?

Common methods include face swaps and facial reenactment, but manipulation can also involve broader frame-level generation or editing.

Does video detection require a face?

Not always. Full-frame analysis can help Reality Defender evaluate videos even when a usable face is not present, though videos with clear human faces generally provide the strongest signal.

What affects accuracy?

Resolution, compression, blur, lighting, face visibility, motion quality, and available visual context can all affect confidence.

Does the model use audio?

No — video and audio detectors run independently. A mismatched voiceover won’t trigger a false positive in the video detector.

Does the system provide localization?

Yes, when localization data is available. Localization details may vary depending on the file and detection path.

Are higher resolutions supported?

Scanning defaults to 360p for performance reasons. 1080p+ scanning is on the roadmap.

Where can I learn more about video results?

For help interpreting video results in the web app, see Understanding Video Results.

For API response fields, model names, and score behavior, see Video Detection: How Results Are Returned.

Did this answer your question?