This is the complete reference for the video rule types in Cloudinary Moderation. Video rules analyze a video over its whole duration - sampling frames from the footage, and, in some cases, the audio - and give the video a single overall result.
All rules share the common settings described in "Create a rule" (name, description, Review threshold, optional rejection reason). Video rules add one recurring setting worth understanding up front:
Frames per second to sample - how many frames per second of footage are analyzed. Higher sampling catches brief appearances more reliably but takes longer to process. The defaults are a good balance for most content.
Video Content Detection
What it checks: whether specific objects, brands, or visual concepts appear anywhere in the video's frames.
Settings:
Labels to detect - a comma-separated list of what to look for, for example: "Bottle, Car, Happy." Leave empty to report everything detected. The Validate and Extend Search Terms AI assist can check and enrich your list.
What do these labels represent? - choose Unwanted content (the video fails if any label is found - the default) or Desired content (the video passes only if the labels are found).
Frames per second to sample
When to use it: spotting competitor products, alcohol, or other unwanted items in footage - or confirming your product actually appears in creator-submitted videos.
Video On-Screen Text Check
What it checks: the text that appears on screen (captions, overlays, signs, packaging), evaluated against a yes/no question you write.
Settings:
Question - a yes/no question about the on-screen text, for example: "Does the video contain profanity or offensive language?"
Pass this rule when the answer is Yes - choose Pass on Yes or Pass on No, depending on whether your question describes acceptable or unacceptable text.
Frames per second to sample
When to use it: policing burned-in captions and overlays - profanity, prohibited claims, outdated promotions, competitor mentions.
Video Transcript Check
What it checks: what is said in the video. The audio is transcribed, and the transcript is checked against a yes/no question you write.
Settings:
Question - a yes/no question about the spoken content, for example: "Does the transcript contain profanity or offensive language?"
Pass this rule when the answer is Yes - Pass on Yes or Pass on No.
When to use it: spoken claims and compliance - required disclaimers, forbidden phrases, offensive language in voiceovers or interviews.
Good to know: this rule listens to the audio; it doesn't look at the visuals. Pair it with Video Content Detection or Video On-Screen Text Check for full coverage.
Video Text Metadata Match
What it checks: whether the text shown on screen matches the asset's own metadata - for example, that the product name burned into the video matches the product name in your metadata.
Settings:
Question - a yes/no question comparing on-screen text to metadata. Insert metadata fields using the field chips (or type
#followed by the field name), for example: "Does the on-screen text match #product_name?"Metadata fields - select which metadata fields Moderation may compare against the video's text.
Pass this rule when the answer is Yes - Pass on Yes or Pass on No.
Asset's metadata field is empty - what should happen when the selected metadata field has no value: Needs Review (the default), Approve, or Reject.
No text is detected on the asset - what should happen when no on-screen text is found in the video: Needs Review (the default), Approve, or Reject.
Frames per second to sample
When to use it: catching mismatches between creative and catalog data at scale - wrong product names, prices, or campaign codes in localized video variants.
Reading video results
Whatever the rule, a video gets one overall result: Approved, Rejected, or Needs review. The rule's explanation tells you what was found, and when it refers to a specific moment it includes a clickable timestamp that jumps the player straight there. A brief appearance is enough to trigger a rule - the result reflects the worst moment in the video, not the average.
Things to keep in mind
Processing time: video is analyzed frame by frame, so runs with many or long videos take noticeably longer than image runs.
Cost: video text rules perform heavier analysis and use more Moderation Actions per video than a typical image check.
Limits: very large or very long videos may be rejected before processing; keep files within your plan's limits.



