Character.ai Tackles AI Video Quality 'Slop'

Mayur Bril from Character.AI discusses the challenges of evaluating AI video quality and introduces 'Judge Judy,' a new harness for better, faster assessments.

Mayur Bril presenting on AI video evaluation at AI Engineer World's Fair
Mayur Bril of Character.AI discusses the challenges and solutions for evaluating AI-generated video quality.· AI Engineer
Visual TL;DR
AI Video Quality 'Slop'Driver
majority of AI video outputs suffer from hallucinations, inconsistent physics, poor audio sync
From the article 6 mentionsTraditional metrics like CLIP Score, while useful for single frames, and metrics for frame-to-frame consistency (like LPIPS) are insufficient for evaluating the holistic quality of a video as a narrative medium.
Evaluation Methods LagDriver
current metrics like CLIP Score and LPIPS are insufficient for holistic video quality
From the articleThe rapid advancements in AI video generation, showcased by models like Sora, Kling, and Veo, have outpaced the development of robust evaluation methods.
Sound/Lip Sync ChallengesDriver
tackling the difficult problem of accurate audio and lip synchronization in AI video
Character.AI's Judge JudyCore
new evaluation harness for better, faster assessments of AI-generated video content
From the article 3 mentionsTo address this, Character.AI has developed an open-source video evaluation harness named 'Judge Judy.' The goal is to move evaluation from an offline, often slow process to an integrated, real-time component within the AI video generation pipeline.
Holistic Video AssessmentEffect
captures complex aspects like storytelling, physics, pacing, beyond individual frames
From the article 2 mentionsThe development of 'Judge Judy' represents a significant step towards more reliable and efficient evaluation of AI-generated video, moving beyond simple metrics to a more holistic understanding of video quality as a storytelling medium.
Manufacturing 'Badness'Context
intentionally creating poor quality examples to improve training and model robustness
From the articleCharacter.AI's approach to training Judge Judy involved intentionally 'manufacturing badness.' This was achieved by corrupting high-quality real footage and pairing it with AI-generated content to create datasets that teach the model to distinguish between good and bad video attributes.
Key Evaluation InsightsContext
understanding what defines good video content for more effective AI model training
Better AI VideoOutcome
From the article 9+ mentionsBril explained that while generating video has become increasingly accessible and affordable, the majority of outputs suffer from common issues like hallucinations, inconsistent physics, and poor audio synchronization.
Contents(6)

The rapid advancements in AI video generation, showcased by models like Sora, Kling, and Veo, have outpaced the development of robust evaluation methods. This gap was highlighted by Mayur Bril from Character.AI during a presentation at the AI Engineer World's Fair. Bril detailed the challenges in assessing the quality of AI-generated video, pointing out that current tools often focus on individual frames or basic prompt adherence, failing to capture the more complex aspects of storytelling, physics, and pacing that define good video content.

StartupHub data

Companies working on this

Profiles of the companies named in this story, with founding year, headquarters, and a short description from our database.

AI-powered text-to-video generation model and social media app.

Founded
2024
Location
San Francisco, United States

Veo Robotics provides a free-space optical communication system for industrial robots.

Founded
2016
Location
Waltham, United States
Funding
$70M

A leading personal AI platform creating revolutionary open-ended conversational applications.

Founded
2021
Location
Menlo Park, United States
Valuation
$1K
Character.ai Tackles AI Video Quality 'Slop' - AI Engineer
Character.ai Tackles AI Video Quality 'Slop', from AI Engineer

The Problem: Evaluating 'Good Enough' Video

Bril explained that while generating video has become increasingly accessible and affordable, the majority of outputs suffer from common issues like hallucinations, inconsistent physics, and poor audio synchronization. Traditional metrics like CLIP Score, while useful for single frames, and metrics for frame-to-frame consistency (like LPIPS) are insufficient for evaluating the holistic quality of a video as a narrative medium. He noted that even advanced LLM-based judges can be slow and highly dependent on prompt phrasing, leading to inconsistent results.

Introducing 'Judge Judy': A New Evaluation Approach

To address this, Character.AI has developed an open-source video evaluation harness named 'Judge Judy.' The goal is to move evaluation from an offline, often slow process to an integrated, real-time component within the AI video generation pipeline. Bril emphasized the importance of making evaluation 'cheap, consistent, and everywhere,' ideally occurring early in the generation process to catch and correct errors more efficiently.

Key Insights for AI Video Evaluation

Bril shared three core takeaways for those working on AI video generation:

  • Go Relative, Not Absolute: Instead of scoring videos on an absolute scale (e.g., 1-10), focus on relative comparisons between outputs. Training models on pairs (A vs. B) leads to more reliable and consistent judgments.
  • Score the Real Axes: Evaluate videos based on the criteria that truly matter for storytelling and coherence: the narrative, physics, pacing, and audio sync, not just pixel-level details or prompt adherence.
  • Put Eval Inside the Loop: Integrate evaluation metrics directly into the generation process. This allows for early intervention, enabling regeneration of short segments rather than entire videos, thus improving efficiency and quality.

Manufacturing 'Badness' for Better Training

Character.AI's approach to training Judge Judy involved intentionally 'manufacturing badness.' This was achieved by corrupting high-quality real footage and pairing it with AI-generated content to create datasets that teach the model to distinguish between good and bad video attributes. The initial version of the model, while confidently scoring videos, was 'confidently wrong' because it focused on superficial 'gloss' rather than underlying quality. By refining the dataset to include comparisons of real footage against generated content, ensuring consistent encoding and annotation across both, the model became more effective at judging actual video quality.

Addressing Sound and Lip Sync Challenges

In the Q&A, Bril addressed the evaluation of sound, explaining that the model uses audio quality metrics and identifies keyframes to correlate sounds with visual events, such as a door slam matching the action. However, he admitted that lip-syncing remains an unsolved problem in AI video generation, particularly for animated characters where the mouth movements may not have a direct correlation to speech patterns.

The development of 'Judge Judy' represents a significant step towards more reliable and efficient evaluation of AI-generated video, moving beyond simple metrics to a more holistic understanding of video quality as a storytelling medium.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer