Character.ai Tackles AI Video Quality 'Slop'

Mayur Bril from Character.AI discusses the challenges of evaluating AI video quality and introduces 'Judge Judy,' a new harness for better, faster assessments.

9 min read
Mayur Bril presenting on AI video evaluation at AI Engineer World's Fair
Mayur Bril of Character.AI discusses the challenges and solutions for evaluating AI-generated video quality.· AI Engineer

Visual TL;DR. AI Video Quality 'Slop' due to Evaluation Methods Lag. Evaluation Methods Lag introduces Character.AI's Judge Judy. Character.AI's Judge Judy enables Holistic Video Assessment. Holistic Video Assessment leads to Better AI Video. AI Video Quality 'Slop' includes Sound/Lip Sync Challenges. Character.AI's Judge Judy uses Manufacturing 'Badness'. Character.AI's Judge Judy provides Key Evaluation Insights. Sound/Lip Sync Challenges improves Better AI Video.

  1. AI Video Quality 'Slop': majority of AI video outputs suffer from hallucinations, inconsistent physics, poor audio sync
  2. Evaluation Methods Lag: current metrics like CLIP Score and LPIPS are insufficient for holistic video quality
  3. Character.AI's Judge Judy: new evaluation harness for better, faster assessments of AI-generated video content
  4. Holistic Video Assessment: captures complex aspects like storytelling, physics, pacing, beyond individual frames
  5. Manufacturing 'Badness': intentionally creating poor quality examples to improve training and model robustness
  6. Better AI Video: addresses common issues like inconsistent physics, hallucinations, and audio synchronization
  7. Sound/Lip Sync Challenges: tackling the difficult problem of accurate audio and lip synchronization in AI video
  8. Key Evaluation Insights: understanding what defines good video content for more effective AI model training
Visual TL;DR
Visual TL;DR, startuphub.ai AI Video Quality 'Slop' due to Evaluation Methods Lag. Evaluation Methods Lag introduces Character.AI's Judge Judy. Character.AI's Judge Judy enables Holistic Video Assessment. Holistic Video Assessment leads to Better AI Video due to introduces enables leads to AI Video Quality 'Slop' Evaluation Methods Lag Character.AI's Judge Judy Holistic Video Assessment Better AI Video From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Video Quality 'Slop' due to Evaluation Methods Lag. Evaluation Methods Lag introduces Character.AI's Judge Judy. Character.AI's Judge Judy enables Holistic Video Assessment. Holistic Video Assessment leads to Better AI Video due to introduces enables leads to AI Video Quality'Slop' EvaluationMethods Lag Character.AI'sJudge Judy Holistic VideoAssessment Better AI Video From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Video Quality 'Slop' due to Evaluation Methods Lag. Evaluation Methods Lag introduces Character.AI's Judge Judy. Character.AI's Judge Judy enables Holistic Video Assessment. Holistic Video Assessment leads to Better AI Video due to introduces enables leads to AI Video Quality 'Slop' majority of AI video outputs suffer fromhallucinations, inconsistent physics, pooraudio sync Evaluation Methods Lag current metrics like CLIP Score and LPIPSare insufficient for holistic videoquality Character.AI's Judge Judy new evaluation harness for better, fasterassessments of AI-generated video content Holistic Video Assessment captures complex aspects likestorytelling, physics, pacing, beyondindividual frames Better AI Video addresses common issues like inconsistentphysics, hallucinations, and audiosynchronization From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Video Quality 'Slop' due to Evaluation Methods Lag. Evaluation Methods Lag introduces Character.AI's Judge Judy. Character.AI's Judge Judy enables Holistic Video Assessment. Holistic Video Assessment leads to Better AI Video due to introduces enables leads to AI Video Quality'Slop' majority of AIvideo outputssuffer from… EvaluationMethods Lag current metricslike CLIP Score andLPIPS are… Character.AI'sJudge Judy new evaluationharness for better,faster assessments… Holistic VideoAssessment captures complexaspects likestorytelling,… Better AI Video addresses commonissues likeinconsistent… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Video Quality 'Slop' due to Evaluation Methods Lag. Evaluation Methods Lag introduces Character.AI's Judge Judy. Character.AI's Judge Judy enables Holistic Video Assessment. Holistic Video Assessment leads to Better AI Video. AI Video Quality 'Slop' includes Sound/Lip Sync Challenges. Character.AI's Judge Judy uses Manufacturing 'Badness'. Character.AI's Judge Judy provides Key Evaluation Insights. Sound/Lip Sync Challenges improves Better AI Video due to introduces enables leads to includes uses provides improves AI Video Quality 'Slop' majority of AI video outputs suffer fromhallucinations, inconsistent physics, pooraudio sync Evaluation Methods Lag current metrics like CLIP Score and LPIPSare insufficient for holistic videoquality Character.AI's Judge Judy new evaluation harness for better, fasterassessments of AI-generated video content Holistic Video Assessment captures complex aspects likestorytelling, physics, pacing, beyondindividual frames Manufacturing 'Badness' intentionally creating poor qualityexamples to improve training and modelrobustness Better AI Video addresses common issues like inconsistentphysics, hallucinations, and audiosynchronization Sound/Lip Sync Challenges tackling the difficult problem of accurateaudio and lip synchronization in AI video Key Evaluation Insights understanding what defines good videocontent for more effective AI modeltraining From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Video Quality 'Slop' due to Evaluation Methods Lag. Evaluation Methods Lag introduces Character.AI's Judge Judy. Character.AI's Judge Judy enables Holistic Video Assessment. Holistic Video Assessment leads to Better AI Video. AI Video Quality 'Slop' includes Sound/Lip Sync Challenges. Character.AI's Judge Judy uses Manufacturing 'Badness'. Character.AI's Judge Judy provides Key Evaluation Insights. Sound/Lip Sync Challenges improves Better AI Video due to introduces enables leads to includes uses provides improves AI Video Quality'Slop' majority of AIvideo outputssuffer from… EvaluationMethods Lag current metricslike CLIP Score andLPIPS are… Character.AI'sJudge Judy new evaluationharness for better,faster assessments… Holistic VideoAssessment captures complexaspects likestorytelling,… Manufacturing'Badness' intentionallycreating poorquality examples to… Better AI Video addresses commonissues likeinconsistent… Sound/Lip SyncChallenges tackling thedifficult problemof accurate audio… Key EvaluationInsights understanding whatdefines good videocontent for more… From startuphub.ai · The publishers behind this format

The rapid advancements in AI video generation, showcased by models like Sora, Kling, and Veo, have outpaced the development of robust evaluation methods. This gap was highlighted by Mayur Bril from Character.AI during a presentation at the AI Engineer World's Fair. Bril detailed the challenges in assessing the quality of AI-generated video, pointing out that current tools often focus on individual frames or basic prompt adherence, failing to capture the more complex aspects of storytelling, physics, and pacing that define good video content.

Character.ai Tackles AI Video Quality 'Slop' - AI Engineer
Character.ai Tackles AI Video Quality 'Slop' — from AI Engineer

The Problem: Evaluating 'Good Enough' Video

Bril explained that while generating video has become increasingly accessible and affordable, the majority of outputs suffer from common issues like hallucinations, inconsistent physics, and poor audio synchronization. Traditional metrics like CLIP Score, while useful for single frames, and metrics for frame-to-frame consistency (like LPIPS) are insufficient for evaluating the holistic quality of a video as a narrative medium. He noted that even advanced LLM-based judges can be slow and highly dependent on prompt phrasing, leading to inconsistent results.

Introducing 'Judge Judy': A New Evaluation Approach

To address this, Character.AI has developed an open-source video evaluation harness named 'Judge Judy.' The goal is to move evaluation from an offline, often slow process to an integrated, real-time component within the AI video generation pipeline. Bril emphasized the importance of making evaluation 'cheap, consistent, and everywhere,' ideally occurring early in the generation process to catch and correct errors more efficiently.

Key Insights for AI Video Evaluation

Bril shared three core takeaways for those working on AI video generation:

  • Go Relative, Not Absolute: Instead of scoring videos on an absolute scale (e.g., 1-10), focus on relative comparisons between outputs. Training models on pairs (A vs. B) leads to more reliable and consistent judgments.
  • Score the Real Axes: Evaluate videos based on the criteria that truly matter for storytelling and coherence: the narrative, physics, pacing, and audio sync, not just pixel-level details or prompt adherence.
  • Put Eval Inside the Loop: Integrate evaluation metrics directly into the generation process. This allows for early intervention, enabling regeneration of short segments rather than entire videos, thus improving efficiency and quality.

Manufacturing 'Badness' for Better Training

Character.AI's approach to training Judge Judy involved intentionally 'manufacturing badness.' This was achieved by corrupting high-quality real footage and pairing it with AI-generated content to create datasets that teach the model to distinguish between good and bad video attributes. The initial version of the model, while confidently scoring videos, was 'confidently wrong' because it focused on superficial 'gloss' rather than underlying quality. By refining the dataset to include comparisons of real footage against generated content, ensuring consistent encoding and annotation across both, the model became more effective at judging actual video quality.

Addressing Sound and Lip Sync Challenges

In the Q&A, Bril addressed the evaluation of sound, explaining that the model uses audio quality metrics and identifies keyframes to correlate sounds with visual events, such as a door slam matching the action. However, he admitted that lip-syncing remains an unsolved problem in AI video generation, particularly for animated characters where the mouth movements may not have a direct correlation to speech patterns.

The development of 'Judge Judy' represents a significant step towards more reliable and efficient evaluation of AI-generated video, moving beyond simple metrics to a more holistic understanding of video quality as a storytelling medium.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.