# Character.ai Tackles AI Video Quality 'Slop' _Mayur Bril from Character.AI discusses the challenges of evaluating AI video quality and introduces 'Judge Judy,' a new harness for better, faster assessments._ **Published:** 2026-07-25 **Source:** https://www.startuphub.ai/ai-news/artificial-intelligence/2026/character-ai-tackles-ai-video-quality-slop --- The rapid advancements in AI video generation, showcased by models like [Sora](/ai-news/claude), Kling, and [Veo](/ai-news/ai-research/2026/dear-upstairs-neighbors-how-controlled-ai-changes-animation), have outpaced the development of robust evaluation methods. This gap was highlighted by Mayur Bril from Character.AI during a presentation at the AI Engineer World's Fair. Bril detailed the challenges in assessing the quality of AI-generated video, pointing out that current tools often focus on individual frames or basic prompt adherence, failing to capture the more complex aspects of storytelling, physics, and pacing that define good video content. AI Video Quality 'Slop'Driver majority of AI video outputs suffer from hallucinations, inconsistent physics, poor audio syncFrom the article 6 mentionsTraditional metrics like CLIP Score, while useful for single frames, and metrics for frame-to-frame consistency (like LPIPS) are insufficient for evaluating the holistic quality of a video as a narrative medium.Evaluation Methods LagDrivercurrent metrics like CLIP Score and LPIPS are insufficient for holistic video qualityFrom the articleThe rapid advancements in AI video generation, showcased by models like Sora, Kling, and Veo, have outpaced the development of robust evaluation methods.Sound/Lip Sync ChallengesDrivertackling the difficult problem of accurate audio and lip synchronization in AI videointroducesCharacter.AI's Judge JudyCorenew evaluation harness for better, faster assessments of AI-generated video contentFrom the article 3 mentionsTo address this, Character.AI has developed an open-source video evaluation harness named 'Judge Judy.' The goal is to move evaluation from an offline, often slow process to an integrated, real-time component within the AI video generation pipeline.Holistic Video AssessmentEffectcaptures complex aspects like storytelling, physics, pacing, beyond individual framesFrom the article 2 mentionsThe development of 'Judge Judy' represents a significant step towards more reliable and efficient evaluation of AI-generated video, moving beyond simple metrics to a more holistic understanding of video quality as a storytelling medium.Manufacturing 'Badness'Contextintentionally creating poor quality examples to improve training and model robustnessFrom the articleCharacter.AI's approach to training Judge Judy involved intentionally 'manufacturing badness.' This was achieved by corrupting high-quality real footage and pairing it with AI-generated content to create datasets that teach the model to distinguish between good and bad video attributes.Key Evaluation InsightsContextunderstanding what defines good video content for more effective AI model trainingleads toBetter AI VideoOutcomeFrom the article 9+ mentionsBril explained that while generating video has become increasingly accessible and affordable, the majority of outputs suffer from common issues like hallucinations, inconsistent physics, and poor audio synchronization. ## The Problem: Evaluating 'Good Enough' Video Bril explained that while generating video has become increasingly accessible and affordable, the majority of outputs suffer from common issues like hallucinations, inconsistent physics, and poor audio synchronization. Traditional metrics like CLIP Score, while useful for single frames, and metrics for frame-to-frame consistency (like LPIPS) are insufficient for evaluating the holistic quality of a video as a narrative medium. He noted that even advanced LLM-based judges can be slow and highly dependent on prompt phrasing, leading to inconsistent results. ## Introducing 'Judge Judy': A New Evaluation Approach To address this, Character.AI has developed an open-source video evaluation harness named 'Judge Judy.' The goal is to move evaluation from an offline, often slow process to an integrated, real-time component within the AI video generation pipeline. Bril emphasized the importance of making evaluation 'cheap, consistent, and everywhere,' ideally occurring early in the generation process to catch and correct errors more efficiently. ## Key Insights for AI Video Evaluation Bril shared three core takeaways for those working on AI video generation: - **Go Relative, Not Absolute:** Instead of scoring videos on an absolute scale (e.g., 1-10), focus on relative comparisons between outputs. Training models on pairs (A vs. B) leads to more reliable and consistent judgments. - **Score the Real Axes:** Evaluate videos based on the criteria that truly matter for storytelling and coherence: the narrative, physics, pacing, and audio sync, not just pixel-level details or prompt adherence. - **Put Eval Inside the Loop:** Integrate evaluation metrics directly into the generation process. This allows for early intervention, enabling regeneration of short segments rather than entire videos, thus improving efficiency and quality. ## Manufacturing 'Badness' for Better Training Character.AI's approach to training Judge Judy involved intentionally 'manufacturing badness.' This was achieved by corrupting high-quality real footage and pairing it with AI-generated content to create datasets that teach the model to distinguish between good and bad video attributes. The initial version of the model, while confidently scoring videos, was 'confidently wrong' because it focused on superficial 'gloss' rather than underlying quality. By refining the dataset to include comparisons of real footage against generated content, ensuring consistent encoding and annotation across both, the model became more effective at judging actual video quality. ## Addressing Sound and Lip Sync Challenges In the Q&A, Bril addressed the evaluation of sound, explaining that the model uses audio quality metrics and identifies keyframes to correlate sounds with visual events, such as a door slam matching the action. However, he admitted that lip-syncing remains an unsolved problem in AI video generation, particularly for animated characters where the mouth movements may not have a direct correlation to speech patterns. The development of 'Judge Judy' represents a significant step towards more reliable and efficient evaluation of AI-generated video, moving beyond simple metrics to a more holistic understanding of video quality as a storytelling medium. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.