Character.ai Tackles AI Video Quality 'Slop'
Mayur Bril from Character.AI discusses the challenges of evaluating AI video quality and introduces 'Judge Judy,' a new harness for better, faster assessments.

Visual TL;DR
majority of AI video outputs suffer from hallucinations, inconsistent physics, poor audio sync
From the article 6 mentionsTraditional metrics like CLIP Score, while useful for single frames, and metrics for frame-to-frame consistency (like LPIPS) are insufficient for evaluating the holistic quality of a video as a narrative medium.
current metrics like CLIP Score and LPIPS are insufficient for holistic video quality
From the articleThe rapid advancements in AI video generation, showcased by models like Sora, Kling, and Veo, have outpaced the development of robust evaluation methods.
tackling the difficult problem of accurate audio and lip synchronization in AI video
new evaluation harness for better, faster assessments of AI-generated video content
From the article 3 mentionsTo address this, Character.AI has developed an open-source video evaluation harness named 'Judge Judy.' The goal is to move evaluation from an offline, often slow process to an integrated, real-time component within the AI video generation pipeline.
captures complex aspects like storytelling, physics, pacing, beyond individual frames
From the article 2 mentionsThe development of 'Judge Judy' represents a significant step towards more reliable and efficient evaluation of AI-generated video, moving beyond simple metrics to a more holistic understanding of video quality as a storytelling medium.
intentionally creating poor quality examples to improve training and model robustness
From the articleCharacter.AI's approach to training Judge Judy involved intentionally 'manufacturing badness.' This was achieved by corrupting high-quality real footage and pairing it with AI-generated content to create datasets that teach the model to distinguish between good and bad video attributes.
understanding what defines good video content for more effective AI model training
From the article 9+ mentionsBril explained that while generating video has become increasingly accessible and affordable, the majority of outputs suffer from common issues like hallucinations, inconsistent physics, and poor audio synchronization.
Contents(6)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.
More from Daniel Singer