State-of-the-art text-to-video model generating high-fidelity, synchronized audio videos from text prompts.