"We knew what we had to do: we had to scale this up." This declaration, made by Greg Brockman, President and Co-founder of OpenAI, during the Developer State Of The Union at OpenAI DevDay [2025], encapsulates a pivotal strategic pivot that has fundamentally reshaped the landscape of artificial intelligence. The event, a series of presentations from OpenAI's leadership, unveiled a suite of advancements and developer tools poised to redefine software creation.
Brockman began by tracing OpenAI’s journey back to its foundational plan in 2015, a simple three-step blueprint: "Solve reinforcement learning, solve unsupervised learning, and gradually learn really complicated things." He recounted early successes, like the breakthrough in Dota 2 in 2017, demonstrating the power of reinforcement learning. Simultaneously, their work on unsupervised learning, particularly with the "unsupervised sentiment neuron" in the same year, revealed that models could spontaneously learn semantics through next-step prediction. This dual progression laid the groundwork for a critical insight: scaling these capabilities could unlock unprecedented utility.
The realization led to a profound shift in OpenAI's strategy. Rather than building specific, vertical AI applications in fields like healthcare or education, a path Brockman described as "tunnel vision" and "exactly backwards the way you're supposed to build a startup", the company chose a different route. In 2020, they introduced the OpenAI API, allowing developers to connect their advanced models, like GPT-3, to real-world applications. This decision, initially fraught with doubt, "felt doomed," Brockman admitted, yet it proved wildly successful. It harnessed the collective ingenuity of developers, transforming a technology in search of a problem into a versatile platform for innovation.
The latest suite of product releases continues this strategic trajectory, emphasizing agentic capabilities and multimodal interaction. Olivier Godement, Head of Platform at OpenAI, detailed GPT-5, describing it as their "most capable and reliable model," specifically designed for agentic tasks, excelling at instruction following and tool use. This iteration pushes the boundaries of coding intelligence, enabling GPT-5 to refactor complex codebases, generate tasteful front-end UIs from single prompts, and even "work autonomously for hours at a time."
Further enhancing multimodal capabilities, Sora 2, the highly anticipated video generation model, was announced for the API. Sora 2 is offered in two versions: a standard model for "fast experimentation" and a "Pro" version that "pays even more closer attention to detail." This allows developers to iterate rapidly on video concepts before committing to higher-fidelity generation. Complementing these, OpenAI also introduced smaller, more cost-effective "mini" versions of their real-time speech-to-speech (GPT-realtime-mini) and image generation (GPT-image-1-mini) models, drastically reducing costs while maintaining quality. This initiative aims to truly democratize access to powerful AI, enabling local and offline applications.
