NVIDIA's AI Factory: Software, Validation, and Internal Use Cases

NVIDIA execs discuss Enterprise Validated Designs, the company's internal AI Factory scale, and the future of secure, autonomous AI agents.

Three men sitting on a couch, discussing AI strategy.
YouTube
Visual TL;DR
AI Factory ComplexityDriver
building and running enterprise-grade AI factories presents significant challenges for organizations
From the article 9+ mentionsIn the latest episode of AI Factory Insider, NVIDIA's Kaushik Shirhatti, Distinguished Solutions Architect, and Jon Fernandez, Director of Compute, delve into the complexities of building and running an enterprise-grade AI factory.
Reference Architectures (ERAs)Context
foundational layers covering energy, chips, and infrastructure for AI factory setup
From the article 2 mentionsShirhatti clarified the distinction between NVIDIA's Enterprise Reference Architectures (ERAs) and Enterprise Validated Designs (EVDs).
Validated Designs (EVDs)Core
address software, Kubernetes, agents, models, and applications for reliable AI workloads
From the article 3 mentionsShirhatti clarified the distinction between NVIDIA's Enterprise Reference Architectures (ERAs) and Enterprise Validated Designs (EVDs).
NVIDIA Internal AI FactoryCore
NVIDIA uses EVD principles to scale its own internal AI factory operations
From the article 7 mentionsBuilding upon the previous episode's focus on hardware-centric reference architectures, this discussion elevates the conversation to software and validated designs, illustrating NVIDIA's internal adoption of these principles.
Kubernetes LayerContext
EVDs establish a reliable Kubernetes layer for scheduling AI workloads and integration
From the article 3 mentionsThe EVD establishes a reliable Kubernetes layer, serving as a platform for scheduling AI workloads and integrating various software components necessary for a complete AI factory.
Secure AI AgentsEffect
future involves autonomous AI agents with security and tokenomics considerations
From the article 5 mentionsIn contrast, EVDs address the layers above, including software, Kubernetes, agents, models, and applications.
Enterprise-Grade AIOutcome
enables robust, scalable, and secure AI solutions for various enterprise use cases
From the articleIn the latest episode of AI Factory Insider, NVIDIA's Kaushik Shirhatti, Distinguished Solutions Architect, and Jon Fernandez, Director of Compute, delve into the complexities of building and running an enterprise-grade AI factory.
Contents(5)

In the latest episode of AI Factory Insider, NVIDIA's Kaushik Shirhatti, Distinguished Solutions Architect, and Jon Fernandez, Director of Compute, delve into the complexities of building and running an enterprise-grade AI factory. Building upon the previous episode's focus on hardware-centric reference architectures, this discussion elevates the conversation to software and validated designs, illustrating NVIDIA's internal adoption of these principles.

NVIDIA's AI Factory: Software, Validation, and Internal Use Cases - YouTube
NVIDIA's AI Factory: Software, Validation, and Internal Use Cases, from YouTube

The Difference Between Reference Architectures and Validated Designs

Shirhatti clarified the distinction between NVIDIA's Enterprise Reference Architectures (ERAs) and Enterprise Validated Designs (EVDs). He explained that ERAs cover the foundational layers of an AI factory, focusing on energy, chips, and infrastructure. In contrast, EVDs address the layers above, including software, Kubernetes, agents, models, and applications. The EVD establishes a reliable Kubernetes layer, serving as a platform for scheduling AI workloads and integrating various software components necessary for a complete AI factory.

Key Components of Validated Designs

Fernandez likened EVDs to an 'iPhone and app store' model, where software is grouped by functionality. He elaborated on the lifecycle of a data flywheel, encompassing data curation, model training and tuning, agent skill harnessing, deployment, monitoring, and observability. This process involves a broad ecosystem of partners, with NVIDIA selecting those known to integrate cleanly with their AI factory solutions. Security is also a paramount dependency, ensuring that all components operate securely and efficiently.

NVIDIA's Internal AI Factory: A Case Study

The conversation then shifted to NVIDIA's own AI Factory, which has been operational for approximately a year. Shirhatti highlighted that the company built its first AI factory using its validated designs and partner solutions, aiming to provide a blueprint for enterprise customers with a low integration risk. He noted that while enterprises may adapt these designs to their specific needs, NVIDIA's internal process serves as a proof of concept. The scale of NVIDIA's internal AI usage has surged, with a 40% month-over-month growth in token consumption. The company is now serving four trillion tokens per month with 99.9% availability, handling approximately 200 million infrastructure inference requests daily.

Use Cases and the Role of Agents

Fernandez shared examples of how NVIDIA employees are benefiting from the AI factory, including the use of AI for ticket deflection and empowering users with self-service capabilities. He emphasized that the AI factory is versatile and supports workloads beyond simple token generation, such as drug discovery, financial modeling, and physical AI simulations. The concept of agentic AI was also explored, with the understanding that while it's a current buzzword, the principles apply across various industries and use cases. The ability for agents to run autonomously, coupled with robust security measures, is seen as a key superpower.

Security, Tokenomics, and the Future of AI Factories

Both speakers touched upon the increasing importance of security as AI agents become more capable. The discussion also highlighted the growing influence of tokenomics on enterprises considering on-premise AI infrastructure. The scalability of the AI factory, leveraging a hybrid cloud and on-prem strategy, is crucial for meeting escalating demand. Looking ahead, Shirhatti expressed excitement about confidential computing for inference workloads, enabling frontier AI labs to deploy on-premise with enhanced security for their intellectual property. The ability to encrypt model weights and enterprise data, while preventing observation by platform operators, is seen as a significant advancement, fostering a new zero-trust model.

The episode concluded with a preview of the next installment, which will delve deeper into agentic AI, long-running agents, investment considerations, and business growth strategies within the AI factory paradigm.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer