Together AI Offers Predictable Inference

Together AI introduces Provisioned Throughput, offering reserved inference capacity for open models with token-based pricing and a 99% uptime SLA.

5 min read
Together AI logo with text 'Provisioned Throughput'
Together AI announces Provisioned Throughput for predictable AI inference.· Together AI
Visual TL;DR
AI Inference Costs RiseDriver
From the article 2 mentionsThis move aims to provide businesses with predictable performance and pricing, a critical factor as AI inference costs become a significant line item for companies.
Inference DilemmaDriver
choose between serverless convenience or dedicated infrastructure
From the article 9 mentionsTogether AI is introducing Provisioned Throughput, a new service designed to offer reserved inference capacity for open-weight frontier models.
Together AI's SolutionCore
introduces Provisioned Throughput service
From the article 4 mentionsFor highly customized needs, dedicated inference solutions are still available.
Reserved CapacityContext
guaranteed inference capacity for open models
From the article 5 mentionsTogether AI is introducing Provisioned Throughput, a new service designed to offer reserved inference capacity for open-weight frontier models.
Predictable PricingContext
token-based pricing similar to proprietary models
From the article 4 mentionsThis move aims to provide businesses with predictable performance and pricing, a critical factor as AI inference costs become a significant line item for companies.
Reliable PerformanceEffect
From the article 4 mentionsThis new offering comes with a 99% uptime SLA and token-based pricing, positioning it as a more reliable option for production workloads than traditional serverless offerings.
Lower CostsOutcome
From the article 4 mentionsCosts are reported to be significantly lower than proprietary alternatives, potentially reducing expenses by up to 90% compared to models like Claude Opus 4.8.

Together AI is introducing Provisioned Throughput, a new service designed to offer reserved inference capacity for open-weight frontier models. This move aims to provide businesses with predictable performance and pricing, a critical factor as AI inference costs become a significant line item for companies.

Historically, organizations have had to choose between the convenience of best-effort serverless inference or the control of dedicated, managed infrastructure. Provisioned Throughput seeks to occupy a middle ground, offering the simplicity of token-based pricing, akin to proprietary model providers, combined with guaranteed capacity and a service level agreement (SLA).

This new offering comes with a 99% uptime SLA and token-based pricing, positioning it as a more reliable option for production workloads than traditional serverless offerings. Costs are reported to be significantly lower than proprietary alternatives, potentially reducing expenses by up to 90% compared to models like Claude Opus 4.8.

Initially available for models such as MiniMax M3 and GLM-5.2, the service is accessible across North America and EMEA. The economics are structured around Provisioned Throughput Units (PTUs), which represent fixed slices of capacity. Each PTU guarantees a specific rate of tokens per minute for a given model, priced at $0.05 per PTU per minute.

The PTUs account for different token types, input, cached input, and output, each consuming capacity at a distinct rate. This allows for optimization based on specific traffic patterns without altering the underlying SLA. For instance, one PTU on MiniMax M3 can handle 138,840 input tokens per minute, 694,200 cached input tokens per minute, or 23,140 output tokens per minute.

This move is particularly relevant as companies increasingly adopt open-weight model inference for various tasks, including coding, finance, and workflow automation. The company notes a substantial shift in traffic, with token volume growing from 30 billion to over 400 trillion tokens per month, much of which has migrated from closed APIs.

Provisioned Throughput is positioned for production workloads requiring guarantees, while serverless inference remains ideal for rapid development. For highly customized needs, dedicated inference solutions are still available. This development provides a clearer migration path for businesses looking to leverage frontier-quality open models with predictable costs and reliability, potentially slashing expenses compared to current proprietary solutions.

The company highlighted that for MiniMax M3 inference, teams can now run these models at scale with commitments that support business operations. This aligns with their efforts in Together AI Masters MiniMax M3 Inference, ensuring robust performance for demanding applications.

This innovation addresses the growing need for cost-effective and reliable AI infrastructure, directly tackling LLM inference costs through a new approach to capacity management.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.