Together AI adds Inkling multimodal model

Together AI integrates Inkling, a new multimodal AI model from Thinking Machines Lab, offering text, image, and audio processing with controllable reasoning.

4 min read
Abstract visualization representing multimodal AI data streams feeding into a central processing core.
Together AI integrates Thinking Machines Lab's Inkling, a multimodal AI model supporting text, image, and audio.· Together AI
Visual TL;DR
Together AI IntegratesCore
integrating a new multimodal AI model from Thinking Machines Lab
From the article 5 mentionsLightweight embedding towers seamlessly integrate image patches and quantized audio into the model's sequence, enabling joint reasoning over diverse data types.
Inkling Multimodal ModelCore
novel model processing text, image, and audio inputs with controllable reasoning
From the article 2 mentionsTogether AI is now offering developers access to Inkling, a novel multimodal model developed by Thinking Machines Lab.
Advanced ArchitectureContext
features query-conditioned relative attention, convolutions, and MoE design
From the article 4 mentionsThis integration marks a significant step in making advanced AI reasoning accessible on a production-ready inference platform.
Text OutputsContext
From the articleIt accepts text, image, and audio inputs, unifying them through a single decoder architecture to produce text outputs.
Efficient ReasoningEffect
From the article 8 mentionsInkling is engineered for efficient reasoning and native understanding across various data types.
Developer AccessOutcome
offering developers access to Inkling on a production-ready inference platform
From the article 4 mentionsDevelopers can access Inkling immediately via Together AI Serverless, eliminating the need for infrastructure provisioning or complex setup.
Fine-tune ReasoningEffect
From the article 8 mentionsDevelopers can fine-tune the model's reasoning depth and resource utilization per task, balancing performance with latency and cost.
Contents(3)

Together AI is now offering developers access to Inkling, a novel multimodal model developed by Thinking Machines Lab. This integration marks a significant step in making advanced AI reasoning accessible on a production-ready inference platform.

Inkling is engineered for efficient reasoning and native understanding across various data types. It accepts text, image, and audio inputs, unifying them through a single decoder architecture to produce text outputs. Developers can fine-tune the model's reasoning depth and resource utilization per task, balancing performance with latency and cost.

Inkling's Architectural Innovations

The Inkling model features architectural advancements beyond standard Transformer models. These include query-conditioned relative attention, short causal convolutions, and a mixture-of-experts (MoE) design with a shared expert sink. These components are intended to enhance both reasoning capabilities and multimodal processing efficiency.

Together AI highlights its optimized inference stack, which specifically supports Inkling's unique attention mechanisms, ensuring practical efficiency gains for production workloads.

Broad Task Versatility and Performance

Inkling has been post-trained across a wide spectrum of tasks, including scientific reasoning, coding, agentic workflows, forecasting, and calibrated prediction. Preliminary evaluations show strong performance on graduate-level scientific reasoning and mathematics benchmarks.

This versatility extends to applications requiring uncertainty representation and calibrated estimation, moving beyond traditional question-answering tasks.

Day-Zero Access and Developer Benefits

Developers can access Inkling immediately via Together AI Serverless, eliminating the need for infrastructure provisioning or complex setup. The platform offers a unified API endpoint for all multimodal inputs, simplifying development pipelines.

The controllable reasoning effort feature allows for per-request tuning of depth, latency, and token expenditure, providing granular control over cost and speed.

Inkling's architecture, boasting 975 billion total parameters with 40 billion active per token and a 1 million token context window, utilizes a novel query-conditioned relative bias for positional encoding. It also incorporates sliding-window and full causal attention layers, alongside a lightweight causal convolution (sconv) for enhanced local context processing.

The MoE architecture with a shared expert sink dynamically routes tokens to specialized experts while allowing shared experts to compete for mixture weight, optimizing computation. Lightweight embedding towers seamlessly integrate image patches and quantized audio into the model's sequence, enabling joint reasoning over diverse data types.

This unified approach powers applications like visual question answering, document analysis, and multimodal agentic systems. Developers can begin building with Inkling on Together AI Serverless today, scaling from experimentation to dedicated production capacity.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.