Thinking Machines Lab cuts costs with Inkling-Small

Thinking Machines Lab launches Inkling-Small, a 276B parameter model that delivers comparable performance to its larger predecessor at a fraction of the cost.

Performance graph showing Inkling-Small efficiency compared to larger models
Inkling-Small demonstrates improved efficiency in reasoning tasks.· Thinking Machines Lab
Visual TL;DR
Cut CostsDriver
Thinking Machines Lab aims to gain ground against competitors by reducing compute requirements
From the articleThe model allows users to adjust reasoning effort, letting them choose between lower costs or higher performance.
Inkling-Small ModelCore
a 276B parameter model designed to compete with larger systems efficiently
From the article 7 mentionsThinking Machines Lab today announced the Inkling-Small release, an efficient open-weights model designed to compete with larger systems while slashing compute requirements.
Mixture-of-ExpertsContext
From the articleThe model functions as a mixture-of-experts transformer, utilizing 12 billion active parameters out of 276 billion total.
Comparable PerformanceEffect
achieves parity with the larger 975B parameter Inkling model on key benchmarks
From the articleThe model allows users to adjust reasoning effort, letting them choose between lower costs or higher performance.
StartupHub.ai ScoreContext
original Inkling scored 54/100, behind Intapp (70/100) but ahead of 360Learning (45/100)
From the articleStartupHub.ai data assigns the original Inkling a score of 54/100.
Exceeds SWEBench-VerifiedEffect
From the articleTesting shows Inkling-Small exceeding 80 percent on SWEBench-Verified.
Retains CapabilitiesEffect
From the articleIt retains the native audio and image processing capabilities found in the larger version.
Fraction of CostOutcome
delivers comparable performance to its larger predecessor at a significantly reduced cost
From the articleThe model allows users to adjust reasoning effort, letting them choose between lower costs or higher performance.

Thinking Machines Lab today announced the Inkling-Small release, an efficient open-weights model designed to compete with larger systems while slashing compute requirements. This Inkling-Small model release arrives as the company attempts to gain ground against competitors.

Architecture and Efficiency

The model functions as a mixture-of-experts transformer, utilizing 12 billion active parameters out of 276 billion total. By training on NVIDIA GB300 hardware, the team claims it achieves parity with the larger 975B parameter Thinking Machines Lab Inkling model on key benchmarks.

StartupHub.ai data assigns the original Inkling a score of 54/100. This puts it behind specialized competitors like Intapp, which holds a 70/100 rating, though it remains ahead of 360Learning at 45/100.

Performance and Capability

Testing shows Inkling-Small exceeding 80 percent on SWEBench-Verified. It retains the native audio and image processing capabilities found in the larger version. The model allows users to adjust reasoning effort, letting them choose between lower costs or higher performance.

The company is providing full weights on Hugging Face. Developers can also access the model via the Tinker platform for fine-tuning.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.