2 min read

Inkling-Small beats its larger sibling on key benchmarks

Thinking Machines Lab released Inkling-Small, a 276B-parameter MoE model with 12B active per token that beats Inkling on three benchmarks.

Image: ITzine

Thinking Machines Lab has released Inkling-Small, a multimodal model whose smaller active footprint outperformed the company’s larger Inkling on several applied benchmarks. The result highlights a practical trade-off in model design: parameter count alone does not determine performance; training methods and the amount of computation used per token also matter.

Inkling-Small architecture and capabilities

Inkling-Small uses a Mixture-of-Experts (MoE) architecture. Its full model contains 276 billion parameters, but only 12 billion parameters are activated for each token. That reduces the computational load while aiming to preserve model quality.

The model supports text, images, and audio, with a context window of up to 1 million tokens. Users can also adjust reasoning depth to balance response speed against quality. Thinking Machines Lab has released the model’s weights, allowing developers to fine-tune it and run it through Tinker.

Benchmark results against the original Inkling

According to the company, Inkling-Small scored higher than the original Inkling on three applied evaluations:

  • Terminal-Bench 2.1: 64.7% versus 63.8%
  • Humanity’s Last Exam: 31.6% versus 29.7%
  • WE-Bench Verified: 80.2% versus 77.6%

Source: Thinking Machines Lab

Recommended reading

Yandex delivery robots launch in Astana

The company says it trained Inkling-Small with on-policy distillation, using the larger Inkling as its teacher. The smaller model then spent another two weeks training on agentic programming tasks with reinforcement learning.

That process appears to have improved performance in narrower, practical tests. However, the original Inkling remains stronger in knowledge coverage and factual accuracy, according to the source. The company did not provide additional figures for those areas.

The release also puts the model’s open weights, multimodal input, adjustable reasoning depth, and fine-tuning support at the center of its proposition. Open-weight systems from companies including Meta* and Mistral have increasingly competed on efficiency per token, multimodality, and customization rather than parameter totals alone.

Whether Inkling-Small’s 1-million-token context window holds up in real-world workloads remains unverified in the information provided. For teams focused on computational cost rather than headline specifications, that validation may matter more than the model’s 276-billion-parameter total.

*Meta is recognized as an extremist organization in Russia, and its activities are prohibited there.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via ITzine

/ Keep reading