


Inkling is Thinking Machines’ first open-weights model, a 975B MoE with 41B active parameters, 1M context, native reasoning across text, images, and audio, and controllable thinking effort. Fine-tune it on Tinker or download the Apache 2.0 weights.
Inkling is Thinking Machines’ first open-weights model, released under the Apache 2.0 license. It is a Mixture-of-Experts (MoE) transformer with 975 billion total parameters and 41 billion active parameters per forward pass. Inkling supports a context window of up to 1 million tokens and was pretrained on 45 trillion tokens spanning text, images, audio, and video. It reasons natively across these modalities and offers controllable thinking effort, allowing users to balance cost and performance. The model is available for fine-tuning on Tinker, and its full weights can be downloaded directly.
Inkling uses a Mixture-of-Experts architecture that activates only a fraction of its total parameters per token. This design delivers the capability of a large model while keeping inference compute efficient and latency manageable.
The model supports up to 1 million tokens of context, making it suitable for long-document analysis, extended conversations, and tasks that require reasoning over large amounts of information in a single pass.
Inkling was pretrained on text, images, audio, and video, and it reasons across these modalities natively. It does not rely on separate encoders or pipelines for different input types.
Users can adjust how much computation the model spends on reasoning before generating an answer. This allows fine-grained control over the trade-off between response quality and cost or latency.
Inkling is not the strongest overall model available today — instead, a combination of qualities makes it a good open-weights base for customization.
This honest positioning sets Inkling apart. Rather than chasing top benchmark scores, Thinking Machines focused on building a balanced, flexible foundation model that excels as a starting point for fine-tuning. The model is available on Tinker for immediate customization, and the team demonstrated this by having Inkling fine-tune itself — writing its own training job, running it, and evaluating the result — all within the Tinker platform.
You are looking for a permissively licensed, open-weights model that balances multimodal capability with efficient inference, and you value the ability to fine-tune and customize the model for your own use cases. Inkling is especially worth exploring if you want to experiment with controllable reasoning effort or need a model that can handle long contexts across text, images, and audio.
Other tools you might consider
Loading comments…
Maker
neon_dev
Visit Website
thinkingmachines.ai/news/introducing-inkling/
Project Info
Product Keywords