An open-weight multimodal mixture-of-experts model from Thinking Machines Lab with 12B active parameters out of 276B total, positioned as the smaller, more efficient member of the Inkling family. Accepts text, image, and audio input with reasoning, tool use, and a 524K-token context window.
Sorted by total cost (input + output per 1M tokens). Select a row to view provider details.
| Provider | Pricing (per 1M) | Rate limits | Regions | Health | Latency |
|---|---|---|---|---|---|
In: $0.45Out: $1.20 | 60 RPM / 200K TPM | us-east-1 | Healthy | 0ms |
Use this model via OpenRouter with an OpenAI-compatible SDK.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://openrouter.ai/api/v1",
apiKey: process.env.OPENROUTER_API_KEY,
});
const response = await client.chat.completions.create({
model: "thinkingmachines/inkling-small",
messages: [
{ role: "user", content: "Hello!" }
],
});
console.log(response.choices[0].message.content);Using OpenRouter API · OpenAI-compatible SDK
Every price recorded for this model, per provider. Prices are per 1M tokens in USD.
No price changes recorded since Oct 1, 2026.
| Date | Provider | Input | Output |
|---|---|---|---|
| Oct 1, 2026Listed · Current | OpenRouter | $0.45 | $1.20 |