Inkling Small

thinkingmachines/inkling-small

lightweight MoE reasoning at lower cost and latency.

cmd --model thinkingmachines/inkling-small
Intelligence index
27.8
Output speed
230.2 tok/s
Input
$0.50 /M
Output
$1.20 /M
Cache read
$0.10 /M
Agent-loop cost
$0.22 /M in
Context window
1M tokens
Released
July 30, 2026
Modalities
→

vs. the lineup

Inkling Small beside its stablemates and nearest rivals. The ◆ marks the best value in each column across every row shown.

pin a rival:
ModelIntelligenceCodingSpeedInput $/MOutput $/MBlended $/MContext
Muse Spark 1.3 Contributor48.1◆75.8◆247.5$0.10◆$0.20◆$0.13◆1.05M◆
Qwen 3.6 Max Preview28.4——$1.30$7.80$2.92200K
GLM-527.9——$1$3.20$1.55200K
Inkling Small ◆27.852.9230.2$0.50$1.20$0.681M
Inkling2552.1109.2$1$4.05$1.76256K
Gemini 3.5 Flash Lite22.249.3370◆$0.30$2.50$0.851M
DeepSeek V4 Pro (latest)pinned3668.878.6$0.66$1.98$0.991M
DeepSeek V4 Flash (latest)pinned3469.1—$0.15$0.60$0.261M

coding performance

The Intelligence Index and its sub-scores, ranked against every scored model in the catalog. A metric that has not been measured for Inkling Small has been left empty.

Coding Index
52.9
#46 of 54 scored
Terminal-Bench
55.1
#48 of 54 scored
Intelligence Index
27.8
#47 of 67 scored
Long-context reasoning
75.7
reasoning across a long context
SciCode
49.7
scientific coding
GPQA Diamond
89.5
graduate-level QA

usage calculator

How far a month of credits goes on Inkling Small.

Input tokensfresh prompt
800
Output tokensmodel reply
180
Cache read tokensre-read context
50K
cost / request $0.0056 · in $0.50 · out $1.20 · cache $0.10 per M
fresh input 7%output 4%cache reads 89%
Requests / 30 days
1.8K
$10 credits ÷ $0.0056 per request
~356 quick fixes~71 bug fixes~12 feature PRs

what real work costs

Real coding tasks priced end to end on Inkling Small, from a quick lookup to a full-repo agent run.

One agent task
$0.05
180K in at 75% cache hit, 12K out
What you pay
$0.26 /M
all-in across every token that task touched
Sticker input
$0.50 /M
cache reads bill at $0.10 /M instead
TaskTokens in · outInkling SmallMuse Spark 1.3 ContributorClaude Haiku 4.5
Quick lookup / one-liner8K · 1K$0.0023$0.0003$0.0065
Review a 500-line PR60K · 4K$0.02$0.0027$0.04
Fix a bug (agent loop)180K · 12K$0.05$0.0072$0.12
Refactor a module320K · 20K$0.09$0.01$0.21
Full-repo agent run900K · 45K$0.22$0.03$0.49

frequently asked

How much does Inkling Small cost?
$0.50/M input and $1.20/M output, cache reads $0.10/M. In an agent loop most input is cache-read, so the effective input rate is about $0.22/M.
Which plan do I need?
Available on Go and above.
How do I switch to it?
Run cmd --model thinkingmachines/inkling-small, or type /model in a session and pick it. You can switch mid-session without losing context.

Ship code that matches your taste

Command Code is the AI coding agent that continuously learns your taste. Start for $1.

Benchmarks from Artificial Analysis (v4.3)commandcode.ai/models/inkling-small