modelbenchmark.io

Ling 3.0 Flash

InclusionAI · ling-3-0-flash

Compare

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.

Specification

most-agreed values

Context
262K
Max output
33K
Released
2026-07-23
Knowledge cutoff
Retires
Open weights
no
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.075
Output
$0.22
Cache read
$0.015

Available from 9 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
openrouter · inclusionai$0.021$0.063$0.0042262K33K
vercel · inclusionai$0.021$0.063$0.0042256K32K
deepinfra$0.06$0.18$0.012131K
kilo · inclusionai$0.06$0.18$0.012262K33K
llmgateway$0.06$0.18$0.012262K262K
llmgateway-providers$0.06$0.18$0.012262K33K
novita$0.06$0.18$0.012262K33K
nano-gpt · inclusionai$0.075$0.22$0.015262K33K
nano-gpt · thinking$0.075$0.22$0.015262K33K

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02vercel · inclusionaiCache read$0.012
2026-09-15vercel · inclusionaiCache read$0.0042
2026-09-02vercel · inclusionaiInput$0.06
2026-09-15vercel · inclusionaiInput$0.021
2026-09-02vercel · inclusionaiOutput$0.18
2026-09-15vercel · inclusionaiOutput$0.063

More from InclusionAI

most-hosted first