Ling 3.0 Flash
InclusionAI · ling-3-0-flash
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.
Specification
most-agreed values
Context
262K
Max output
33K
Released
2026-07-23
Knowledge cutoff
—
Retires
—
Open weights
no
Input
text
Output
text
Price
US dollars per million tokens · most-agreed
Input
$0.075
Output
$0.22
Cache read
$0.015
Available from 9 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| openrouter · inclusionai | $0.021 | $0.063 | $0.0042 | — | 262K | 33K | — |
| vercel · inclusionai | $0.021 | $0.063 | $0.0042 | — | 256K | 32K | — |
| deepinfra | $0.06 | $0.18 | $0.012 | — | 131K | — | — |
| kilo · inclusionai | $0.06 | $0.18 | $0.012 | — | 262K | 33K | — |
| llmgateway | $0.06 | $0.18 | $0.012 | — | 262K | 262K | — |
| llmgateway-providers | $0.06 | $0.18 | $0.012 | — | 262K | 33K | — |
| novita | $0.06 | $0.18 | $0.012 | — | 262K | 33K | — |
| nano-gpt · inclusionai | $0.075 | $0.22 | $0.015 | — | 262K | 33K | — |
| nano-gpt · thinking | $0.075 | $0.22 | $0.015 | — | 262K | 33K | — |
Price history
append-only observations · a listing writes a row only when its price moves
| Date | Host | Field | $/M |
|---|---|---|---|
| 2026-09-02 | vercel · inclusionai | Cache read | $0.012 |
| 2026-09-15 | vercel · inclusionai | Cache read | $0.0042 |
| 2026-09-02 | vercel · inclusionai | Input | $0.06 |
| 2026-09-15 | vercel · inclusionai | Input | $0.021 |
| 2026-09-02 | vercel · inclusionai | Output | $0.18 |
| 2026-09-15 | vercel · inclusionai | Output | $0.063 |
More from InclusionAI
most-hosted first