modelbenchmark.io

Qwen3.8 2.4T A95B

Alibaba · qwen3-8-2-4t-a95b

Compare

Open-weight sparse MoE (2.4T total, 95B active), the open-weight twin of Qwen3.8 Max for coding, research, complex reasoning, and agentic workflows

Specification

most-agreed values

Context
262K
Max output
262K
Released
2026-08-12
Knowledge cutoff
Retires
Open weights
yes
Input
text, image
Output
text

Price

US dollars per million tokens · most-agreed

Input
$2.00
Output
$6.00
Cache read
$0.20
Cache write
$2.50

Available from 14 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
aihubmix$2.00$6.00$0.50262K262K
deepinfra$2.00$6.00$0.20262K131K
edenai$2.00$6.00$0.25$2.501M131K
fireworks-ai · accounts$2.00$6.00$0.25262K131K
hyper$2.00$6.00$0.251M128K
kilo$2.00$6.00$0.25$2.501M131K
modal$2.00$6.00$0.251M131K
openrouter$2.00$6.00$0.251M
131K
262K
requesty$2.00$6.00$0.20262K262K
together_ai$2.00$6.00$0.251M
vercel$2.00$6.00$0.25262K128K
cortecs$2.50$6.00$0.625262K262K
huggingface$2.50$6.25262K131K
merge-gateway$2.50$6.25$0.50262K1M

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02openrouterCache read$0.20
2026-09-03openrouterCache read$0.25
2026-09-06openrouterCache read$0.25
2026-09-02together_aiCache read$0.50
2026-09-04together_aiCache read$0.25
2026-09-02openrouterInput$2.00
2026-09-06openrouterInput$2.00
2026-09-02together_aiInput$2.50
2026-09-04together_aiInput$2.00
2026-09-02openrouterOutput$6.00
2026-09-06openrouterOutput$6.00
2026-09-02together_aiOutput$6.25
2026-09-04together_aiOutput$6.00

More from Alibaba

most-hosted first