modelbenchmark.io

Qwen3.8 Flash Next

Alibaba · qwen3-8-flash-next

Compare

Open-weight experimental preview of the Qwen4 architecture: hybrid-attention MoE (125B total, 6B active) with vision encoder for coding, agent tasks, and image and video understanding

Specification

most-agreed values

Context
262K
Max output
262K
Released
2026-08-27
Knowledge cutoff
Retires
Open weights
yes
Input
text, image, video
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.20
Output
$0.50
Cache read
$0.05

Quality

1 benchmark · 1 source

Composite
76th
percentile of 334 scored models
Rank
81 / 334
±0.53 sd
Evidence
1 × 1
too little evidence to place with confidence
Effort range
one configuration only
general61st1/1 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
LiveBench77.3default · 23 runsLiveBench2026-06-25

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 3 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
amd$0.15$0.47$0.016262K131K
requesty$0.20$0.50$0.05262K262K
cortecs$0.201$0.50$0.05262K64K

More from Alibaba

most-hosted first