modelbenchmark.io

Qwen3.5 Flash

Alibaba · qwen3-5-flash

Compare

Qwen vision-language model for visual reasoning, documents, and agent tasks

Specification

alibaba-cn listing

Context
1M
Max output
66K
Released
2026-02-23
Knowledge cutoff
2025-04
Retires
Open weights
no
Input
text, image, video
Output
text

Price

US dollars per million tokens · alibaba-cn list price

Input
$0.172
Output
$1.72
Cache read
$0.05
Cache write
$0.125

Quality

8 benchmarks · 1 source

Composite
47th
percentile of 334 scored models
Rank
178 / 334
±0.06 sd
Evidence
8 × 1
one source, or fewer than 3 benchmarks
Effort range
one configuration only
reasoning75th3/4 bench
math38th4/7 bench
knowledge17th1/1 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
OTIS Mock AIME 2024-202584.4 ±5.5defaultEpoch AI2026-08-07
GPQA diamond82.3 ±2.7defaultEpoch AI2026-08-07
Chess Puzzles21.0 ±4.1defaultEpoch AI2026-08-07
Mystery Game Puzzles20.0 ±4.0defaultEpoch AI2026-08-27
FrontierMath-Tier-4-2025-07-010.0 ±0.0defaultEpoch AI2026-05-12
FrontierMath-v18.1 ±1.4default · 2 runsEpoch AI2026-05-12
SimpleQA Verified20.3 ±1.3defaultEpoch AI2026-08-27
FrontierMath-Tiers-1-3-v29.5 ±1.7defaultEpoch AI2026-08-28

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 10 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
alibaba-cnvendor$0.172$1.721M66K
merge-gateway$0.029$0.287$0.00581M250K
empiriolabs$0.09$0.368$0.091M33K
nano-gpt$0.10$0.40$0.05992K66K
nano-gpt · thinking$0.10$0.40$0.05992K66K
ofox$0.10$0.40$0.01$0.1251M64K
ofox · bailian$0.10$0.40$0.01$0.1251M66K
orcarouter$0.10$0.401M66K
vercel$0.10$0.40$0.01$0.1251M64K
zenmux$0.10$0.401M1M

Price history

append-only observations · a listing writes a row only when its price moves

DateHostField$/M
2026-09-02vercelCache read$0.001
2026-09-15vercelCache read$0.01

More from Alibaba

most-hosted first