modelbenchmark.io

Holo3-35B-A3B

holo3-35b-a3b

Compare

Compact GPT model for low-latency assistance and high-volume workloads

Specification

most-agreed values

Context
66K
Max output
8K
Released
2024-01-01
Knowledge cutoff
Retires
Open weights
yes
Input
text, image
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.25
Output
$1.80
Cache read
$0.125

Available from 2 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
nano-gpt$0.25$1.80$0.12566K8K
nano-gpt · thinking$0.25$1.80$0.12566K8K