modelbenchmark.io

llama-3-8b

Meta · llama-3-8b

Compare

Specification

most-agreed values

Context
8K
Max output
8K
Released
Knowledge cutoff
Retires
Open weights
Input
Output

Price

US dollars per million tokens · most-agreed

Input
$0.05
Output
$0.25

Quality

4 benchmarks · 1 source

Composite
7th
percentile of 334 scored models
Rank
311 / 334
±0.07 sd
Evidence
4 × 1
one source, or fewer than 3 benchmarks
Effort range
one configuration only
reasoning9th2/4 bench
math2nd2/7 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
Chess Puzzles0.0 ±0.0defaultEpoch AI2026-08-28
OTIS Mock AIME 2024-20251.9 ±1.0defaultEpoch AI2026-08-30
MATH level 56.1 ±0.4defaultEpoch AI2025-01-27
GPQA diamond26.1 ±1.7defaultEpoch AI2025-01-27

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 2 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
replicate · replicate$0.05$0.258K8K
vercel_ai_gateway · vercel-ai-gateway$0.05$0.088K8K

Other listings of this model

same model, different key

More from Meta

most-hosted first