Open Llama multimodal model for image understanding and text reasoning
Specification
most-agreed values
Context
128K
Max output
4K
Released
2024-09-18
Knowledge cutoff
2023-12
Retires
2026-06-13
Open weights
yes
Input
text, image
Output
text
Price
US dollars per million tokens · most-agreed
Input
free
Output
free
Rate limits
Requests per minute
300
Available from 9 hosts
| Host | In $/M | Out $/M | Cache rd | Cache wr | Context | Output | Retires |
|---|---|---|---|---|---|---|---|
| nvidia | free | free | — | — | 128K | 4K | — |
| cloudflare | $0.0485 | $0.676 | — | — | 128K | 128K | — |
| cloudflare-workers-ai · cf | $0.0485 | $0.676 | — | — | 128K | 128K | — |
| deepinfra | $0.049 | $0.049 | — | — | 131K | 131K | — |
| inference | $0.055 | $0.055 | — | — | 16K | 4K | — |
| edenai | $0.345 | $0.345 | — | — | 131K | 4K | — |
| watsonx | $0.35 | $0.35 | — | — | 128K | 128K | — |
| azure_ai | $0.37 | $0.37 | — | — | 128K | 2K | 2026-06-13 |
| oci · oci | $2.00 | $2.00 | — | — | 128K | 4K | — |
Other listings of this model
same model, different key
More from Meta
most-hosted first