modelbenchmark.io

Schematron V2 Turbo

schematron-v2-turbo

Compare

Inference.net's 3B-parameter HTML-to-JSON extraction model, optimized for throughput and low cost on high-volume workloads. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.

Specification

most-agreed values

Context
128K
Max output
8K
Released
2026-09-12
Knowledge cutoff
Retires
Open weights
no
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.03
Output
$0.15
Cache read
$0.015

Available from 3 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
kilo · inference-net$0.03$0.15$0.03128K8K
nano-gpt · inference-net$0.03$0.15$0.015128K8K
openrouter · inference-net$0.03$0.15$0.03128K8K