Schematron V2 Turbo
AvailableA 3B-parameter HTML-to-JSON extraction model from Inference.net, listed Sep 12 2026 (inference-net/schematron-v2-turbo). Part of the company's "workhorse model" line — small, purpose-built LLMs sold on cost per unit of work rather than general capability — and tuned for maximum throughput on high-volume scraping and ingestion pipelines, reported at ~4.14 requests/second on a single H100. It turns messy HTML into clean structured JSON, and is schema-driven rather than prompt-driven: the extraction target is supplied as a JSON schema in the response_format parameter, not in the system or user prompt. 128K-token context, 8,192 max output tokens, text in and text out. Proprietary and API-only via Inference.net and OpenRouter, at $0.03 input / $0.15 output per Mtok. Inference.net reports an LLM-as-judge quality score of 4.039 and 79.42 on SimpleQA — positioning it as the cheaper, faster half of the V2 pair against the higher-quality Small. In scope as a narrow document/structured-extraction LLM, consistent with the Cohere Parse 5 and North Micro Vision precedent.
Specifications
- License
- Proprietary
- Weights
- Not released
- Architecture
- unknown
- Parameters
- 3B
- Context window
- 128K tokens
- Max output
- 8K tokens
- Knowledge cutoff
- —
- Price (in / out, $/M)
- $0.03 / $0.15
- Modalities
- Text
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.