Schematron V2 Small
AvailableThe quality-oriented half of Inference.net's Schematron V2 pair, listed Sep 12 2026 (inference-net/schematron-v2-small). Also a 3B-parameter HTML-to-JSON extraction model, but tuned to hold up on complex schemas and long pages where the throughput-optimized Turbo variant degrades. Like Turbo it is schema-driven — the extraction target goes in a JSON schema via response_format rather than the prompt — with a 128K-token context and 4,096 max output tokens, text in and text out. Proprietary and API-only via Inference.net and OpenRouter, at $0.05 input / $0.23 output per Mtok with cache reads at $0.05; OpenRouter reported ~1.74s P50 latency at listing. Inference.net reports an LLM-as-judge quality score of 4.060 and 83.10 on SimpleQA, both slightly ahead of Turbo's 4.039 / 79.42 — the trade the two variants are meant to express. In scope as a narrow document/structured-extraction LLM, consistent with the Cohere Parse 5 and North Micro Vision precedent.
Specifications
- License
- Proprietary
- Weights
- Not released
- Architecture
- unknown
- Parameters
- 3B
- Context window
- 128K tokens
- Max output
- 4K tokens
- Knowledge cutoff
- —
- Price (in / out, $/M)
- $0.05 / $0.23
- Modalities
- Text
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.