ZDTaichu5.0-9B
AvailableZDTaichu5.0-9B is an open-weights spatial vision-language model from the Taichu (ZiDong Taichu) series, published ~2026-09-24 under the NVIDIA Open Model License (retaining Qwen3.5 Apache-2.0 and third-party notices). ~10B parameters combining a Qwen3.5-9B decoder with an NVIDIA C-RADIOv4-H vision encoder. Multimodal understanding: accepts text, single or multiple images and video at any resolution and outputs text, with a focus on spatial reasoning. Up to 128K-token context. Reported benchmarks include ViewSpatial 62.5, MMSI 47.2, MindCube-tiny 78.3, TAU2 87.7, Claw-Eval 71.4 and IFEval 93.7. A derivative built on a Qwen backbone, added for parity with other tracked multimodal-understanding VLMs.
Specifications
- License
- Open weights · NVIDIA Open Model License
- Weights
- Downloadable
- Architecture
- Dense
- Parameters
- 10B
- Context window
- 131K tokens
- Max output
- —
- Knowledge cutoff
- —
- Price (in / out, $/M)
- —
- Modalities
- TextVisionVideo
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.