Member of Inception Program

Apertus v1.5 70B on a single RTX PRO 6000 Blackwell

Apertus v1.5 70B ships as approximately 145 GB (135 GiB) of BF16 weights that officially need two H200 datacenter GPUs. We show how to run it on a single NVIDIA RTX PRO 6000 Blackwell with 96 GB VRAM, with vision input, native tool calling, and a 192k context profile (229k measured maximum). Using three-tier mixed-precision quantization (NVFP4 + FP8) it decodes 55% faster than FP8 while using 32% less weight storage, with 99% of FP8 MMLU quality.