We build on NVIDIA GPUs, from the enterprise server to the AI data center. The onprem.ai software supports the full spectrum:
- NVIDIA RTX PRO 6000 Blackwell: Blackwell GPU with an excellent price-performance ratio for inference, the basis of the configurations S1, S2, DC4, and DC8
- NVIDIA H100: proven Hopper data center GPU for demanding inference and fine-tuning workloads
- NVIDIA H200: Hopper GPU with extended memory and higher bandwidth for more context and larger models per GPU
- NVIDIA GH200: Grace Hopper superchip for the highest demands
Why this selection?
Professional LLM inference needs above all memory capacity and memory bandwidth. The GPUs we deploy offer both, across a spectrum from cost-efficient team servers to maximum-performance scenarios, with mature drivers and broad software support.
Our configurations
Entry starts with the Server S1 and grows in 2-GPU steps: S2 (2x RTX PRO 6000 Blackwell, 192 GB VRAM) for whole teams, DC4 and DC8 for data centers. All details on the server page.
Note: GPU prices are currently volatile. What your configuration costs is shown by the cost calculator, binding figures come with your quote.
Next steps
- See the software with all supported GPUs
- See the servers with all configurations
- Contact us for hardware advice