Swiss Engineering Server S2: enterprise AI for your whole team
Our most popular server: 2x NVIDIA RTX PRO 6000 Blackwell with 192 GB VRAM, preinstalled onprem.ai software, and the latest LLM models. Runs standalone, scales as a cluster.
One server that serves an entire team
Designed for teams in regulated industries such as legal, finance, and healthcare. All data is processed exclusively on premises, fully air-gapped if required. Thanks to OpenAI-compatible APIs, the S2 replaces cloud AI services without code changes.
- GPUs
- 2x NVIDIA RTX PRO 6000 Blackwell
- VRAM
- 192 GB @ 1.8 TB/s
- AI performance
- 8000 AI TOPS
- Power draw
- ~1.8 kW max
- Form factor
- 4U rack
- Users
- 5 - 20 concurrent
Built for speed
A single S2 runs DeepSeek v4 flash with the full 1M context and serves an entire team at once. The key measurements:
Text generation under parallel requests, 200 tokens/s single-stream
Input processing under parallel requests
Context window, with a 1.4M token KV cache
Concurrent in office use, plus 10 parallel coding agents
Document classification in continuous operation
Generated tokens, plus 26 bn processed input tokens
Preliminary measurements with DeepSeek v4 flash on the Server S2, as of August 2026. Verification in progress.
More configurations
All configurations run the same software and can operate standalone or join a cluster. Expand capacity as soon as your needs grow.
- 1x Blackwell GPU
- Nvidia RTX6000WS
- 96 GB VRAM @ 1.8 TB/s
- 4000 AI TOPS
- ~1 KW max
The compact entry point: one RTX PRO 6000 Blackwell in a 2U chassis, ideal for first production workloads and pilot projects.
- 4x Blackwell GPU
- Nvidia RTX6000S
- 384 GB VRAM @ 1.6 TB/s
- 16000 AI TOPS
- ~3 KW max
Four GPUs in one system for departments with several parallel workloads, from RAG and OCR to coding agents.
- 8x Blackwell GPU
- Nvidia MGX RTX6000S
- 768 GB VRAM @ 1.6 TB/s
- 32000 AI TOPS
- ~5.4 KW max
Eight GPUs for company-wide use: large models, high concurrency, and headroom for new use cases.
- 8x Blackwell GPU
- Nvidia DGX B200
- 1440 GB VRAM @ 8 TB/s
- 144000 AI TOPS
- ~14.3 KW max
NVIDIA's reference system for the datacenter: eight B200 GPUs with NVLink, for training, fine-tuning, and inference of the largest models.
- 2U rack
- Up to 3 GPUs
- 2-GPU configuration: Server S2
The compact 2U chassis behind S1 and S2: fitted with one or two RTX PRO 6000 Blackwell and quiet enough to run in existing server rooms.
- 4U rack
- Up to 8 GPUs
- 8-GPU configuration: DC8
The large 4U chassis behind DC4 and DC8: up to eight GPUs per system, built for high density in the datacenter and growth into an AI cluster.