With the right choices, any SME can run a professional AI server on‑premise. We explain what matters and give concrete recommendations based on our practical experience.
Introductions
Many customers ask us: Are the top‑tier AIs only available in the cloud? Aren’t there affordable on‑premise options? While cloud giants bombard us daily with tempting AI offers, on‑prem solutions for SMEs still seem to lag behind.
For Swiss companies in fields like private banking, fiduciary services, law, or medical technology, a professional on‑premise infrastructure is often the only viable option: Risks in the cloud, such as data leaks or industrial espionage, can devastate a hard‑won reputation. Understandably, many SMEs prefer not to take that risk.
Uncertainty is high: Which concrete on‑prem solutions for SMEs are succeeding today? How do they compare to the cloud in performance, and what do they cost? We see too many offerings that run on unsuitable hardware. These lead to disappointment and fuel the misconception that professional AI cannot be run on‑premise.
In this article, we want to bring clarity and, based on our many years of practical experience, spell out exactly what matters for a successful on‑premise AI solution.
Contents
- Choosing the right AI for professional on‑premise applications
- AI server vs. running AI directly on your workstation
- Matching hardware with the ideal price‑performance ratio
- Modular server software for smooth operations
Choosing the Right AI for Professional On‑Premise Applications
Are the best AIs only in the cloud? It can seem that way when listening to cloud providers. The fact is: Many of today’s leading cloud AIs are open source, with licenses that even allow commercial use at no cost. The selection is almost too large, and new open‑source models appear continuously. So which model should you choose? What should you optimize for?
One factor has proven to be the most reliable starting point in practice: model size in GB (checkpoint). Important: model size in GB is not the same as parameter count. Parameter counts can be misleading, for example when models are heavily quantized. Model size in GB is a decisive factor both for a model’s intelligence and for hardware selection.
Model sizes at a glance
The following four models are not intended as a permanent leaderboard. They are a practical snapshot. The number of active parameters is not the only factor that matters. With Mixture of Experts models, inactive experts must also remain in memory. Hardware planning therefore needs to account for the specific checkpoint format, context memory, and inference server headroom.
| Category | Compact | Small | Medium | Large |
|---|---|---|---|---|
| Example model | Granite 4.1 8B | Gemma 4 31B | DeepSeek V4 Flash | GLM 5.2 |
| Architecture | 8B dense | 31B dense | 284B MoE, 13B active | approximately 743B MoE, 39B active |
| Typical use | Tests and technical experiments | General assistance, RAG, multimodal tasks | Demanding agents, coding, long context | Complex long running tasks and parallel enterprise operations |
| Typical server class | Compact system or UMA | 1 × RTX PRO 6000 Blackwell | 2 × RTX PRO 6000 Blackwell | 8 × RTX PRO 6000 Blackwell |
| Practical assessment | unsuitable for business applications | strong SME all rounder | very high quality with good efficiency | top tier performance with substantial infrastructure requirements |
Granite 4.1 8B from IBM is a modern Apache 2.0 licensed model with support for German, RAG, extraction, and function calling. Its limited model size restricts task comprehension, reliability, and subject matter depth. We therefore do not consider this class a viable foundation for business applications.
Gemma 4 31B is a dense multimodal model from Google DeepMind. It supports more than 140 languages and up to 256,000 context tokens. An RTX PRO 6000 Blackwell with 96 GB of VRAM provides enough memory for a production oriented deployment, including reasonable headroom for context and runtime requirements.
DeepSeek V4 Flash is a Mixture of Experts model with 284 billion total parameters, of which only 13 billion are active per token. Depending on the variant, the published checkpoint requires approximately 160 GB of memory. Two RTX PRO 6000 Blackwell GPUs with a combined 192 GB of VRAM are therefore an attractive server class. The usable context length depends heavily on quantization, KV cache, and concurrency. The theoretically supported one million tokens do not automatically constitute a sensible production configuration.
GLM 5.2 is designed for long running agent, coding, and reasoning workloads. The model contains approximately 743 billion parameters, with around 39 billion active. The official FP8 checkpoint targets eight GPUs with more memory per GPU. Running it on eight RTX PRO 6000 Blackwell GPUs therefore requires a suitable NVFP4 quantization. This class demands careful validation of the inference engine, PCIe topology, cooling, and power supply.
Models that are too small
Even modern 8B models are too small for production business applications. Their answers can appear convincing in simple demonstrations. In daily work, however, they are more likely to misunderstand tasks, miss important relationships, and produce unreliable results for demanding professional questions. The resulting review and correction effort eliminates the apparent cost advantage.
We therefore advise against using this model class as a business assistant or as the foundation for automated processes. It is suitable for technical experiments and noncritical tests, but not for production SME applications. For a professional starting point, we recommend at least the next class at approximately 31 billion parameters.
Small models
Models around 31 billion parameters now provide a compelling entry point for professional on premise AI. They offer substantially more capacity for reasoning, RAG, multilingual work, and multimodal tasks than the 8B class. Gemma 4 31B fits on one professional 96 GB GPU, making it a strong SME all rounder with manageable infrastructure requirements.
Medium sized models
Mixture of Experts models such as DeepSeek V4 Flash combine a large knowledge capacity with relatively few active parameters per token. This enables high quality and speed without computing every parameter for each output token. The full checkpoint must still fit in GPU memory. Two professional 96 GB GPUs form a sensible performance class for demanding agents, software development, and extensive document context.
Large models
For demanding domains such as private banking, fiduciary services, medical technology, or defense, large models and powerful multi GPU servers are appropriate. GLM 5.2 targets complex long running tasks, agents, and coding workloads. A configuration with eight RTX PRO 6000 Blackwell GPUs provides 768 GB of VRAM, but this model still requires a suitable quantization. In this class, testing with your own data, a coordinated security concept, and a professional operating environment are essential.
Model size alone does not guarantee factual accuracy. Large models can also make mistakes. Business critical applications therefore require source citations, access controls, quality assurance, and human approval for consequential decisions.
AI server vs. running AI on your workstation?
Professional AI use in a company requires a stable and high‑performance infrastructure. We therefore recommend running your AI workloads on a dedicated server configured specifically for these compute‑intensive tasks and operating independently of everyday work devices.
Separation of concerns
Stability is crucial in professional environments: That’s why we advise our clients to run their AI on a dedicated server with suitable hardware. This reduces risks such as system overloads and ensures other business‑critical applications can continue running undisturbed.
Office PCs or laptops are designed for user interfaces and low‑compute software such as office applications. Large AI models, by contrast, present immense computational demands and require a different software and hardware environment.
A clean separation is both more stable and more flexible. If the AI runs on its own server, updates, maintenance, and backups can be performed centrally and independently of the workstation. You keep control over office computers while the AI continues to work reliably.
Noise and heat
Compute means heat, which means cooling, which in most cases means noise. Similar to gaming PCs, medium and large AI models produce a lot of heat and therefore require noticeable cooling. Gamers with headphones may not mind, but in an office you need to concentrate. Cooling a professional 600 W AI system is akin to a small vacuum cleaner and can be distracting.
Ideally, you have a server room or a lockable room such as a basement or even a storage closet to place the AI server. Smaller server clusters often get by without room air‑conditioning. If you must place an AI server in a work area, water cooling is advisable. It still uses fans and is somewhat more expensive, but is significantly quieter at the same performance because the heat can be dissipated over a larger surface area.
Background tasks
A dedicated AI server can continuously do preparatory work in the background. There are various types of such background tasks. One is document indexing, which can make search via AI significantly more efficient and successful. Documents are embedded into a vector space so that semantic similarities can be recognized and queries can be answered more precisely.
Another form of background processing that boosts SME efficiency is automated document handling. For example, incoming emails can be analyzed to automatically generate structured reports. Ongoing quality and compliance checks can also run in the background, saving resources while ensuring standards are met.
Matching hardware with the ideal price to performance ratio
Professional AI models primarily require enough fast memory. The model checkpoint, the KV cache for context, runtime buffers, and concurrent requests must all fit at the same time. Parameter count alone is not sufficient for sizing a system.
There are two sensible platform classes for on premise AI: compact systems with shared memory and dedicated GPU servers. The first class offers a large amount of addressable memory with low power consumption. The second provides substantially more compute performance, higher memory bandwidth, and the more mature NVIDIA software stack for production inference.
Systems with shared memory
| NVIDIA DGX Spark | AMD Ryzen™ AI Halo | Apple Mac Studio | |
|---|---|---|---|
| Processor platform | NVIDIA GB10 Grace Blackwell | Ryzen AI Max+ 395 | M4 Max or M3 Ultra |
| Shared memory | 128 GB | up to 128 GB | up to 128 GB with M4 Max, up to 512 GB with M3 Ultra |
| Memory bandwidth | 273 GB/s | 256 GB/s | up to 546 GB/s with M4 Max, 819 GB/s with M3 Ultra |
| AI software | CUDA, NVIDIA AI stack | ROCm | MLX, Metal |
| Platform power | 240 W power supply | up to 120 W TDP | depends on configuration |
| Strength | CUDA in a compact system | strong balance of memory, price, and efficiency | very large memory capacity and quiet desktop operation |
| Limitation | much lower memory bandwidth than a dedicated RTX PRO GPU | verify software compatibility in advance | different software stack from typical Linux GPU servers |
These systems are particularly interesting for development, pilot projects, quiet single user solutions, and large quantized models with moderate throughput. A large shared memory pool does not automatically deliver high speed.
Professional GPU servers
| 1 GPU | 2 GPUs | 8 GPUs | |
|---|---|---|---|
| GPU | 1 × RTX PRO 6000 Blackwell Server Edition | 2 × RTX PRO 6000 Blackwell Server Edition | 8 × RTX PRO 6000 Blackwell Server Edition |
| Total VRAM | 96 GB GDDR7 with ECC | 192 GB GDDR7 with ECC | 768 GB GDDR7 with ECC |
| Theoretical aggregate memory bandwidth | 1.6 TB/s | 3.2 TB/s | 12.8 TB/s |
| Maximum GPU power | 600 W | 1,200 W | 4,800 W |
| Suitable example model | Gemma 4 31B | DeepSeek V4 Flash | GLM 5.2 with suitable NVFP4 quantization |
| Typical role | powerful SME inference server | demanding agents and higher throughput | central AI platform for complex models and many users |
The RTX PRO 6000 Blackwell Server Edition is a professional data center GPU with 96 GB of GDDR7 ECC memory and up to 1.6 TB/s of memory bandwidth. Unlike consumer GPUs, it is designed for dense server configurations, continuous operation, and enterprise workloads. Whether a specific configuration reaches its full performance depends on the server chassis, power supply, cooling, PCIe topology, and inference software.
Prices and throughput change quickly and depend heavily on model format, context length, and concurrency. We therefore do not recommend a general words per second figure. Reliable procurement starts with a benchmark of the desired model on the intended server configuration.
Graphics cards
Dedicated GPUs remain the most powerful standard solution for interactive and parallel AI inference. Their strength comes from the combination of high memory bandwidth, specialized tensor compute units, and mature inference frameworks.
For professional on premise systems, we recommend server GPUs with ECC memory, documented cooling requirements, and a clear support path. Consumer graphics cards can appear attractive in laboratory settings, but their smaller memory capacity, chassis requirements, and lack of suitability for dense multi GPU servers make them an unsuitable foundation for the SME systems recommended here.
Unified Memory Architecture (UMA)
In a conventional GPU server, the CPU has its system memory and every GPU has its own VRAM. The model must be loaded into GPU memory and divided across multiple GPUs where necessary. This design provides very high bandwidth and compute performance, but imposes clear limits based on the available VRAM on each card and the connections between GPUs.
With a Unified Memory Architecture, the CPU and integrated accelerators access a shared memory pool. This removes the need to copy large volumes of data between separate RAM and VRAM pools. More importantly, a large share of system memory can be made available to the model. This makes compact systems with 128 GB or more of shared memory attractive for large quantized checkpoints.
The three platforms take different approaches. NVIDIA DGX Spark combines 128 GB of coherent memory with CUDA and the NVIDIA software stack. AMD Ryzen™ AI Halo offers up to 128 GB of shared LPDDR5x memory, ROCm, and support for Linux and Windows. Apple Mac Studio provides up to 512 GB of Unified Memory with M3 Ultra and high memory bandwidth, but primarily uses MLX and Metal rather than CUDA for local AI.
UMA is not automatically faster. DGX Spark and Ryzen AI Halo reach approximately 273 GB/s and 256 GB/s respectively. A single RTX PRO 6000 Blackwell reaches up to 1.6 TB/s. Mac Studio with M3 Ultra sits between them at 819 GB/s. Shared memory systems therefore excel primarily in capacity, efficiency, noise, and compact design. Dedicated GPU servers excel in throughput, concurrency, and the breadth of production inference software.
A useful rule is straightforward: If a model, its context, and sufficient headroom fit on a professional GPU, that GPU is generally the faster server solution. If very large memory capacity is needed at moderate throughput, UMA can be more economical and simpler.
Clusters
In multi GPU systems, model weights and computation are distributed across several accelerators. The available memory only combines effectively when the inference software parallelizes the model correctly. Communication between GPUs also introduces overhead.
The RTX PRO 6000 Blackwell Server Edition connects through PCIe and does not include NVLink. Systems with two or eight GPUs therefore require careful planning of motherboard topology, PCIe lanes, NUMA allocation, and NCCL configuration. Eight cards installed in eight arbitrary computers are not automatically comparable to a coordinated eight GPU server.
Additional GPUs can serve two different goals. A large model can be divided through tensor or expert parallelism, or multiple model instances can increase throughput for concurrent users. The second approach generally scales more easily. Reliable planning requires combined testing of time to first token, output tokens per second, context length, and concurrent requests.
Modular server software for smooth operations
Everyone knows the problem: Windows needs another update and then suddenly something stops working. The reason is that programs from different developers must coexist on a computer without interfering with each other. Dependencies can get out of sync during updates.
Operating systems like Windows, Linux, or macOS have been continuously improved for decades. In professional environments, however, there is generally zero tolerance for incompatibilities. Additional measures are therefore used to maximize stability and process safety, and thereby reduce maintenance effort and outages.
Docker Compose
Photo by Tom Fisk
Probably the best‑known and most widely used platform for modular operation of complex software is “Docker Compose.” Docker is based on the concept of “containers,” analogous to shipping containers in global trade. Thanks to strict standards for form factors and static requirements, containers fit on different ships, trains, and trucks worldwide and can be stacked regardless of contents. They are deployable anywhere.
Software containers are likewise independent of their environment. For companies, this means different applications can run isolated and safely without conflicts. Docker Compose is particularly suitable for smaller SME setups and for quickly entering the world of professional AI.
With a compact docker-compose.yml file, multiple services, such as an AI API, a database, and a web frontend, can be started with a single command. This significantly reduces complexity and enables IT teams without deep DevOps experience to set up a stable on‑premise infrastructure. Particularly convincing: Updates and rollbacks are very easy with Compose, which significantly increases operational reliability in SMEs.
Typical use cases in SMEs include:
- Provisioning single AI instances for internal use
- Document search with Elasticsearch or MeiliSearch combined with an AI interface
- Small pilot projects that can later scale to Kubernetes
Kubernetes
As soon as dozens of users access an AI infrastructure simultaneously, or different models need to run in parallel, Kubernetes is the professional solution. Kubernetes provides automatic scaling, load distribution, and self‑healing when a service fails. This allows stable operation of larger setups with multiple GPU servers in a cluster.
For SMEs, getting started with Kubernetes is more complex, but it brings clear advantages:
- Central management of clusters with multiple AI servers
- Rolling out updates without downtime
- Integration of load balancers, secrets, and monitoring tools
- Flexible extension with additional services such as vector databases or API gateways
Especially in sensitive industries like banking and healthcare, where high availability is mandatory, Kubernetes is a solid foundation. In smaller environments, a hybrid approach often makes sense: start with Docker Compose during the pilot phase, then scale to Kubernetes later.
Remote maintenance
For on‑premise systems, professional maintenance is highly recommended. Outages can massively disrupt business operations. Remote maintenance tools ensure that system administrators always have an overview of performance and can intervene early in the event of problems.
Important building blocks for remote maintenance are:
- Prometheus + Grafana for detailed metrics (GPU utilization, memory usage, network load)
- Alertmanager with email, SMS, or Teams/Slack notifications for critical states
- Remote logging solutions such as Loki or the ELK stack to analyze root causes after the fact
- VPN tunnels for secure access to dashboards and maintenance systems, even outside the office
A modular maintenance setup provides not only security but also transparency: Decision‑makers can track the efficiency of their on‑premise AI at any time and thus make a stronger case to leadership that the investment is worthwhile.