Member of Inception Program

Schweizer FlaggeSwiss Engineering The operating system for AI in your local datacenter

Our software connects servers, GPUs, models, APIs, monitoring, and updates in one platform. Fully on premises, even in air-gapped environments.

Put models into operation safely

The model catalog only shows configurations that match the hardware you have. Before a launch, the platform checks free VRAM and blocks what does not fit, before a mistake reaches production.

  • Servers: Nodes with Kubernetes runtime, VRAM allocation, and performance and temperature timelines in a single view.
  • Model catalog: Multimodal from text through vision and OCR to embeddings, with clear maturity levels from stable to unsupported.
  • Deployment: Tested memory profiles, versioned manifests, and readiness probes make every deployment reproducible.

Protected interfaces for your applications

Every model is served through an authenticated gateway, including OpenAI-compatible APIs as a drop-in replacement for cloud services.

  • API gateway: Authentication, routing, usage attribution, and overload protection in one place.
  • API keys: Named keys with last usage and token consumption, instantly renewable or revocable.
  • Playground: Tests through the real production path and delivers the integration code along with it.

Keep operations in view at all times

Status, alerts, and usage put API behavior, pod events, and GPU activity on the same timeline. That shortens root-cause analysis and gives incident reviews a shared factual basis.

  • Status: Endpoint history, time to first token, and cluster events correlated on one timeline.
  • Alerts: Grouped by AI/GPU, Kubernetes, and host, with evidence and handover to the DevOps agent.
  • Usage: Tokens, KV cache, GPU energy, and rejection reasons, groupable by model and API key.

Full control, even without an internet connection

Security controls are visible in the product and mapped to ISO 27001 and SOC 2; open items stay clearly marked. Updates are always started by an operator: online or via USB package in an air gap.

  • Governance: Control status for identity, RBAC, TLS, audit logging, and high availability.
  • Identity: OAuth2 and OIDC via Keycloak connect your existing enterprise identities.
  • Updates: Versioned releases with history and logs, for connected and air-gapped sites.

Verified hardware for reliable operations

Our software is hardware-aware: the model catalog, the memory profiles, and the deployments refer to concrete GPUs, standalone or clustered across multiple nodes. These GPU families and Dell chassis are approved for operation:

NVIDIA RTX PRO 6000 Blackwell

NVIDIA RTX PRO 6000 Blackwell

96 GB GDDR7

Blackwell GPU with an excellent price-performance ratio for inference. The basis of the S1, S2, DC4, and DC8 configurations.

NVIDIA H100

NVIDIA H100

80 GB HBM3

Proven Hopper datacenter GPU for demanding inference and fine-tuning workloads.

NVIDIA H200

NVIDIA H200

141 GB HBM3e

Hopper GPU with expanded memory and higher bandwidth: more context and larger models per GPU.

NVIDIA GH200

NVIDIA GH200

Up to 144 GB HBM3e + 480 GB LPDDR5X

Grace Hopper superchip: CPU and GPU with shared memory access, ideal for very large models and memory-intensive pipelines.

NVIDIA B200

NVIDIA B200

180 GB HBM3e

Blackwell datacenter GPU for maximum inference throughput, the basis of the DGX B200.

NVIDIA B300

NVIDIA B300

288 GB HBM3e

Blackwell Ultra GPU with expanded memory for the largest models and longest contexts.