Executive Summary: SUSE (Rancher Prime + SUSE AI + SUSE AI Factory with NVIDIA + SLES 16)

SUSE is an open-source infrastructure vendor that turned its Kubernetes into an AI platform by curating what runs on it, and the map reads it strongest where Rancher always was. Layer 2A is strong on multi-cluster Kubernetes with AI-conformant RKE2, virtual clusters that hand out GPU slices under quotas, and an Apache 2.0 AI Factory operator; six layers are moderate (a hardware-agnostic abstraction with SUSE Linux Enterprise Server (SLES) 16 and Multi-Instance GPU (MIG) capable virtualization at Layer 0; Longhorn and MinIO as platform storage at 1A; supported open-source vector stores plus NVIDIA's retrieval blueprint at 1B; NVIDIA's fixed retrieval-augmented generation (RAG) ingest pipeline at 1C; open inference engines plus NVIDIA AI Enterprise through the Factory at 2B; observability and Rancher AI's supervisor and human-validation gates at 2C; Liz, the Rancher Prime MCP server, and SUSE Industrial Edge at Layer 3), and none is a gap.

The capture is the lightest on the map, because SUSE sells support for software the enterprise could run without it. Twenty-eight components read Retained 13, Delegated 3, Ceded 12: SLES, Harvester, RKE2, K3s, K3k, Fleet, Longhorn, MinIO, Milvus, Qdrant, OpenSearch, Open WebUI, vLLM, Ollama, LiteLLM, NeuVector, and the AI Factory operator are Apache 2.0 or GPL projects the enterprise operates itself; what's SUSE's is Rancher Prime's extensions, the Application Collection's hardened images and lifecycle, the AI Factory's subscription catalog and blueprints, SUSE Observability, Rancher AI, and Industrial Edge, and what's NVIDIA's through SUSE's paper (NIM, NeMo, Nemotron, the RAG blueprint) is Ceded to NVIDIA except for model access through the standard interface. The decision-authority readings are Delegated at five layers because SUSE's engines document the pins, per-object policies, and confirmation prompts the override rule asks for.

The buyer's trade: a sovereign, air-gappable AI stack on their own hardware with the fewest captive opinions of any row, in exchange for assembling retrieval and pipelines from parts, taking the agent runtime and the retrieval models from NVIDIA, and waiting on the pieces still in preview (Universal Proxy, the Multi-Linux Manager MCP server, the SLES 16 MCP host, OpenShell and NemoClaw, Longhorn v2, Run:ai in the Factory's docs). What would move a cell: a SUSE retrieval engine (1B), Universal Proxy reaching GA with tool-call policy (2C), or the AI Factory's docs naming Run:ai and the NVIDIA agent runtime as supported deployables (2A, 2B).

Layer-by-layer status: Layer 0 (Hardware-Agnostic Abstraction: SLES 16, SUSE Virtualization with MIG, Any Vendor's Metal), Layer 1A (Platform Storage (Longhorn) and S3-Compatible Objects (MinIO), Not a Data Foundation), Layer 1B (Supported Open-Source Vector Stores and RAG in the AI Library; NVIDIA's RAG Blueprint Through the Factory), Layer 1C (NVIDIA's Fixed RAG Ingest Pipeline Through the Factory; No General Pipeline Layer), Layer 2A (Rancher Heritage Strength: Multi-Cluster Kubernetes, AI-Conformant RKE2, Virtual Clusters with GPU Quotas, an Apache 2.0 AI Factory Operator), Layer 2B (Supported Open Inference (vLLM, Ollama) + NVIDIA AI Enterprise (NIM, NeMo, Nemotron) Through the Factory; No SUSE Runtime), Layer 2C (Observability for AI Workloads + Rancher AI's Supervisor, Agent Configs, and Human-Validation Gates; MCP Gateway in Preview), Layer 3 (+1) (Liz and the Crew for Infrastructure Operations, SUSE Industrial Edge for the Plant Floor; Platform-Enabled Beyond).

Assessment framework: 4+1 Layer AI Infrastructure Model. Scoring model: Decision Authority Placement Model (DAPM) — Retained, Delegated, or Ceded. Published by The CTO Advisor LLC (DBA The Advisor Bench). Author: Keith Townsend. Date assessed: September 5, 2026. Version: v1.0 - 4+1 v2: Authority Split.

SUSE (Rancher Prime + SUSE AI + SUSE AI Factory with NVIDIA + SLES 16)

Mapped to the 4+1 Layer AI Infrastructure Model

v1.0 - 4+1 v2: Authority SplitAssessed September 5, 2026Source: SUSE documentation (SUSE AI 1.0 deployment and AI Library; SUSE AI Factory 1.0 monitoring; SUSE AI Factory with NVIDIA overview, building blocks, pre-built blueprints, lifecycle appendix, and release notes; Rancher AI architecture, user, and administrator guides; RKE2 GPU operators and AI conformance; SUSE Virtualization VM creation; SUSE Storage StorageClass parameters and settings; Longhorn eviction; SUSE Virtual Clusters; Vulnerability Scanner MCP server; SLES 16.0 release notes; Application Collection MinIO reference and subscriptions) and product pages (Rancher Prime, SUSE AI, SUSE AI Factory, SUSE Observability, SUSE Virtualization); SUSE press releases and blogs (SUSE AI Factory with NVIDIA generally available, July 9, 2026, and announced April 21, 2026; SUSE Industrial Edge, April 22, 2026; Losant acquisition, February 19, 2026; intelligent infrastructure management update, March 24, 2026; KubeCon EU 2026 posts on the agentic ecosystem, MCP plug-and-play, AI crews, virtual clusters and GPU sharing, MIG in SUSE Virtualization 1.7, SUSE Observability Hosted; secure agentic AI for infrastructure management with partners, April 2026; simplicity, security, and observability for AI workloads, November 11, 2025; AI observability dashboards GA, May 28, 2025; Universal Proxy first release, January 30, 2026; SUSE AI GA at KubeCon North America 2024; SLES 16 launch, November 4, 2025); the SUSE/suse-ai-up, SUSE/aif, rancher/k3k, and harvester/harvester repositories; NVIDIA's NIM Operator, NeMo Retriever, RAG reference architecture, and OpenShell documentation for the NVIDIA components. Peer-reviewed cell by cell through the labs claims ledger (suse-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (suse-row-agy); totals and escalated items in labs/reviews/suse-judgment.md.
ACTIVE ASSESSMENT
Strength
Moderate
Gap
Partner
Layer 0 · ComputeCompute & Network FabricHardware-Agnostic Abstraction: SLES 16, SUSE Virtualization with MIG, Any Vendor's Metaldecides: vendor · Delegated

Raw compute, networking, and acceleration fabric

Vendor-Provided

SUSE Linux Enterprise Server 16 (GA November 4, 2025; Agama, SELinux, 16-Year Lifecycle)Retained

The enterprise operating system under every SUSE path, open source the enterprise runs itself with SUSE's subscription for updates and support. Retained on the open-source seam.

SUSE Virtualization (Harvester 1.7, Apache 2.0; KubeVirt; NVIDIA MIG vGPU GA; Per-VM Node Scheduling)Retained

The hyperconverged virtualization layer running VMs and containers on one Kubernetes, with GPU multi-tenancy and per-VM placement. Open source the enterprise operates: Retained on the open-source seam.

Multi-Vendor Hardware Support (x86, Arm; Fsas, Dell, HPE, Lenovo, Supermicro Servers)Retained

SUSE runs on any certified server the enterprise buys. The enterprise's choice: Retained.

NVIDIA GPU Integration (MIG, vGPU, GPU Operator) Through SUSE's PaperDelegated

NVIDIA's partitioning, drivers, and operator packaged and supported by SUSE. NVIDIA's technology, substitutable at the accelerator: Delegated.

NVIDIA-Provided

NVIDIA GPU Integration in SUSE Virtualization and RKE2; NVIDIA AI Enterprise in the AI Factory

NVIDIA Multi-Instance GPU (MIG) vGPU support is generally available in SUSE Virtualization 1.7 (January 2026, announced March 24, 2026); RKE2 meets the Cloud Native Computing Foundation (CNCF) AI conformance with the NVIDIA GPU Operator installed per NVIDIA's docs; SUSE AI Factory with NVIDIA embeds NVIDIA AI Enterprise, NIM, NeMo, and Nemotron. SUSE sells none of the hardware.

Gap Analysis

SUSE sells the software that makes any vendor's hardware an AI substrate, and nothing below it. SUSE Linux Enterprise Server 16 (generally available November 4, 2025) is the operating system, with the Agama installer, SELinux by default, a sixteen-year lifecycle, and a technology-preview Model Context Protocol (MCP) host; SUSE Virtualization (the KubeVirt-based Harvester, Apache 2.0, 1.7 in January 2026) runs virtual machines and containers side by side on the same Kubernetes with NVIDIA MIG vGPU multi-tenancy, per-VM node scheduling (any node, a specific node, or affinity rules), upgrade control, and VM auto-balancing and live storage migration in early access; both run on whatever x86 or Arm servers the enterprise buys, and the Openchip partnership (2026) points a RISC-V stack at European sovereignty. Fsas Technologies (Fujitsu) is the launch hardware partner for the AI Factory and Vultr the co-engineered cloud. The buyer gets the Nutanix and VMware value proposition from an open-source vendor: an abstraction over their own metal, with the GPU slicing handled. Applying the exposure test: the substrate is the customer's and the abstraction is SUSE's; the MIG partitioning and VM placement the enterprise administers are real Layer 0 controls, on hardware SUSE never touches. Calibration: Nutanix and VMware read moderate on hardware-agnostic abstractions with NVIDIA GPU integration; Supermicro and CoreWeave strong on hardware they sell; Cloudera gap because it ships no abstraction at this layer. SUSE is the Nutanix and VMware case with an open-source hypervisor and operating system. Moderate.

Borrowed Judgment

Bounded and pinnable, on the open-source seam. SLES and SUSE Virtualization are open source the enterprise operates (Retained; the subscription buys updates and support). Multi-vendor hardware support is the enterprise's choice: Retained. NVIDIA's MIG and vGPU integration is NVIDIA's through SUSE's paper: Delegated. The scheduler places VMs and GPU slices within the enterprise's policies, and a per-VM nodeSelector or required affinity rule is a pin the engine must honor: vendor decides, visible, overridable, Delegated under the override rule (the Nutanix and VMware Ceded readings predate the rule and are logged for the reconcile).

Working Notes

Watch-list, notes only: VM auto-balancing and live storage migration (early access); the SLES 16 MCP host (technology preview); the Openchip RISC-V stack (partnership, no product); SUSE as launch partner for the AWS European Sovereign Cloud. Public evidence that moves the cell: SUSE-sold hardware.

Layer 1A · StorageData Storage & GovernancePlatform Storage (Longhorn) and S3-Compatible Objects (MinIO), Not a Data Foundationdecides: vendor · Delegated

Durable, governed data foundation — the Governance Catalog that Layer 2C queries

Vendor-Provided

SUSE Storage (Longhorn, Apache 2.0 CNCF Project; Replicated Block Volumes, Snapshots, Backups, DR, Per-Volume Placement)Retained

Rancher-originated distributed block storage for Kubernetes, bundled with Rancher Prime. Open source the enterprise operates: Retained on the open-source seam.

MinIO via the SUSE Application Collection (S3-Compatible Object Storage, Distributed Mode)Retained

SUSE-built, SUSE-supported MinIO charts installed from the Application Collection or the AI Factory catalog. Self-run open source behind the S3 standard: Retained.

NVIDIA-Provided

No NVIDIA Layer 1A Dependency

Nothing at this layer runs on accelerators.

Gap Analysis

SUSE's data layer is the storage under its Kubernetes. SUSE Storage is Longhorn, the Apache 2.0 Cloud Native Computing Foundation project Rancher Labs created and SUSE now delivers: distributed block storage with replication across nodes, snapshots, backups to S3 or NFS, disaster recovery volumes, per-volume placement (node and disk selectors, data locality, replica auto-balance overriding the global setting) and documented eviction with cancel, bundled with Rancher Prime; the Longhorn v2 data engine is a technology preview. Objects come from the SUSE Application Collection, which builds, tests, and distributes MinIO (S3-compatible object storage, distributed by default) under the Rancher Prime and SUSE AI subscriptions, and the AI Factory installs from that catalog. That's durability for the AI stack's state and an S3 endpoint for its data, and it's where SUSE's data story stops: no native object tier since SUSE Enterprise Storage (the Ceph product) was retired, no catalog, no lineage, no classification, no policy engine over data as an estate. The buyer gets storage they trust under the platform, and brings the data foundation. Applying the test: durable stores, yes, block and object; a governed data foundation Layer 2C could query, no. Calibration: VMware reads moderate on vSAN as platform storage that isn't AI-native; Nutanix moderate on Files, Objects, and Data Lens; NetApp strong on a data foundation with a metadata catalog; Elastic moderate on a tiered store with access governance; CoreWeave moderate on open default storage. SUSE is the VMware case with open-source engines and an S3 endpoint. Moderate, thin.

Borrowed Judgment

Low. Longhorn and MinIO are open source the enterprise runs itself: Retained on the open-source seam, with SUSE's support subscription the only thing that leaves, and MinIO is consumed through the S3 standard. The storage engine places replicas within the enterprise's per-volume selectors and locality settings, which are per-object policies it must honor, and eviction is a documented, cancelable action: vendor decides, visible, overridable, Delegated under the override rule.

Working Notes

Watch-list, notes only: Longhorn v2 data engine (technology preview). Public evidence that moves the cell: a governance catalog on SUSE's paper.

Layer 1B · RetrievalContext Management & RetrievalSupported Open-Source Vector Stores and RAG in the AI Library; NVIDIA's RAG Blueprint Through the Factorydecides: vendor · Delegated

Low-latency retrieval for RAG — vector/hybrid search, context windows

Vendor-Provided

Vector Stores and RAG Front End in the SUSE AI Library (Milvus, Qdrant, OpenSearch, Open WebUI Pipelines)Retained

Open-source retrieval components installed from the AI Library and run by the enterprise, OpenSearch's hybrid query included. Retained on the open-source seam.

SUSE Application Collection Curation (Signed, Hardened Images; SBOMs; Lifecycle; Support)Ceded

SUSE's packaging and support opinions around the open components, the subscription's value. SUSE's: Ceded.

NVIDIA AI-Q with RAG Blueprint (RAG Server, NeMo Retriever Embedding, Nemotron) via SUSE AI Factory with NVIDIACeded

NVIDIA's generally available RAG stack under NVIDIA AI Enterprise on the customer's GPUs through SUSE's paper. Proprietary third-party software: Ceded to NVIDIA.

NVIDIA-Provided

NVIDIA RAG Blueprint, NeMo Retriever Embedding, and Nemotron Models in SUSE AI Factory with NVIDIA

The AI-Q with RAG blueprint deploys NVIDIA's RAG server, ingestor, and the nemotron-vlm-embedding microservice with the Nemotron Super 49B model under NVIDIA AI Enterprise on the customer's GPUs; the SUSE AI Library path runs on any GPU the enterprise has.

Gap Analysis

SUSE's retrieval layer is curated open source with SUSE support, plus NVIDIA's blueprint. The SUSE AI Library, delivered through the SUSE Application Collection as signed, hardened images with software bills of materials, installs Milvus (the default vector database for the stack), Qdrant, and OpenSearch as vector stores, Open WebUI with its pipelines as the retrieval-augmented generation (RAG) front end, and PyTorch (LiteLLM, the model router, is scored at 2B), all deployable air-gapped, with documented OpenTelemetry integration patterns for Milvus and Open WebUI once the applications are instrumented. SUSE AI Factory with NVIDIA (generally available July 9, 2026, co-engineered with Vultr) ships the NVIDIA AI-Q with RAG blueprint as a pre-built, generally available stack: NVIDIA's RAG server, ingestion service, and NeMo Retriever embedding microservice with the Nemotron Super 49B model, plus a minimal low-GPU variant the docs say isn't for production. SUSE ships no retrieval engine of its own. The buyer gets a supported, secured open-source retrieval stack they'd otherwise assemble, on their own hardware, and NVIDIA's RAG through one contract. Applying the test: the capability is real and generally available on SUSE's paper and general in kind; it's the enterprise's to compose (OpenSearch's own hybrid query included) and there's no SUSE-integrated hybrid or permission-aware retrieval service. Calibration: Kamiwaza reads strong on a context manager and retrieval service of its own; Elastic strong on a native engine; Cloudera gap with its studio in preview; Cloudflare moderate on an index plus open models. SUSE is a curated kit plus a partner blueprint, shipped and supported. Moderate.

Borrowed Judgment

Low on the open path, NVIDIA's on the other. Milvus, Qdrant, OpenSearch, and Open WebUI are open projects the enterprise runs itself: Retained; the vectors are captive to whichever embedding model was chosen, under the carve-out. SUSE's contribution, the Application Collection's curated images, SBOMs, lifecycle, and support, is SUSE's and Ceded, and it's what a customer pays for. The NVIDIA RAG blueprint and NeMo Retriever microservices are NVIDIA AI Enterprise software with no source license, Ceded to NVIDIA through SUSE's paper on the channel-substitution rule. The enterprise composes the pipeline: vendor decides at the store, visible, overridable, Delegated, the Elastic reading for a surface the enterprise configures end to end.

Working Notes

Inference flagged: the current SUSE AI version (the documentation set is 1.0; the November 2025 release added components without a new major version). Public evidence that moves the cell: a SUSE retrieval engine, or hybrid and permission-aware retrieval on SUSE's paper.

Layer 1C · PipelinesData Movement & PipelinesNVIDIA's Fixed RAG Ingest Pipeline Through the Factory; No General Pipeline Layerdecides: vendor · Ceded

Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering

Vendor-Provided

NVIDIA RAG Blueprint Ingestion (NV-Ingest: Parsing, OCR, Chunking, Embedding, Bulk and Streaming Ingest) via SUSE AI Factory with NVIDIACeded

NVIDIA's fixed-function ingest pipeline into the RAG vector store, generally available in the Factory's pre-built blueprint. Proprietary third-party software through SUSE's paper: Ceded to NVIDIA.

NVIDIA-Provided

NVIDIA RAG Blueprint Ingestion (NV-Ingest) in SUSE AI Factory with NVIDIA

The AI-Q with RAG blueprint's ingestor-server is NVIDIA's NV-Ingest pipeline (parsing, OCR and table extraction, chunking, embedding, bulk and streaming ingestion into the vector database) under NVIDIA AI Enterprise; the minimal variant disables OCR and table and chart extraction.

Gap Analysis

SUSE moves data into one place: the vector database under NVIDIA's RAG. SUSE AI Factory with NVIDIA deploys the AI-Q with RAG blueprint as a generally available stack, and its ingestor is NVIDIA's NV-Ingest: document parsing with optical character recognition and table and chart extraction, chunking, embedding, and bulk or streaming ingestion into the vector store, a fixed-function pipeline with one destination. Beyond it, SUSE moves software rather than data: Fleet delivers manifests and Helm charts to thousands of clusters, the Factory promotes AI workloads from development to production, and Kubeflow Pipelines (in the AI Library as a technology preview, with its own ML Metadata lineage) orchestrate training steps. There's no extract-transform-load engine, no change data capture, no generally available lineage on SUSE's paper, no cost-aware movement, and no KV-cache tiering. The buyer gets NVIDIA's ingest for NVIDIA's RAG, and brings their own pipelines. Applying rule 4: a fixed-function AI-ingest pipeline that terminates in the vendor's retrieval surface, on SUSE's paper through the NVIDIA bundle. That's the NetApp AI Data Engine shape (moderate), not the absence of a layer. Calibration: NetApp reads moderate on a fixed sync-classify-vectorize pipeline; Nutanix and VMware gap with no pipeline layer; Cloudera and Elastic score real movement products. SUSE is NetApp's shape with NVIDIA's pipeline. Moderate.

Borrowed Judgment

NVIDIA's. The ingest pipeline's parsing, chunking, and embedding opinions are NVIDIA's inside a blueprint SUSE ships: Ceded to NVIDIA through the channel. The pipeline runs as configured with no per-document override documented: vendor decides, visible, not overridable, Ceded, the NetApp reading.

Working Notes

Grade moved during review: drafted gap, moved to moderate when the AI Factory docs showed the NVIDIA RAG blueprint's ingestor as a generally available deployable on SUSE's paper. Watch-list, notes only: Kubeflow Pipelines (technology preview in the AI Library; ML Metadata lineage inside). Public evidence that moves the cell: a data pipeline or movement product on SUSE's paper.

Layer 2A · OrchestrationInfrastructure OrchestrationRancher Heritage Strength: Multi-Cluster Kubernetes, AI-Conformant RKE2, Virtual Clusters with GPU Quotas, an Apache 2.0 AI Factory Operatordecides: vendor · Ceded

GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization

Vendor-Provided

Rancher Multi-Cluster Management Core (Apache 2.0; Cluster and Fleet Configuration the Enterprise Runs)Retained

The open-source Rancher the enterprise operates, whose cluster definitions and GitOps configuration lift to community Rancher. Retained on the open-source seam.

SUSE Rancher Prime Extensions (AI Assistant, GPU Multi-Tenancy Packaging, Support Entitlements)Ceded

The proprietary extensions and entitlements Prime adds over the open core, whose configuration doesn't lift to community Rancher. SUSE's: Ceded.

RKE2 + K3s (Apache 2.0; RKE2 CNCF Kubernetes AI Conformant with the NVIDIA GPU Operator and a Gang Scheduler Installed)Retained

The supported Kubernetes distributions the enterprise runs, conformant to the standard interface. Retained on the open-source seam.

Virtual Clusters via K3k (Apache 2.0) + Virtual Cluster Policies (GPU Quotas, Live Updates)Retained

Isolated K3s control planes on shared GPU infrastructure with tenant quotas. Open source the enterprise runs; the policy packaging is Prime's. Retained on the open-source seam, flagged.

Fleet (GitOps at Scale, Apache 2.0)Retained

Continuous delivery of manifests and charts to fleets of clusters. Open source the enterprise runs: Retained.

SUSE AI Factory Operator + Rancher UI Extension (Apache 2.0)Retained

The open-source control plane that installs, manages, and monitors AI applications and blueprints on Rancher-managed clusters. Self-run open source: Retained.

SUSE AI Factory Subscription Catalog + Pre-Validated BlueprintsCeded

The subscription-gated application catalog and SUSE's immutable, version-controlled blueprint content. SUSE's: Ceded.

NVIDIA-Provided

NVIDIA GPU Operator on RKE2; NVIDIA Applications and Blueprints from NGC Through the AI Factory

RKE2 meets CNCF AI conformance with the NVIDIA GPU Operator installed per NVIDIA's docs and a gang scheduler (Volcano or Kueue) installed alongside; the AI Factory operator discovers and installs NVIDIA applications and blueprints from the NGC catalog. NVIDIA Run:ai appears in the April announcement and not in the product docs.

Gap Analysis

This is the layer SUSE bought Rancher for, and it's where the AI additions landed. Rancher Prime manages fleets of Kubernetes clusters (fifty thousand teams run Rancher) with RKE2 and K3s as the supported distributions; RKE2 is generally available as a Cloud Native Computing Foundation (CNCF) Certified Kubernetes AI Conformant distribution, meeting the conformance with the NVIDIA GPU Operator installed per NVIDIA's documentation and a gang scheduler such as Volcano or Kueue installed for distributed training. Virtual Clusters through K3k (Apache 2.0) are generally available: hundreds of isolated K3s control planes on one physical cluster, with Virtual Cluster Policies that hand GPU slices to tenants under quotas, updated live, so teams self-serve GPUs without owning nodes. Fleet does GitOps at scale; SUSE Virtualization schedules VMs beside containers with MIG multi-tenancy. SUSE AI Factory with NVIDIA (generally available July 9, 2026) is a Rancher UI extension plus Kubernetes operator (Apache 2.0) that installs, manages, and monitors AI applications from a subscription catalog and the NGC catalog through immutable, version-controlled blueprints, from ClickOps to GitOps. The buyer gets fleet-scale Kubernetes with GPU multi-tenancy from a vendor whose scheduler is the upstream one. The architect's concern is which opinions lift, and the answer is nearly all of them. RKE2, K3s, K3k, Fleet, Rancher, and the Factory operator are Apache 2.0; what's SUSE's is Prime (support, extensions, the AI assistant, GPU multi-tenancy packaging) and the Factory's subscription catalog and pre-validated blueprints. The scheduling judgment is Kubernetes', bounded by the quotas the enterprise writes. Calibration: Nutanix and VMware read strong on orchestration heritage (Prism, VCF Automation); CoreWeave strong on open-substrate GPU orchestration (Kubernetes plus Slurm); Cloudera and Elastic moderate on platform-scoped orchestration. SUSE is CoreWeave's shape sold as software. Strong.

Borrowed Judgment

Bounded, and mostly on the open-source seam. RKE2, K3s, K3k, Fleet, and the AI Factory operator and UI extension are Apache 2.0 the enterprise runs itself: Retained; the Rancher core the enterprise runs is Retained, and Prime's proprietary extensions (the AI assistant, GPU multi-tenancy packaging, entitlements) are Ceded. The Factory's subscription catalog and pre-validated blueprints are SUSE's: Ceded. The scheduler places pods and GPU slices within the quotas and policies the enterprise set, which are configuration: vendor decides, visible, not overridable, Ceded, the Nutanix and VMware reading.

Working Notes

Watch-list, notes only: NVIDIA Run:ai for GPU orchestration (named in the April 21, 2026 announcement of SUSE AI Factory with NVIDIA; not named in the product's overview, building blocks, blueprints, or release notes; it would read Ceded to NVIDIA as a proprietary scheduler if scored); SUSE Observability Hosted (KubeCon EU 2026). Inference flagged: whether Virtual Cluster Policies' GPU quotas are Prime-only packaging or in open-source K3k (the product page lists GPU multi-tenancy as a Prime feature; the SUSE Virtual Clusters docs cover the VirtualClusterPolicy resource). Public evidence that moves the cell: nothing upward from strong.

Layer 2B · RuntimeApplication Runtime & ExecutionSupported Open Inference (vLLM, Ollama) + NVIDIA AI Enterprise (NIM, NeMo, Nemotron) Through the Factory; No SUSE Runtimedecides: model · Delegated

Model serving, agent execution, inference APIs, distributed inference

Vendor-Provided

Open Inference Engines and Tooling in the SUSE AI Library (vLLM, Ollama, Open WebUI, LiteLLM, MLflow, PyTorch, mcpo)Retained

Supported open-source engines and tools the enterprise runs on its own GPUs. Retained on the open-source seam.

Model Access via OpenAI-Compatible Interfaces (vLLM, LiteLLM, and NIM Endpoints)Delegated

Standard-shaped inference interfaces over the open engines and NVIDIA's NIM microservices. Delegated on the inference-interface ruling.

NVIDIA AI Enterprise via SUSE AI Factory with NVIDIA (NIM Microservices, NeMo Containers, RAG Blueprint Services)Ceded

NVIDIA's inference and agent-building stack on the customer's GPUs through SUSE's paper (GA July 9, 2026), beyond the standard interface. Proprietary third-party software: Ceded to NVIDIA.

NVIDIA Nemotron Open Models via SUSE AI Factory with NVIDIADelegated

Open-weights models the enterprise can serve itself, bundled in the Factory. Delegated on the open-weights ruling.

SUSE Application Collection Curation (Hardened Images, SBOMs, Lifecycle, Support)Ceded

SUSE's packaging and support opinions around the open components. SUSE's: Ceded.

Customer-Created and Managed Tools (Function Calling / Client-Side)Retained

Tools the enterprise writes and runs in its own processes, called by models served on the stack. The enterprise's: Retained.

NVIDIA-Provided

NIM, NeMo, and Nemotron in SUSE AI Factory with NVIDIA; OpenShell and NemoClaw Announced, Not Documented

The Factory's GA release embeds NVIDIA AI Enterprise with NIM microservices, the NeMo framework, and Nemotron open models; the AI-Q blueprint's deep-research agent runs on them. OpenShell (alpha on NVIDIA's own README) and NemoClaw appear in the April announcement and not in the product docs. The SUSE AI Library path (vLLM, Ollama) runs on any GPU.

Gap Analysis

SUSE serves models with open engines and takes the rest from NVIDIA. Serving: the SUSE AI Library installs vLLM (added as the inference engine in November 2025) and Ollama as supported, hardened images, with Open WebUI as the chat front end, LiteLLM as the OpenAI-compatible router, MLflow for the model lifecycle, Kubeflow (technology preview) for ML pipelines, PyTorch, and mcpo (an MCP-to-OpenAPI bridge) for tool integration, air-gapped if required, on the customer's GPUs; SUSE AI became a CNCF-conformant AI platform in November 2025. SUSE AI Factory with NVIDIA (generally available July 9, 2026) embeds NVIDIA AI Enterprise: NIM microservices as containerized inference engines running entirely on premises, the NeMo framework, and the Nemotron open models, with the AI-Q with RAG blueprint's deep-research agent as the shipped agentic example. Universal Proxy, SUSE's Apache 2.0 MCP proxy (first release January 30, 2026), fronts MCP servers with discovery, a registry, session management, and authentication, and isn't in the Factory's documentation; NVIDIA's OpenShell agent runtime and NemoClaw were announced with the Factory in April and aren't in its docs either. Liz, Rancher Prime's AI agent, is scored at Layer 3. The buyer gets a supported open inference stack they'd otherwise assemble, and NVIDIA's inference stack through one contract. The architect's concern: SUSE ships no runtime of its own. The serving engines are open source (the enterprise's), the NVIDIA stack is NVIDIA's (through SUSE's paper), and SUSE's value is curation, security, observability, and support. That's a real dependence and a thin one, and the agent runtime is still on the announcement side of the line. Calibration: Nutanix reads moderate on platform-native serving with agents pre-GA; VMware moderate on a model runtime plus an agent builder; Dell's 2B is two components (blueprints and accelerator services) with NVIDIA's runtime at the center, and it watch-lists NemoClaw and OpenShell; Kamiwaza moderate on its own runtime; Cloudera strong on any-model serving plus a GA agent builder of its own. SUSE is Dell's shape sold as software. Moderate.

Borrowed Judgment

Low for the engines, NVIDIA's for the models it packages. vLLM, Ollama, Open WebUI, LiteLLM, MLflow, and mcpo are open source the enterprise runs: Retained; model access through vLLM's and LiteLLM's OpenAI-compatible interfaces, and through NIM's, is Delegated on the inference-interface ruling, and the models the enterprise brings are its own. NVIDIA AI Enterprise's blueprints, RAG services, and NeMo containers beyond the interface are NVIDIA's: Ceded to NVIDIA through SUSE's paper; the Nemotron open models are Delegated on the open-weights ruling. The Application Collection's curated images are SUSE's: Ceded. Customer-created tools running in the enterprise's own processes are Retained. The model decides which tool to call; the enterprise's client-side tools gate the effect: model decides, visible, overridable, Delegated, the Nutanix and VMware reading.

Working Notes

Watch-list, notes only: OpenShell and NemoClaw (announced April 21, 2026; OpenShell is alpha per NVIDIA's README; not in the Factory's product docs; the Dell row watch-lists the same runtime), Universal Proxy (Apache 2.0; first release January 30, 2026; no GA badge; not in the Factory docs), the SLES 16 MCP host (technology preview), Kubeflow (technology preview), SUSE Observability Hosted. Public evidence that moves the cell: a SUSE-owned agent runtime, or the Factory's docs naming OpenShell and NemoClaw as supported deployables.

Layer 2C · ReasoningAgentic Infrastructure — The Reasoning PlaneObservability for AI Workloads + Rancher AI's Supervisor, Agent Configs, and Human-Validation Gates; MCP Gateway in Previewdecides: model · Delegated

Policy-driven placement and resource coordination — the Autonomy Layer

Vendor-Provided

SUSE Observability for AI Workloads (AI Observability Dashboards GA May 2025; OpenTelemetry Operator, GenAI StackPack, Health Monitors for vLLM, Ollama, Milvus, OpenSearch, GPUs, KServe, Kubeflow)Ceded

The OpenTelemetry-native observability platform with AI-workload bindings, in Rancher Prime and as a hosted edition. SUSE's platform: Ceded.

Rancher AI Orchestration (Liz Supervisor, AIAgentConfig, Built-In and Custom MCP Agents, Human Validation Tools, Rancher MCP Server with Read-Only Mode)Ceded

The generally available multi-agent supervisor, agent configuration, confirmation gates, and controlled API gateway inside Rancher Prime. SUSE's: Ceded.

NVIDIA-Provided

No NVIDIA Layer 2C Dependency Scored

NeMo Guardrails is an Apache 2.0 toolkit not named in the AI Factory's product docs; SUSE Observability, Rancher AI, and SUSE Security run on CPU.

Gap Analysis

SUSE's reasoning plane has two shipped legs and one in preview. Observability: SUSE Observability (OpenTelemetry-native, integrated with Rancher Prime and offered hosted, with agentic root-cause analysis) ships the AI Observability Dashboards (generally available since May 2025) and, since November 2025, the OpenTelemetry operator and metric bindings and health monitors for vLLM, Ollama, Milvus, OpenSearch, GPU infrastructure, KServe, and Kubeflow, with a GenAI StackPack that turns LLM telemetry (prompts, tokens, response quality) into topology, dashboards, and health monitors, once applications are instrumented. Orchestration and policy inside Rancher AI (generally available March 2026): Liz is a supervisor that routes each request across specialized agents (Rancher, Continuous Delivery, Provisioning, Application Collection, Observability, Security, CloudCasa, and custom MCP agents the administrator adds) selected by AIAgentConfig metadata, each agent reasoning with an LLM the administrator configures (Ollama, OpenAI, Gemini, Bedrock, or any OpenAI-compatible endpoint), acting through a Rancher MCP server that the docs call the secure, controlled gateway to the Rancher and Kubernetes APIs with a read-only mode, and gated by Human Validation Tools, the administrator's list of tools that require explicit user confirmation before Liz executes them. SUSE Security (NeuVector) is container security, a state guardrail scored below this layer; its Vulnerability Scanner MCP server is experimental. The gateway leg for the estate is Universal Proxy, one point of access for agent interactions with MCP servers, with discovery, a curated registry, security hardening, and telemetry, Apache 2.0 since its first release on January 30, 2026 and absent from the Factory's docs, so it's watch-listed. The buyer gets to see what their AI workloads are doing across the estate, and a governed multi-agent operator whose actions a human confirms. Applying the test: Intelligence-2C has an observability leg (GA, cross-workload), an orchestration leg with agent configuration and a confirmation gate (GA, scoped to Rancher AI's own agents); no agent registry across systems, no agent identity distinct from the user, and the estate-wide tool gateway in preview. Infrastructure-2C, placement across model, cost, and compliance tiers, is absent. Calibration: Nutanix reads moderate on a single gateway leg; VMware moderate on AgentMinder's identity, gateway, and observability; Elastic moderate on registry, orchestration, and observability; Cohere gap with North's in-product policies alone; NetApp gap on data guardrails. SUSE has Elastic's observability and orchestration legs (the second inside one product), with the estate-wide gateway pending. Moderate.

Borrowed Judgment

Split. SUSE Observability's monitors and the StackPack are SUSE's (Ceded); Rancher AI's supervisor, agent configurations, and validation lists are SUSE's (Ceded). The layer's decision, which agent acts and with which tool, is Liz's LLM routing, visible in the UI, gated by the administrator's Human Validation Tools and the MCP server's read-only mode before the effect: model decides, visible, overridable, Delegated.

Working Notes

Watch-list, notes only, and these are what move the cell: Universal Proxy (Apache 2.0; first release January 30, 2026; not in the Factory docs), the SUSE Security Vulnerability Scanner MCP server (experimental; NeuVector itself is container security below this layer), the Multi-Linux Manager MCP server (technology preview), the SLES 16 MCP host (technology preview), NeMo Guardrails (an Apache 2.0 toolkit the enterprise could run, Retained if it does; not named in the Factory's docs). Public evidence that moves the cell: Universal Proxy GA with tool-call policy, or an agent registry or identity on SUSE's paper.

Layer 3 (+1) · ApplicationsAI Application Layer — The Value PlaneLiz and the Crew for Infrastructure Operations, SUSE Industrial Edge for the Plant Floor; Platform-Enabled Beyonddecides: vendor · Delegated

AI-powered business capabilities — business logic, workflow automation

Vendor-Provided

Liz + the Crew (Rancher Prime AI Assistant and Specialist Agents; GA March 2026; External MCP Integration; Confirmation Before Changes)Ceded

The first-party agentic SRE for Rancher-managed estates, extended through MCP servers the enterprise plugs in, with a confirmation prompt before any modification. SUSE's: Ceded.

Rancher Prime MCP Server (SUSE Products as Tools for External Agents)Ceded

The MCP interface external agents use to drive Rancher Prime. The tools are SUSE's objects behind the standard protocol: Ceded, the Elastic MCP-server reading.

SUSE Industrial Edge (Losant; Available April 22, 2026; Visual Workflow Engine, Dashboards, Templates, Anomaly Alerts, AI-Driven Quality Checks)Ceded

The first-party industrial IoT application platform for the Tiny Edge. SUSE's engine, with open sourcing announced as intent: Ceded.

NVIDIA-Provided

No NVIDIA Layer 3 Dependency

Liz runs against the models the enterprise configures; Industrial Edge runs on the Tiny Edge.

Gap Analysis

SUSE's own applications are an agent for the people who run SUSE and a platform for the people who run plants. Liz, the AI assistant in Rancher Prime, went from tech preview (November 2025) to generally available (March 2026) as a context-aware agent that coordinates a Crew of specialists for Linux, observability, security, provisioning, and fleet management, detects issues, explains them, and, through external MCP servers the enterprise plugs in (Atlassian, a firewall, documentation), creates tickets and takes actions across the operational stack, presenting a confirmation prompt the user must accept before any change is applied; Rancher Prime exposes an MCP server of its own so external agents can drive it, with n8n and Revenium as the documented examples and Amazon Web Services, Fsas Technologies, and Stacklok as named relationships, while the Multi-Linux Manager MCP server and the SLES 16 MCP host are technology previews. SUSE Industrial Edge (available April 22, 2026, on the Losant platform acquired February 19, 2026) is a first-party industrial IoT application: protocol-agnostic device connectivity (Siemens, Beckhoff, OPC UA), a no-code visual workflow engine, multi-level dashboards, pre-built application templates, anomaly alerts, and AI-driven quality checks, with SUSE AI endpoints on the roadmap and the technology slated for open sourcing. Everything else at this layer is what customers build on SUSE AI or buy from certified partners. The buyer gets an agentic SRE for their infrastructure, a workflow platform for their factories, and a platform for the rest. Applying the test: the layer asks for AI-powered business capabilities, and Liz is an AI-powered operations capability while Industrial Edge is a business-process platform with AI features inside it; both are first-party and generally available, and neither is the breadth of a security or analytics estate. Calibration: Nutanix and VMware read moderate on platform-enabled applications; Kamiwaza moderate on a platform plus its Kaizen agent; CoreWeave moderate on a developer plane; Elastic strong on first-party security and observability applications; Qlik strong on an analytics estate. SUSE is Kamiwaza's case plus a vertical application. Moderate, with the boundary named: Industrial Edge maturing into an AI-driven process platform would argue for strong.

Borrowed Judgment

Low and split. Liz and the Crew are SUSE's (Ceded) and their opinions about the enterprise's clusters accumulate in Rancher Prime; Industrial Edge's workflows and dashboards are the enterprise's designs on SUSE's engine (Ceded until it's open-sourced); the MCP servers the enterprise plugs in are its own. Liz acts on the environment only after a confirmation prompt the user accepts, a persisted gate before the effect: vendor decides, visible, overridable, Delegated. (The MongoDB assistant's approval on a read-only query tool didn't move that row's reading; Liz's function is acting, so the gate carries this one.)

Working Notes

Escalated: the grade (moderate written; strong argued on Industrial Edge as a first-party process platform; the whole-row reviewer's rule 4 reading is that a domain-specific platform is a fixed-function slice that anchors moderate). Watch-list, notes only: the Multi-Linux Manager MCP server (technology preview), the SLES 16 MCP host (technology preview), SUSE AI endpoints on Industrial Edge (roadmap), the open-sourcing of Losant's technology (announced intent), SUSE AI certified partner applications. Public evidence that moves the cell: Industrial Edge's AI features documented as generally available product capabilities.

Summary Finding

SUSE is an open-source infrastructure vendor that turned its Kubernetes into an AI platform by curating what runs on it, and the map reads it strongest where Rancher always was. Layer 2A is strong on multi-cluster Kubernetes with AI-conformant RKE2, virtual clusters that hand out GPU slices under quotas, and an Apache 2.0 AI Factory operator; six layers are moderate (a hardware-agnostic abstraction with SUSE Linux Enterprise Server (SLES) 16 and Multi-Instance GPU (MIG) capable virtualization at Layer 0; Longhorn and MinIO as platform storage at 1A; supported open-source vector stores plus NVIDIA's retrieval blueprint at 1B; NVIDIA's fixed retrieval-augmented generation (RAG) ingest pipeline at 1C; open inference engines plus NVIDIA AI Enterprise through the Factory at 2B; observability and Rancher AI's supervisor and human-validation gates at 2C; Liz, the Rancher Prime MCP server, and SUSE Industrial Edge at Layer 3), and none is a gap.

The capture is the lightest on the map, because SUSE sells support for software the enterprise could run without it. Twenty-eight components read Retained 13, Delegated 3, Ceded 12: SLES, Harvester, RKE2, K3s, K3k, Fleet, Longhorn, MinIO, Milvus, Qdrant, OpenSearch, Open WebUI, vLLM, Ollama, LiteLLM, NeuVector, and the AI Factory operator are Apache 2.0 or GPL projects the enterprise operates itself; what's SUSE's is Rancher Prime's extensions, the Application Collection's hardened images and lifecycle, the AI Factory's subscription catalog and blueprints, SUSE Observability, Rancher AI, and Industrial Edge, and what's NVIDIA's through SUSE's paper (NIM, NeMo, Nemotron, the RAG blueprint) is Ceded to NVIDIA except for model access through the standard interface. The decision-authority readings are Delegated at five layers because SUSE's engines document the pins, per-object policies, and confirmation prompts the override rule asks for.

The buyer's trade: a sovereign, air-gappable AI stack on their own hardware with the fewest captive opinions of any row, in exchange for assembling retrieval and pipelines from parts, taking the agent runtime and the retrieval models from NVIDIA, and waiting on the pieces still in preview (Universal Proxy, the Multi-Linux Manager MCP server, the SLES 16 MCP host, OpenShell and NemoClaw, Longhorn v2, Run:ai in the Factory's docs). What would move a cell: a SUSE retrieval engine (1B), Universal Proxy reaching GA with tool-call policy (2C), or the AI Factory's docs naming Run:ai and the NVIDIA agent runtime as supported deployables (2A, 2B).

4+1 Layer AI Infrastructure Model · Vendor Assessment Series · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com