# IBM Cloud — 4+1 Layer AI Infrastructure Assessment

> Mapped to the 4+1 Layer AI Infrastructure Model  
> Version: v1.0 - IBM Cloud as a Cloud, Split from the IBM Row · Date: October 8, 2026  
> Source: IBM Cloud documentation (VPC accelerated profiles and release notes, cluster networks, IBM Cloud release levels, Cloud Object Storage S3 compatibility, Code Engine fleets, jobs, events, and release notes, IBM Cloud Kubernetes Service GPU driver migration for 1.36, Red Hat OpenShift on IBM Cloud and its OpenShift AI add-on, IBM Cloud Satellite), watsonx.data and watsonx.data intelligence SaaS documentation (catalog, classifications, data quality, lineage, Managed MCP Server, OpenRAG provisioning, unstructured data integration), watsonx.ai SaaS documentation (deployment methods, chat tool calling, rerank, Model Gateway), watsonx.data integration on IBM Cloud, watsonx Orchestrate documentation and announcements (Agentic Control Plane, AI Gateway, AgentOps, Agent Identity previews, prebuilt agents), IBM Verify Agent Identity and HashiCorp Vault agentic IAM announcements, IBM product and announcement pages (OpenRAG GA June 18, 2026; IBM Confluent Cloud GA March 17, 2026; Event Streams deprecation; watsonx BI GA), Cognos Analytics offerings, and published lab results (Labs 002, 008, 017 on relationship-gated GPU access across clouds). First-hand: Keith Townsend, 'When Small Isn't Simple: Lessons from Deploying Granite 13B on IBM Cloud,' The CTO Advisor, October 9, 2025. v1.0 (October 8, 2026): new row, the follow-up from the IBM v2.0 rescope (October 5, 2026), which kept IBM's owned software and systems on the IBM row and named IBM Cloud a separate cloud shape. Boundary ruled October 8: this row grades what IBM sells as a cloud; Layer 3 grades the business applications IBM Cloud runs as services, under rule 8. Rulings made on this row: select availability is a documented, for-sale release level and scores, with relationship-gated GPU access named as universal (Layer 0, calibrated to OCI); building blocks count at 1B as at 2C (Layer 1B strong). Each cell was peer-checked by ChatGPT before review; ledgers in reviews/ibm-cloud-layer*-chatgpt.json.  
> Published by: The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com  
> Author: Keith Townsend

[Full interactive assessment](https://layer2c.com/assessment/ibm-cloud) · [Methodology](https://layer2c.com/methodology) · [What Is Layer 2C?](https://layer2c.com/what-is-layer-2c)

## Executive Summary

IBM Cloud is the cloud that sells you the parts. It is strong at the data and runtime layers (storage and catalog, retrieval, pipelines, orchestration, and serving) and moderate at the edges: a GPU fleet a generation behind the frontier at Layer 0, an agent control plane missing first-class agent identity at 2C, and a narrower first-party application estate at Layer 3. Decision authority reads Delegated through most of the stack, Retained in the pipelines where the enterprise's own flows and code decide, and Ceded only at Layer 0 and Layer 3.

Much of the strength is assembled rather than packaged. Retrieval is OpenRAG, managed Milvus, Elasticsearch, permission-carrying ingestion, and a watsonx.ai rerank API; orchestration is managed Kubernetes and OpenShift with Kueue beside Code Engine's serverless GPU fleets; serving is watsonx.ai with the enterprise's own agent loop in Code Engine or OpenShift AI. The instrument credits a capability an architect can assemble from generally available parts on the vendor's paper, and it names the seams as the enterprise's: wiring rerank and permission-aware ingestion into the retrieval path, and routing serving traffic through a door IBM hasn't yet made generally available.

That construction makes IBM Cloud the least captive cloud on the instrument by its own chips: 46 percent of its components read Ceded, against 57 to 76 percent for CoreWeave, AWS, Azure, Google Cloud, and OCI. The parts underneath are open (OpenRAG under Apache 2.0, Milvus, Kubernetes, Knative, Kueue, Granite's open weights), so indexes, workloads, and models lift. The capture that remains is decoupled and easy to miss: the classifications, lineage, and governance rules in watsonx.data intelligence, the DataStage and StreamSets flows, and the Orchestrate control plane don't leave with the data.

The edges are where a buyer feels the difference. Most of IBM Cloud's accelerators are sold at select availability through a support case, the relationship gate every cloud keeps at the top of its GPU menu, but its newest NVIDIA parts sit in one zone with no Blackwell cluster network and no rack-scale system, where OCI, the closest peer by shape, runs NVL72 at supercluster scale. And the company with one of the deepest identity stacks among these vendors is shipping agent identity last: IBM Verify and IBM Cloud IAM are generally available, and the product that makes an agent its own principal is still in preview.

The trade is open parts for owned seams. An architect who wants a data and runtime stack that leaves cleanly, and who will staff the integration work, gets the least captive cloud stack on the map. One who needs frontier GPUs at scale or a finished, pre-wired agent plane will find it elsewhere today. IBM's owned software and systems (watsonx software, z17, Power, FlashSystem, Confluent Platform, Vault) are read on the IBM row.

## Layer Status

| Layer | Status | Classification |
|---|---|---|
| Layer 0 · Compute | ◑ Three-Vendor GPU Cloud, a Generation Behind the Frontier | Compute & Network Fabric |
| Layer 1A · Storage | ● Object Storage + a Catalog Agents Can Query | Data Storage & Governance |
| Layer 1B · Retrieval | ● Retrieval Kit Assembled from GA Parts: OpenRAG, Milvus, Elasticsearch, Rerank | Context Management & Retrieval |
| Layer 1C · Pipelines | ● One Integration Plane (Batch, Streaming, Replication) + Event-Driven Code Engine Jobs | Data Movement & Pipelines |
| Layer 2A · Orchestration | ● Managed Kubernetes and OpenShift with Kueue + Serverless GPU Fleets + Satellite | Infrastructure Orchestration |
| Layer 2B · Runtime | ● Tool-Calling Serving on watsonx.ai + Open Serving on OpenShift AI; Gateway Not Yet GA | Application Runtime & Execution |
| Layer 2C · Reasoning | ◑ Orchestrate Control Plane on IBM Cloud; Agent Identity in Preview | Agentic Infrastructure — The Reasoning Plane |
| Layer 3 (+1) · Applications | ◑ Domain Agents and AI Analytics as IBM Cloud Services; Models and Builders Not Graded | AI Application Layer — The Value Plane |

## DAPM Portability Profile (components)

| Classification | Count | Meaning |
|---|---|---|
| Retained | 5 | I possess the capability and can operate it independently of this provider |
| Delegated | 11 | Someone else provides the capability, but I can substitute that provider without reconstructing my accumulated opinions |
| Ceded | 14 | Changing providers requires reconstructing those opinions |

**Decision authority (per layer, gaps included)**

| Reading | Layers | Meaning |
|---|---|---|
| Retained | 1 | The enterprise, or code it writes or controls, decides |
| Delegated | 5 | Vendor or model decides; the enterprise can see and override |
| Ceded | 2 | Vendor or model decides; no override, often invisible |
| Absent | 0 | Nothing offered, nothing inherited |

## Strongest Layers

- **Layer 1A** (Data Storage & Governance) — Object Storage + a Catalog Agents Can Query
- **Layer 1B** (Context Management & Retrieval) — Retrieval Kit Assembled from GA Parts: OpenRAG, Milvus, Elasticsearch, Rerank
- **Layer 1C** (Data Movement & Pipelines) — One Integration Plane (Batch, Streaming, Replication) + Event-Driven Code Engine Jobs
- **Layer 2A** (Infrastructure Orchestration) — Managed Kubernetes and OpenShift with Kueue + Serverless GPU Fleets + Satellite
- **Layer 2B** (Application Runtime & Execution) — Tool-Calling Serving on watsonx.ai + Open Serving on OpenShift AI; Gateway Not Yet GA

## Layer-by-Layer Detail

### ◑ Layer 0 · Compute: Compute & Network Fabric

*Raw compute, networking, and acceleration fabric*  
**Status:** Three-Vendor GPU Cloud, a Generation Behind the Frontier

**Decision authority:** Ceded (decides: vendor; visible: true; overridable: false; boundary: vendor)

**NVIDIA GPU Instances on IBM Cloud VPC (L4, L40S, A100, H100 and H200 HGX, B300 HGX)** [DAPM: Retained]  
L4 (up to four) and L40S (up to two) on gx3, generally available; A100, eight-GPU H100 and H200 HGX on gx3d and eight-GPU B300 HGX on gx4d at select availability, bought through a support case in listed regions. Standard NVIDIA instances: a CUDA or PyTorch workload moves to another cloud's or an OEM's NVIDIA systems without rebuilding.

**AMD Instinct MI300X on IBM Cloud (Virtual Server and Bare Metal)** [DAPM: Retained]  
Eight-GPU MI300X virtual server (select availability) and, since July 31, 2026, a bare metal server in Washington DC for single-node inference with no cluster network. ROCm is an open runtime: a ROCm or PyTorch workload moves to another cloud or an OEM server without rebuilding, as on Azure's ND MI300X and the AMD row.

**Intel Gaudi 3 on IBM Cloud (gx3d)** [DAPM: Retained]  
Eight-accelerator Gaudi 3 virtual servers at select availability in Washington DC, Dallas, and Frankfurt; IBM Cloud was the first cloud to deploy Gaudi 3. Read as the Intel row reads Gaudi 3: standard Ethernet, PyTorch workloads that move to OEM Gaudi systems.

**Hopper 1 Cluster Network (VPC; 3.2 Tbps per Node, RoCE v2, RDMA)** [DAPM: Ceded]  
Generally available since December 10, 2024: eight 400 Gbps RoCE v2 links per node with RDMA, joining H100 and H200 nodes on a separate, externally unrouted IPv4 space, over NVIDIA ConnectX-family adapters. IBM's VPC fabric, configured through IBM's VPC APIs; multi-node layouts tuned to it rebuild on another provider's fabric. No Blackwell cluster network is documented.

**Gap Analysis:** IBM Cloud rents accelerators from three vendors as virtual servers in its VPC: NVIDIA L4 and L40S on gx3; A100, eight-GPU H100 and H200 HGX nodes, and Intel Gaudi 3 on gx3d; eight-GPU B300 HGX nodes on gx4d; and AMD Instinct MI300X as a virtual server and, since July 31, 2026, as a bare metal server. The Hopper 1 cluster network (generally available December 10, 2024) joins H100 and H200 nodes at 3.2 Tbps per node (eight 400 Gbps RoCE v2 links) with RDMA.

IBM sells most of that menu at select availability, which it defines as 'production-ready, available for sale, and accessible to select customers': L4 and L40S are generally available, and the rest is bought through a support case. That is a documented, for-sale product, not a preview (IBM reserves beta and experimental for evaluation), so it scores. The gate is not unique to IBM. Lab 002 found scarce accelerators relationship-gated on every cloud it tried ('quota at zero, a 60-second soft-deny, thousand-dollar-a-day minimums, a support ticket'), Lab 008 found Google's Trillium capacity dry in every catalogued zone with quota in hand, and Lab 017 found Azure's MI300X offered in two of seventeen priced regions behind a quota of zero. Access friction at the top of the GPU menu is a market fact, named here and charged to no one vendor.

The grade is moderate on thin deployment, calibrated to OCI, the closest peer by shape: an enterprise-heritage vendor renting GPUs beside a large software estate. OCI rents B200 and B300 broadly, runs GB200 NVL72 racks, scales a single cluster to 131,072 B200 GPUs on its own Acceleron fabric, and sells AMD's MI355X. IBM Cloud matches OCI's breadth (NVIDIA, AMD, and Intel, and it was the first cloud to deploy Gaudi 3) and trails it everywhere else: B300 in one Dallas zone for select customers, no cluster network for Blackwell so B300 can't scale out, no rack-scale system, and an AMD part a generation older. The capability is real and deployable, and it stands a generation behind the strong rows (AWS, Azure, Google Cloud, OCI, CoreWeave), each of which carries rack-scale NVL72 capacity.

As on every cloud, the enterprise decides nothing at Layer 0 beyond the profile it rents: placement, fabric, and the silicon roadmap are IBM's.

**Borrowed Judgment:** The enterprise borrows IBM's judgment on which accelerators to stock, where, and who may buy them. The accelerators themselves are portable from IBM Cloud: a CUDA or PyTorch workload on NVIDIA, a ROCm or PyTorch workload on MI300X, or a PyTorch workload on Gaudi 3 moves to another cloud or an OEM server without rebuilding, so the chips read Retained relative to IBM Cloud; CUDA adoption cedes to NVIDIA, read on NVIDIA's row. The cluster network is IBM's VPC fabric, and RDMA job layouts tuned to it rebuild elsewhere.

### ● Layer 1A · Storage: Data Storage & Governance

*Durable, governed data foundation — the Governance Catalog that Layer 2C queries*  
**Status:** Object Storage + a Catalog Agents Can Query

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**IBM Cloud Object Storage (S3-Compatible Subset)** [DAPM: Delegated]  
Object storage behind a subset of the S3 API; IBM documents the supported operations and limits (1,000 objects per list call, 1,000 buckets per instance). A managed service behind a multi-vendor standard interface: applications that use IBM's supported S3 subset move to another S3-compatible implementation.

**watsonx.data on IBM Cloud (Managed Lakehouse)** [DAPM: Ceded]  
Managed open lakehouse over Apache Iceberg tables with Presto and Spark engines, available in IBM Cloud regions including Dallas, Washington, Toronto, Frankfurt, London, Tokyo, and Sydney. Tables are Iceberg; the lakehouse's engine configuration, metastore, and data-services opinions are IBM's.

**watsonx.data intelligence on IBM Cloud (Catalog, Classification, Lineage, Data Quality, Managed MCP Server)** [DAPM: Ceded]  
SaaS catalog: automated classification and sensitive-data discovery, data protection rules, data quality, Data Lineage across 50+ sources with OpenLineage export, data products, and a Managed MCP Server (documented; announced available April 23, 2026) exposing catalog search, governance, quality, lineage, enrichment, and data products to external AI agents. IBM's catalog structure, business glossary, and governance rules don't lift.

**Gap Analysis:** IBM Cloud pairs object storage with a catalog that answers what the data is, as managed services. Cloud Object Storage holds the data behind a subset of the S3 API. watsonx.data runs as a managed lakehouse on IBM Cloud over Apache Iceberg tables with Presto and Spark engines. watsonx.data intelligence, the SaaS form of the catalog IBM sells as software (formerly IBM Knowledge Catalog, with Data Lineage from the Manta acquisition), classifies data for sensitivity and confidentiality with predefined PI, PII, and SPI classifications, profiles data classes, scores data quality on dimensions that include timeliness, maps lineage across 50+ sources with OpenLineage export, and publishes governed data products.

The rule 9 pull-through is explicit. watsonx.data intelligence SaaS on IBM Cloud includes a Managed MCP Server, documented and included with the entitlement, that exposes the catalog to external AI agents through tools for search and discovery, semantic search and query, data protection and governance, data quality, lineage, metadata enrichment, and data products. IBM introduced it as a SaaS tech preview on April 7, 2026 and announced it 'now available' on April 23; its documentation carries no preview label, and the documentation governs. The catalog answers the questions rule 9 asks: sensitivity and classification, lineage, and freshness through quality and timeliness signals. That is a catalog a reasoning plane can ask what the data is, which reads strong, calibrated to Azure (Purview), Google Cloud (Knowledge Catalog), AWS (Glue Data Catalog, Lake Formation, SageMaker Catalog), and OCI.

The capture is the governance layer, not the bytes. Data in Cloud Object Storage and Iceberg tables leaves; the classifications, business terms, quality rules, and lineage maps built in watsonx.data intelligence don't.

**Borrowed Judgment:** Discovery, classification, and quality scoring are IBM's judgment, and the enterprise sees the results and can change them: classifications show on the asset overview and the asset owner, an editor, or a catalog admin can change them, and stewards manage the term, data class, and classification assignments metadata enrichment proposes. Data protection rules are written by the enterprise and enforced as written. The grade-carrying half is the catalog's description of the data, so the layer reads vendor-decides, visible, overridable: Delegated, the Purview, Knowledge Catalog, Horizon, and Unity Catalog reading.

The IBM row reads its own 1A Ceded on the storage half it owns (FlashSystem and Storage Fusion placing data, the Dell and NetApp reading). The IBM Cloud row carries no owned storage system at 1A, so its reading follows the catalog.

### ● Layer 1B · Retrieval: Context Management & Retrieval

*Low-latency retrieval for RAG — vector/hybrid search, context windows*  
**Status:** Retrieval Kit Assembled from GA Parts: OpenRAG, Milvus, Elasticsearch, Rerank

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**OpenRAG on watsonx.data (IBM Cloud, Generally Available June 18, 2026)** [DAPM: Delegated]  
Managed agentic RAG: Docling ingestion, OpenSearch indexing and hybrid search, Langflow orchestration, connectors to SharePoint, OneDrive, Google Drive, and S3, 20+ file types, Medium and Large sizes, embeddings from OpenAI or watsonx.ai. Built on the open-source OpenRAG framework (Apache 2.0): flows and indexes lift to a self-run OpenRAG or another provider's OpenSearch and Langflow.

**watsonx.data Milvus Service on IBM Cloud** [DAPM: Delegated]  
Managed Milvus (2.5 line) for vector similarity and hybrid search with scalar filtering, up to three billion vectors at 1,024 dimensions. Open-source Milvus: collections and queries move to a self-run Milvus or another Milvus provider.

**IBM Cloud Databases for Elasticsearch (Vector Search; Platinum Semantic Search and Document-Level Security)** [DAPM: Delegated]  
Managed Elasticsearch: the Enterprise plan deploys Elastic Basic and the Platinum plan Elastic Platinum; vector search on both, semantic search and document- and field-level security on Platinum. Indexes and queries move to Elastic Cloud or a self-managed Elasticsearch.

**watsonx.data Unstructured Data Integration on IBM Cloud** [DAPM: Ceded]  
SaaS pipelines that ingest, cleanse, transform, chunk, and embed documents into vector stores such as Milvus, kept live as sources change, with source access control lists carried through the pipeline and PII redaction. IBM's flows and operators: Ceded.

**watsonx.ai Embeddings and Rerank on IBM Cloud** [DAPM: Ceded]  
Embedding models and the text rerank API (/ml/v1/text/rerank) with cross-encoder rerankers, the relevance stage an architect adds after vector retrieval. Embeddings are model-captive under the embedding carve-out, and rerank scores depend on the model chosen: Ceded.

**Gap Analysis:** IBM Cloud sells retrieval as parts, and the parts cover the layer. OpenRAG on watsonx.data, generally available on IBM Cloud since June 18, 2026, packages Docling ingestion and preparation, OpenSearch indexing and hybrid search, and Langflow agentic RAG orchestration over SharePoint, OneDrive, Google Drive, S3, and more than 20 file types. watsonx.data's managed Milvus serves vector and hybrid search with scalar filters at up to three billion vectors. IBM Cloud Databases for Elasticsearch serves vector search, and on the Platinum plan semantic search with document- and field-level security. watsonx.data's unstructured data integration ingests, cleanses, chunks, and embeds documents into Milvus on live pipelines and carries source access control lists through the pipeline. watsonx.ai supplies embedding models and a text rerank API with cross-encoder rerankers, the relevance stage.

That reads strong under the building-blocks reading ruled October 5, 2026 at 2C and applied here October 8: each function a frontier retrieval service performs (ingestion that keeps permissions, hybrid search, a relevance stage, managed indexes at scale) exists as a GA part on IBM's paper. Calibration: level with Azure AI Search, Google Cloud's Agent Search, and Amazon Bedrock Knowledge Bases, which ship those functions inside one mature service.

The seam is the enterprise's. OpenRAG, the one pre-wired path, is a first release, and its IBM documentation names neither the rerank stage nor permission-trimmed results; IBM documents ACL preservation in its data integration product pages, not as an enforcement model at query time. The architect wires watsonx.ai rerank and ACL-aware ingestion into the retrieval path and verifies that permissions survive into results.

The capture is low for a cloud. OpenRAG is the open-source OpenRAG framework (Apache 2.0) on open-source Langflow, Docling, and OpenSearch, and Milvus is open source; the flows, indexes, and collections lift to a self-run stack or another provider. Elasticsearch moves to Elastic Cloud or a self-managed Elasticsearch, under Elastic's license. The ingestion operators, embeddings, and rerankers are where the capture sits.

**Borrowed Judgment:** The engines decide ranking and chunking at runtime, and the enterprise sets them per request and per flow: query parameters, hybrid weights, rerank calls, Langflow flow definitions, and Docling and ingestion settings are the enterprise's to change. That reads vendor-decides, visible, overridable: Delegated, the reading every cloud's managed retrieval gets.

Embeddings are model-captive: vectors are useful only with the model that produced them. An index built with an open embedding model moves with it; one built with a model served only by IBM doesn't.

### ● Layer 1C · Pipelines: Data Movement & Pipelines

*Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering*  
**Status:** One Integration Plane (Batch, Streaming, Replication) + Event-Driven Code Engine Jobs

**Decision authority:** Retained (decides: code; visible: true; overridable: true; boundary: vendor)

**watsonx.data integration on IBM Cloud (DataStage, StreamSets, Data Replication, Databand)** [DAPM: Ceded]  
Generally available IBM Cloud service (Trial and Essentials plans) with one control plane for batch ETL and ELT (DataStage), real-time streaming pipelines (StreamSets), continuous replication for supported database pairs (Data Replication), pipeline observability (Databand), and unstructured data integration. Flow and pipeline definitions are IBM's formats: Ceded.

**Code Engine Jobs with Cloud Object Storage Event Subscriptions** [DAPM: Delegated]  
Run-to-completion container jobs on open-source Kubernetes and Knative, triggered on a schedule, on demand, or by a Cloud Object Storage subscription on each write or delete (up to 100 subscriptions per project). The job code is the enterprise's container and moves to any Kubernetes or Knative platform; the event wiring is IBM's.

**Gap Analysis:** IBM Cloud moves data two ways. watsonx.data integration, a generally available IBM Cloud service with Trial and Essentials plans, puts DataStage batch ETL and ELT, StreamSets streaming pipelines, continuous, CDC-style Data Replication for supported database pairs, Databand pipeline observability, and unstructured data integration under one control plane. Code Engine, IBM Cloud's serverless container platform on open-source Kubernetes and Knative, runs run-to-completion jobs for ETL on a schedule, on demand, or when an event fires: a Cloud Object Storage subscription triggers a job on each write or delete in a bucket, the pattern IBM documents for processing data where it lands. Lineage for the pipelines is read at 1A, in watsonx.data intelligence.

That reads strong on watsonx.data integration and Code Engine. The pipelines are general: DataStage runs arbitrary flows, StreamSets arbitrary streams, and Code Engine arbitrary code in containers the enterprise writes, so the enterprise's second requirement runs on the same platform. Calibration: level with Azure (Data Factory and Fabric pipelines), Google Cloud (BigQuery-native pipelines with Dataflow and Datastream), and AWS (Glue, managed Airflow, Step Functions); above OCI, whose 1C reads moderate on database-centric movement.

Streaming is the one thin spot on IBM Cloud itself. IBM Event Streams, IBM's managed Kafka, is deprecated with no further enhancements, and IBM Confluent Cloud, generally available on IBM paper since March 17, 2026, runs on the hyperscalers rather than in IBM Cloud regions; it is credited on the IBM row. Streaming on IBM Cloud runs through StreamSets in watsonx.data integration.

**Borrowed Judgment:** The engines execute what the enterprise wrote: DataStage flows, StreamSets pipelines, replication subscriptions, and Code Engine job code. StreamSets drift handling and DataStage runtime column propagation, which adapt to columns that appear at runtime, are opt-in settings the author enables per pipeline or job, configuration under the defaults clause, the reading Azure's opt-in schema drift gets. The layer reads code-decides, Retained, as Azure, Google Cloud, and AWS do at 1C.

Portability splits. Code Engine jobs are containers on Knative, so the code moves to any Kubernetes or Knative platform; the Cloud Object Storage event subscriptions that trigger them are IBM's. DataStage and StreamSets definitions are IBM's formats: StreamSets pipelines export and import within StreamSets Control Hub, not as a neutral pipeline standard, and rebuild on another engine.

### ● Layer 2A · Orchestration: Infrastructure Orchestration

*GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization*  
**Status:** Managed Kubernetes and OpenShift with Kueue + Serverless GPU Fleets + Satellite

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**IBM Cloud Kubernetes Service + Red Hat OpenShift on IBM Cloud (GPU Worker Pools)** [DAPM: Delegated]  
Managed Kubernetes and managed OpenShift with GPU worker pools on VPC (L4, L40S, H100, H200 by zone). IBM-managed GPU drivers on IKS through 1.35, customer-managed from 1.36. Manifests and operators lift to any conformant Kubernetes or OpenShift cluster: a managed service behind a standard interface.

**Red Hat OpenShift AI on IBM Cloud + Red Hat Build of Kueue** [DAPM: Delegated]  
OpenShift AI through the managed cluster add-on on OpenShift on IBM Cloud, with Kueue for quota management, resource allocation, prioritized and gang scheduling, and heterogeneous accelerators across Ray and PyTorch jobs, notebooks, and inference services. Kueue is an open-source Kubernetes project and the queues lift to any OpenShift AI or Kueue cluster.

**Code Engine Serverless Fleets (GPU Task Scheduling)** [DAPM: Ceded]  
One endpoint takes a batch of tasks; Code Engine queues them and, for the GPU family or worker profile the user names, provisions and scales single-tenant workers within the fleet's scaling settings, then releases them. Fleets, tasks, and workers can be cancelled or deleted, with replacement workers on a hard stop. The task containers are the enterprise's; fleet definitions, queueing, and scaling configuration are Code Engine's API and rebuild elsewhere.

**IBM Cloud Satellite (Managed OpenShift and Services on Customer Hosts)** [DAPM: Delegated]  
Runs IBM Cloud's managed OpenShift and selected services on hosts the enterprise supplies, on premises or in another cloud, managed from IBM Cloud. The workloads are OpenShift and lift to any OpenShift; the managed control plane is IBM's to substitute.

**Gap Analysis:** IBM Cloud orchestrates GPU work three ways. IBM Cloud Kubernetes Service and Red Hat OpenShift on IBM Cloud run managed clusters with GPU worker pools (NVIDIA L4, L40S, H100, and H200 flavors on VPC, by zone), and the OpenShift AI cluster add-on (patch updates through June 2026) brings Red Hat OpenShift AI with the Red Hat build of Kueue: quota management, prioritized and gang scheduling, and heterogeneous accelerators for Ray and PyTorch jobs, notebooks, and inference services. Code Engine Serverless Fleets, released September 26, 2025, take a batch of tasks at one endpoint and queue them; once the user names a GPU family or worker profile (L4, L40S, V100, A100, H100, H200, B300), Code Engine provisions and scales single-tenant workers within the fleet's scaling settings and releases them, and tasks can be added after a fleet starts. IBM Cloud Satellite extends IBM Cloud's managed OpenShift and services to hosts the enterprise supplies on premises or in another cloud.

That reads strong. Calibration: AWS (EKS with Karpenter), Azure (AKS with Arc), Google Cloud (GKE with Fluid Compute), OCI (OKE with instance pools), and CoreWeave (CKS with Slurm) each pair managed Kubernetes with GPU-aware scheduling. On IBM Cloud, OpenShift AI with Kueue carries the fair-share, quota, and gang-scheduling comparison; Code Engine fleets are IBM's own serverless batch scheduler beside it, not a substitute for it; Satellite matches Arc and Google Distributed Cloud as the hybrid reach.

The architect's concern: Kubernetes manifests, OpenShift operators, and Kueue queues lift to any conformant cluster, but fleet definitions, task queues, and scaling configuration are Code Engine's API.

**Borrowed Judgment:** IBM's schedulers decide where and when work runs inside the bounds the enterprise sets (Kueue quotas and priorities, fleet scaling limits, GPU per task), and the enterprise can reverse specific calls: Code Engine documents cancelling a fleet with a soft or hard stop, cancelling pending and running tasks, and deleting individual workers, with a replacement provisioned on a hard stop. Those are documented manual actions under the override rule, so the layer reads vendor-decides, visible, overridable: Delegated, the Azure and CoreWeave reading. On Satellite the enterprise owns the hosts and IBM's control plane decides what runs where on them.

### ● Layer 2B · Runtime: Application Runtime & Execution

*Model serving, agent execution, inference APIs, distributed inference*  
**Status:** Tool-Calling Serving on watsonx.ai + Open Serving on OpenShift AI; Gateway Not Yet GA

**Decision authority:** Delegated (decides: model; visible: true; overridable: true; boundary: model)

**watsonx.ai Inference on IBM Cloud (Shared, Deploy on Demand, Custom Foundation Models)** [DAPM: Ceded]  
IBM-provided models including Granite and gpt-oss-120b on shared hardware billed per token; deploy on demand (including gpt-oss-20b and gpt-oss-120b) on hardware dedicated to the enterprise, billed hourly, without rate limits; custom foundation models uploaded as safetensors and served on dedicated GPU hardware specifications; chat API tool calling for models that support function calling. The chat API, deployment spaces, and hardware specifications are IBM's shape, and hosted Granite serving sits here: Ceded.

**OpenShift AI Model Serving on IBM Cloud (KServe, ModelMesh)** [DAPM: Delegated]  
KServe and ModelMesh Serving through the OpenShift AI add-on on Red Hat OpenShift on IBM Cloud (supported add-on versions 416 to 420). Open-source serving on OpenShift: models and serving configuration lift to any OpenShift AI or KServe cluster.

**Customer-Created and Managed Tools (Function Calling / Client-Side)** [DAPM: Retained]  
GA through watsonx.ai chat tool calling for supported models. The enterprise owns and operates the business logic, APIs, and deterministic validators the model's tool calls invoke; they move with the enterprise.

**Granite Model Family (Open Weights, Apache 2.0)** [DAPM: Retained]  
IBM's Granite models, published as open weights under Apache 2.0 (Granite 4.0 among them): the enterprise can pull and run the same weights on any runtime. Hosted Granite serving on watsonx.ai is read in the watsonx.ai chip.

**Code Engine Apps and Jobs (Enterprise Agent Loops on Knative)** [DAPM: Delegated]  
The enterprise's own agent loop as a container on Code Engine, with scale to zero on Knative Serving. The code moves to any Kubernetes or Knative platform.

**watsonx Orchestrate Agent Runtime (Hosted)** [DAPM: Ceded]  
IBM-hosted agent runtime on IBM Cloud for agents built in Orchestrate or its Agent Development Kit. The hosted loop and its configuration are IBM's: Ceded.

**Gap Analysis:** IBM Cloud serves models through watsonx.ai in three modes: IBM-provided models on shared hardware billed per token (the Granite family and OpenAI's gpt-oss-120b among them), deploy on demand on hardware dedicated to the enterprise and billed hourly without rate limits (gpt-oss-20b and gpt-oss-120b included), and custom foundation models the enterprise uploads (safetensors, supported architectures) and serves on dedicated GPU hardware specifications. The watsonx.ai chat API carries tool calling for the models IBM lists as supporting function calling, with reasoning controls on gpt-oss. Beside the managed path, the OpenShift AI add-on on Red Hat OpenShift on IBM Cloud serves models with KServe and ModelMesh Serving, Code Engine runs the enterprise's own agent loop as containers on Knative, and watsonx Orchestrate supplies a hosted agent runtime.

That reads strong under the 2B agent-execution reading: tool calling on the chat API, the enterprise's own loop in Code Engine or OpenShift, and dedicated or bring-your-own hosting carry the loop, and the hosted Orchestrate runtime is a component, not the price of the grade. Tool support varies by model, as it does on every cloud, and IBM documents which models support it. Calibration: level with Azure, Google Cloud, AWS, and OCI on serving that carries the loop.

The difference from those peers is the door. Each of them scores an OpenAI-compatible endpoint on its managed service. IBM's equivalent, the watsonx.ai Model Gateway, documents an OpenAI-compatible unified API but launched as a public preview and has no general-availability note for IBM Cloud, so it isn't scored, and watsonx.ai's own API reads Ceded. The open door on IBM Cloud is OpenShift AI's serving path.

Keith ran this path first-hand: in October 2025 he served Granite 13B through watsonx on IBM Cloud for $69 over a few days, and found the smaller model needed explicit prompting and external retrieval to perform ('When Small Isn't Simple,' The CTO Advisor, October 9, 2025). That is a model finding, not a platform one, and Granite 13B has since been succeeded by Granite 4.0.

**Borrowed Judgment:** A model decides when to act. The enterprise keeps the deterministic gate before the effect where its own code executes the tool calls the model proposes (client-side function calling) and where its own loop runs in Code Engine or OpenShift, so the layer reads model-decides, visible, overridable: Delegated, the reading every cloud's serving gets. In watsonx Orchestrate's hosted runtime the loop is IBM's, read as a Ceded component.

### ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane

*Policy-driven placement and resource coordination — the Autonomy Layer*  
**Status:** Orchestrate Control Plane on IBM Cloud; Agent Identity in Preview

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**watsonx Orchestrate Agentic Control Plane on IBM Cloud (AI Gateway, Registry, Orchestration, AgentOps)** [DAPM: Ceded]  
Introduced on IBM Cloud and AWS in June 2026, with deployment-specific availability: the AI Gateway's controls apply at runtime to agent inputs, outputs, and MCP traffic for agents connected through it; the registry takes agents discovered and registered from other platforms (Amazon Bedrock GA August 31, 2026); multi-agent orchestration across IBM, LangGraph, LangFlow, and A2A agents; observability through the AgentOps Agent, runtime agent monitoring, and LLM-as-a-judge evaluators (GA September 30, 2026). IBM's control plane and its policy model: Ceded.

**watsonx.governance on IBM Cloud (Cross-Platform AI Governance)** [DAPM: Ceded]  
SaaS governance mapping AI assets through policies, risks, and regulatory requirements, with continuous monitoring of models and agents including those IBM doesn't run. IBM's governance model: Ceded.

**Gap Analysis:** watsonx Orchestrate runs on IBM Cloud as the agent control plane IBM sells (generally available May 2026). Its Agentic Control Plane, introduced on AWS and IBM Cloud in June 2026 with availability that varies by deployment (some June governance capabilities were AWS-only), carries four of the five legs: an AI Gateway whose controls apply at runtime to agent inputs, outputs, and MCP traffic for agents connected through it; an agent registry fed by discovery and registration from other platforms (Amazon Bedrock generally available August 31, 2026); multi-agent orchestration; and observability through the AgentOps Agent and runtime agent monitoring, with LLM-as-a-judge evaluators generally available September 30, 2026. watsonx.governance on IBM Cloud maps AI assets through policies, risks, and regulations and monitors models IBM doesn't run.

The missing leg is first-class agent identity, and only that. Identity is generally available across IBM's portfolio: IBM Verify for workforce and customer identity, IBM Cloud IAM service IDs, trusted profiles, and access groups for any workload on IBM Cloud, and Orchestrate connections that call tools with OAuth credentials, including on behalf of the signed-in user. What the leg asks for is the agent as its own principal, registered with an owner and a lifecycle so policy and audit can tell it apart from the user it acts for and the workload it runs in. Generic workload identity hasn't counted as the leg on any row, and acting on behalf of a user is delegated user identity. The product that packages agent identity, IBM Agent Identity on IBM Verify, is in preview: in Orchestrate (private preview September 21, 2026; 'now in preview' October 2) and in Verify (public preview September 1). HashiCorp Vault's agentic IAM is generally available on IBM paper (September 9, 2026) but runs self-managed or as HCP on other clouds; IBM Cloud Secrets Manager is built on Vault and doesn't expose its agentic IAM. Vault is read on the IBM row.

A partial plane with the missing leg named reads moderate. The gateway acts at runtime on connected agents, which clears the 2C action floor. Calibration: level with Snowflake, Salesforce, VMware, and Nutanix, which hold moderate with their missing leg named, and with the IBM row, which reads the same plane the same way; below Azure, Google Cloud, and Databricks, whose planes are complete and GA.

Infrastructure-2C (per-request placement across cost, residency, and capacity) is absent here as everywhere on the instrument, and the deterministic outcome-validator is likewise universal: both are noted, not charged.

**Borrowed Judgment:** IBM's engines evaluate the gateway's controls, the registry's governance rules, and the orchestration flows; the AgentOps Agent and evaluators judge agent behavior with models the enterprise can see. The enterprise can change specific outcomes: it changes gateway controls, policies, evaluator settings, and connections, and suspends or removes registered agents from the managed estate. That reads vendor-decides, visible, overridable: Delegated, the IBM row's reading. The enforcement machinery is IBM's proprietary control plane, Ceded on portability.

### ◑ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane

*AI-powered business capabilities — business logic, workflow automation*  
**Status:** Domain Agents and AI Analytics as IBM Cloud Services; Models and Builders Not Graded

**Decision authority:** Ceded (decides: model; visible: true; overridable: false; boundary: model)

**watsonx Orchestrate Prebuilt Domain Agents on IBM Cloud (Sales, Procurement, HR)** [DAPM: Ceded]  
Production prebuilt agents with domain logic and application integrations (SAP SuccessFactors, CRM and procurement systems): Sales and Procurement Agents generally available and included with Orchestrate Standard and Premium; HR agents documented as available on IBM Cloud. IBM's agents and their logic: Ceded.

**watsonx BI on IBM Cloud** [DAPM: Ceded]  
Generally available: natural-language business questions answered over a governed semantic model with the organization's definitions, calculations, and KPIs, with suggested metric definitions for analysts. IBM's application and semantic layer: Ceded.

**Cognos Analytics on Cloud (Hosted and On-Demand)** [DAPM: Ceded]  
IBM's analytics application as SaaS on IBM Cloud infrastructure: single-tenant hosted, or multitenant on-demand at the latest release. Reports and models are Cognos's formats: Ceded.

**IBM Cloud Catalog: Partner Applications on IBM Paper** [DAPM: Delegated]  
Partner software procured, billed, and supported through the IBM Cloud account. Distribution, not graded capability under rule 8; the relationship is substitutable: Delegated.

**Gap Analysis:** IBM Cloud runs first-party business applications as services. watsonx Orchestrate's prebuilt domain agents do sales and procurement work out of the box (generally available, included with every Orchestrate Standard and Premium subscription), and IBM documents its HR agents as available on IBM Cloud; the customer care entries in the agent catalog mix IBM and partner agents. watsonx BI, generally available on IBM Cloud, answers business questions in natural language over a governed semantic model with the organization's own metric definitions. Cognos Analytics on Cloud runs on IBM Cloud infrastructure as a single-tenant hosted service or a multitenant on-demand service.

Under rule 8 as amended October 5, 2026, the models, catalogues, builders, and developer tools IBM Cloud also offers earn no Layer 3 grade: the watsonx.ai model catalog (Granite and third-party models), Watson AI service APIs, Orchestrate's Agent Builder and Agent Development Kit, IBM's coding assistants, and the IBM Cloud catalog of partner software. They are named, not graded; partner applications bought through IBM Cloud carry the commercial relationship on the DAPM axis.

That reads moderate. Calibration: below OCI (the Fusion application suite), Azure (Microsoft 365 and Dynamics), and the IBM row (which adds Maximo and Planning Analytics from IBM's wider portfolio), whose first-party application estates are broader; above rows with no first-party business application. The domain agents and analytics are real business capability, and IBM Cloud's own application estate is narrower than IBM's.

**Borrowed Judgment:** In the domain agents and watsonx BI, a model decides: which action an agent takes and how a question becomes a query. The enterprise supplies the definitions the model works within (the semantic model and its SQL expressions, metrics, connections), and IBM documents configurable confirmation mechanics: agent guidelines that ask for confirmation before a tool fires, and human-task pauses in traces. Guidelines are instructions the model interprets, not a deterministic gate before the effect, and no mandatory approval step was found for the prebuilt agents, so the layer reads model-decides, visible, not overridable: Ceded. A documented persisted approval gate before agent actions would move it to Delegated.

---
*Layer2C · AI Infrastructure Decision Intelligence · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com*
