# Layer2C — Complete Vendor Assessment Data > The CTO Advisor LLC · thectoadvisor.com · Updated July 17, 2026 Complete 4+1 Layer AI Infrastructure Model assessments for 23 vendors. Each layer lists all components with their DAPM classification (Retained / Delegated / Ceded / Absent), followed by gap analysis, borrowed judgment assessment, and working notes. ════════════════════════════════════════════════════════════════════════════════ # Articul8 AI Platform Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.1 - Reconcile: 2C Determinism Rule **Date:** July 17, 2026 **Source:** Articul8 product pages, Series B announcement (Jan 2026), The CTO Advisor Enterprise Reasoning Plane whitepaper (Jan 2026), AWS Startups case study, SiliconANGLE coverage, VentureBeat coverage, Adara Ventures investment thesis, G2 profile, ZenML LLMOps database, Tracxn profile ## Summary Finding Articul8 is a domain-specific GenAI platform that enters the buyer conversation at Layers 2B/2C/3 — the execution, reasoning, and application layers — and works downward into the data foundation (L1A/1B/1C). It is the heaviest top-of-stack vendor in the instrument: three strong layers (L1B, L2B, L3), three moderate (L1A, L1C, L2C), and two gaps (L0, L2A). Every layer it provides is Ceded. Articul8's distinguishing architectural claim is the Enterprise Reasoning Plane — an explicit implementation of Intelligence Layer 2C that externalizes reasoning from both the model runtime (Layer 2B) and the application layer (Layer 3). The reasoning plane interprets enterprise objectives, decomposes missions into sub-tasks, selects appropriate domain-specific intelligence, applies governance and policy constraints, and produces traceable execution plans. Models become interchangeable execution resources within a governed plan rather than the locus of enterprise intelligence. The capture mechanism is deep-stack application capture. Articul8 captures more of the enterprise's AI investment than most vendors because the domain-specific models (fine-tuned Llama 4 with injected domain knowledge), the reasoning engine (ModelMesh), the Knowledge Graph, and the packaged domain agents are all proprietary and mutually reinforcing. The base models are open (Llama 4). The trained domain expertise is captive. The reasoning logic that orchestrates it is captive. The applications that consume it are captive. Each layer deepens the commitment to the layers above and below it. Articul8 explicitly acknowledges the Infrastructure Layer 2C gap — policy-driven placement decisions about where workloads run (region, cluster, cost, compliance) are 'provided by hyperscalers, platform vendors, or custom enterprise control planes.' This is architecturally honest and shared by every non-hyperscaler vendor in the instrument. The Intelligence Layer 2C that Articul8 does provide — mission decomposition, domain-specific agent selection, policy-constrained execution planning — is real and production-validated, and scores moderate rather than strong: its decomposition core is an LLM-driven planner (probabilistic), not the deterministic reasoning mechanism the frontier strongs carry. The structural comparison to Kamiwaza is instructive: both are software-only platforms absent at L0 and L2A, both are strong at L1B. Articul8 is stronger at L2B (owns the models and the runtime) and L3 (packaged domain agents on four cloud marketplaces). At 2C they now separate: Kamiwaza holds strong on a deterministic, code-based governance mechanism (ReBAC) plus the Inference Mesh's partial live placement, while Articul8 is moderate because its mission-decomposition core is a probabilistic planner. Articul8 asks 'which intelligence should handle this mission?' Kamiwaza asks 'may this agent act on this data under this policy?' Both are Intelligence-2C; the determinism rule distinguishes the code-based mechanism from the prompt-shaped one. The buyer's trade: production-grade domain-specific AI with expert-level accuracy in regulated industries — semiconductor, energy, manufacturing, financial services — in exchange for Ceding the reasoning engine, domain models, Knowledge Graph, and application layer to Articul8's proprietary platform. The data stays in the enterprise's perimeter. The intelligence built from that data is captive. A closed system is a closed system. ## ○ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Enterprise Responsibility ### NVIDIA-Provided Components **NVIDIA GPU Training Infrastructure (via AWS)** All documented training runs on NVIDIA GPUs via AWS SageMaker HyperPod. Fine-tunes Llama-4-Maverick on EKS with HyperPod managing distributed training clusters. An Intel spinout that runs its production training on NVIDIA. **Intel Heritage — Gaudi Support Claimed** Platform launched and optimized on Intel Xeon Scalable processors and Intel Gaudi accelerators. Claims hardware-agnostic but no documented AMD MI300X support. Intel Gaudi is the heritage platform; NVIDIA is the production platform. ### Gap Analysis Articul8 provides no compute hardware, networking, or acceleration fabric. Training runs on AWS SageMaker HyperPod with NVIDIA GPUs. Inference deploys into customer VPC, on-prem, or A8-hosted environments — but Articul8 does not own or sell the infrastructure in any deployment model. The A8 Hosted deployment option is notable: Articul8 operates infrastructure on the customer's behalf. This changes the operational model but does not constitute L0 capability — Articul8 is operating cloud infrastructure, not providing hardware. The Intel heritage is architecturally interesting. Articul8 spun out of Intel in 2024, launched on Intel Xeon and Gaudi, but migrated production training to NVIDIA GPUs via AWS. Claims hardware-agnostic but no evidence of AMD MI300X support. The hardware-agnostic claim is aspirational rather than demonstrated. The enterprise retains full responsibility for Layer 0. ### Working Notes Intel and DigitalBridge are strategic investors. SoftBank is acquiring DigitalBridge. The Intel connection is financial and historical, not architectural — Articul8's production stack runs on NVIDIA/AWS. ## ◑ Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Derived Semantic Layer ### Vendor-Provided Components **Knowledge Graph** [DAPM: Ceded] Advanced Knowledge Graph linking entities, topics, and relationships for unified context across enterprise data sources. Governed semantic layer that the Intelligence Layer 2C queries for mission context. Hybrid graph + vector architecture. Proprietary — graph structure and query interface captive to Articul8. **Autonomous Data Perception** [DAPM: Ceded] Self-evolving hybrid graph + vector architecture that processes structured and unstructured data — text, tables, images, PDFs, streams, 3D meshes, engineering drawings. Uncovers hidden patterns and relationships to enhance semantic understanding. Proprietary perception engine captive to Articul8. **Security & Governance Model** [DAPM: Ceded] SOC 2 Type II compliance. Data remains within customer security perimeter (VPC, on-prem, air-gapped). Enterprise-grade governance with policy-driven access and audit controls. Proprietary governance model captive to Articul8. ### Gap Analysis Articul8 builds a governed Knowledge Graph over enterprise data — entities, topics, and relationships linked for unified context. The platform ingests from siloed enterprise systems (sensor data, process parameters, defect logs, CAD drawings, configuration data, topology diagrams) and creates a semantic governance layer that the reasoning plane queries. This is a derived semantic layer, not a source data platform. The enterprise's data physically lives in existing databases, file systems, S3 buckets, and manufacturing systems. Articul8 governs the semantic representation; the source data stores are governed by whatever already governs them. Compare to VAST (strong) which IS the data substrate — data physically lives in VAST DataStore/DataBase/Element Store with ACID guarantees, multiprotocol access, and integrated security. Articul8 builds understanding over someone else's data foundation. The strong bar at L1A requires being the data substrate or providing a complete, standalone governance layer. Articul8 complements the enterprise's existing data platform rather than replacing it. SOC 2 Type II certified. Data stays within customer security perimeter in all deployment models (VPC, on-prem, air-gapped). ### Working Notes The Knowledge Graph spans L1A and L1B as an integrated system. The L1A/L1B boundary is somewhat artificial for Articul8's architecture — the Knowledge Graph is simultaneously the governed data artifact (L1A) and the context retrieval surface (L1B). ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Articul8 Differentiator ### Vendor-Provided Components **Knowledge Graph (Retrieval Surface)** [DAPM: Ceded] The Knowledge Graph serves dual duty as L1A governed storage and L1B retrieval surface. Context retrieval is domain-aware — the system understands engineering drawings, schematics, sensor data, and manufacturing logs as structured domain objects, not flat documents. Proprietary retrieval surface captive to Articul8. **Multimodal Domain Context Engine** [DAPM: Ceded] Processes complex multimodal formats including 3D meshes, engineering drawings, tables, schematics, sensor data, and long-form interdependent workflows where steps and sequencing matter. Grounds reasoning in real-world physical systems. Proprietary context engine captive to Articul8. **Domain-Specific Retrieval Intelligence** [DAPM: Ceded] Context retrieval is not generic document search — it draws on Articul8's domain-specific models to understand the semantic meaning of industrial data types. A8-Semicon understands Verilog and chip design workflows. A8-Energy understands grid topology and environmental data. The retrieval layer and the model layer are mutually reinforcing. Proprietary intelligence captive to Articul8. ### Gap Analysis Articul8's context management goes well beyond simple RAG. The whitepaper draws the distinction explicitly: simple RAG treats intelligence as a monolithic process of fetching documents for a single general-purpose model. Articul8's context layer handles multimodal domain-specific data — schematics, sensor data, maintenance logs, 3D meshes, engineering drawings — that requires specialized reasoning to even parse. The Knowledge Graph provides the context foundation that the Intelligence Layer 2C queries. When the reasoning plane decomposes a mission, it draws context from the Knowledge Graph to determine which domain-specific agents and models should handle each sub-task. Production evidence: • Semiconductor: siloed sensor data, process parameters, and defect logs correlated across systems for root cause analysis • Mechanical engineering: CAD drawings and electrical schematics analyzed with geometric reasoning agents — 93% accuracy in anomaly detection • Network operations: configuration data, logs, and topology diagrams synthesized into a semantic network graph Comparable to Kamiwaza's strong at L1B but for different reasons. Kamiwaza's strength is the living ontology maintained across distributed sources without data movement. Articul8's strength is multimodal domain-specific context across complex industrial data types. Both are strong; different architectures serving different buyer profiles. ### Working Notes The multimodal capability — reasoning across text, tables, images, 3D meshes, engineering drawings, schematics, sensor data — is genuinely differentiated. Most L1B implementations handle text and maybe tabular data. Articul8 handles data types that require domain-specific intelligence to parse, which is why L1B and L2C are tightly coupled in this architecture. ## ◑ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Autonomous Ingestion ### Vendor-Provided Components **Autonomous Data Ingestion** [DAPM: Ceded] Connector-driven ingestion from siloed enterprise systems — manufacturing data, engineering systems, financial documents, network infrastructure. Handles multimodal formats (text, tables, images, 3D meshes, schematics, sensor streams). Feeds the Knowledge Graph. Proprietary ingestion pipeline captive to Articul8. **Synthetic & Augmented Data Generation** [DAPM: Ceded] Train, fine-tune, and continuously improve domain-specific models using synthetic and augmented data derived from enterprise sources. Supports the model training pipeline, not general-purpose data movement. Proprietary capability captive to Articul8. ### Gap Analysis Articul8 explicitly claims L1C — the reference stack maps it to 'autonomous ingestion.' The platform ingests from siloed enterprise systems (sensor data, process parameters, defect logs, CAD drawings, configuration data, topology diagrams) into the governed Knowledge Graph. The ingestion is multimodal and handles complex industrial data types that most ingestion pipelines cannot parse. Production evidence: • Semiconductor: autonomous ingestion across siloed sensor data, process parameters, and defect logs • Network operations: configuration data, logs, and topology diagrams ingested and synthesized • Supports synthetic and augmented data generation for training and fine-tuning What Articul8's L1C does not provide: general-purpose ETL/ELT pipeline orchestration, data lineage graphs, cost-aware data movement, or cross-system data movement beyond what feeds the Knowledge Graph. The autonomous ingestion serves the intelligence platform — it is ingestion-for-intelligence, not general-purpose data movement infrastructure. Compare to VAST (strong) where DataEngine provides native pipeline capability with lineage, cost-aware movement, and real-time streams. Articul8's L1C is narrower in scope — it serves L1A/1B, not general-purpose enterprise data movement. ### Working Notes The whitepaper explicitly lists L1C as an Articul8-implemented layer. The capability is real but specialized — autonomous ingestion of complex multimodal industrial data, not general-purpose pipeline orchestration. ## ○ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Enterprise Responsibility ### Gap Analysis Articul8 does not provide infrastructure orchestration. GPU scheduling, quota management, cluster provisioning, and infrastructure lifecycle are provided by the underlying platform — AWS SageMaker HyperPod for training, customer VPC/on-prem orchestration for inference, or A8-hosted infrastructure managed by Articul8. The whitepaper is explicit: Infrastructure Layer 2C (policy-driven placement) is 'provided by hyperscalers, platform vendors, or custom enterprise control planes.' If Articul8 doesn't provide Infrastructure 2C, it certainly doesn't provide 2A. The boundary is clear — Articul8 orchestrates intelligence, not infrastructure. The enterprise retains full responsibility for Layer 2A. Same architectural reason as Kamiwaza — software-only platform that operates above the infrastructure layer. ### Working Notes In the A8 Hosted deployment model, Articul8 operates cloud infrastructure on the customer's behalf. This is operational delegation, not infrastructure orchestration capability. Articul8 is consuming cloud orchestration (AWS), not providing it. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** ModelMesh Runtime ### Vendor-Provided Components **ModelMesh™ Agentic Reasoning Engine** [DAPM: Ceded] Autonomous agent-of-agents that dynamically orchestrates data, models, and tools — domain-specific, general-purpose, and non-LLMs — to deliver contextualized outcomes without static rules. Optimizes resources in real-time. The core runtime for all Articul8 workloads. Proprietary reasoning engine captive to Articul8. **LLM-IQ™ Model Evaluation & Routing** [DAPM: Ceded] Model evaluation and dynamic routing engine. Assesses model fitness for task requirements and routes execution to optimal models. Proprietary evaluation and routing logic captive to Articul8. **Domain-Specific Models (A8-Semicon, A8-Energy, A8-SupplyChain, A8-Finance)** [DAPM: Ceded] Proprietary domain-specific models fine-tuned from Meta Llama-4-Maverick with domain knowledge injected via continued pre-training on curated technical and scientific corpora. A8-Semicon: 2x performance over SOTA, Verilog-capable. A8-Energy: 96.9% accuracy (developed with EPRI). A8-SupplyChain: 92% accuracy, autonomous technical documentation reasoning. Base models are open (Llama 4); trained domain expertise is captive to Articul8. **Agent Factory & Multi-Agent Squads** [DAPM: Ceded] Enterprise teams create domain-specific agents via Agent Factory. Multi-Agent Squads autonomously execute complex missions at scale. Google A2A protocol integration for cross-agent interoperability. MCP support via ModelMesh Dock & InterLock. Proprietary agent creation and orchestration framework captive to Articul8. ### NVIDIA-Provided Components **NVIDIA GPU Inference** Inference execution runs on NVIDIA GPUs in all documented deployments. No evidence of inference on Intel Gaudi or AMD accelerators in production. ### Gap Analysis Articul8's primary product layer. ModelMesh™ is the agentic reasoning engine that orchestrates model execution — domain-specific, general-purpose, and non-LLM models as decision and action nodes. The whitepaper frames it precisely: 'Models become interchangeable execution resources within a governed plan rather than the locus of enterprise intelligence.' Key 2B capabilities: • ModelMesh™ — agent-of-agents runtime orchestrating specialized models and agents • LLM-IQ™ — model evaluation and dynamic routing • Agent Factory — enterprise teams create domain-specific agents • Multi-Agent Squads — autonomous mission execution at scale • Domain-specific models (A8-Semicon, A8-Energy, A8-SupplyChain, A8-Finance) — fine-tuned Llama 4 with domain knowledge injected • Google A2A protocol integration for cross-agent interoperability • MCP support via ModelMesh Dock & InterLock • Deployment: A8 Hosted, Customer VPC, On-Prem, air-gapped The combination of owning the runtime AND the models is distinctive. Most vendors provide model serving infrastructure without models (VAST, Kamiwaza) or models without a runtime (NVIDIA open models). Articul8 provides both as an integrated system. Stronger than Kamiwaza at L2B because Articul8 owns both the orchestration engine and the domain-specific models it orchestrates. Kamiwaza's Inference Mesh is model-agnostic with thinner model serving. Production validation: semiconductor root cause analysis (days to hours), 93% CAD anomaly detection accuracy, network topology reasoning. These are real 2B outcomes. ### Working Notes The A2A protocol integration and MCP support create interoperability seams — Articul8 agents can communicate with non-Articul8 agents via open protocols. But interoperability is not portability. The enterprise can reach Articul8 agents from outside; they cannot lift ModelMesh out of Articul8. Open protocols at the boundary, proprietary reasoning inside. ## ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Governed Mission Dispatch — Probabilistic Core ### Vendor-Provided Components **ModelMesh™ Intelligence Reasoning Plane** [DAPM: Ceded] Mission-aware dispatch: decomposes enterprise missions into sub-tasks, selects domain-specific intelligence, applies governance constraints, produces traceable execution plans. Externalizes reasoning from both the model runtime (Layer 2B) and the application layer (Layer 3). The core Intelligence Layer 2C implementation. Proprietary reasoning plane captive to Articul8. **Mission Decomposition & Agent Routing** [DAPM: Ceded] Autonomously decomposes complex missions into sub-tasks routed to specialized agents. Not simple model routing — evaluates domain context, task complexity, data characteristics, and institutional knowledge to select intelligence. Produces compound, auditable results from multi-agent execution. Proprietary decomposition and routing logic captive to Articul8. **Observability & Auditability** [DAPM: Ceded] Real-time visualization of reasoning flows and agent decisions. Full decision traceability — every step in mission decomposition, agent selection, and execution is logged and auditable. 100% decision traceability claimed. Policy-driven access and audit controls. Proprietary observability framework captive to Articul8. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** Articul8's reasoning plane has no NVIDIA dependency. ModelMesh reasoning logic, mission decomposition, and agent routing are entirely Articul8 IP. ### Gap Analysis Articul8 provides real, production-validated Intelligence Layer 2C: mission decomposition, domain-specific agent selection, policy-constrained execution planning, and full decision traceability. It correctly claims Intelligence-2C and disclaims Infrastructure-2C (where inference physically runs), which the whitepaper assigns to 'hyperscalers, platform vendors, or custom enterprise control planes.' That boundary is honest and the capability is genuinely deployed. The grade is moderate, not strong, under the determinism rule (ratified Salesforce v1.0). Decompose the cell by mechanism. The headline capability — autonomously decomposing a mission and selecting intelligence 'in full domain context' — is LLM-driven planning: a model reads the objective and plans the breakdown. That is prompt-shaped and probabilistic, not a deterministic reasoning mechanism. The genuinely deterministic part is narrower: the governance and policy constraints applied to execution, plus policy-driven access and audit. Observability is decision tracing, not outcome validation. So the reasoning the buyer inherits is the model's judgment inside a governed plan, not a code-based reasoning plane. Frontier check (rule 6): strong at 2C stands with the frontier — the comprehensive productized control planes (Google, Azure: identity, registry, gateway, orchestration, observability) and Palantir's Apollo, the one code-based constraint solver. An LLM-driven mission dispatcher plus policy enforcement stands with neither. It is real, deployable, and partial, which is the definition of moderate, and it calibrates cleanly to the moderate cohort: the multi-agent decomposition is comparable to Databricks' Supervisor Agent, and the policy-plus-audit surface matches Salesforce's Agent Fabric and IBM's governance. The universal finding applies here as everywhere: the deterministic outcome-validator (did the decision achieve intent, under a graduated escalation policy) is absent, and its absence is noted as universal across the instrument rather than charged to Articul8. The live-placement (Infrastructure-2C) gap is likewise universal among non-hyperscalers. ### Borrowed Judgment Split. The reasoning the buyer inherits — how a mission is decomposed and which intelligence is selected — is the model's, probabilistic and inherited, even though Articul8 governs the plan around it. The governance and policy-constraint logic, the routing framework, and the observability are Articul8 IP and Ceded to Articul8: the ModelMesh opinions do not lift to another reasoning plane without rebuilding. The enterprise gets a governed, auditable place to run domain missions, and cedes the orchestration logic while renting the reasoning itself. ### Working Notes The whitepaper makes the most architecturally rigorous case for the Intelligence-2C / Infrastructure-2C distinction of any vendor document in the assessment set. This distinction — first introduced in the Palantir assessment — is now validated by a vendor explicitly building to one half and acknowledging the other as someone else's responsibility. Reconcile (July 2026): 2C moved strong→moderate under the determinism rule — the mission-decomposition core is a probabilistic planner, not a deterministic reasoning mechanism, so it reads real-but-partial rather than frontier. Instrument follow-up (Keith's ruling, open): Kamiwaza 2C paired re-read — its strong rests on ReBAC (a deterministic, code-based access-enforcement mechanism) plus the Inference Mesh (genuine partial Infrastructure-2C), which the determinism rule rewards rather than discounts, so it is not automatically dragged down with Articul8. ## ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Domain Applications ### Vendor-Provided Components **Semiconductor Expert Agent** [DAPM: Ceded] Domain-specific agent for semiconductor engineering. 15% productivity boost, 89% test compilation success rate. Verilog-capable via A8-Semicon DSM. Integrates domain knowledge with reasoning for complex chip design workflows. Deployed at Intel for fab root cause analysis and at a leading Santa Clara semiconductor company for product release acceleration. Proprietary agent captive to Articul8. **Energy Expert Agent** [DAPM: Ceded] Domain-specific agent developed with EPRI. 97% accuracy across diverse energy tasks. 96.9% average accuracy across 10 specialized energy topics vs. 71.3% for GPT-OSS-20B. Enables breakthroughs in grid optimization and environmental monitoring. Proprietary agent captive to Articul8. **Finance Expert Agent** [DAPM: Ceded] Domain-specific agent for financial analysts and investment researchers. Deep understanding of finance domain, market events, and company insights. Production deployment processing 2M financial documents with multimodal insights. Proprietary agent captive to Articul8. **Telco Expert Agent / Weave** [DAPM: Ceded] Network topology intelligence as a service. Transforms raw network logs and topology diagrams into a live, queryable graph. Isolated filegroup architecture for security. Supports visibility, change detection, scenario planning, and policy enforcement across clouds, data centers, and distributed environments. Proprietary agent captive to Articul8. **A8 Essential** [DAPM: Ceded] Self-service GenAI product for domain experts. No-code, no data science experience required. Abstracts complexity of building GenAI for real-world applications. Intuitive UI designed for domain experts to have immersive collaborative experience. Powered by ModelMesh. Proprietary application captive to Articul8. **Agent Factory & MicroApp Templates** [DAPM: Ceded] Enterprise teams create domain-specific agents via Agent Factory. Industry-tailored MicroApp templates for rapid time-to-value. No-code MicroApps with industry templates. Custom outputs including reports, presentations, and apps grounded in enterprise data. Proprietary platform captive to Articul8. ### Gap Analysis Articul8 is more application vendor than platform vendor at Layer 3 — multiple packaged domain agents with production deployments, marketplace distribution, and quantified outcomes. This is the strongest L3 of any software-only vendor in the instrument. Domain-specific agents (available on AWS Marketplace, Google Cloud, Microsoft, Databricks): • Semiconductor Expert Agent — 15% engineering productivity boost, 89% test compilation success rate • Energy Expert Agent — 97% accuracy across energy tasks, developed with EPRI • Finance Expert Agent — financial analysis and investment research • Telco Expert Agent / Weave — network topology intelligence as a service, live queryable graph Applications: • A8 Essential — self-service GenAI product for domain experts, no-code, abstracts complexity of building GenAI • MicroApp templates — industry-tailored application templates for rapid time-to-value • Agent Factory — enterprise teams create their own domain-specific agents Production case studies: • Intel fab root cause analysis — saving millions, investigation time from days to hours • BCG knowledge discovery — 39% work completion uplift, 27% search relevance improvement • Financial research — multimodal insights across 2M financial documents • Leading semiconductor company — product release cycle acceleration on Google Cloud The domain-specific models (A8-Semicon, A8-Energy, A8-SupplyChain) are themselves L3 assets — trained expertise packaged for consumption. The combination of domain agents + domain models + A8 Essential + MicroApp templates constitutes a complete L3 offering for industrial and regulated verticals. Stronger than Kamiwaza at L3 (moderate — Kaizen agent plus templates). Articul8 has multiple packaged domain agents with published benchmarks and named customer deployments. Four cloud marketplace distribution vs. Kamiwaza's HPE partnership channel. ### Working Notes $500M+ valuation (Jan 2026 Series B). ~65 employees. $70M Series B led by Adara Ventures. Strategic investors: DigitalBridge (SoftBank acquisition pending), Intel. Eyes $250M ARR in four years. Available on AWS Marketplace, Google Cloud Marketplace, Microsoft, and Databricks. ════════════════════════════════════════════════════════════════════════════════ # AWS AI Infrastructure Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.4 - Single-Vendor API Re-Lens & Label Reconciliation **Date:** July 12, 2026 **Source:** re:Invent 2025, GTC 2026, Bedrock AgentCore GA, AgentCore Policy GA (Mar 2026), SageMaker Unified Studio, AWS/NVIDIA collaboration, OpenAI/AWS partnership, analyst coverage. v1.2 (instrument reconciliation): 1A S3 and 2A EKS Retained→Delegated — a managed service behind a multi-vendor standard interface is Delegated; Retained is reserved for an open substrate the enterprise operates. Corrects the prior over-application of Retained. Also aligned the VAST Layer 2C cross-reference to VAST's current gap status (PolicyEngine GA end 2026). v1.3 (lab-validated against the TFD corpus RAG build): added Amazon S3 Vectors (GA Dec 2025) to Layer 1B as a Ceded component (single-vendor vector API, no open exit), and added the self-orchestrate-versus-Bedrock-KB Retain/Delegate fork to the 1B and 1C narration, mirroring 2A/2B. v1.4 (July 12, 2026, /reconcile): the hyperscaler lens drift resolved — AWS's single-vendor managed services re-lensed under the litmus's own rule (single-vendor API = Ceded; Delegated requires a multi-vendor standard or open substrate), aligning the row with the Azure/GCP/OCI treatment of architecturally identical services. Chips moved D→C: Glue Catalog + Lake Formation, SageMaker Catalog (1A); Bedrock Knowledge Bases (1B); Bedrock, AgentCore Runtime (2B); Guardrails, AgentCore Policy, Evaluations + Memory (2C); Amazon Q, Kiro, Bedrock Agents (L3). Invalid dual-enum chips split into per-mode components (SageMaker managed Ceded / self-hosted Retained; Bedrock Agents Ceded / Strands apps Retained; Glue ETL Delegated / Unified Studio Ceded; MWAA Delegated / Step Functions Ceded). Agent Registry (Preview) descored to the watch-list per the GA-gate. statusLabels moved to capability vocabulary. No capability grades changed. ## Summary Finding AWS is the first vendor in this assessment series that makes a credible claim across every layer of the 4+1 model — including Layer 2C. The structural difference between AWS and every on-prem vendor (Dell, HPE, VAST) is the direction of authority. On-prem vendors build upward from hardware, attempting to extend authority into orchestration and runtime layers. AWS builds downward from managed services, extending authority into custom silicon (Trainium, Inferentia, Graviton), custom networking (EFA/SRD, Nitro), and now on-prem infrastructure (AWS AI Factories). The DAPM classification for AWS is structurally inverted compared to on-prem vendors. The enterprise architect using AWS retains less direct authority at every layer — but gains operational leverage that on-prem vendors cannot match. The question is not whether AWS has the capabilities. The question is whether the enterprise architect has made the authority delegation explicit, and whether they understand what borrowed judgment they inherit when they adopt AWS’s Reasoning Plane as their own. AWS already has the control plane everyone else is trying to build. The problem is that customers do not always see where AWS’s control plane ends and their own authority begins. When Bedrock routes inference, SageMaker auto-scales, or Karpenter provisions nodes — those are Layer 2C functions operating invisibly inside managed services. This is the ‘DGX Realization’ that birthed the 4+1 model: the cloud operates an invisible Reasoning Plane that becomes visible only when you try to replicate it on bare metal. The OpenAI partnership (2GW of Trainium capacity, Stateful Runtime on Bedrock) and NVIDIA deepened collaboration (1M+ GPUs including Blackwell and Rubin) demonstrate AWS positioning as the substrate on which multiple AI ecosystems converge — creating one of the broadest Layer 3 ecosystems and one of the most complex borrowed judgment landscapes. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Custom Silicon Full Stack ### Vendor-Provided Components **AWS Custom Silicon (Annapurna Labs)** [DAPM: Ceded] Trainium3 GA: EC2 Trn3 UltraServers, 144 chips/UltraServer, 4.4x compute vs Trn2, 4x energy efficiency, 4x memory bandwidth. Sub-10µs chip-to-chip latency. Designed for agentic AI, MoE models, large-scale RL. Trainium4 expected 2027. Inferentia2 for inference. Graviton for ARM CPU. AWS-owned silicon IP. **NVIDIA GPUs on AWS** [DAPM: Ceded] Broadest NVIDIA GPU collection of any cloud. P5 (H100), P5e (H200), P6 (B200), P6e (GB200). 1M+ GPUs added 2026 including Blackwell and Rubin. **Nitro System + EFA + SRD** [DAPM: Ceded] Custom hardware/firmware for I/O offload. Hardware-enforced security isolation. EFA: OS-bypass with 3,200 Gbps bandwidth. SRD: AWS custom multi-path, fault-tolerant transport. EC2 UltraClusters: petabit-scale, 20,000 GPUs, 16% latency reduction (v2.0). **AWS AI Factories (On-Prem)** [DAPM: Ceded] Dedicated on-prem environments as private AWS Region. Customer provides space/power; AWS deploys and manages Trainium, NVIDIA GPUs, networking, storage, and full managed services (Bedrock, SageMaker). Inverted vs Dell/HPE: AWS operates infrastructure the customer houses. ### NVIDIA-Provided Components **NVIDIA GPU Silicon + NIXL** 1M+ GPUs including Blackwell and Rubin. NIXL support with EFA for disaggregated LLM inference. ### Gap Analysis The enterprise has no authority over Layer 0 hardware beyond choosing instance types. The multi-accelerator marketplace (Trainium, NVIDIA, AMD, Intel) creates a workload-to-silicon matching problem that is itself a Layer 2C function. This problem doesn’t exist in on-prem (accelerator choice made once at procurement) but recurs with every cloud workload placement decision. AWS is the only vendor that owns accelerator silicon IP (Annapurna Labs). Dell and HPE brand third-party silicon. VAST has no Layer 0 silicon. AWS AI Factories invert the on-prem model: Ceded infrastructure even when physically in the customer’s facility. ### Borrowed Judgment Inverted. The enterprise Cedes Layer 0 entirely — AWS makes all silicon, networking, and infrastructure decisions. The enterprise selects from AWS’s menu but does not influence underlying hardware design, networking topology, or physical infrastructure. The trade-off: loss of direct hardware authority in exchange for operational leverage (no procurement lead time, per-workload silicon selection, managed scaling). ### Working Notes The switching cost / decision frequency distinction between cloud and on-prem at Layer 0 is structural. On-prem: capital decision at procurement. Cloud: per-workload decision at runtime. That fluidity itself requires Layer 2C. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Managed Data Foundation ### Vendor-Provided Components **Amazon S3 + S3 Tables** [DAPM: Delegated] De facto object storage standard. S3 Tables (re:Invent 2025): Iceberg-native table storage. S3 Express One Zone: single-digit ms latency. S3 is a vendor-managed service behind the multi-vendor S3 standard — object opinions lift to any S3-compatible platform, so Delegated (operation delegated to AWS; the standard interface keeps the opinions portable). Reserve Retained for an open substrate the enterprise operates (e.g. self-run Ceph/MinIO). **AWS Glue Data Catalog + Lake Formation** [DAPM: Ceded] Centralized metadata with catalog federation to remote Iceberg catalogs. Fine-grained access control, cross-account sharing, column/row-level security. SageMaker and Bedrock inherit IAM/Lake Formation context. The primitives for 1A→2C exist; the composition is customer-built. Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. **SageMaker Catalog** [DAPM: Ceded] Discovery, subscription, governed sharing of data assets within SageMaker Unified Studio. Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. ### Gap Analysis Most mature governance catalog in this assessment. Glue + Lake Formation metadata is API-accessible to higher layers. The 1A→2C connection (Reasoning Plane querying governance metadata for placement decisions) is not a product today — the primitives exist, composition is customer-built. Catalog federation to remote Iceberg catalogs is unmatched within this series. Hybrid gap: federated catalog covers S3 and Iceberg-compatible catalogs but not proprietary on-prem storage metadata. ### Borrowed Judgment Delegated with customer-retained policy. Lake Formation policies are customer-defined; enforcement is AWS-managed. Cleaner DAPM than Dell (MetadataIQ indexes Dell-only) or VAST (governance catalog is proprietary). ### Working Notes The ‘Governance Enables Autonomy’ principle from the 4+1 model is achievable on AWS but requires the enterprise architect to build the governance-to-placement linkage. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Managed Retrieval Portfolio ### Vendor-Provided Components **Amazon Bedrock Knowledge Bases** [DAPM: Ceded] Managed RAG: ingest → chunk → embed → index → retrieve. Supports OpenSearch, Aurora/pgvector, Pinecone, Redis. Vertically integrates 1B+1C within managed boundary. Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. **Amazon OpenSearch Serverless** [DAPM: Delegated] Vector search with HNSW/FAISS. Default Bedrock Knowledge Bases backend. Serverless scaling. **Amazon Neptune** [DAPM: Delegated] Graph database for relationship-aware retrieval. Multi-hop reasoning for agentic workloads. Delegated survives on the consumed interface: Gremlin and openCypher are multi-vendor query standards — the weakest survivor of the July 2026 re-lens, flagged as such. **Amazon S3 Vectors** [DAPM: Ceded] Native vector storage in S3 buckets. GA Dec 2025. Up to 2 billion vectors per index, sub-100ms on frequent queries, roughly 90 percent cheaper than specialized vector databases. A metadata filter on query-vectors gives hybrid semantic-plus-structured retrieval in one call. Usable standalone or as a Bedrock Knowledge Bases storage engine. The put/query interface is single-vendor AWS with no open implementation elsewhere, so the store is captive even though the vectors re-embed cheaply: Ceded. The cheapest 1B option is also the most captive interface. ### Gap Analysis Bedrock Knowledge Bases erases the 1B/1C boundary within its managed surface — borrowed judgment, not a gap. AWS makes chunking, embedding, retrieval strategy decisions on the customer’s behalf. The enterprise should ask whether defaults suit their domain. The governance fork mirrors 2A and 2B: consume Bedrock Knowledge Bases and the chunk, embed, and retrieve decisions are AWS's (Delegated), or self-orchestrate the pipeline on primitives (InvokeModel embeddings into S3 Vectors or OpenSearch) and Retain the retrieval-strategy opinions. Self-orchestration Retains the pipeline logic, not the store: OpenSearch stays portable, S3 Vectors is captive to a single-vendor API. Interoperability gap: no unified retrieval abstraction across Bedrock Knowledge Bases + self-hosted Weaviate + Neptune. Routing logic between backends is a Layer 2C function living in application code. ### Borrowed Judgment Moderate, with a Retain path. Self-orchestrating the embed-and-index pipeline on primitives (InvokeModel embeddings into S3 Vectors or OpenSearch) Retains the retrieval-strategy opinions, though the chosen store keeps its own authority. Consume Bedrock Knowledge Bases and AWS makes retrieval quality decisions the enterprise inherits without explicit governance. Compare to VAST (InsightEngine — tighter but VAST-controlled) or Dell (Elastic — separate ISV). ### Working Notes Neptune graph-based retrieval is increasingly relevant for agentic workloads needing relationship-aware context — a pattern neither Dell’s Elastic nor VAST’s InsightEngine natively provides. ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Managed Pipelines (Spark/Airflow Substrates) ### Vendor-Provided Components **AWS Glue ETL (Spark Substrate)** [DAPM: Delegated] Glue: serverless ETL with Spark 3.5.6, Iceberg 1.10. SageMaker Unified Studio: horizontal integration across 1A/1C/2B with one-click onboarding. Single governed environment collapsing organizational boundaries across data engineering, data science, ML engineering. Chip scoped to the ETL engine: Glue jobs are Spark code — an open substrate; job logic ports to any Spark. The catalog surface is scored Ceded at 1A; the Studio console is split below. **SageMaker Unified Studio** [DAPM: Ceded] Single-vendor development and pipeline console — workspace and pipeline-definition opinions accumulate against an AWS-only surface. Ceded, split from the Spark-substrate ETL chip. **Amazon MWAA (Managed Airflow)** [DAPM: Delegated] Managed Airflow for complex DAGs. Step Functions for serverless multi-step workflows in ML pipeline reference architectures. Chip scoped to MWAA: managed open-source Airflow — DAGs port to any Airflow. Step Functions split below. **AWS Step Functions** [DAPM: Ceded] Single-vendor state-machine language — workflow definitions have nowhere to run but AWS. Ceded, split from the Airflow chip. ### Gap Analysis AWS horizontal integration vs VAST vertical integration: same functional coverage, different authority models. VAST = one authority boundary, fewer choices, fewer seams. AWS = many services sharing governance via Lake Formation/IAM, more policy control, more operational complexity. SageMaker Unified Studio addresses fragmentation but underlying services remain distinct. ### Borrowed Judgment Delegated with customer-retained configuration. AWS provides pipeline services; customer defines transformations and flows. The same fork as 1B applies to the embed-and-index pipeline: consume Bedrock Knowledge Bases and the chunk, embed, and index movement is collapsed and managed (Delegated), or self-orchestrate it on primitives (InvokeModel embeddings feeding S3 Vectors or OpenSearch) and Retain the pipeline opinions, with the chosen store carrying its own authority. Operational complexity of maintaining consistent governance across many AWS accounts is the trade-off for flexibility. ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** EKS + Capacity Management ### Vendor-Provided Components **Amazon EKS Auto Mode + Karpenter** [DAPM: Delegated] Managed K8s with GPU-aware scheduling. EKS Auto Mode automates cluster/compute management. Karpenter: open-source autoscaler provisioning exact instance types. Mixed compute (NVIDIA, Trainium, Inferentia, Graviton). Note: no DRA support, Capacity Blocks negate scale-to-zero. EKS is a managed service behind the standard Kubernetes API — manifests lift to another conformant cluster, so Delegated (operation delegated to AWS; the standard interface keeps the opinions portable). **Capacity Management** [DAPM: Ceded] Capacity Block Reservations, Flex Start, Savings Plans, Spot. All AWS-controlled allocation. These are capacity acquisition mechanisms, not workload placement reasoning. ### NVIDIA-Provided Components **NVIDIA GPU Operator (on EKS)** Available but optional. AWS controls GPU scheduling through EKS Auto Mode and Karpenter. Run:ai available but not required. ### Gap Analysis For Dell/HPE, Layer 2A is where authority slips to NVIDIA. For AWS, Layer 2A is where AWS retains authority through managed services while integrating NVIDIA optionally. GPU scheduling primitives are AWS-controlled. Governance choice: Retain 2A by running self-managed EKS, or Cede 2A by consuming Bedrock (no EKS, no Karpenter — AWS handles 2A invisibly). Both legitimate; the choice is a governance decision with DAPM implications. ### Borrowed Judgment Low to moderate depending on path. EKS (Kubernetes interface): Delegated — managed K8s, manifests lift to another conformant cluster. Bedrock consumption: Ceded. NVIDIA dependency is optional, structurally different from Dell/HPE where Run:ai is the primary GPU scheduler. ### Working Notes Capacity Blocks sometimes described as proto-2C but cost-optimized capacity acquisition is not multi-objective placement reasoning. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Bedrock + Open Paths ### Vendor-Provided Components **Amazon Bedrock** [DAPM: Ceded] Foundation model access: Anthropic Claude, Amazon Nova, Meta Llama, OpenAI (Stateful Runtime), Mistral, Cohere, NVIDIA Nemotron. Unified API. Fine-tuning including RFT. Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. **Bedrock AgentCore Runtime** [DAPM: Ceded] Serverless agent runtime. Framework-agnostic (Strands, LangChain, CrewAI). Protocol-agnostic (MCP, A2A). Model-agnostic. MicroVM session isolation. 2M+ SDK downloads in 5 months. Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. **SageMaker AI (Managed)** [DAPM: Ceded] Training, fine-tuning, inference endpoints. LMI containers with vLLM. Multi-LoRA. Supports Trainium + NVIDIA. vLLM on EKS and Ray on EKS for fully self-hosted (Retained). Chip scoped to the managed platform: SageMaker training/serving APIs are single-vendor — Ceded, matching Vertex on the GCP row. The self-hosted path is split below. **Self-Hosted OSS Runtimes on EC2/EKS** [DAPM: Retained] vLLM, Ray, KServe-class serving the customer operates on rented compute — open substrates; the opinions lift to any infrastructure. The Retained escape valve of the AWS runtime story. **Strands Agents SDK** [DAPM: Retained] AWS open-source agentic framework. Model-first, native AgentCore/Guardrails/OpenTelemetry integration. Multi-agent patterns with A2A. ### NVIDIA-Provided Components **NVIDIA GPU Instances + NIM** P5/P6 instances. NVIDIA Nemotron via Bedrock. NVIDIA dependency optional — Trainium-only inference is architecturally possible. ### Gap Analysis AWS owns multiple 2B surfaces (Bedrock, SageMaker, AgentCore, EKS). NVIDIA dependency is optional in a way it’s not for Dell/HPE. Agent frameworks blur 2B/2C/3 boundaries — AgentCore bundles Runtime (2B) + Policy (2C) + agent logic (3). Product boundary ≠ architectural boundary. Borrowed judgment: using Bedrock to access Anthropic Claude or Meta Llama means the model provider’s alignment decisions become part of the enterprise’s AI system. Guardrails constrain output but reasoning in model weights is not customer-configurable. ### Borrowed Judgment Varies by path. Bedrock: Delegated + model provider borrowed judgment. SageMaker self-hosted: Retained. AgentCore: Delegated. Self-hosted EKS: fully Retained. Runtime proliferation is itself a 2C decision AWS doesn’t automate. ### Working Notes Product boundary (AgentCore = Runtime + Policy + Evaluations + Memory + Registry) doesn’t align with 4+1 architectural boundary (Runtime = 2B, Policy = 2C, agent logic = 3). Same cross-layer bundling seen in Google’s and VAST’s products. ## ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Intelligence-2C Toolkit; Placement Implicit ### Vendor-Provided Components **Bedrock Guardrails** [DAPM: Ceded] Content filtering, PII protection, topic blocking. ApplyGuardrail API works with any model. Cross-account organizational safeguards (GA Apr 2026). Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. **AgentCore Policy (GA Mar 2026)** [DAPM: Ceded] Centralized governance outside agent code. Natural language → Cedar policy. Intercepts every tool call before execution. 13 AWS regions. Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. **AgentCore Evaluations + Memory** [DAPM: Ceded] Built-in evaluators for correctness, safety, adherence, consistency. Episodic Memory for stateful reasoning across sessions. Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** All Layer 2C components are AWS IP. NVIDIA does not control governance, policy, or reasoning in the AWS stack. ### Gap Analysis Intelligence Layer 2C (partially present): AgentCore Policy + Guardrails + Evaluations govern agent behavior. Real and productized. Natural-language-to-Cedar conversion is the most accessible policy authoring in this assessment. Infrastructure Layer 2C (not built): No service answers ‘given data residency, cost, latency, and compliance, should this run on Trainium in us-east-1 or NVIDIA in eu-west-1?’ Capacity primitives are building blocks, not a policy-driven placement engine querying 1A governance metadata. The structural insight: AWS already has the control plane everyone else is trying to build — but it’s implicit. Dell’s 2C gap is product absence. AWS’s 2C gap is visibility and authority — the capability exists but is implicit, managed, and Ceded. Five-vendor Layer 2C comparison: • Dell: Absent. • HPE: Retained (IT ops) + Delegated (Kamiwaza). • VAST: Gap, emerging (Polaris ships as placement abstraction — routing, not reasoning; PolicyEngine GA end 2026). • AWS: Intelligence 2C Delegated (productized). Infrastructure 2C implicit (inside managed services). The question is not ‘Does AWS have Layer 2C?’ but ‘How much can the enterprise configure, audit, and control — and how much has been Ceded without explicit classification?’ ### Borrowed Judgment Intelligence 2C: Low — AgentCore Policy, Guardrails, Evaluations are AWS IP. Customer defines policies; AWS enforces. Infrastructure 2C: Ceded (implicit) — placement decisions inside managed services without explicit customer policy input. When SageMaker auto-scales or Bedrock routes, those are 2C functions the enterprise has Ceded without classification. DAPM discipline demands: for every managed service placement decision, classify as Delegated (customer sets policy) or Ceded (AWS decides). ### Working Notes The re:Invent 2025 and AgentCore announcements are the strongest vendor validation of the 4+1 model’s Layer 2C thesis. The ‘invisible Reasoning Plane’ observation is the conceptual foundation of the 4+1 model itself. Watch-list (descored July 12, 2026, /reconcile GA-gate): AWS Agent Registry — in preview; a preview is never a scored component on any row. Scores at doc-confirmed GA. ## ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Broadest Ecosystem ### Vendor-Provided Components **Bedrock Agents (Managed)** [DAPM: Ceded] No-code (Bedrock Agents) and full-code (Strands) agent construction. Bedrock Agents: Delegated behavior. Strands: Retained authority. Both span 2B/2C/3. Chip scoped to the managed agent service: single-vendor agent definitions and tool bindings — Ceded. Strands-built applications split below. **Strands SDK Applications** [DAPM: Retained] Agents built on the open-source Strands SDK — model-agnostic, self-hostable; application opinions lift out. Retained. **Amazon Q** [DAPM: Ceded] AWS AI assistant for business and development. Enterprise Delegates application behavior to AWS. Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. **Model + ISV Ecosystem** [DAPM: Delegated] Anthropic Claude, Amazon Nova, Meta Llama, OpenAI, Mistral, Cohere, NVIDIA Nemotron. Thousands of ISV applications. 11,000+ government agencies. **AWS Kiro (Agentic IDE Platform)** [DAPM: Ceded] Spec-driven agentic development platform replacing Amazon Q Developer (new signups ended May 2026). Three surfaces: VS Code-compatible IDE, CLI, and autonomous cloud agent. Spec-driven development generates requirements.md, design.md, and tasks.md before code — specs are source-of-truth, code is build artifact. Hooks system: 17 automated quality gates (security, linting, testing, validation) firing on file save and PR events. Multi-model routing: Claude Sonnet for reasoning-heavy specs, Amazon Nova for high-throughput code generation, Bedrock as unified model plane. 50+ Powers (MCP integrations: Figma, Terraform, Stripe, Datadog). Autonomous agent executes backlog tasks and opens PRs without developer in the loop. Deep AWS context: native Powers for AWS pricing, docs, Well-Architected, cost analysis. The most opinionated developer AI surface from any cloud vendor — enforces structured development discipline rather than freeform 'vibe coding.' Single-vendor API — the litmus names this class Ceded: the opinions accumulate against a surface only AWS implements, with nowhere to take them. Re-lensed July 2026 to match the Azure/GCP/OCI treatment of the architecturally identical service. ### NVIDIA-Provided Components **NVIDIA NIM on Bedrock** NVIDIA models via Bedrock API alongside all other providers. ### Gap Analysis Broadest Layer 3 in this assessment — different category than Dell ISV partnerships, HPE Unleash AI, or VAST Cosmos. Each Layer 3 application brings its own governance domain. AgentCore Policy and Guardrails provide cross-agent governance primitives; whether they compose into enterprise-wide agent governance remains an implementation question. The Retained/Delegated boundary is not uniform. Custom Strands on self-hosted EKS: fully Retained. Bedrock Agents / Q / partner apps: substantially Delegated. Same enterprise may have both patterns simultaneously. Kiro represents AWS's strongest Layer 3 opinion: spec-driven development enforces structured requirements before code generation. This is an opinionated development methodology embedded in tooling — the enterprise Delegates development workflow decisions to AWS's architectural opinions about how AI-assisted software should be built. The autonomous agent (cloud agent executing tasks and opening PRs without human in the loop) creates a new DAPM question: when Kiro's agent writes and ships code autonomously, who owns the judgment embedded in that code? The developer who assigned the task, or Kiro's multi-model routing logic that chose which model to apply? Compare to Google Antigravity 2.0 (agent orchestration platform, multi-agent parallel execution, Gemini-native) and GitHub Copilot (IDE-embedded coding agent with cloud agent for autonomous PR creation). All three clouds now have agentic developer surfaces that span Layer 2B (execution) and Layer 3 (application). The competitive dynamics are shifting from 'which cloud has the best models' to 'which cloud has the most productive developer surface.' ### Borrowed Judgment Distributed and complex. Model providers bring training data, alignment, safety decisions as inherited borrowed judgment. AWS platform defaults shape application behavior. DAPM Action 3 applies with force: when you move off AWS, what judgment doesn’t move with you? Answer: almost everything above Layer 0. ### Working Notes OpenAI partnership (Stateful Runtime on Bedrock) + NVIDIA (1M+ GPUs) position AWS as convergence substrate for multiple AI ecosystems. Broadest possibilities, most complex borrowed judgment landscape. The Q Developer → Kiro transition (new Q Developer signups ended May 2026) is a significant strategic signal. AWS is consolidating its developer AI surface into a single opinionated platform rather than maintaining parallel tools. Kiro's spec-driven approach is the inverse of 'vibe coding' — it imposes engineering discipline through AI tooling. Whether enterprise development teams accept this opinionated workflow or prefer the freeform approach of Cursor/Copilot is the adoption question. Red Hat Summit 2026 announced OpenShift Dev Spaces support for Kiro alongside Microsoft Copilot, Claude CLI, Cline, Continue, and Roo — meaning Kiro can run inside IBM's governed platform. This cross-vendor interoperability matters for the 4+1 model: the developer tool (Layer 3) can be decoupled from the infrastructure platform (Layer 2A). An enterprise could run Kiro on OpenShift on Dell hardware — three vendors' authority at three different layers. ════════════════════════════════════════════════════════════════════════════════ # Microsoft Azure AI Infrastructure Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.3 - Label Vocabulary & Channel-Substitution Reconciliation **Date:** July 12, 2026 **Source:** Build 2025, GTC 2026, KubeCon Europe 2026, FabCon/SQLCon 2026, Ignite 2024, Microsoft/OpenAI restructured agreement (April 2026), Entra Agent ID GA (April 2026), Agent Governance Toolkit (April 2026), Foundry Agent Service, analyst coverage. v1.2 (instrument reconciliation): 2A AKS Retained→Delegated — managed K8s behind a standard interface is Delegated, not Retained. Azure Blob remains Ceded (single-vendor API). v1.3 (July 12, 2026, /reconcile): statusLabels moved to capability vocabulary (authority lives in the DAPM chips); Azure Databricks Delegated→Ceded under the channel-substitution rule (proprietary platform Ceded-to-owner through any paper, matching the Databricks row). ## Summary Finding Microsoft Azure is the most structurally complex vendor in this assessment series because it operates three distinct authority systems simultaneously: a massive cloud infrastructure platform (Azure), a deeply integrated but newly non-exclusive frontier model partnership (OpenAI), and the largest enterprise software installed base on earth (Microsoft 365, Entra ID, Purview, Fabric). No other assessed vendor straddles all three domains. Google Cloud owns its frontier model outright. AWS partners with model providers at arm's length. Dell, HPE, and VAST operate below the model layer entirely. Microsoft is the only vendor that must coordinate authority across infrastructure, model intelligence, and enterprise application identity — and the April 2026 OpenAI restructuring has made that coordination both more flexible and more visible. The DAPM classification for Azure reveals a paradox unique among the assessed vendors: Microsoft has more productized Layer 2C capability than any vendor except Google — Agent Governance Toolkit (open-source, sub-millisecond policy enforcement), Entra Agent ID (GA April 2026, agent identity as first-class Entra citizens), Microsoft Agent 365 (unified agent registry and control plane), Foundry Control Plane (centralized observability for agents across frameworks) — yet the enterprise Cedes the most judgment to consume it. The Layer 2C surface is real. The authority delegation is also real. Both statements are true simultaneously. The OpenAI restructuring (April 27, 2026) is the most significant borrowed judgment event in this assessment series. Azure exclusivity is gone — OpenAI models now ship on AWS Bedrock the next day. Microsoft retains a four-month first-mover window on new frontier models, IP license through 2032, ~27% equity stake, and OpenAI's commitment to $250B in Azure consumption. But the structural dependency has shifted from contractual lock-in to commercial preference. Microsoft is simultaneously scaling its own model development (MAI-1, speech/image models via Mustafa Suleyman's CoreAI division) — hedging the borrowed judgment it once embraced unconditionally. The enterprise identity story is Azure's most underappreciated differentiator. No other vendor has extended enterprise identity governance to AI agents. Entra Agent ID treats agents as identity citizens alongside humans and workloads — same Conditional Access, same lifecycle management, same risk detection. This is Layer 2C infrastructure that every other vendor will eventually need to build or integrate. Microsoft has it because it already owns the enterprise identity plane. Identity is a cross-cutting concern not fully addressed in earlier assessments in this series — the Azure assessment surfaces it as a structural dimension that should be retroactively evaluated across all vendors. Azure's structural question is not capability — the capabilities span every layer of the 4+1 model. The question is authority composition: when the enterprise adopts Azure AI Foundry + OpenAI models + Entra Agent ID + Fabric data governance + Purview compliance, how many independent judgment systems has it inherited, and has it classified each delegation explicitly? The 4+1 model exists to make that composition visible. Azure is the vendor where the composition is most complex. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Hyperscale Compute Stack ### Vendor-Provided Components **Azure Custom Silicon (Maia + Cobalt)** [DAPM: Ceded] Maia 100: AI accelerator on TSMC 5nm, 105B transistors. Designed for LLM training and inference. Optimized for Azure OpenAI Service and Copilot workloads. Second-generation Maia 200 ('Braga') in development — reported design revisions pushed to 2026. Cobalt 100: ARM-based CPU for general cloud workloads. Both designed in-house by Azure hardware teams. OpenAI partnership now includes rights to OpenAI's custom chip designs for integration into Maia/Cobalt roadmap. **GPU Accelerators on Azure (NVIDIA + AMD)** [DAPM: Ceded] Among the first hyperscalers to deploy Vera Rubin NVL72 (via Azure Local + Foundry Local). ND-series VMs: H100, H200, B200/B300. 1M+ NVIDIA GPUs. Fractional GPU via Azure Kubernetes Service. AMD Instinct MI300X via ND MI300X v5 VMs for inference workloads. The multi-accelerator marketplace (Maia, NVIDIA, AMD) creates a workload-to-silicon matching problem that is itself a Layer 2C function. **Azure Networking (SONiC + Accelerated Networking)** [DAPM: Ceded] Microsoft created SONiC (Software for Open Networking in the Cloud) and open-sourced it through OCP. SONiC is now the de facto open-source network OS for hyperscale, running on switches from Broadcom, NVIDIA, Intel, and others — Microsoft shaped the networking substrate and gave it away. Azure Accelerated Networking with hardware-level SR-IOV. InfiniBand interconnect for GPU clusters. RDMA for distributed training. Microsoft-designed rack architecture, power, and cooling. 80+ regions, 500+ datacenters, 800,000+ km of fiber. **Azure Local (On-Prem)** [DAPM: Delegated] Hyper-converged, customer-owned cluster running Azure services on-premises. Windows/Linux VMs, AKS containers, Azure Virtual Desktop. Connected and fully disconnected modes (Feb 2026). Validated OEM hardware from Dell, HPE, Lenovo, and others. Azure Arc extends management, governance, and security. Foundry Local for on-prem AI inference. $10/core/month + optional services. Sovereign Private Cloud (Azure Local + Microsoft 365 Local) for air-gapped environments. ### NVIDIA-Provided Components **NVIDIA GPU Silicon + Networking** Vera Rubin NVL72, Blackwell B200/B300, H100/H200. InfiniBand for GPU cluster interconnect. ConnectX/BlueField NICs. Microsoft manages the NVIDIA integration and instance types. **NVIDIA NIM on Azure** NVIDIA inference microservices available through Microsoft Foundry model catalog alongside OpenAI, open-source, and Microsoft models. ### Gap Analysis Azure's Layer 0 follows the same structural pattern as AWS and GCP: the enterprise Cedes all infrastructure authority in exchange for operational leverage. The multi-accelerator marketplace (Maia, NVIDIA, AMD) creates the same workload-to-silicon matching problem identified in the AWS assessment — a Layer 2C function that no hyperscaler yet automates with policy-driven placement. Azure's custom silicon is less mature than AWS's (Trainium is in production at scale; Maia 100 powers internal services but Maia 200 has slipped) and less differentiated than Google's (TPUs are architecturally distinct; Maia is NVIDIA-competitive). The OpenAI chip design rights add an interesting dimension — Azure could incorporate inference-optimized design ideas from OpenAI into future Maia generations, creating a silicon-model co-optimization loop unique to Microsoft. SONiC is an underappreciated Layer 0 authority claim. Microsoft designed the network OS that runs hyperscale data centers globally — including competitors' — and open-sourced it. This is a different networking authority model than any other vendor: Dell brands NVIDIA switches, HPE acquired Juniper ($14B), Google built Virgo (proprietary), AWS built SRD (proprietary). Microsoft built SONiC and made it public infrastructure. The strategic value is ecosystem shaping, not proprietary control. Azure Local inverts the on-prem model differently than AWS AI Factories: Azure Local is customer-operated on customer-owned hardware with Azure management plane (Delegated). AWS AI Factories are AWS-operated on AWS-owned hardware in customer facilities (Ceded). Azure Local also runs on multi-vendor OEM hardware (Dell, HPE, Lenovo) while AWS AI Factories run on AWS hardware only — creating cross-OEM visibility similar to VMware's hardware-agnostic model. ### Borrowed Judgment Inverted, same as AWS and GCP. The enterprise Cedes Layer 0 entirely. The trade-off: loss of direct hardware authority in exchange for operational leverage and multi-accelerator choice. The OpenAI silicon co-design relationship is a unique form of borrowed judgment: Microsoft can incorporate OpenAI's hardware ideas but inherits OpenAI's optimization priorities (inference-first, GPT-family architectures). Whether that alignment holds as Microsoft scales its own model development (MAI-1, CoreAI) is an open question. ### Working Notes Microsoft's data center capex run rate exceeds $150B annually (2026), with 1 GW of additional capacity added in Q3 FY2026 alone — among the largest infrastructure investments in corporate history. Custom server boards, racks, and cooling designed for Maia and GPU density. The Stargate project (OpenAI/SoftBank/Oracle JV) is related but distinct — Stargate involves Oracle infrastructure and SoftBank capital, a shared infrastructure authority model with DAPM implications of its own. The Maia 200 slip is worth tracking: AWS shipped Trainium3 on schedule; Google shipped TPU 8t/8i on schedule; Microsoft's second-generation AI accelerator is delayed. Microsoft's near-term answer is massive NVIDIA GPU deployment — deepening the same NVIDIA dependency that Dell and HPE face, just at cloud scale. Azure Local's multi-vendor hardware support makes it the only hyperscaler on-prem offering that runs across OEM boundaries. This parallels VMware's hardware-agnostic model and creates the same potential for a multi-vendor reasoning plane. The difference: Azure Local is managed by Azure Arc (Microsoft's control plane); VMware is managed by VCF (Broadcom's control plane). Both see across OEM boundaries. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Unified Data Foundation (Fabric/OneLake) ### Vendor-Provided Components **Azure Blob Storage + ADLS Gen2** [DAPM: Ceded] Object and data lake storage. Hierarchical namespace for analytics workloads. Hot/Cool/Cold/Archive tiers. Immutable blobs, versioning, lifecycle management. S3-compatible API. The storage substrate for OneLake, Fabric, and AI training data pipelines. **Microsoft Fabric / OneLake** [DAPM: Ceded] Unified SaaS data platform: data engineering, warehousing, real-time analytics, data science, Power BI — all on a single data lake (OneLake). Zero-copy access patterns — data remains in place while being reused across experiences. FabCon/SQLCon 2026 positioned Fabric as 'the operating system for enterprise data' and the central control plane for OneLake data management. EXPOSE TO FABRIC T-SQL extension virtualizes database objects in OneLake without moving data. **Microsoft Purview (Governance)** [DAPM: Ceded] Unified data governance across clouds, data types, and the full data estate. Built into Fabric. Automated sensitivity labeling on OneLake assets. Data Loss Prevention policies detect and restrict sensitive data uploads. Audit logs capture all Fabric user activities including AI interactions. Classification, lineage, compliance (GDPR, HIPAA, PCI DSS). Sensitivity labels, access policies, and compliance controls extend to data shared across tenants via OneLake data sharing. Purview extends beyond Azure into Microsoft 365 (SharePoint, Teams, Exchange), on-premises SQL Server, and multi-cloud environments. ### Gap Analysis Azure's Layer 1A is among the most expansive in this assessment series because Microsoft controls the cloud storage infrastructure (Blob, ADLS Gen2), the enterprise data governance plane (Purview), AND the unified data platform (Fabric/OneLake). Microsoft's structural advantage: Purview governance extends beyond Azure into Microsoft 365 (SharePoint, Teams, Exchange), on-premises SQL Server, and multi-cloud environments. The governance catalog that a Layer 2C reasoning plane would query already contains metadata about the enterprise's entire data estate — not just cloud-resident data. The Fabric evolution from analytics platform to 'operating system for enterprise data' (FabCon/SQLCon 2026) is an explicit control plane claim. The unified data catalog spanning all Fabric workloads with automatic governance inheritance is the closest thing in this assessment to a data-layer reasoning plane. The gap: Purview's governance metadata is rich but it's unclear whether it's API-accessible in a way that a Layer 2C placement engine could query programmatically for real-time decisions. Governance as compliance reporting vs. governance as runtime policy input are different functions. ### Borrowed Judgment Ceded with customer-retained policy. Purview policies are customer-defined; enforcement is Microsoft-managed. The enterprise defines what's sensitive and who can access it; Microsoft enforces across the data estate. Fabric introduces a form of borrowed judgment through embedded Copilot: AI assistance for authoring, exploration, and development across Fabric workloads respects tenant, data, and permission boundaries — but the AI assistance logic is Microsoft's. ### Working Notes The Purview-to-Layer-2C connection is the most compelling governance-to-reasoning pathway in the assessment. If Microsoft builds a reasoning plane that queries Purview classification metadata, sensitivity labels, and compliance policies to make placement decisions about AI workloads, it would have the richest governance input of any vendor — because Purview already sees the enterprise's data across Azure, Microsoft 365, and on-premises. The OneLake 'single logical data lake across the tenant' vision parallels VAST's DataSpace 'global namespace' — both attempt to make data location transparent. OneLake is a cloud-native abstraction within Microsoft's platform; DataSpace is an infrastructure-level abstraction across physical sites. Different layers of the stack, same architectural ambition. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Managed Retrieval (AI Search) ### Vendor-Provided Components **Azure AI Search** [DAPM: Ceded] Enterprise retrieval engine combining full-text search (BM25/Lucene), vector search (HNSW), hybrid search with Reciprocal Rank Fusion (RRF), and transformer-based semantic ranker — all in a single managed platform. Integrated vectorization with built-in chunking and embedding. Agentic retrieval (preview): LLM-assisted query planning, multi-source access, structured responses optimized for agent consumption. **Enterprise + Web Knowledge Grounding** [DAPM: Ceded] Foundry Agent Service agents access SharePoint, Microsoft Fabric, Azure AI Search, Azure Blob Storage, and Bing as knowledge sources. The Microsoft 365 corpus (SharePoint, Teams, Exchange) provides enterprise context that is already semantically rich — documents, conversations, and emails created in business workflows carry business meaning natively. Bing adds web-scale grounding. No other vendor has both enterprise productivity data and a web search engine integrated as agent context sources. This is the structural contrast to Google's Knowledge Catalog approach: Google uses Gemini to derive business semantics from raw data; Microsoft retrieves from a corpus where business semantics were created by humans in the course of work. ### Gap Analysis Azure AI Search is one of the strongest Layer 1B offerings in this assessment. The hybrid search + semantic ranker + agentic retrieval combination addresses the retrieval quality problem comprehensively. The structural contrast with Google's Knowledge Catalog is the key 1B finding. Google's approach is model-integrated: Gemini enriches raw data on arrival, extracting business semantics and building a context graph that didn't exist before. Microsoft's approach is corpus-integrated: the M365 corpus already contains business context because humans created it in business workflows. A SharePoint document about Q3 revenue already carries the business semantics that Knowledge Catalog would need Gemini to extract from a raw CSV. Microsoft doesn't need a model to derive meaning because the data was born semantic. Both approaches have trade-offs. Google's model-derived semantics are consistent and exhaustive — every data asset gets enriched. Microsoft's human-created semantics are richer but inconsistent — the quality depends on how well the enterprise organizes its M365 content. Google's approach works on any data. Microsoft's advantage depends on the enterprise already having its knowledge in M365. The agentic retrieval mode (preview) blurs the boundary between retrieval (1B) and reasoning (2C) — the retrieval engine uses an LLM to decompose complex queries into sub-queries. Same cross-layer blurring observed in other vendors' products. Bing grounding introduces a unique borrowed judgment: the enterprise inherits Bing's web index, coverage, biases, and ranking decisions as part of agent context. No other vendor has this dependency because no other vendor owns a web search engine. ### Borrowed Judgment Ceded with high integration value. Azure AI Search is Microsoft IP. The semantic ranker is a Microsoft model. Agentic retrieval uses Microsoft's LLM for query decomposition. The enterprise Cedes retrieval intelligence to Microsoft but gains the integrated M365 corpus and Bing web index as context. The M365 corpus as borrowed judgment: the enterprise's own data is the context source, but Microsoft controls how it's indexed, chunked, embedded, and served to agents. The enterprise created the content; Microsoft controls the retrieval path. ### Working Notes The 'retrieval as reasoning' evolution (agentic retrieval mode) is worth tracking as a 4+1 model observation. If the retrieval engine uses LLM reasoning to plan queries, where does Layer 1B end and Layer 2C begin? The product boundary (Azure AI Search) doesn't align with the architectural boundary (retrieval vs. reasoning). The Google Knowledge Catalog contrast deserves tracking as both approaches mature. If Google's Gemini-derived semantics prove more reliable than human-created M365 content for agent grounding, the model-integrated approach wins. If enterprise-specific context (internal jargon, organizational knowledge, relationship context) proves more valuable than machine-extracted semantics, the corpus-integrated approach wins. The answer likely varies by use case. ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Managed Pipelines (Fabric/ADF) ### Vendor-Provided Components **Azure Data Factory + Fabric Data Pipelines** [DAPM: Ceded] Cloud-based ETL/ELT orchestration available as a standalone Azure service (Data Factory) and within the unified Fabric experience (Data Factory in Fabric). 200+ connectors in Fabric. Visual pipeline builder. Gartner Leader for Data Integration Tools (5 consecutive years). Fabric version adds Dataflows Gen2 for self-service data preparation with Power Query (no-code), Medallion architecture (Bronze → Silver → Gold) with Delta Lake, ACID transactions, schema enforcement, time-travel queries, Data Activator for event-driven actions, and Copilot for natural-language pipeline authoring. Scheduled and event-based triggers. **Azure Databricks (Partnership)** [DAPM: Ceded] Apache Spark-based analytics platform. Delta Lake as the transactional storage layer. Data engineering, data science, ML. Azure-optimized but Databricks-owned IP — and under the channel-substitution rule that decides the chip: Databricks is a proprietary single-implementation platform whose pipeline and workspace opinions are Databricks-captive through any cloud. Multi-cloud availability is a choice of invoice, not an exit. Ceded-to-Databricks, matching the Databricks row. ### Gap Analysis Azure's Layer 1C is comprehensive and mature. Data Factory's 5-year Gartner leadership position reflects enterprise-grade data movement capability. The Fabric integration collapses what were previously separate services (Data Factory, Synapse, Power BI) into a unified pipeline-to-analytics experience. Azure's Layer 1C advantage: Fabric's unified data platform means the pipeline, storage, analytics, and governance are the same system. Data moves through experiences within OneLake rather than between services. This architectural philosophy parallels vertically integrated approaches in other assessed platforms but at a different layer of the stack. The Databricks dependency is worth noting: many enterprise Azure customers use Databricks rather than native Fabric pipelines for complex data engineering. This creates a Delegated component within an otherwise Ceded layer — the enterprise's data pipeline intelligence is Databricks' IP, running on Azure's infrastructure. No KV cache tiering story is evident on Azure. Dell has validated NVIDIA CMX (19x TTFT improvement), HPE has Alletra X10000 KV cache support, VAST collocates cache and compute in CNode-X. This gap matters as inference workloads scale and KV cache management becomes a data movement problem. ### Borrowed Judgment Ceded for native services. Delegated for Databricks. Copilot for Data Factory adds a dimension: natural-language pipeline authoring means the enterprise inherits Microsoft's LLM's understanding of data engineering patterns. Whether Copilot-authored pipelines match the quality of expert-authored ones is an open question with borrowed judgment implications. ### Working Notes The Fabric positioning as 'operating system for enterprise data' (FabCon/SQLCon 2026) makes Layer 1C the layer where Microsoft's data platform ambitions are most visible. If Fabric succeeds as the unified data control plane, it collapses 1A (storage) + 1B (retrieval) + 1C (movement) into a single authority boundary — which simplifies DAPM classification but concentrates authority. ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** AKS + Arc Hybrid Orchestration ### Vendor-Provided Components **Azure Kubernetes Service (AKS)** [DAPM: Delegated] Managed Kubernetes with GPU-aware scheduling. Dynamic Resource Allocation (DRA) GA at KubeCon Europe 2026 — fine-grained GPU resource allocation enabling multiple AI workloads to share GPU resources, reducing idle GPU time from typical 30-40% rates. AKS-managed GPU node pools automate NVIDIA driver, device plugin, and DCGM metrics. MIG, MPS, and time-slicing for GPU sharing. Kueue for fair queuing and priority. Cilium mTLS encryption in public preview — sidecarless pod-to-pod security eliminating sidecar proxy overhead for AI workloads (built with Isovalent/Cisco). AI Runway (preview): unified inference API positioning Kubernetes as the AI infrastructure operating system, with cross-cloud GPU scheduling previewed. Microsoft contributed DRA to upstream Kubernetes — these GPU scheduling primitives are available to all Kubernetes users, not just AKS. AKS is a managed service behind the standard Kubernetes API — manifests lift to another conformant cluster, so Delegated (operation delegated to Azure; the standard interface keeps the opinions portable). **Azure Arc (Hybrid Orchestration)** [DAPM: Delegated] Extends Azure management to any infrastructure — on-prem, edge, multi-cloud. Projects external resources into Azure Resource Manager. Unified RBAC, policy, monitoring. AKS enabled by Azure Arc: AI inference on hybrid Kubernetes clusters with centralized governance. Connected and disconnected operation modes. The only hyperscaler management plane designed to orchestrate across non-native infrastructure — AWS manages AWS; Google manages GCP and GDC; Azure Arc manages anything. ### NVIDIA-Provided Components **NVIDIA GPU Operator + DRA** NVIDIA GPU Operator manages GPU drivers and device plugins on AKS. DRA GA provides Kubernetes-native GPU scheduling primitives. Microsoft contributed DRA to upstream Kubernetes. ### Gap Analysis Azure's Layer 2A is mature and comprehensive. AKS is the most widely deployed managed Kubernetes service for AI workloads on Azure, and the KubeCon 2026 DRA GA + AI Runway announcements extend its capabilities specifically for AI scheduling. Azure Arc is the hybrid orchestration differentiator: it extends Azure's management plane to any infrastructure, including competitor hardware. An enterprise running Azure Arc on Dell, HPE, and Lenovo hardware gets a unified orchestration plane across OEM boundaries — managed by Microsoft rather than Broadcom (VMware) or the OEM itself. The AI Runway + cross-cloud GPU scheduling preview is the most aggressive multi-cloud orchestration claim in this assessment. If Azure can schedule workloads across its own GPUs, AWS GPUs, and GCP GPUs based on availability and pricing, it becomes the first cross-hyperscaler Layer 2A orchestration plane. This is aspirational — the preview was just announced — but architecturally significant. DRA contribution to upstream Kubernetes means these GPU scheduling primitives are available to all Kubernetes users, not just AKS. Microsoft is building the open-source foundation that competitors also benefit from — a strategic choice that prioritizes ecosystem leadership over proprietary advantage. Brendan Burns' authorship (Kubernetes co-creator, Microsoft employee) gives Azure unique credibility in shaping Kubernetes' evolution. ### Borrowed Judgment Ceded for cloud workloads. Delegated for Azure Arc-managed on-prem infrastructure. Microsoft's Kubernetes contributions (DRA, AI Runway) are open-source — the enterprise can run them on any Kubernetes. But operational maturity (managed upgrades, monitoring, scaling) is Azure-specific. The Kubernetes interface keeps the enterprise's opinions portable to any K8s (Delegated — managed K8s service; the enterprise could switch without rebuilding); Azure-specific operational maturity is convenience, not capture. ### Working Notes Microsoft's KubeCon 2026 framing of 'Kubernetes as the AI Infrastructure OS' is a Layer 2A statement, not a Layer 2C statement. The distinction matters: Kubernetes schedules and manages resources. A Reasoning Plane governs them with policy-driven intelligence. Same distinction applies to VMware's framing of VCF as 'the permanent abstraction layer.' ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Foundry + OpenAI Runtime ### Vendor-Provided Components **Microsoft Foundry + Azure OpenAI Service** [DAPM: Ceded] Unified platform-as-a-service for enterprise AI operations. 11,000+ models in catalog including OpenAI (GPT-5.5, o-series — four-month first-mover window per April 2026 restructuring), open-source, Microsoft (MAI-1, Phi), and NVIDIA models. Foundry Agent Service: fully managed agent runtime supporting no-code prompt agents and code-based agents (Agent Framework, LangGraph, custom). Handles hosting, scaling, identity, observability, enterprise security. OpenResponses, Activity, Invocations, and A2A protocols for agent distribution through M365 Copilot, Teams, and Entra Agent Registry. Foundry Control Plane centralizes observability for agents across frameworks — including agents NOT running on the platform (register custom LangGraph, A2A, or HTTP-based agents, route through AI Gateway, send OTel traces to Application Insights). Batch evaluations for third-party agents using built-in evaluators for safety, fluency, and task adherence. OpenAI models are now also available on AWS Bedrock — the model is no longer an Azure differentiator; the platform integration is. **Microsoft Agent Framework (Open-Source)** [DAPM: Retained] Production framework for building multi-agent systems. Orchestration patterns: sequential, concurrent, handoff, group chat (Magentic One). OpenAPI integration, A2A protocol, MCP support. Local development → Azure deployment with observability and compliance. Open-source (MIT license). KPMG Clara AI built on Agent Framework. ### NVIDIA-Provided Components **NVIDIA Nemotron + NIM on Foundry** NVIDIA models available through Foundry model catalog. Vera Rubin support via Azure Local + Foundry Local for on-prem inference. **NVIDIA NeMo Data Designer** Integration through Foundry for model training and fine-tuning. Same NVIDIA training dependency seen across multiple assessed vendors. ### Gap Analysis Azure's Layer 2B is one of the broadest — 11,000+ models, managed agent runtime, open-source agent framework, cross-framework observability. The breadth creates complexity: the enterprise must choose between Foundry Agent Service (Ceded, managed), Agent Framework on AKS (Retained, self-hosted), Azure OpenAI direct (Ceded, model-specific), and fully self-hosted options. The April 2026 Custom Agent Monitoring is architecturally significant: Foundry extends governance to agents it doesn't host. This is a Layer 2B/2C crossover — the runtime observability reaches beyond the runtime boundary. The OpenAI restructuring creates a unique Layer 2B dynamic: OpenAI models are Azure's flagship capability AND are now available on AWS Bedrock. The model is commodity; the platform around the model is the lock-in. The Agent Framework being open-source (MIT) means the enterprise Retains the code and can run it anywhere. But the operational envelope (Foundry hosting, Entra identity, Purview compliance) is Azure-specific. Code portability vs. operational portability — same distinction identified in the Google Cloud assessment with ADK. ### Borrowed Judgment Multi-layered. OpenAI models: borrowed alignment, training data, and safety decisions — the most significant model-provider borrowed judgment in this assessment because OpenAI is simultaneously a partner, a competitor (ChatGPT Enterprise vs. Microsoft 365 Copilot), and a platform (OpenAI API vs. Azure OpenAI Service). Microsoft's own models (MAI-1, CoreAI): borrowed judgment shifts to Microsoft's model team. Open-source models: community-borrowed judgment. The enterprise using Azure OpenAI Service inherits three judgment systems simultaneously: Microsoft's platform decisions (Foundry defaults, content filtering), OpenAI's model decisions (alignment, capabilities, safety), and NVIDIA's silicon decisions (GPU scheduling, driver behavior). This is the most complex borrowed judgment stack in the assessment. ### Working Notes The April 2026 restructuring eliminated Azure exclusivity for OpenAI models. GPT-5.5 appeared on AWS Bedrock the next day. This validates the 4+1 model's prediction that model-layer lock-in is transient while platform-layer lock-in (2A/2B/2C) is durable. Microsoft's response is right: invest in Foundry, Entra Agent ID, and governance infrastructure that doesn't move with the model. The CoreAI division under Mustafa Suleyman signals Microsoft is building its own model capability to reduce OpenAI dependency. The borrowed judgment composition changes as Microsoft's own models mature. Today it's Microsoft platform + OpenAI models. Tomorrow it could be Microsoft platform + Microsoft models — a vertically integrated model similar to Google's. ## ● Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Intelligence 2C: Productized | Infra 2C: Emerging ### Vendor-Provided Components **Agent Governance Toolkit (April 2026, Open-Source)** [DAPM: Delegated] Seven-package, multi-language (Python, TypeScript, Rust, Go, .NET) system for governing autonomous AI agents. Agent OS: stateless policy engine, sub-millisecond enforcement (<0.1ms p99). Addresses all 10 OWASP agentic AI risks. 9,500+ tests. Deterministic, not LLM-based — 0.43 seconds total overhead across 11 agents over 11 days in Microsoft's internal deployment. Intent-based authorization: Declare → Approve → Execute → Verify lifecycle. Drift detection with soft-block, hard-block, or log-only responses. Deploy as AKS sidecar, Foundry middleware, or Azure Container Apps. **Microsoft Entra Agent ID (GA April 2026)** [DAPM: Ceded] Agent identity as first-class Entra citizens. Same Conditional Access, lifecycle management, risk detection, and governance as human identities. Agent blueprints: reusable identity templates for consistent governance. Shadow AI detection: discover unsanctioned agents. Sponsor lifecycle management. Four Conditional Access policy templates for agents. ID Protection extends anomaly detection to agent identities. Part of Microsoft Agent 365. Identity is a cross-cutting concern not fully assessed at this layer for other vendors in the series. **Microsoft Agent 365 (Control Plane)** [DAPM: Ceded] Unified agent registry and distributed control plane. Single inventory of all agents — Microsoft and non-Microsoft — operating in the organization. Agent Card Manifests provide rich metadata. Collection-based policies for discovery governance. SDK for third-party agent platforms to register agents. Converging the Entra Agent Registry under Agent 365 for simplified management. **Foundry AI Gateway** [DAPM: Ceded] Secures and manages MCP tools with policies and observability. Routes agent traffic through a governed gateway. Content Safety in Foundry Control Plane provides guardrails. Cross-prompt injection attack (XPIA) protection. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** All Layer 2C components are Microsoft IP or open-source. NVIDIA does not provide or control governance, policy, agent identity, or reasoning in the Azure stack. ### Gap Analysis Microsoft has the most productized Intelligence Layer 2C alongside Google. Four distinct 2C capabilities are GA or recently shipped: (1) Agent Governance Toolkit: Open-source, deterministic policy enforcement addressing all 10 OWASP agentic AI risks with sub-millisecond enforcement. Being open-source (MIT) means it's available to every vendor's customers — Microsoft built a governance standard others can adopt. (2) Entra Agent ID: The only vendor that has extended enterprise identity governance to AI agents as first-class identity citizens. Conditional Access for agents is GA. Shadow AI detection for unsanctioned agents is uniquely valuable for enterprises that don't yet know what agents are running. Identity as a governance dimension is not fully assessed across other vendors in this series — the Azure assessment surfaces it as a structural concern that warrants cross-vendor evaluation. (3) Agent 365: Unified agent registry covering Microsoft and non-Microsoft agents. The SDK for third-party platforms to register means Microsoft is building the universal agent inventory — even for agents that don't run on Azure. (4) Agent Governance Toolkit + Agent Framework integration: The Declare → Approve → Execute → Verify lifecycle with drift detection is the most structured agent governance protocol in this assessment. Infrastructure Layer 2C (emerging): Same gap as AWS — no service answers 'given data residency, cost, latency, and compliance, should this workload run on Maia, NVIDIA, or AMD in which region?' The policy-driven infrastructure placement engine does not exist as a product. Azure Arc + cross-cloud GPU scheduling (previewed at KubeCon) are building blocks, not a composed reasoning plane. The key differentiator vs. Google's Layer 2C: Azure's Intelligence 2C is model-agnostic — it governs agents regardless of which model powers them. Google's is Gemini-integrated. For enterprises running multi-model strategies, Azure's approach provides governance without model lock-in. ### Borrowed Judgment Intelligence 2C: Low — Agent Governance Toolkit is open-source, Entra Agent ID is Microsoft IP, Agent 365 is Microsoft IP. The enterprise defines governance policies; Microsoft provides the enforcement infrastructure. The Entra Agent ID dependency is worth flagging: extending agent identity into Entra means agent governance is tied to Microsoft's identity plane. An enterprise running agents on AWS that are governed by Entra Agent ID has a cross-cloud identity dependency on Microsoft. This is a deliberate strategic move — Microsoft is positioning Entra as the universal agent identity standard, as Active Directory became the universal enterprise identity standard. Infrastructure 2C: Not yet built. The building blocks (Arc, DRA, cross-cloud scheduling) exist but have not been composed into a policy-driven placement engine. ### Working Notes The Agent Governance Toolkit validates the 4+1 model's Layer 2C thesis directly. The OWASP Agentic Top 10 alignment, intent-based authorization lifecycle, and deterministic enforcement model all map precisely to what the 4+1 model describes as the Reasoning Plane's governance function. Microsoft's internal deployment data (11 agents, 7,000+ decisions, 0.43 seconds total governance overhead over 11 days) provides the first production evidence that Layer 2C governance can operate at negligible performance cost. Microsoft's Cloud Adoption Framework for agent governance (April 2026) provides a prescriptive four-layer composition model: data governance/compliance (Purview), agent observability (Agent 365, Defender, Log Analytics), agent security (Defender AI threat protection, Content Safety, AI Red Teaming Agent, RBAC, Sentinel), and agent development (Agent Framework, Foundry SDK, MCP, A2A). This is not a product but a reference architecture showing how the productized 2C components compose into an enterprise governance posture. The strategic play: Microsoft is building Layer 2C as an identity and governance story (Entra Agent ID + AGT). Google is building Layer 2C as a model intelligence story (Gemini integrated into infrastructure). VAST is building Layer 2C as a data platform story (PolicyEngine + Polaris). Three different vectors converging on the same architectural function. The 4+1 model predicted this convergence. ## ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Broadest Enterprise Ecosystem ### Vendor-Provided Components **Microsoft Foundry Model Catalog (11,000+ Models)** [DAPM: Delegated] OpenAI (GPT-5.5, o-series), Meta (Llama), Mistral, Cohere, NVIDIA Nemotron, Microsoft (MAI-1, Phi), plus thousands of open-source and industry-specific models. One of the broadest model marketplaces available. **Copilot Studio** [DAPM: Ceded] Low-code/no-code platform for building custom Copilot agents and extensions. Agents publish through Microsoft 365 Copilot and Teams. Enterprise-accessible without developer expertise. Managed MCP tool governance. **ISV + Partner Ecosystem** [DAPM: Delegated] Azure Marketplace with thousands of AI applications. SI partnerships (Accenture, Deloitte, KPMG, Infosys). ISV integrations across every industry. The Microsoft partner network is the largest in enterprise technology. **GitHub Copilot (Agentic Developer Platform)** [DAPM: Ceded] GitHub Copilot has evolved from code completion to a full agentic development platform with three surfaces: IDE agent mode (VS Code, Visual Studio 2026 with cloud agent integration), Copilot CLI (GA for all paid subscribers — Plan mode, Autopilot mode, dynamic agent delegation), and cloud agent (autonomous coding agent that researches repos, creates plans, makes code changes, opens PRs without developer in the loop via GitHub Actions). Multi-model: Claude, GPT, and OpenAI Codex models selectable per task. GitHub Copilot SDK enables building custom agents using Copilot's orchestration runtime — planning, tool invocation, streaming, MCP server integration. Foundry Local integration enables fully local, air-gapped agentic coding with data sovereignty. Microsoft Agent Framework supports Copilot SDK as agent backend. Visual Studio 2026 adds Debugger Agent that validates fixes against real runtime behavior. The most pervasive developer AI surface by installed base — integrated into the largest code hosting platform (GitHub) and the most widely used enterprise IDE ecosystem (VS Code + Visual Studio). ### NVIDIA-Provided Components **NVIDIA NIM + Blueprints on Foundry** NVIDIA models and application patterns available through Foundry. Same blueprints available across other assessed vendors — non-differentiating. ### Gap Analysis Azure's Layer 3 is structurally unique because Microsoft owns both the infrastructure platform AND the largest enterprise application estate in market. GitHub Copilot (developer AI), Microsoft 365 Copilot (knowledge worker AI, 3.3% paid adoption as of early 2026), Power Platform AI (business process automation), and Dynamics 365 Copilot (CRM/ERP AI) are Microsoft application products — not Azure services — but they consume Azure AI infrastructure at Layers 0–2C. This creates a dynamic no other vendor has: the enterprise's Layer 3 application decision often drives the Layer 0 infrastructure decision rather than vice versa. On-prem vendors sell infrastructure first; Azure sells applications first. This application estate context does not appear as assessed components because these are Microsoft products, not Azure services. The parallel exists in the Google Cloud assessment where Gemini's consumer properties are acknowledged as context without being assessed as GCP Layer 3 components. The GitHub Copilot vs. Microsoft 365 Copilot adoption contrast has implications beyond Azure. GitHub Copilot succeeds (high adoption, measurable productivity impact in a specific workflow). M365 Copilot struggles (3.3% conversion on the largest enterprise installed base). This suggests Layer 3 AI applications succeed when targeting specific professional workflows rather than augmenting general knowledge work — an observation relevant to every vendor's Layer 3 strategy. Copilot Studio is the Azure-native Layer 3 capability: low-code/no-code agent building with distribution through M365 and Teams. This is the bridge between the Microsoft application estate and the Azure AI platform — agents built in Copilot Studio consume Foundry models, are governed by Entra Agent ID, and are distributed through the M365 surface. GitHub Copilot's evolution from code completion to autonomous cloud agent represents the most significant Layer 3 shift in the Azure/Microsoft ecosystem. The cloud agent (coding autonomously in GitHub Actions, opening PRs without developer presence) creates a new category of AI-generated code flowing through enterprise repositories. The DAPM question: when Copilot's cloud agent writes production code autonomously, who owns the engineering judgment? The developer who assigned the issue, the model that generated the code, or GitHub's agent orchestration logic? The GitHub Copilot SDK is strategically important: it exposes Copilot's production agent runtime as a programmable API, enabling enterprises to build custom agents on top of GitHub's orchestration engine. Combined with Foundry Local for air-gapped on-device inference, this creates a Retained-authority path for enterprises that need agentic development without cloud dependency — the only assessed cloud vendor offering a fully local agentic developer tool. Three-cloud comparison of agentic developer surfaces: • AWS Kiro: Spec-driven, methodology-opinionated. Enforces structured requirements → design → implementation. Replaces Q Developer. Deep AWS integration (pricing, Well-Architected, Bedrock). Most prescriptive. • Google Antigravity 2.0: Agent orchestration platform. Multi-agent parallel execution, scheduled background tasks. Desktop + CLI + SDK. Replaces Gemini CLI. Gemini-native. Most ambitious multi-agent vision. • GitHub Copilot: IDE-embedded + CLI + cloud agent. Multi-model (Claude, GPT, Codex). GitHub-native (repos, issues, PRs, Actions). Copilot SDK for custom agents. Foundry Local for air-gapped. Most pervasive installed base. All three are Ceded or Delegated — the enterprise adopts the vendor's opinions about how AI-assisted development should work. The competitive differentiation is in the development philosophy: AWS imposes discipline (specs first), Google enables parallelism (multiple agents), Microsoft/GitHub enables delegation (assign and forget). ### Borrowed Judgment The most complex borrowed judgment landscape in this assessment. The enterprise using Azure AI inherits judgment from: Microsoft's platform decisions (Foundry defaults, content filtering), OpenAI's model decisions (alignment, capabilities, safety), NVIDIA's silicon decisions (GPU scheduling, driver behavior), ISV application decisions, and Microsoft's enterprise application decisions (Copilot integration points, Power Platform automation patterns). The unique risk: Microsoft's Layer 3 applications are also the enterprise's productivity tools. If Copilot in Word makes a poor suggestion, it affects the document. If Copilot in Dynamics 365 makes a poor recommendation, it affects the sales pipeline. The blast radius of borrowed judgment at Layer 3 is larger for Microsoft than for other vendors because the applications are mission-critical business tools, not standalone AI applications. ### Working Notes The Microsoft 365 Copilot adoption data is the enterprise AI reality check this assessment series benefits from. At 3.3% conversion on the largest installed base in enterprise software, the question is whether the industry's Layer 3 ambitions are ahead of enterprise readiness — and whether that readiness gap affects infrastructure investment decisions at Layers 0–2. GitHub Copilot's success vs. Microsoft 365 Copilot's adoption challenge has implications for every vendor's Layer 3 strategy: AI applications may succeed faster when they target specific professional workflows (coding, design, data engineering) than when they target general knowledge work. The GitHub Copilot SDK's availability as a programmable agent runtime backend via Microsoft Agent Framework creates a developer platform play that spans Layers 2B and 3. Enterprises building custom agents on the Copilot SDK inherit GitHub's orchestration opinions — planning, tool invocation, context management — as borrowed judgment. The SDK is the distribution mechanism for Microsoft's agent architecture opinions into enterprise codebases. Foundry Local + Copilot SDK for air-gapped agentic development is a unique capability in this assessment. No other cloud vendor provides a fully local, data-sovereign agentic developer tool. AWS Kiro requires Bedrock connectivity. Google Antigravity requires Gemini API. GitHub Copilot with Foundry Local runs entirely on-device. This matters for defense, financial services, and government enterprises where source code cannot leave the local environment. The agentic developer tools market is consolidating around three philosophies: structured discipline (Kiro), parallel orchestration (Antigravity), and delegation-first autonomy (Copilot). Red Hat's OpenShift Dev Spaces supporting Kiro, Copilot, Claude CLI, and others from a single governed runtime suggests the enterprise will run multiple agentic developer tools simultaneously — governed by the platform layer (2A) rather than choosing a single tool at Layer 3. ════════════════════════════════════════════════════════════════════════════════ # Cisco Secure AI Factory with NVIDIA Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.5 - VAST GPL Re-read & DAPM Reconciliation **Date:** July 12, 2026 **Source:** Cisco Live EMEA 2026, GTC 2026, RSA Conference 2026, Cisco Q3 FY2026 earnings, Galileo acquisition announcement (Apr 2026), AGNTCY/Linux Foundation donation (Jul 2025), DefenseClaw open source, Cisco/VAST partnership, analyst coverage, published 4+1 model. v1.5: two instrument rulings applied. (1) Whose-paper capability rule (the GPU parallel): the VAST AI Operating System is on Cisco's Global Price List with full Cisco support — the full platform deploys on Cisco paper, exposed and customer-administered, passing the exposure test; 1A/1B/1C promote moderate→strong at parity with the Supermicro row's identical delivery of the same engine. (2) Retraction of the OEM-channel DAPM rule: VAST components at 1A/1B/1C move Delegated→Ceded — a proprietary single-implementation platform's opinions are captive to its owner through any paper, matching VAST's own row. statusLabels moved to capability vocabulary. Capability grades and authority chips now read on their own axes. ## Summary Finding Cisco is the only vendor in this assessment series whose AI infrastructure authority is anchored by networking and security rather than compute, storage, or cloud services. Where Dell builds up from servers, HPE from sovereign compute, VAST from storage, and hyperscalers from managed services, Cisco builds outward from the network fabric and security posture that connects everything else. The Secure AI Factory with NVIDIA is a reference architecture, not a vertically integrated platform — and that distinction defines Cisco's entire 4+1 profile. Layer 0 is Cisco's strongest position with depth across both networking and compute. Cisco owns the network fabric (Silicon One G300, Nexus 9000/8000, Nexus Hyperfabric) and has built a three-tier GPU compute portfolio (C885A dense HGX for training, C845A modular MGX for inference, X-Series X580p for composable blade AI) that competes directly with Dell PowerEdge and HPE ProLiant. The X-Series disaggregated GPU architecture — independently managing CPU and GPU lifecycles via X-Fabric — is an architectural differentiator no other blade vendor matches. Cisco does not own storage; the AI POD storage story is partner-delivered (VAST Data primary — now on Cisco's Global Price List with full Cisco support — plus NetApp, Pure Storage, Hitachi Vantara, Nutanix as validated options), which reads as strong capability whose authority cedes to the chosen partner. This makes Cisco's Layer 0 the inverse of Dell's: Dell owns compute and storage but brands NVIDIA networking silicon. Cisco owns networking silicon and has purpose-built AI compute but depends on partners for the data foundation. The security and observability layers are where Cisco makes its most differentiated claim. AI Defense, Duo Agentic Identity, DefenseClaw, Splunk Observability Cloud with AI Agent Monitoring, and the Galileo acquisition together represent the most comprehensive agent security and observability portfolio of any infrastructure vendor assessed. These capabilities span Layers 2A through 2C — and they constitute genuine Cisco-owned IP, not rebranded NVIDIA or partner technology. The structural question is whether security and observability are sufficient to constitute a control plane, or whether they remain constraint enforcement without placement reasoning. The AGNTCY initiative — open-sourced by Cisco's Outshift incubator, donated to the Linux Foundation with Dell, Google Cloud, Oracle, and Red Hat as formative members — contributes infrastructure primitives (discovery, identity, messaging, observability) to the broader agentic standards ecosystem. AGNTCY sits within the Agentic AI Foundation (AAIF), the fastest-growing project in Linux Foundation history, where MCP (Anthropic), A2A (Google), AGENTS.md (OpenAI) define the foundational protocols. Cisco is a Gold member of AAIF, not a Platinum founder — the ecosystem-defining standards were originated by Anthropic, Google, and OpenAI, not Cisco. AGNTCY contributes essential plumbing beneath those protocols, and Cisco's commercial products (AI Defense, Duo, Splunk) could become the enterprise implementation layer if these open standards achieve adoption. But Cisco's position is contributor, not definer. Cisco's ~$9B in projected FY2026 AI infrastructure orders, record $15.8B quarterly revenue, and hyperscaler design wins (Silicon One P200, G200) validate market traction. The Galileo acquisition and DefenseClaw open-source release signal an intent to own the AI agent trust layer. But the 4+1 assessment reveals a vendor whose authority is wide but thin: strong at the network layer and security perimeter, Delegated or Absent at the data plane and runtime layers. Cisco secures the AI Factory. It does not yet govern it. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Networking Strength — Asymmetric ### Vendor-Provided Components **Silicon One G300 (AI Networking Silicon)** [DAPM: Ceded] 102.4 Tbps switching silicon for massive AI cluster buildouts. Intelligent Collective Networking delivers 33% increase in network utilization and 28% improvement in job completion time. Powers gigawatt-scale AI clusters for training, inference, and real-time agentic workloads. Cisco-designed silicon — not NVIDIA-branded, not repackaged. This is Cisco's most significant Layer 0 differentiator: owned networking silicon IP at the same architectural level as AWS Nitro/EFA/SRD, Google Virgo, and Microsoft SONiC. Proprietary Cisco platform — opinions captive, no open exit. G300 silicon ships H2 2026 (announced Cisco Live EMEA, Feb 2026); current-gen Silicon One and the Nexus platform ship today, which is what anchors this layer. **Nexus 9000 / 8000 Systems (G300-powered)** [DAPM: Ceded] 102.4 Tbps switching speeds in liquid-cooled and air-cooled designs. Nearly 70% energy efficiency improvement with liquid cooling and advanced optics. Common hardware for diverse fabric types including Nexus Hyperfabric. Designed for hyperscalers, neoclouds, sovereign clouds, service providers, and enterprises. First P200 design wins confirmed with hyperscalers; third P200 win in early Q4 2026. Proprietary Cisco platform — opinions captive, no open exit. Existing Nexus ships today; G300-powered systems H2 2026. **Nexus Hyperfabric (Cloud-Managed AI Fabric)** [DAPM: Ceded] Cloud-managed network fabric integrating networking, GPU, and storage into a unified infrastructure. Fabric-as-a-Service model with automated deployment and operations. NVIDIA NCP (Networking Connectivity Platform) design validation. Sharon AI selected Hyperfabric for Australia's first Cisco Secure AI Factory deployment (1,024 NVIDIA Blackwell Ultra GPUs). Proprietary Cisco platform — opinions captive, no open exit. **Nexus One (Unified Management Plane)** [DAPM: Ceded] Unified management across on-premises and cloud-based data center deployments. AI Job Observability provides job-aware, network-to-GPU visibility correlating network telemetry with AI workload behavior. Native Splunk Platform integration for in-place telemetry analysis without data movement — essential for sovereign cloud and compliance-sensitive environments. API-driven automation and customization built-in. Proprietary Cisco platform — opinions captive, no open exit. **Cisco UCS C885A M8 (Dense GPU — HGX Platform)** [DAPM: Retained] Cisco's first 8-way accelerated computing system. Built on NVIDIA HGX platform with 8x NVIDIA H100 or H200 Tensor Core GPUs, OR 8x AMD MI300X/MI350X OAM GPUs — multi-vendor GPU support in the same dense server. One ConnectX-7 NIC or BlueField-3 SuperNIC per GPU for cluster-scale training. BlueField-3 DPUs for accelerated GPU-to-data access and zero-trust security. Designed for large LLM training, fine-tuning, and large model inference. MLPerf Training benchmarks published January 2026. This is Cisco's direct competitor to Dell PowerEdge XE9680 and HPE ProLiant DL380a — a genuine leadership-class AI server, not a rebadged reference design. **Cisco UCS C845A M8 (Modular GPU — MGX Platform)** [DAPM: Retained] Flexible, scalable AI server based on NVIDIA MGX modular reference design. Configurable from 2 to 8 NVIDIA or AMD PCIe GPUs including RTX PRO 6000 and RTX PRO 4500 Blackwell Server Edition in a 4RU chassis. Enhanced power delivery, fewer PCBs, improved cable routing for optimal airflow and thermal management. GPU hot-swap for faster replacement. E1.S SSDs for increased storage density. First VAST Certified CNode-X platform in market (UCS C845A M8 with RTX PRO 6000). 'Start small and scale up' positioning for enterprises ramping AI from inference to fine-tuning. Day-1 Intersight management for unified AI and traditional workload operations. **Cisco UCS X-Series + X580p GPU Node (Modular Blade AI)** [DAPM: Ceded] Disaggregated, composable AI compute within the UCS X9508 blade chassis. X580p PCIe Node adds up to 4 NVIDIA GPUs (RTX PRO 6000/4500 Blackwell, H200 NVL, L40S) to X210c/X215c compute nodes via X-Fabric Technology. X9516 X-Fabric Module provides PCIe Gen 5 switching with ultra-low-latency, NVLink Bridge support, and dynamic GPU-to-host provisioning. Independent CPU and GPU lifecycle management — upgrade GPUs without replacing compute nodes. MLPerf Inference benchmarks published April 2026. Up to 24 GPUs per 7RU chassis. This is architecturally distinctive: no other vendor offers a modular blade system with disaggregated GPU composition for AI workloads. Dell PowerEdge MX and HPE Synergy do not have equivalent GPU composability. **Cisco UCS C240 M8 / C225 M8 (Mainstream + Storage)** [DAPM: Retained] C240 M8: 2U rack server with Intel Xeon 6 processors, up to 2 double-width NVIDIA GPUs (RTX PRO 6000 Blackwell, H200 NVL). Balanced compute for distributed AI, analytics, edge AI, and vision workloads. Up to 28 NVMe/SAS/SATA drives for high-speed data access. C225 M8: designated as VAST Data Platform EBOX (storage servers) in AI POD architecture — the persistent storage foundation for Cisco AI PODs. Both managed through Cisco Intersight. **Cisco AI PODs (Pre-Validated Full-Stack Infrastructure)** [DAPM: Ceded] Modular, full-stack AI infrastructure platform combining UCS C885A/C845A compute, Nexus 9000 networking (up to 800G), NVIDIA GPUs, and partner storage (VAST Data, NetApp, Pure Storage, Hitachi Vantara, Nutanix). Scalable from 32 to 128+ GPUs. Cisco Validated Designs (CVDs) and NVIDIA Enterprise Reference Architectures (ERAs) for training, fine-tuning, inference, and RAG workloads. Pre-validated designs reduce setup time by up to 50%. Integrated management through Cisco Intersight and Nexus Dashboard. Modular scale-unit design enables growth without full infrastructure overhaul. Proprietary Cisco platform — opinions captive, no open exit. **Cisco Unified Edge (Edge AI Compute)** [DAPM: Retained] Edge compute platform supporting NVIDIA RTX PRO 4500 and 6000 Blackwell Server Edition GPUs for mission-critical AI at the edge. Also supports NVIDIA L4 GPUs. Zero-touch deployment with pre-validated blueprints. Centralized management via Intersight with Splunk and ThousandEyes integrations for end-to-end edge observability. Multi-layered zero-trust security with tamper-proof features, deep telemetry, drift-free configurations. Cisco AI Grid reference design extends edge AI to service providers via Cisco Mobility Services Platform — a unique go-to-market that no other assessed vendor offers. ### NVIDIA-Provided Components **NVIDIA GPU Silicon (HGX, MGX, RTX PRO Blackwell)** All GPU acceleration in Cisco UCS servers depends on NVIDIA silicon — same structural dependency as Dell and HPE. But Cisco adds genuine compute engineering above the GPU: the X-Series X580p disaggregated GPU composition, X-Fabric PCIe Gen 5 switching, MGX reference design improvements (enhanced power delivery, PCB reduction, cable routing for thermal management), and multi-vendor GPU support (AMD MI300X/MI350X on C885A alongside NVIDIA HGX). Cisco's compute differentiation is modularity and composability, not thermal engineering (Dell) or sovereign heritage (HPE). **NVIDIA Spectrum-X Switch Silicon** Cisco offers BOTH Silicon One-powered switches (Cisco-designed silicon) AND Spectrum-X-powered switches (NVIDIA silicon). This dual-silicon networking strategy is unique in the assessment series — Dell only brands NVIDIA Spectrum, HPE only uses Juniper/Aruba/Slingshot, VAST depends on OEM networking. Cisco is the only vendor that competes with NVIDIA in switching silicon while also offering NVIDIA's switching silicon. **NVIDIA BlueField DPUs** Cisco extends Hybrid Mesh Firewall policy enforcement to BlueField DPUs embedded in GPU servers, enabling threat mitigation at the server level before reaching sensitive data. This is Cisco adding security value on top of NVIDIA hardware — a pattern consistent across the assessment. **NVIDIA AI Enterprise + NIM** Cloud-native software tools, libraries, frameworks, dynamic GPU resource allocation, AI workload scheduling, and production-ready models. Same NVIDIA software dependency as Dell and HPE at this layer. ### Gap Analysis Layer 0 is Cisco's strongest position with genuine depth across both networking AND compute — not just networking. Networking authority is unmatched among on-prem vendors. Silicon One is proprietary switching silicon comparable in strategic significance to AWS's Nitro/EFA, Google's Virgo, or Microsoft's SONiC. No other on-prem infrastructure vendor designs their own switching silicon for AI networking. Dell brands NVIDIA Spectrum. HPE acquired Juniper for networking IP but doesn't design switching ASICs. VAST depends entirely on OEM networking. The compute portfolio is broader and more architecturally deliberate than initial assessment suggested. The three-tier GPU server strategy (C885A for dense 8-GPU HGX training, C845A for flexible 2-8 GPU MGX inference and fine-tuning, X-Series X580p for composable blade AI) covers the full spectrum of enterprise AI workloads from leadership-class training to distributed inference. The C885A competes directly with Dell PowerEdge XE9680 and HPE ProLiant DL380a as a genuine 8-GPU dense server with MLPerf benchmarks published. The X-Series X580p is an architectural differentiator that deserves specific recognition: disaggregated GPU composition via X-Fabric allows enterprises to independently manage CPU and GPU lifecycles, dynamically provision GPU resources to compute nodes, and scale GPU density within a blade chassis (up to 24 GPUs per 7RU). No other vendor offers modular blade-based GPU composability at this level. Dell PowerEdge MX and HPE Synergy do not have equivalent GPU disaggregation. This is Cisco applying its composable infrastructure heritage (UCS X-Series has always been about disaggregation) to the AI compute problem. Multi-vendor GPU support (AMD MI300X/MI350X alongside NVIDIA HGX on C885A) gives Cisco the same silicon optionality that HPE offers with GX5000 (NVIDIA Rubin + AMD MI430X). Dell's AI Factory is NVIDIA-only (AMD under separate branding). The storage gap remains the most significant Layer 0 structural difference from Dell and HPE. Cisco does not manufacture or sell storage. The AI POD storage story depends on partners: VAST Data (primary), plus NetApp, Pure Storage, Hitachi Vantara, and Nutanix as validated options. This multi-vendor storage flexibility is arguably a strength for the reference-architecture model — the enterprise retains storage vendor choice — but it means Cisco cannot build a vertically integrated data plane. The dual-silicon networking strategy (Silicon One + Spectrum-X) is strategically unique. Cisco gives customers a choice: Cisco-designed silicon or NVIDIA silicon, both managed through Nexus One. This preserves the NVIDIA partnership while maintaining Cisco's networking authority. If NVIDIA changes its Spectrum roadmap, Cisco's customers have an alternative that Dell's customers do not. The ~$9B in FY2026 AI infrastructure orders and hyperscaler design wins (P200, G200) validate that the networking + compute + reference architecture approach has market traction at the highest scale. ### Borrowed Judgment Moderate but structurally different from Dell's or HPE's. Cisco borrows GPU silicon judgment from NVIDIA (same as everyone) but retains genuine compute platform engineering judgment: X-Series composable architecture, X-Fabric GPU disaggregation, MGX reference design improvements, multi-vendor GPU support, and the three-tier server strategy. Cisco also retains networking judgment entirely — Silicon One, Nexus 9000/8000, Hyperfabric, and Nexus One are all Cisco IP. Storage judgment is borrowed from VAST Data and other partners. The comparison: • Dell retains compute packaging (thermal, mechanical, rack-scale) and storage judgment, borrows networking silicon from NVIDIA (Spectrum). • HPE retains compute judgment (Cray heritage, ProLiant) and networking judgment (Juniper/Aruba/Slingshot), borrows runtime from NVIDIA. • Cisco retains networking judgment AND compute platform engineering (UCS, X-Series composability), borrows GPU silicon (NVIDIA) and storage (partners). Each on-prem vendor retains authority in their heritage domain. Cisco's heritage is networking, but UCS — now in its M8 generation with purpose-built AI servers — has matured from 'networking company does compute' to a genuine multi-tier AI compute platform. The X-Series disaggregated GPU architecture is Cisco's compute contribution that has no direct equivalent from Dell or HPE. ### Working Notes The UCS X-Series composable architecture is Cisco's most underappreciated Layer 0 capability. The X580p + X-Fabric model — dynamically provisioning GPU resources to compute nodes, independently managing CPU and GPU lifecycles, NVLink Bridge support within a blade chassis — applies the same disaggregation principle that defined UCS from inception. If GPU lifecycle velocity continues to outpace CPU lifecycle velocity (which it will), the ability to upgrade GPUs without replacing the compute node is a genuine operational and CapEx advantage. No other blade vendor offers this. The C885A M8 with AMD MI300X/MI350X support is worth tracking. If AMD Instinct gains enterprise traction, Cisco is one of two on-prem vendors (alongside HPE) positioned to offer both NVIDIA and AMD in the same dense-GPU server platform. Dell's AMD support is under separate 'Dell AI Platform with AMD' branding — a different SKU family, not GPU optionality within the same chassis. The Cisco AI Grid reference design for service providers (Cisco Mobility Services Platform + NVIDIA RTX PRO Blackwell GPUs) is a unique go-to-market that no other assessed vendor offers. Dell, HPE, and VAST target enterprises and neoclouds. Cisco targets enterprises, neoclouds, AND the service provider edge — leveraging telco relationships that predate the AI era. If edge inference becomes a significant workload category, Cisco's service provider footprint is a distribution advantage that pure-compute vendors cannot match. The energy efficiency story is worth noting: nearly 70% improvement with liquid cooling and advanced optics. At hyperscale, energy efficiency is a buying criterion that often outranks raw performance — and it plays to Cisco's infrastructure engineering strengths. Cross-row finding (July 2026): Lenovo's Hybrid AI Factory second phase integrates Cisco 800GbE Silicon One switches — Cisco fabric selling through a competitor OEM's AI factory. Cisco's networking IP is becoming a component in other vendors' stacks, a distribution surface no other assessed vendor's fabric has. The multi-vendor storage partner model (VAST, NetApp, Pure Storage, Hitachi Vantara, Nutanix) within AI PODs is structurally different from Dell's single-vendor storage story (PowerScale/ObjectScale/Exascale). The reference architecture model gives enterprises storage vendor choice — but it also means Cisco cannot optimize the compute-to-storage integration path the way Dell can with Exascale or VAST can with DASE. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Partner-Delivered Data Platform ### Vendor-Provided Components **VAST Data Platform on Cisco AI PODs (EBOX)** [DAPM: Ceded] VAST Data Platform running on Cisco UCS C225 M8 servers (designated EBOX). VAST DASE architecture provides shared-everything model with NVMe-over-Fabrics, global namespace, and ACID guarantees. Managed through Cisco Intersight alongside compute and networking. VAST was recognized at VAST Forward 2026 for its Cisco partnership. The customer buying a Cisco Secure AI Factory gets VAST storage as a validated, integrated component — and the VAST AI OS rides Cisco's Global Price List with full Cisco support. Cisco delivers the capability; VAST's opinions do the capturing: the platform's data-architecture opinions are VAST-captive through any paper, as on VAST's own row. **Multi-Vendor Storage Options (AI POD Validated)** [DAPM: Delegated] Cisco AI PODs validate multiple storage partners: VAST Data (primary/deepest integration), NetApp, Pure Storage, Hitachi Vantara, and Nutanix. The enterprise retains storage vendor choice — a structural advantage of the reference architecture model over Dell's single-vendor storage story (PowerScale/ObjectScale) or VAST's vertically integrated approach. Each storage partner brings its own governance metadata model, meaning Layer 1A governance capabilities vary by storage choice. Menu-altitude Delegated: the choice is real before purchase; once deployed, each chosen platform captures per its own terms. **Cisco Hypershield + Isovalent (Data Security)** [DAPM: Ceded] Zero-trust, ransomware-resilient storage security and inline network security. Cisco Secure AI Factory security principles applied to the data layer. Isovalent provides eBPF-based runtime security for containerized AI workloads. Hybrid Mesh Firewall extends policy enforcement to BlueField DPUs at the storage server level. Security is Cisco-owned and layered on top of whichever storage partner the enterprise selects — consistent security posture regardless of storage choice. Proprietary Cisco platform — opinions captive, no open exit. ### NVIDIA-Provided Components **NVIDIA AI Data Platform Reference Design** Cisco + NVIDIA + VAST validated as one of the first NVIDIA AI Data Platform reference designs. NVIDIA provides the reference architecture standard; Cisco provides networking and compute; VAST provides the data platform. ### Gap Analysis Applying the litmus test — what does the enterprise get when it asks Cisco for a 4+1 AI infrastructure? — the answer is clear: the Cisco Secure AI Factory delivers a complete data foundation. VAST Data Platform provides AI-optimized storage with DASE architecture, global namespace, and ACID guarantees. NetApp, Pure Storage, Hitachi Vantara, and Nutanix are validated alternatives. The customer gets Layer 1A capability through Cisco's go-to-market. The two axes read separately, which is the point: the capability is strong — the full VAST platform on Cisco's Global Price List with Cisco support, plus a validated multi-vendor menu — and the authority on the deployed primary is Ceded to VAST, whose opinions are captive through any paper. The customer buying Cisco doesn't experience VAST storage as a gap; they experience it as strength whose capture belongs to the partner they chose at menu time. Where Cisco's Layer 1A is genuinely thinner than Dell's or HPE's is governance metadata. Dell's MetadataIQ indexes billions of files across PowerScale/ObjectScale with automated classification, tagging, and metadata enrichment — Dell-owned IP. HPE's Data Fabric provides policy-based data placement with lineage tracking — HPE-owned IP. Cisco has no equivalent Cisco-owned metadata or governance capability. The governance metadata available depends entirely on which storage partner the enterprise selects. With VAST, the enterprise gets VAST Catalog. With NetApp, different governance primitives. Cisco adds consistent security across all of them (Hypershield, Isovalent) but does not add a Cisco-owned governance layer above the storage partner. The 4+1 model defines Layer 1A as the 'Governance Catalog that Layer 2C queries.' The catalog exists in the Cisco Secure AI Factory — it's provided by the storage partner. Cisco does not own that catalog, which means Cisco's future Layer 2C ambitions depend on a partner's metadata model. Dell and HPE can build from their own metadata to their own control plane. Cisco would need to build from VAST's metadata to Cisco's control plane — a cross-vendor integration that neither party has announced. The multi-vendor storage model is both a strength and a structural constraint. Strength: the enterprise retains storage vendor choice, avoiding single-vendor lock-in at the data layer. Constraint: Cisco cannot optimize the compute-to-storage-to-governance integration path the way Dell can with Exascale + MetadataIQ or VAST can with DASE + Catalog. Each storage partner brings its own data architecture, its own metadata model, and its own governance surface — Cisco must integrate across all of them rather than optimizing for one. ### Borrowed Judgment High for storage platform and governance metadata, low for data security. The enterprise buying a Cisco Secure AI Factory inherits the storage partner's data architecture decisions — VAST's DASE model, NetApp's ONTAP model, or Pure's Purity model depending on selection. Cisco contributes consistent security posture across all storage choices (Hypershield, Isovalent, Hybrid Mesh Firewall) but does not contribute governance metadata, data classification, or compliance tagging above the storage partner. Comparison: Dell borrows Layer 1A governance judgment from Trust3 AI (a specific partner function) but retains storage platform judgment (PowerScale, ObjectScale, MetadataIQ are Dell IP). HPE retains both storage platform (Alletra) and governance (Data Fabric) judgment. Cisco borrows the storage platform from partners but retains the security layer — structurally higher borrowed judgment than Dell or HPE at Layer 1A, but the customer still receives a complete solution. ### Working Notes The strategic question is whether Cisco intends to remain storage-agnostic (a networking company that partners with storage vendors) or will eventually acquire or build storage capability. The VAST partnership depth suggests the former — but the competitive pressure from Dell's Exascale and HPE's Alletra may eventually force the question. Cisco's absence from the storage market is also why VAST ships CNode-X through Cisco as an OEM partner. The relationship is complementary, not competitive: Cisco needs VAST for data, VAST needs Cisco for networking and enterprise distribution. Dell is notably absent from VAST's OEM partner list precisely because Dell and VAST compete at Layer 1A. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Partner-Delivered Retrieval ### Vendor-Provided Components **VAST InsightEngine on Cisco AI PODs** [DAPM: Ceded] VAST InsightEngine on Cisco UCS C845A M8 servers (first VAST Certified CNode-X platform). Real-time vector embedding and retrieval for RAG and agentic workflows. Integrates with NVIDIA NIM microservices for AI-native retrieval. Automates embedding, indexing, and retrieval pipelines. Marketed as reducing RAG pipeline latency from minutes to seconds. This is a validated component of the Cisco Secure AI Factory — the customer asking Cisco for RAG capability receives InsightEngine as part of the delivered solution, on Cisco paper via the GPL. Lift-to-leave: the retrieval and indexing opinions are VAST-captive through any paper, as on VAST's own row. **Cisco AI Networking for Retrieval Performance** [DAPM: Ceded] Silicon One G300 Intelligent Collective Networking directly impacts retrieval latency and throughput. The 33% network utilization improvement and 28% job completion time improvement are not just training metrics — they affect how quickly GPU compute accesses storage-side vector indexes. Cisco's networking fabric is the enabling layer that makes VAST's retrieval performance achievable at scale. Lossless, low-latency Nexus 9000 fabric with up to 800G bandwidth between compute and storage tiers. Proprietary Cisco platform — opinions captive, no open exit. ### NVIDIA-Provided Components **NVIDIA AI Data Platform + NeMo Retriever** NVIDIA AI Data Platform reference design validates the Cisco + VAST retrieval architecture. NeMo Retriever provides embedding models and reranking. Same NVIDIA retrieval stack available to Dell and HPE — the differentiation is in the validated integration, not the NVIDIA components. ### Gap Analysis Applying the customer litmus test: the enterprise asking Cisco for RAG and retrieval capability receives VAST InsightEngine as a validated component of the Secure AI Factory, on Cisco paper. The retrieval function is delivered at the strength it carries on VAST's own row — that is the capability grade. The authority chip says the rest: the retrieval opinions are Ceded to VAST through any paper. Cisco's owned contribution at Layer 1B is the networking fabric that makes high-performance retrieval possible — lossless connectivity between GPU compute nodes and VAST storage with predictable latency. This is not the retrieval capability itself, but it is the enabling condition. The Cisco + NVIDIA + VAST data platform is marketed as 'the first enterprise architecture unifying compute, fabric, and storage into a single, validated platform to accelerate RAG' — and the validated integration is genuine, even though each component has a different owner. Where Cisco is thinner than Dell at Layer 1B: Dell has Data Analytics Engine with its own MCP Server (Feb 2026), blurring the 1A/1B boundary with search, analytics, and orchestration surfaced as a single queryable service — Dell-owned IP. HPE has Data Fabric with integrated vector search capabilities. Cisco's Layer 1B retrieval is entirely partner-provided. If the enterprise selects a storage partner other than VAST (NetApp, Pure Storage, etc.), the Layer 1B retrieval story changes entirely — each partner brings different retrieval capabilities, and Cisco provides no Cisco-owned retrieval abstraction above them. The network-to-retrieval performance correlation is an underappreciated Cisco contribution. When Nexus One's AI Job Observability shows that retrieval latency degraded because of network congestion on a specific path between compute and storage tiers, that's retrieval-relevant intelligence that no storage vendor can provide independently. Cisco sees the network between the GPU and the data — VAST sees the data, NVIDIA sees the GPU, Cisco sees the fabric connecting them. ### Borrowed Judgment High for retrieval logic, low for retrieval-enabling networking. All retrieval and context management intelligence is provided by VAST (InsightEngine) and NVIDIA (NeMo Retriever, AI Enterprise). Cisco provides the networking substrate that determines retrieval performance characteristics — and that substrate is Cisco-owned IP with genuine impact on retrieval latency and throughput. The structural comparison: Dell borrows retrieval acceleration from NVIDIA (cuVS, NeMo Retriever) but retains the storage platform on which retrieval operates (PowerScale). Cisco borrows both retrieval logic (VAST InsightEngine) and retrieval acceleration (NVIDIA) but retains the networking fabric that connects them. ### Working Notes The network-as-retrieval-enabler framing is worth developing further. In a disaggregated architecture where GPU compute and vector storage are on separate server tiers connected by fabric, the network IS part of the retrieval path. Cisco's ability to correlate network telemetry with retrieval latency via Nexus One and Splunk is a genuine observability advantage — one that could feed Layer 2C placement decisions about which compute-to-storage path to use for a given retrieval workload. The VAST InsightEngine as first VAST Certified CNode-X platform on Cisco UCS C845A M8 (with RTX PRO 6000 acceleration) represents the deepest Cisco-VAST integration point. This is where the partnership moves from 'VAST storage on Cisco servers' to 'VAST compute-storage fusion on Cisco infrastructure' — closer to a joint product than a validated configuration. ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Partner-Delivered Pipeline Engines ### Vendor-Provided Components **VAST DataEngine / SyncEngine (on Cisco AI PODs)** [DAPM: Ceded] VAST DataEngine executes serverless functions directly where data lives — 'bringing compute to the data.' SyncEngine indexes and synchronizes data from external sources, triggering enrichment pipelines automatically. Both are delivered as part of the Cisco Secure AI Factory with VAST Data, on Cisco paper via the GPL. The customer asking Cisco for data pipeline capability receives these as validated components — a pipeline capability set rated strong on VAST's own row. Lift-to-leave: DataEngine functions and SyncEngine flows are VAST-captive through any paper. **Cisco Networking Fabric (Data Movement Substrate)** [DAPM: Ceded] Silicon One G300 Intelligent Collective Networking optimizes data flow across the AI cluster. The network fabric is the physical data movement layer — every byte of training data, every embedding pipeline, every model checkpoint traverses Cisco's switching fabric. Lossless, low-latency networking with up to 800G bandwidth is a prerequisite for high-throughput data pipelines. Cisco's contribution to Layer 1C is the movement infrastructure, not the pipeline logic. Proprietary Cisco platform — opinions captive, no open exit. ### Gap Analysis Applying the customer litmus test: the enterprise asking Cisco for data pipeline capability receives VAST DataEngine and SyncEngine as validated components of the Secure AI Factory, on Cisco paper via the GPL. The pipeline function is delivered at the strength it carries on VAST's own row — the capability grade credits it fully, the GPU-parallel rule. The authority chip says the rest: pipeline opinions are Ceded to VAST through any paper. Cisco has no proprietary data pipeline, ETL, or feature engineering tools — and this is consistent with Cisco's architectural identity. Cisco has never been a data management software company. Dell's Dataloop-powered Data Orchestration Engine is notable precisely because it's Dell's first proprietary software asset in the data lifecycle. HPE's Data Fabric provides policy-based data placement. IBM has watsonx.data with Confluent streaming. Cisco's equivalent is the validated partner integration. Cisco's owned contribution at Layer 1C is the networking fabric that data pipelines traverse. In a disaggregated AI architecture, data movement between storage tiers, GPU compute, and model serving endpoints is constrained by network bandwidth and latency. Silicon One G300's Intelligent Collective Networking — the 33% utilization improvement and 28% job completion time improvement — directly accelerates data pipeline throughput. This is an infrastructure contribution, not a software contribution, but it's a real one. The Splunk data ingestion and processing capabilities could theoretically extend into data pipeline territory — Splunk already handles high-volume data streams for observability. But Splunk is positioned as an observability and security platform, not an AI data pipeline tool. No Cisco signals suggest this expansion. ### Borrowed Judgment High for pipeline logic, low for data movement infrastructure. All data pipeline orchestration, serverless execution, and data synchronization judgment is provided by VAST (DataEngine, SyncEngine) or by the enterprise's own tooling. Cisco provides the networking fabric that determines data movement performance — Cisco-owned IP that directly impacts pipeline throughput. If the enterprise selects a storage partner other than VAST, the Layer 1C story changes significantly. NetApp, Pure Storage, and Hitachi Vantara each have their own data movement capabilities, none validated at the same depth as VAST within the Cisco AI POD architecture. ### Working Notes The networking-as-data-movement framing is structurally consistent across Layers 1A, 1B, and 1C. At each data layer, Cisco's Retained contribution is the networking fabric that connects and enables partner-provided data capabilities. This is not a gap in the customer experience — the customer gets a complete data plane. It is a structural characteristic of Cisco's reference architecture model: Cisco owns the connectivity, partners own the data logic. This pattern has a Layer 2C implication. If Cisco ever builds a control plane that makes placement decisions about data (where should this data live, how should it move, which pipeline should process it), that control plane would need to integrate with multiple storage partners' data APIs. Dell's hypothetical control plane only needs to integrate with Dell's own storage. Cisco's hypothetical control plane needs to be multi-vendor by design — which is harder to build but potentially more valuable in heterogeneous enterprise environments. ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Networking-Centric Orchestration ### Vendor-Provided Components **Cisco Intersight (Unified Fleet Management)** [DAPM: Ceded] Cloud-based infrastructure management for UCS compute, Nexus networking, and VAST storage lifecycle. Policy-driven server profiles, firmware management, and workload optimization. Manages the chassis and networking — comparable to Dell's OpenManage Enterprise but with stronger networking integration. Does not manage GPU workload scheduling directly. Proprietary Cisco platform — opinions captive, no open exit. **Nexus One (AI Network Operations)** [DAPM: Ceded] Unified management plane for data center networking. AI Job Observability correlates network telemetry with AI workload behavior. AgenticOps-powered autonomous network operations. The AI-aware network management capability is Cisco-unique — no other vendor correlates network health with GPU job completion at this level of integration. Proprietary Cisco platform — opinions captive, no open exit. **AgenticOps (Agent-First IT Operating Model)** [DAPM: Ceded] Agent-driven IT operations across networking, security, and observability. Cross-domain telemetry from Cisco Networking, Security Cloud Control, Nexus One, Splunk, and ThousandEyes. Agentic Workflows and AI Canvas for troubleshooting and automation. Deep Network Model provides system-wide awareness. Extends from cloud to on-premises to air-gapped industrial environments. Proprietary Cisco platform — opinions captive, no open exit. **Cisco Hybrid Mesh Firewall** [DAPM: Ceded] Policy enforcement across network switches, workloads, and NVIDIA BlueField DPUs. Extends security policy to the GPU server level. This is a Layer 0/2A security function — infrastructure-level policy enforcement that operates below the application runtime. Proprietary Cisco platform — opinions captive, no open exit. ### NVIDIA-Provided Components **NVIDIA GPU Operator + Run:ai** GPU scheduling, resource allocation, and workload orchestration. Same NVIDIA dependency as Dell and HPE at this layer. Cisco does not own GPU-aware scheduling primitives. **NVIDIA AI Enterprise** Enterprise AI software platform providing GPU drivers, container runtime, and validated AI frameworks. Licensed separately from Cisco infrastructure. ### Gap Analysis Cisco's Layer 2A is networking-centric: Intersight manages the infrastructure fleet, Nexus One manages the AI network, AgenticOps provides agentic IT operations. These are genuine Cisco-owned capabilities with no equivalent in Dell's or HPE's stack in terms of network-to-GPU observability correlation. But GPU-aware scheduling and workload orchestration — the core Layer 2A functions — are NVIDIA-controlled (GPU Operator, Run:ai, AI Enterprise). This is the same gap Dell and HPE face. The enterprise using Cisco AI PODs schedules GPU workloads through NVIDIA's stack, not through Cisco's. The AgenticOps framework is architecturally interesting because it applies agentic AI to IT operations itself — using AI agents to manage the infrastructure that runs AI agents. No other on-prem vendor has an equivalent agent-driven IT operations model. But AgenticOps manages infrastructure operations (networking, security, observability), not AI workload placement. It's a Layer 2A operational capability, not a Layer 2C governance capability. Nexus One's AI Job Observability deserves specific note: correlating network telemetry with AI job behavior is a genuine signal that could feed Layer 2C placement decisions. If the network can tell you that a training job's completion time degraded because of network congestion on a specific spine switch, that's placement-relevant intelligence. Whether this signal is consumed by any placement logic today is not evident. ### Borrowed Judgment Moderate. Cisco retains infrastructure management and network orchestration judgment (Intersight, Nexus One, AgenticOps). GPU workload scheduling judgment is borrowed from NVIDIA (Run:ai, GPU Operator), same as Dell and HPE. The AI Job Observability capability reduces borrowed judgment marginally — Cisco can see network-to-GPU correlations that NVIDIA's orchestration layer may not surface. But seeing the problem and acting on it are different functions. Cisco sees; NVIDIA schedules. ### Working Notes AgenticOps' cross-domain telemetry — ingesting signals from networking, security, and observability into a single agentic decision surface — is the closest thing to a unified operations control plane in the on-prem assessment series. HPE's GreenLake Intelligence is comparable but serves a different function (infrastructure operations vs. networking operations). The question is whether AgenticOps evolves from operational automation into policy-driven governance. ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Security Layer — Not Runtime ### Vendor-Provided Components **Cisco AI Defense** [DAPM: Ceded] Industry-first AI security solution for securing AI models, agents, applications, and infrastructure. Integrates NVIDIA NeMo Guardrails for AI application security. Secures NVIDIA OpenShell agent development platform with controls and guardrails to govern agent and claw actions. Model validation, prompt injection defense, data leakage prevention, bias detection. Spans Layer 2B (runtime security) and Layer 2C (agent governance). This is Cisco's most differentiated Layer 2B contribution — no other infrastructure vendor has an equivalent AI-specific security product. Proprietary Cisco platform — opinions captive, no open exit. **DefenseClaw (Open Source Secure Agent Framework)** [DAPM: Retained] Open-source framework that automates security governance for agentic AI. Admission control: scans skills, MCP servers, plugins, and code before they run. Observe mode (log without blocking) and action mode (block HIGH/CRITICAL findings). Integrates with NVIDIA OpenShell as the sandbox runtime. Splunk Observability Cloud dashboards for monitoring. Apache 2.0 licensed. Plans to integrate with NVIDIA OpenShell as the sandbox to eliminate manual steps and accelerate secure agent deployment. **Cisco Isovalent Runtime Security** [DAPM: Ceded] eBPF-based runtime security for containerized AI workloads. Deep kernel-level visibility into container behavior, network flows, and system calls. Part of the Secure AI Factory security stack. Runtime enforcement, not runtime execution — secures the container environment, does not provide the model-serving or agent execution runtime. Proprietary Cisco platform — opinions captive, no open exit. **Splunk AI Agent Monitoring** [DAPM: Ceded] Tracks performance, cost, quality, and behavior of LLM and agentic applications in Splunk Observability Cloud. Visualizes agent workflows. Integrating with Cisco AI Defense for risk mitigation (bias, hallucinations, data leakage, prompt injection). GA February 2026. Galileo acquisition (expected Q4 FY2026) will add real-time guardrails, 20+ evaluation metrics including hallucination detection and context adherence, full ADLC coverage. Proprietary Cisco platform — opinions captive, no open exit. ### NVIDIA-Provided Components **NVIDIA NemoClaw / OpenShell** Agent runtime and sandboxed execution environment. Cisco AI Defense integrates with OpenShell to add security governance. The runtime is NVIDIA's; the security layer is Cisco's. **NVIDIA NIM + AI Enterprise** Containerized model serving and commercial AI platform. Same dependency as Dell and HPE. **NVIDIA Dynamo** Distributed inference framework with KV-aware routing. Performance optimization, not policy-driven placement. **NVIDIA NeMo Guardrails** Runtime safety boundaries integrated into Cisco AI Defense. Cisco extends NeMo Guardrails with its own AI Defense policy enforcement. ### Gap Analysis Cisco does not own the core agent runtime, model-serving runtime, or distributed inference framework — same structural position as Dell and HPE at Layer 2B. The NVIDIA NemoClaw/OpenShell stack provides execution; Cisco provides security governance around it. But Cisco's security contribution at Layer 2B is the most comprehensive of any infrastructure vendor assessed. AI Defense is not a rebranded NVIDIA capability — it's Cisco-developed AI-specific security that integrates with NVIDIA's runtime. DefenseClaw is open-source admission control for agent capabilities. Isovalent provides kernel-level container security. Splunk AI Agent Monitoring provides behavioral observability. Together, these constitute a 'trust layer' for AI execution that no other infrastructure vendor matches. The Galileo acquisition is strategically significant: it extends Splunk from infrastructure observability into AI agent evaluation, covering the full Agent Development Lifecycle (ADLC) — prompt optimization, model selection, production monitoring, and guardrail enforcement. Post-Galileo, Cisco will have the most comprehensive AI observability stack of any infrastructure vendor. The distinction the 4+1 model draws: security constrains what agents CAN'T do. Runtime governs what agents DO. Governance determines what agents SHOULD do. Cisco has the first. NVIDIA has the second. Nobody fully has the third. ### Borrowed Judgment Moderate but inverted from Dell's pattern. Dell borrows runtime judgment from NVIDIA and security judgment from partners (CrowdStrike, Fortanix, F5). Cisco borrows runtime judgment from NVIDIA but retains security judgment entirely — AI Defense, DefenseClaw, Isovalent, and Splunk AI Agent Monitoring are all Cisco IP. The enterprise inherits NVIDIA's runtime decisions but Cisco's security decisions. Post-Galileo, the observability judgment becomes even more Retained: Cisco will own infrastructure observability (Splunk), network observability (ThousandEyes, Nexus One), and AI agent observability (Galileo + AI Agent Monitoring) in a single platform. ### Working Notes Peter Bailey (SVP/GM Cisco Security): 'We have this opportunity to be a trust layer, not just for network activity, but actually what's happening at the application layer, at the workload layer, between agents, between workloads, between data.' This is a Layer 2B/2C strategic claim — Cisco as the trust layer that wraps around other vendors' execution layers. DefenseClaw's open-source model is strategically similar to AGNTCY: Cisco open-sources the agent security framework, building community adoption and standards influence, while retaining the commercial AI Defense product. Open-source for influence, commercial for revenue. ## ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Security + Identity — Not Yet Governance ### Vendor-Provided Components **Duo Agentic Identity** [DAPM: Ceded] Agent identity as first-class non-human identities. Duo Directory registers agents as distinct identity objects mapped to human owners with group-based policy enforcement. Per-action least-privilege enforcement. Lifecycle visibility for agent onboarding and decommissioning. Cisco Identity Intelligence provides continuous inventory of active AI agents — including shadow agents never registered with IdP. Architectural advantage: Cisco spans both identity and network, surfacing agents that communicate across infrastructure even without IdP registration. Proprietary Cisco platform — opinions captive, no open exit. **AGNTCY (Linux Foundation — Infrastructure Layer)** [DAPM: Retained] Open-source infrastructure for multi-agent systems: agent discovery (Open Agent Schema Framework / DNS-like agent directory), agent identity (cryptographic verification across organizational boundaries), agent messaging (SLIM — Secure Low-Latency Interactive Messaging, quantum-safe), agent observability (end-to-end across multi-agent, multi-vendor workflows). Originally open-sourced by Cisco Outshift (March 2025), donated to Linux Foundation (July 2025) with Dell, Google Cloud, Oracle, Red Hat as formative members. 75+ supporting companies. Critical context: AGNTCY is one component within a larger standards convergence, not the defining standard itself. The Agentic AI Foundation (AAIF), formed December 2025, is the umbrella — analogous to CNCF for cloud-native. AAIF's founding projects are MCP (Anthropic), AGENTS.md (OpenAI), and goose (Block). Platinum members are AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft, OpenAI. Cisco is a Gold member of AAIF, not Platinum. A2A (Google) reached v1.0 and joined the broader LF agentic ecosystem. AGNTCY sits as complementary infrastructure beneath these protocols — the plumbing (discovery, identity, messaging, observability) that MCP and A2A need but don't provide themselves. AAIF is already called 'the fastest-growing project in Linux Foundation history' with 190 members by May 2026 — more than double CNCF's membership at the same stage. Microsoft explicitly draws the Kubernetes parallel: 'Just as Kubernetes needed RBAC and admission controllers to be enterprise-ready, agentic systems need governance primitives.' AGNTCY contributes some of those primitives. It does not define the ecosystem. **Cisco AI Defense (Agent Governance Functions)** [DAPM: Ceded] Secures multi-agent systems with controls for agent discovery, behavioral guardrails, and policy enforcement. Integrates with NVIDIA OpenShell for sandbox governance. Combined with Duo Agentic Identity, provides: which agents exist (discovery), who they are (identity), what they can access (authorization), what they do (monitoring), and what they can't do (guardrails). This is the closest thing to an agent governance stack in Cisco's portfolio. Proprietary Cisco platform — opinions captive, no open exit. **Splunk Observability + Galileo (Agent Evaluation)** [DAPM: Ceded] AI Agent Monitoring in Splunk Observability Cloud for production agent behavior tracking. Galileo acquisition adds real-time guardrails, hallucination detection, context adherence, chunk attribution — 20+ evaluation metrics across the full ADLC. The Futurum Group positioned Splunk post-Galileo as 'an AI-era control plane candidate, concentrating network, security, and AI agent behavior telemetry in a single vendor.' Whether telemetry concentration constitutes a control plane is the open question. Proprietary Cisco platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** All Layer 2C components are Cisco IP or Cisco-originated open standards. NVIDIA does not control agent identity, governance policy, or observability in Cisco's stack. This is the same pattern as Google's Layer 2C — the intelligence governance layer is vendor-owned, not NVIDIA-dependent. ### Gap Analysis Applying the 'Routing Is Not Reasoning' test from the 4+1 model: • Duo Agentic Identity = identity management (which agents exist and what they can access) • AI Defense = constraint enforcement (what agents cannot do) • DefenseClaw = admission control (what agent capabilities are approved to run) • AGNTCY = infrastructure primitives (how agents discover, authenticate, and message each other) • Splunk + Galileo = observability and evaluation (what agents are doing and how well) None of these provides policy-driven decisions about where compute runs relative to data, which model serves which request, or how cost/compliance/latency are arbitrated in real time. Cisco's Layer 2C is the most comprehensive SECURITY and OBSERVABILITY story of any vendor assessed — but security and observability are not the same as governance and placement. The AGNTCY positioning requires careful calibration. The K8s analogy applies to the broader Agentic AI Foundation (AAIF) — the Linux Foundation umbrella with MCP (Anthropic), AGENTS.md (OpenAI), goose (Block) as founding projects, and AWS/Google/Microsoft/Anthropic/OpenAI as Platinum members. AAIF is 'the fastest-growing project in Linux Foundation history' with 190 members and more than double CNCF's early-stage membership. AGNTCY is one project within this ecosystem — important infrastructure (discovery, identity, messaging, observability) that MCP and A2A need but don't provide. But Cisco is a Gold member of AAIF, not Platinum. The ecosystem-defining standards (MCP, A2A) were originated by Anthropic and Google respectively, not Cisco. AGNTCY contributes the plumbing beneath the protocols — valuable, but not the protocols themselves. The structural comparison: • Google has a productized Layer 2C control plane (Agent Identity + Gateway + Registry + Orchestration + Observability + Memory Bank) AND originated A2A • IBM has cross-framework agent governance (watsonx Orchestrate) + cross-platform model governance (watsonx.governance) • VAST has PolicyEngine + Polaris for data-plane governance • Cisco has agent identity + agent security + agent observability + open infrastructure primitives (AGNTCY) — but no agent orchestration, no model routing, no placement reasoning Cisco is building the trust layer and contributing to the infrastructure standards layer. The question the 4+1 model poses is whether trust (can this agent be trusted to act?) plus interoperability standards (can agents discover and talk to each other?) is sufficient without governance (should this agent act here, now, with this data, on this model?). Trust and interoperability are necessary for governance. They are not governance. However, Cisco's Layer 2C position is stronger than Dell's (Absent) and arguably stronger than HPE's (Delegated to Kamiwaza). Cisco owns genuine Layer 2C primitives — they just don't compose into a placement engine yet. ### Borrowed Judgment Low for the functions Cisco provides. All Layer 2C components are Cisco IP or Cisco-originated open standards. No NVIDIA dependency, no partner dependency for identity, security, or observability at this layer. But 'low borrowed judgment for partial Layer 2C' is structurally different from 'low borrowed judgment for complete Layer 2C.' VAST has low borrowed judgment for a comprehensive (if captive) Layer 2C. Cisco has low borrowed judgment for agent trust functions — but the placement, routing, and governance functions are Absent, not borrowed. The enterprise architect using Cisco for Layer 2C gets strong agent identity and security with zero borrowed judgment. They do not get agent orchestration, model routing, or policy-driven placement from anyone — Cisco or otherwise. ### Working Notes The Futurum Group's framing of post-Galileo Splunk as 'an AI-era control plane candidate' is the most explicit analyst validation of Cisco's Layer 2C potential. The path from candidate to actual control plane requires composing identity (Duo) + security (AI Defense) + observability (Splunk/Galileo) + networking intelligence (Nexus One AI Job Observability) into a single decision surface that makes placement and governance decisions. The pieces exist. The composition does not. This is the inverse of Dell's position: Dell has no pieces. Cisco has pieces without composition. VAST has composition without openness. Google has composition with captivity. The AAIF ecosystem is the right frame for understanding where agentic governance standards are heading. Microsoft's explicit Kubernetes parallel — 'Just as Kubernetes needed RBAC and admission controllers to be enterprise-ready, agentic systems need governance primitives and those primitives belong in the open' — validates the 4+1 model's Layer 2C thesis. The governance primitives Microsoft references are exactly what Layer 2C defines. AAIF is building the open-standards foundation for Layer 2C; the question is whether any vendor productizes it as a coherent control plane. Cisco's strategic position within AAIF: contributor of infrastructure plumbing (AGNTCY), commercial implementer of trust functions (AI Defense, Duo, Splunk), Gold-tier member. This is a meaningful position — but it's one contributor among 190 members, not the ecosystem definer. Anthropic (MCP) and Google (A2A) contributed the foundational protocols. Cisco contributed the infrastructure layer beneath them. The analogy: Google contributed Kubernetes; Cisco contributed something more like the CNI (Container Network Interface) or service mesh layer — essential infrastructure, but not the orchestration standard itself. The commercial bet: if AAIF standards mature into the enterprise agentic governance stack, Cisco's commercial products (AI Defense for security, Duo for identity, Splunk for observability, DefenseClaw for admission control) become the enterprise implementation layer on top of open standards — the Red Hat model applied to agentic infrastructure. That's a credible business model if the standards achieve adoption. Cisco survey data: 55% of organizations have agentic AI running as pilots or in production (Jan 2026). Only 4% are fully confident in full-scale deployment. 59% cite security concerns as the biggest barrier. Cisco is positioned to address the security barrier. The governance barrier — which the 4+1 model argues is structurally different — is what AAIF and AGNTCY are attempting to address at the standards level. ## ◇ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Partner Ecosystem ### Vendor-Provided Components **Cisco Secure AI Factory Partner Ecosystem** [DAPM: Delegated] Reference architecture that partners build on. VAST Data (storage), NVIDIA (compute/runtime), and a growing ecosystem of security partners integrated into AI Defense. The Secure AI Factory is a validated framework, not a walled garden — partners provide application logic, Cisco provides infrastructure and security. **Splunk Platform (Analytics + Security Applications)** [DAPM: Ceded] Splunk Enterprise Security TDIR platform. Splunk Observability Cloud. These are Layer 3 security applications that run on Cisco's own infrastructure. Splunk's installed base provides a distribution channel for AI capabilities — existing Splunk customers can adopt AI Agent Monitoring without new vendor relationships. **Cisco Security Portfolio (Security Applications)** [DAPM: Ceded] AI Defense, Duo, SecureX, Umbrella, Talos threat intelligence — the security product portfolio that Cisco positions as the 'trust layer' for enterprise AI. Each product addresses a specific security function; together they constitute a security application suite that sits alongside (not above) AI applications from other vendors. Proprietary Cisco platform — opinions captive, no open exit. ### NVIDIA-Provided Components **NVIDIA NIM + NemoClaw / OpenShell Runtime** Execution surface for AI applications on Cisco infrastructure. NVIDIA provides the application runtime; Cisco provides the infrastructure substrate and security layer. ### Gap Analysis Cisco's Layer 3 is structurally different from Dell's, HPE's, or the hyperscalers'. Dell has a broad ISV ecosystem (OpenAI, Palantir, Google, ServiceNow, SpaceXAI). HPE has Unleash AI with 26+ curated ISV partners. AWS and Google have thousands of ISV applications. Cisco's Layer 3 is narrower — the Secure AI Factory is a reference architecture that partners build on, not an application marketplace. Cisco's strongest Layer 3 asset is Splunk — an established platform with deep enterprise penetration that is being extended into AI observability and security. Splunk's competitive advantage at Layer 3 is distribution: enterprises already running Splunk can adopt AI Agent Monitoring, AI Defense integration, and Galileo evaluation capabilities within their existing observability investment. The strategic comparison with Dell at Layer 3: Dell's ecosystem is load-bearing (ISV partners provide infrastructure-level functions the platform lacks). Cisco's ecosystem is enabling (partners build applications on Cisco's infrastructure substrate). Cisco's Layer 3 partnerships are less about filling platform gaps and more about extending the reference architecture to specific use cases. The Cisco 360 Partner Program (launched January 2026) structures partner engagement around the Secure AI Factory, with role-based training paths, dCloud demo environments, and NVIDIA compute training alongside Cisco networking and security training. This is a go-to-market capability, not a technology capability. ### Borrowed Judgment Distributed across partners, which is architecturally appropriate at Layer 3. The structural observation: Cisco's borrowed judgment at Layer 3 is concentrated in two domains — AI runtime (NVIDIA) and data platform (VAST). The security and observability applications are Retained. The enterprise application use cases are partner-provided. The Splunk ecosystem (tens of thousands of enterprise customers, extensive app marketplace, active developer community) provides a distribution advantage for AI capabilities that pure-infrastructure vendors lack. Dell and HPE must sell AI capabilities to new buyer personas. Cisco can extend AI capabilities to existing Splunk and security customers. ### Working Notes The service provider go-to-market is unique to Cisco at Layer 3. Cisco AI Grid enables telcos to offer managed AI services to their enterprise customers using Cisco infrastructure. No other assessed vendor has a comparable service provider distribution channel for AI. If edge inference becomes a significant market, Cisco's telco relationships position it differently from every other vendor in the assessment. The $4,000 employee reduction (5% of ~86,000) alongside record revenue signals active reallocation toward AI infrastructure and security — Cisco is restructuring its workforce to match its strategic pivot, not cutting due to weakness. ════════════════════════════════════════════════════════════════════════════════ # CoreWeave AI Cloud The Open-Substrate Neocloud — Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.4 - Label Vocabulary & Managed-Service Chip Reconciliation **Date:** July 12, 2026 **Source:** CoreWeave Platform and Storage docs, NVIDIA GTC 2026 (HGX B300 GA, Vera Rubin NVL72 H2 2026), SUNK / CKS / Mission Control product pages and docs, CoreWeave AI Object Storage (CAIOS) + LOTA, Distributed File Storage, Dedicated storage tiers (VAST/WEKA/DDN/IBM Spectrum Scale/Pure), CoreWeave Sandboxes, Weights & Biases acquisition (closed 2025) and unified agentic AI launch (May 2026 — Serverless RL, CoreWeave Inference, W&B Weave, W&B Skills), NVIDIA $2B investment (Jan 2026), SEC filings, analyst coverage. v1.1 added Dedicated storage tier + Sandboxes (CoreWeave AR, May 29 2026). v1.2 corrects the 1A authority reasoning: the architectural guarantee is CoreWeave's SLA on a CoreWeave service — the storage substrate (VAST/WEKA/etc.) is an implementation detail invisible to the customer and is removed from the DAPM call, consistent with not scoring a hyperscaler's object store on its drive supplier. v1.4 (July 12, 2026, /reconcile): statusLabels moved to capability vocabulary; CKS+SUNK (2A) and CKS open serving (2B) Retained→Delegated — CoreWeave-operated managed services behind standard interfaces take the EKS-class treatment; Retained is reserved for substrates the customer operates (SUNK Anywhere self-host noted as that form). ## Summary Finding CoreWeave is a pure-play NVIDIA GPU cloud whose entire differentiation lives at three layers — the compute fabric (0), infrastructure orchestration (2A), and application runtime (2B) — and which is, unusually for a cloud in this assessment, built deliberately on open substrate. Its orchestration opinions rest on Kubernetes (CKS) and Slurm (SUNK), not a proprietary managed scheduler, and SUNK Anywhere runs those same workflows on non-CoreWeave and on-prem clusters with few configuration changes. That makes CoreWeave's orchestration layer the rare cloud cell that reads as Retained rather than Ceded. The capture mechanism is decoupled and invisible. The substrate the buyer sees — Kubernetes, Slurm, an S3-compatible object store with no egress fees — is genuinely open and portable, which is reassuring at purchase. The value that actually accumulates sits in two captive places the buyer underprices: the silicon and fabric (wholly NVIDIA, with no x86-OEM substitutability because you rent CoreWeave's fleet rather than buy swappable hardware), and the Weights & Biases opinion layer at the top of the stack — experiment history, evaluation frameworks, and the closed training-to-inference 'Superintelligence Loop' — which is proprietary and does not move when the substrate does. The data layers are where CoreWeave's openness gets its sharpest test, and the reading is substrate-agnostic: the customer contracts with CoreWeave for a CoreWeave service backed by a CoreWeave SLA, so whatever storage platform sits underneath (VAST, WEKA, DDN, and others all appear in the dedicated tier) is an implementation detail and irrelevant to authority — the same way a hyperscaler's object store is not scored on its drive supplier. What matters is the interface and the lift-to-leave. CoreWeave's default data services — AI Object Storage (CAIOS) via an S3-compatible API, and Distributed File Storage via POSIX — sit behind open, portable interfaces: a customer can re-point and leave, so authority is Delegated (a CoreWeave-managed service behind an open interface — the customer could switch without rebuilding; Retained is reserved for an open substrate the enterprise operates itself). The governance surface (catalog, lineage, advanced data services) exists only on the optional Dedicated tier, which CoreWeave's own docs steer most workloads away from; those governed artifacts are a closed proprietary platform with no open exit and are Ceded. So Layer 1A reads moderate — a real governance capability is available and CoreWeave-delivered, but it is optional and the only captive piece, while the default experience stays open. Context/retrieval (1B) and pipelines (1C) have no CoreWeave product at all and remain the enterprise's to provide. The buyer's trade: CoreWeave offers the latest NVIDIA silicon, best-in-class GPU-cluster orchestration on open standards, and a credible agentic development loop via W&B — with a far more favorable orchestration-layer authority profile than any hyperscaler. In exchange it cedes total dependence on NVIDIA at Layer 0 (supplier, strategic partner, and equity holder after the January 2026 $2B investment) and accumulates captive value in the W&B value plane. The instrument's reading: open where it is cheap to be open, captive where the value compounds. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** NVIDIA-Reference Cloud Compute ### Vendor-Provided Components **GPU compute (bare-metal CKS nodes)** [DAPM: Ceded] Single-GPU through 8× NVLink systems and multi-node InfiniBand clusters; HGX B300 with 2.1 TB HBM3e; GB200/GB300 rack-scale; Vera Rubin NVL72 expected H2 2026. Rented capacity, not owned hardware. **Network fabric** [DAPM: Ceded] NVIDIA Quantum-X800 InfiniBand, ConnectX, BlueField DPUs; multi-cloud backbone with private interconnects, direct cloud peering, 400 Gbps-capable ports. No commodity-substrate equivalent. ### NVIDIA-Provided Components **NVIDIA GPU fleet** First-to-market NVIDIA generations: HGX B300 GA, GB200/GB300 rack-scale, Hopper/Ada; among first to deploy Vera Rubin NVL72 in H2 2026. **NVIDIA networking** Quantum-X800 InfiniBand, ConnectX NICs, BlueField DPUs for node/resource offload — the entire fabric is NVIDIA silicon. ### Gap Analysis CoreWeave's Layer 0 is differentiated and complete — the freshest NVIDIA fleet of any cloud, 40+ data centers, ultra-low-latency fiber, BlueField-offloaded bare-metal nodes, and MLPerf-leading utilization. Calibrates with AWS and NVIDIA at this layer: strong capability, zero buyer authority over silicon or fabric. The authority call is harder than it looks. Dell and HPE score Retained at Layer 0 because x86/NVIDIA hardware is substitutable across OEMs — the enterprise owns the box and can swap the vendor. CoreWeave is not that: the enterprise rents CoreWeave's NVIDIA-only fleet, so there is no commodity substrate it controls and no swap that does not mean leaving CoreWeave. The dependence is also uniquely total — NVIDIA is supplier, Elite Cloud Partner, and (after the January 2026 $2B investment) a significant equity holder. There is no alternative-accelerator path here as there is on AWS (Trainium) or Google (TPU). Ceded. ### Working Notes The NVIDIA column is the densest in this row. Unlike the hyperscalers, CoreWeave has no own-silicon hedge — NVIDIA dependence at Layer 0 is the structural fact of the business. ## ◑ Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Open Default Storage + Dedicated Tier ### Vendor-Provided Components **CoreWeave AI Object Storage (CAIOS) + LOTA** [DAPM: Delegated] CoreWeave's own S3-compatible object service; LOTA caches on GPU/CPU nodes for up to 7 GB/s per GPU, cross-region/cross-cloud reach, no egress/request fees, >75% lower cost; versioning, lifecycle, read-after-write consistency, security policies. Open S3 interface — the customer can leave by re-pointing workloads. The storage backend is CoreWeave's implementation detail, not part of the authority call. A managed CoreWeave service behind the multi-vendor S3 standard: object opinions lift to any S3 platform, so Delegated, not Retained (Retained is reserved for an open substrate the enterprise operates). **Distributed File Storage** [DAPM: Delegated] POSIX-compliant shared filesystem for multi-node training synchronization; snapshots, async deletion. A CoreWeave-managed service behind the open POSIX/NFS interface: workloads lift to another POSIX filesystem, so Delegated. Performance, not governance. **Dedicated tier governance services** [DAPM: Ceded] Optional single-tenant dedicated clusters expose a full data-services suite (catalog, database, data-engine, audit logging, cross-cluster replication) — the only path to a real 1A governance surface. CoreWeave-delivered and SLA-backed, but a closed proprietary data platform with no open exit; governed artifacts can't be lifted without rebuilding. Defaulted-away-from for most workloads. ### Gap Analysis This cell is scored on the customer's contractual boundary, not on the storage substrate. The customer contracts with CoreWeave for a CoreWeave service under a CoreWeave SLA; the platform underneath (CAIOS's backend, or the VAST/WEKA/DDN/Spectrum Scale/Pure options in the dedicated tier) is an implementation detail CoreWeave manages and is invisible to the customer — so it does not enter the authority call, the same way a hyperscaler's object store is not scored on who makes its drives. What governs is the interface and the lift-to-leave. Default data services are open at the interface and Delegated. AI Object Storage (CAIOS) is CoreWeave's own service, LOTA-accelerated (up to 7 GB/s per GPU, cross-cloud single-dataset reach, no egress, >75% lower cost), and it speaks an S3-compatible API with versioning, lifecycle, read-after-write consistency, and security policies. Distributed File Storage presents a POSIX shared filesystem for training synchronization. Both sit behind open, portable interfaces — a customer can re-point S3 or POSIX workloads to another provider without rebuilding the data layer. Delegated — a CoreWeave-managed service behind an open interface; the customer could switch without rebuilding, but does not operate the substrate (which would be Retained). A real governance surface exists, but only on the optional Dedicated tier and only as a closed platform. That tier exposes a full data-services suite (catalog, database, data-engine functions, audit logging, cross-cluster replication) — a genuine Layer 1A governance capability that CoreWeave delivers and supports. But CoreWeave's own docs name the default tiers (CAIOS, Distributed File Storage) as the recommended starting point for most AI workloads, so the governance-rich path is the exception, not the norm. And those governed artifacts — the catalog metadata, the database logic, the data-engine pipelines — are a proprietary platform with no open exit: the customer cannot lift them to another vendor without rebuilding. Ceded, on the closed-system litmus, not because a partner supplies the engine. Net: moderate. A governance capability is present and CoreWeave-delivered (so not gap), but it is optional, defaulted-away-from, and not a differentiated CoreWeave opinion set the way AWS Lake Formation is AWS's own (so not strong). Authority splits: Delegated on the open-interface default services, Ceded on the proprietary governance tier. Note the standing decoy: 'S3-compatible, no egress, runs anywhere' describes portability of bytes through an open interface — which is the legitimate basis for the Delegated default here — but says nothing about the separate governance opinions, which are where the capture sits. ### Working Notes v1.2 (May 29 2026): corrected the authority reasoning. v1.0 scored 1A gap on default storage alone; v1.1 moved it to moderate but mis-attributed the authority to VAST as a third party. The customer's guarantee is with CoreWeave, not the substrate vendor, so VAST/WEKA/etc. are removed from the DAPM call. Status stays moderate; the split is Delegated (open S3/POSIX interfaces — CoreWeave-managed services the customer could switch) vs. Ceded (proprietary governance services on the optional dedicated tier). v1.3 (Jun 20 2026): default-storage authority Retained→Delegated under the instrument-wide managed-service rule (a managed service behind a multi-vendor standard interface is Delegated; Retained is reserved for an open substrate the enterprise operates). ## ○ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Enterprise-provided ### Gap Analysis No native vector database, hybrid search, or managed RAG context service. W&B Weave provides tracing and evaluation for agent and LLM behavior, not retrieval. Tensorizer accelerates model loading, not context assembly. The enterprise brings its own retrieval stack (pgvector, Milvus, Weaviate, etc.) and runs it on CoreWeave compute. The same boundary that governs 1A governs here: a capability counts only if CoreWeave exposes it as a service you buy and they SLA. The underlying storage platform on the dedicated tier may have similarity/indexing services, but CoreWeave does not surface a retrieval product, so there is nothing for the customer to consume as a CoreWeave 1B capability. An underlay capability that is not exposed to the customer is not a CoreWeave capability — it is a gap. Calibrates with NVIDIA's 1B gap: CoreWeave accelerates retrieval workloads at Layer 0/2B but does not provide the retrieval capability itself. Function remains the enterprise's — gap, DAPM Retained. ### Working Notes Boundary rule: scored on what CoreWeave exposes and SLAs, not on what the storage substrate can technically do. CoreWeave sells storage tiers, not a retrieval product, so this is a gap regardless of substrate. ## ○ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Fast Transport, No Pipeline Layer ### Gap Analysis LOTA moves bytes fast and across clouds, and AI Object Storage gives a single dataset global reach without replication. But fast data transport is not a data pipeline. There is no ETL/ELT, no transformation framework, no lineage tracking, and no cost-aware movement orchestration product. The enterprise composes this from open tooling (Airflow, Ray Data, dlt, Spark) on CoreWeave compute. Same boundary as 1A and 1B: the dedicated tier's data-engine can do pipeline-style work, but CoreWeave does not expose an ETL/pipeline/lineage product the customer can buy and hold CoreWeave to. An underlay capability CoreWeave does not surface as a service is not a CoreWeave capability. Gap. Data movement ≠ data pipelines. Cross-cloud reach is a Layer 0/1A throughput property, not a Layer 1C pipeline capability. Function remains the enterprise's — gap, DAPM Retained. ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Open-Substrate GPU Orchestration ### Vendor-Provided Components **CKS + SUNK (Slurm on Kubernetes)** [DAPM: Delegated] Managed Kubernetes with Slurm integrated as a K8s scheduler; login/compute/controller nodes as Pods; preemption across Slurm and K8s workloads. Built on CNCF Kubernetes and open-source Slurm. Consumed as a CoreWeave-operated managed service: Delegated per the managed-service-behind-a-standard-interface rule (the EKS/AKS/GKE/OKE treatment) — manifests and Slurm definitions port to any K8s/Slurm environment. SUNK Anywhere extends to non-CoreWeave and on-prem clusters; a customer self-operating that open substrate is the Retained form, scored here as the service actually consumed. **Mission Control** [DAPM: Ceded] Proprietary cluster health, straggler detection, silent-fault mitigation, and node lifecycle management; ~50% fewer interruptions claimed. Operational intelligence layer captive to CoreWeave. ### NVIDIA-Provided Components **Topology awareness on NVIDIA fabric** SUNK topology-aware scheduling tuned to GB200 rack-scale NVLink/InfiniBand layouts; Mission Control detects GPU stragglers and silent hardware faults. ### Gap Analysis This is CoreWeave's signature layer and the most consequential authority call in the row. CKS (CoreWeave Kubernetes Service) plus SUNK (Slurm on Kubernetes) plus Mission Control deliver differentiated, complete GPU-cluster orchestration: unified Slurm batch scheduling and Kubernetes container orchestration on one cluster, topology-aware placement, preemption logic across both schedulers, automated node lifecycle and straggler mitigation, and self-service cluster provisioning. Strong capability — peer to the hyperscalers' managed orchestration and, by several customer accounts, ahead of it for large training clusters. The authority reading is where CoreWeave diverges from every hyperscaler in this assessment. The hyperscalers' managed orchestration is scored present-and-Ceded because the scheduling opinions (Karpenter, proprietary fair-share) are captive. CoreWeave's orchestration opinions rest on open substrate — Kubernetes (CNCF) and Slurm (open-source, SchedMD) — and SUNK Anywhere explicitly runs the same workflows on non-CoreWeave clusters and on-prem 'with very few configuration changes,' confirmed by customers running it across providers. The enterprise can lift its Slurm/K8s scheduling opinions and operate them elsewhere without rebuilding. By the litmus, that is Retained, and it calibrates near the IBM/Red Hat standard. Two proprietary slivers sit inside the otherwise-Retained layer: Mission Control (the health/straggler-detection and lifecycle service) and the SUNK self-service operator/console are CoreWeave IP. A buyer who depends on Mission Control's operational intelligence specifically inherits a Ceded dependency — but the core scheduling abstraction they would carry to another platform is open. ### Working Notes Pressure-tested against the tidy-story risk: it would be neat to call every cloud's 2A Ceded. CoreWeave genuinely breaks that pattern because it ships the open schedulers as the product rather than hiding a proprietary one behind a managed service. The Retained call is evidence-driven (SUNK Anywhere portability), not generous. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Strong — authority splits by what you use ### Vendor-Provided Components **CKS open serving (KServe / KubeFlow) + Tensorizer** [DAPM: Delegated] Open model-serving (KServe/Kubeflow) on CoreWeave-operated managed Kubernetes; Tensorizer accelerates model load from S3. Runtime opinions portable off CoreWeave — Delegated per the managed-service rule; self-operating the same open stack elsewhere is the Retained form. **CoreWeave Inference + W&B Serverless RL** [DAPM: Ceded] Always-on production inference with integrated monitoring; environment-free Serverless RL for agentic post-training (~40% cost reduction, ~1.4× faster claimed). Proprietary runtime. **CoreWeave Sandboxes** [DAPM: Ceded] Productized isolated execution environment for RL, agent, and evaluation workloads, integrated into the inference service. The execution surface the Weave-driven improvement loop writes fine-tuning, eval, and RL changes into. Proprietary. ### NVIDIA-Provided Components **Optimized runtime on NVIDIA** Inference and training tuned to NVIDIA GPUs; integrates NVIDIA stack but does not require NVIDIA AI Enterprise as the only runtime path. ### Gap Analysis Strong, complete runtime capability across the training-to-inference span: CKS natively integrates open serving stacks (KServe, KubeFlow), Tensorizer streams serialized models from S3/HTTPS for fast cold-start (5× faster downloads claimed), and CoreWeave Inference runs continuously-on production workloads with built-in scaling and health monitoring. W&B Serverless RL adds post-training (RL) execution without provisioning, and CoreWeave Sandboxes provide a productized isolated execution environment for agent, RL, and evaluation workloads — the runtime substrate the improvement loop (Layer 2C) drives changes into. Authority is genuinely split and depends on the presented path the buyer chooses — flagged here rather than averaged away: • Open serving (KServe/KubeFlow on CKS) — the runtime opinions are open and portable. Retained. • CoreWeave Inference / W&B Serverless RL / Sandboxes — proprietary CoreWeave/W&B runtime; adopting it means the execution opinions are captive. Ceded where used. Calibrates with NVIDIA 2B (NIM/Dynamo runtime, strong/Ceded) on the proprietary path, but CoreWeave — unlike NVIDIA — also offers a fully open serving path that keeps the layer Retained. Strong capability, mixed authority. ### Working Notes This is the one cell resting partly on inference: the DAPM depends on which runtime a given customer adopts. Scored as presented architecture (both paths shipped), with the split named rather than collapsed. v1.1: added CoreWeave Sandboxes (confirmed by CoreWeave AR, May 29 2026) as the isolated RL/agent/eval execution environment inside the inference service. ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Gap: Improvement Loop, Not a Reasoning Plane ### Gap Analysis CoreWeave's May 2026 'Superintelligence Loop' — Serverless RL + CoreWeave Inference + W&B Weave observability + W&B Skills — is real Intelligence-2C: a closed feedback loop that governs which agents improve, evaluates agent actions against custom signals, surfaces failure modes in multi-agent workflows, and (via Skills + MCP server) turns coding agents into autonomous agent-improvers. It is productized and adopted, not a slideware claim. But routing is not reasoning, and improvement is not placement. This is closed-loop agent operations and continuous improvement, not live Infrastructure-2C: no policy engine answers 'given data residency, cost, latency, and compliance, run this inference on B300 in region X versus region Y at request time.' The physical placement underneath is just SUNK/Kubernetes scheduling. Calibrates with AWS — which has productized Intelligence-2C (AgentCore Policy/Guardrails) and absent Infrastructure-2C — but CoreWeave's Intelligence-2C is narrower: it governs agent improvement and evaluation, not a general action-authorization policy plane like Cedar-based AgentCore Policy. As a single agent-improvement-and-evaluation loop, it sits below the bar for a scored Layer 2C capability: the same line applied to Nutanix's Agent Gateway and VMware's MCP governance, both treated as sub-threshold building blocks at gap, versus the productized multi-component agent-governance/policy planes that earn moderate (AWS AgentCore Policy/Guardrails/Evaluations/Registry; Cisco identity + defense + observability). So this is a gap, not moderate: the enterprise retains policy-driven placement and a general action-authorization plane. The live per-inference placement gap is universal across this assessment, not specific to CoreWeave; it is noted, not penalized as a unique defect. Authority: the W&B improvement/observability loop is proprietary IP — named here as an emerging signal rather than scored; the enterprise retains policy-driven placement, and the underlying scheduling is the Retained SUNK/K8s layer. ### Working Notes Pressure-tested 'routing is not reasoning': the Superintelligence Loop is genuinely agentic but operates on the model/agent-quality axis, not the request-placement axis. Scored as a gap on the same basis as the other infrastructure rows whose single 2C signals (Nutanix Agent Gateway, VMware MCP governance) are sub-threshold. ## ◑ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** W&B developer plane — captive opinion layer ### Vendor-Provided Components **W&B Models + Registry** [DAPM: Ceded] Experiment tracking, model versioning, artifact lineage, lifecycle promotion. The accumulated experiment and lineage history is the captive value. **W&B Weave + Skills** [DAPM: Ceded] Agent/LLM tracing, evaluation framework, production monitoring, Playground; Skills + MCP server for autonomous agent improvement. Proprietary developer value plane. ### Gap Analysis Through Weights & Biases (acquired 2025), CoreWeave owns a genuine, widely-adopted value plane: W&B Models + Registry (experiment tracking, model versioning, lineage, lifecycle promotion), W&B Weave (tracing, evaluation, production monitoring, Playground), and W&B Skills. W&B powers 1,500+ organizations including 30+ foundation-model builders. This is a real Layer 3 opinion the vendor owns and operates, not merely an ISV catalog — so it is not scored 'partner.' It is narrower than the hyperscaler value planes, which is why it is moderate rather than strong. W&B is a developer/MLOps and agent-ops surface, not a breadth of business-workflow applications comparable to AWS Bedrock/Q/Kiro or Google's application stack. There is no business-process application ecosystem here — the value plane is for the people who build AI, not the people who consume it in line-of-business workflows. This is the decoupled capture, and it is the most important reading in the row. The substrate the buyer chose CoreWeave for — open Kubernetes, open Slurm, S3-API storage, no egress — is reassuringly portable. The value that actually compounds — years of experiment history, evaluation frameworks, the improvement loop, the model registry's lineage — accumulates inside W&B, which is proprietary with no open exit. The buyer feels free because the visible layer is open, and underprices the W&B commitment precisely because everything beneath it can move. Ceded. ### Working Notes The W&B capture is the inverse of the hyperscalers' coupled capture: there the data lives visibly in the vendor's namespace; here the data stays open while the opinion layer above it is captive. More reassuring at purchase, which is what makes it worth naming. ════════════════════════════════════════════════════════════════════════════════ # Databricks Data Intelligence Platform Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.0 - Initial Assessment **Date:** June 21, 2026 **Source:** docs.databricks.com, Data + AI Summit 2025 & 2026 announcements, Databricks blog/newsroom (Lakeflow GA, Lakebase GA, Agent Bricks, Unity AI Gateway, Unity Catalog open-source), MLflow 3 release notes, analyst coverage, published 4+1 model ## Summary Finding Databricks is a data + AI lakehouse platform that floats above Layer 0 — it owns no silicon and runs on AWS, Azure, and GCP — and is strong across the entire data and application stack: storage and governance (1A), retrieval (1B), pipelines (1C), runtime (2B), and the value plane (3). It moderates only where it does not own the layer: infrastructure orchestration (2A), where it rents cloud compute and its GPU scheduling is still pre-GA, and the reasoning plane (2C), where it provides governance and multi-agent orchestration but not live placement. Its authority sits in the opinion layer, never the substrate. The capture mechanism is decoupled and invisible, and Databricks is its purest expression in this series. Every openness claim is true: Delta Lake is open source (Linux Foundation), the data sits in the customer's own cloud bucket in open formats (Delta, Iceberg, Parquet), Spark and MLflow are open source, Unity Catalog was donated to open source, and Delta Sharing is an open protocol. But the value the enterprise accumulates lives in the managed services those open pieces sit beneath: managed Unity Catalog governance (lineage, audit, RBAC, masking — none of which the open-source catalog has), the closed-source Photon engine, Lakeflow declarative pipelines, Mosaic AI serving and agents, and the notebooks and workspace. The open formats keep the bytes portable; the opinions — governance policies, pipelines, agents, dashboards — do not lift. The DAPM profile makes this exact. Of 25 scored components, 20 are Ceded, 5 are Delegated, and none are Retained. The five Delegated components are precisely the open-source and open-protocol surfaces (open lakehouse storage, Spark, MLflow, Marketplace via Delta Sharing, and Lakebase's PostgreSQL interface); everything that carries a Databricks opinion is Ceded. The enterprise never operates an open substrate itself within Databricks — even the open formats are read and written through Databricks' managed runtime. This places Databricks in the proprietary-captive cluster with Palantir and VAST, and it shares Palantir's specific mechanism: open storage, captive opinion layer — so the commitment is invisible until you try to leave, and it compounds with every pipeline, governed asset, and agent built. Databricks is the strongest data-plane vendor in the series — Layers 1A, 1B, and 1C are all strong, with 1C (Lakeflow + Spark + Photon) the most mature data-engineering layer assessed. Its Layer 2B runtime is strong and, unlike the on-prem platform vendors (VAST, Nutanix, VMware, Dell — all moderate at 2B), its agent runtime is generally available rather than preview, which calibrates it to the hyperscalers. The ceiling: it owns no Layer 0 (a gap, like Palantir); its orchestration is scoped to its own workloads on rented cloud compute with GPU scheduling pre-GA (2A moderate); and its reasoning plane is governance plus multi-agent orchestration without live placement — routing is not reasoning, and the agent/LLM control-plane gateway is still Beta (2C moderate). The buyer's trade is architectural coherence and a best-in-class data + AI platform in exchange for ceding the entire opinion layer to Databricks under an open-formats banner. You keep your data in open formats in your own cloud — genuinely — and you accumulate governance, pipelines, retrieval, models, agents, and dashboards that are Databricks-shaped and Databricks-bound. Databricks owns the lakehouse and the AI stack on top of it. It does not own the silicon beneath, and it does not yet reason about where inference runs. ## ○ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Not Databricks' Layer (By Design) ### NVIDIA-Provided Components **GPU Compute via Hyperscaler IaaS** ML training and inference run on NVIDIA GPUs, but only as mediated by the hyperscaler's IaaS (AWS/Azure/GCP). There is no structural Databricks-NVIDIA dependency the way the on-prem vendors (Dell, VAST, Nutanix) have — the GPU relationship belongs to the cloud. Serverless GPU ('AI Runtime') is Public Preview, not GA. ### Gap Analysis Layer 0 is not Databricks' layer, by design. Databricks owns no silicon, no networking, and no datacenters — it is a software platform layered entirely on AWS, Azure, and GCP IaaS, split into a control plane (Databricks' own cloud accounts) and a compute plane (classic clusters in the customer's own VPC, or serverless in Databricks' account). The buyer never thinks about Layer 0, and that is the value proposition. The consequence for the 4+1 model is the same as Palantir's: a Databricks adoption decision resolves no Layer 0 authority question. Whatever capture exists at the silicon and fabric layer belongs to the chosen hyperscaler — a different row on this map — not to Databricks. The contrast with the infrastructure vendors is clean: Dell and Cisco are strong here because they own or design hardware; VAST and Nutanix are moderate as software-defined/HCI abstraction layers; Databricks, like Palantir, simply floats above. The one captive compute-adjacent asset, Photon (the closed-source vectorized execution engine), is a query/processing engine rather than silicon or fabric, and is scored at Layer 1C where it does its work — not here. ### Borrowed Judgment Total at Layer 0, and irrelevant to the value proposition by design. Databricks inherits all silicon, networking, and acceleration judgment from the host cloud. The enterprise's Layer 0 authority position is set by its hyperscaler choice (a different vendor's row), and adopting Databricks does not, by itself, resolve it. ### Working Notes Watch-list (Preview, not scored): serverless GPU / 'AI Runtime' (A10/H100) — Public Preview. Classic GPU clusters are long-GA but run on the cloud's GPU instances. Photon is named and scored at Layer 1C, not here. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Lakehouse Strength — Open Formats, Captive Governance ### Vendor-Provided Components **Unity Catalog (Managed Governance)** [DAPM: Ceded] Single governance layer across data and AI assets: column-level lineage, audit logging, RBAC, ABAC, row-filtering and column-masking, data products, and the catalog namespace. The governance opinions are a proprietary managed-service surface and do not lift — the open-source Unity Catalog donation lacks lineage, audit, RBAC, and masking. Proprietary Databricks platform — opinions captive, no open exit. **Open Lakehouse Storage (Delta Lake + Managed Iceberg + UniForm + Delta Sharing)** [DAPM: Delegated] Open-source Delta Lake (Linux Foundation), native managed Iceberg read/write plus an Iceberg REST catalog and UniForm interop, and the Delta Sharing open protocol — with data physically in the customer's own cloud object storage. The consumed interface is a genuine multi-vendor standard: the customer can repoint another engine (Spark, Trino, Snowflake, DuckDB) at their Delta/Iceberg tables, so the data-layer opinions lift. Delegated (a multi-vendor standard consumed through Databricks, not an open substrate the enterprise operates itself). **Lakebase (Managed PostgreSQL OLTP)** [DAPM: Delegated] Serverless PostgreSQL (built on Neon) integrated with the lakehouse — database branching and sync to Delta. GA on AWS; Azure Beta; GCP not yet. The consumed interface is standard PostgreSQL, so the database opinions lift to any Postgres; the branching/sync convenience is the Databricks-specific layer. Delegated. ### NVIDIA-Provided Components **No NVIDIA Layer 1A Dependency** Unity Catalog, Delta Lake, and lineage are Databricks IP or open source. NVIDIA contributes nothing to the governance layer. ### Gap Analysis This is the Lakehouse, and it is Databricks' center of gravity. The buyer gets a unified governed data foundation: Delta Lake tables in their own cloud bucket, Unity Catalog governing every data and AI asset (lineage, audit, access control, ABAC), native managed Iceberg read/write plus an Iceberg REST catalog and UniForm interop, and Delta Sharing for cross-organization sharing. It feels maximally open — your data, open formats, your cloud account, share with anyone, no proprietary storage. This is the textbook decoupled capture, and every openness claim is true and is exactly what hides the commitment. The data is genuinely open — Delta is open source (Linux Foundation), Iceberg is an open standard, Parquet sits in the customer's own S3/ADLS/GCS. But the governance is captive: managed Unity Catalog is where the value accumulates (column-level lineage, audit, RBAC, row-filtering and column-masking, data products, the catalog namespace), and none of that ships in the open-source Unity Catalog donation, which has the APIs but not the lineage, audit, RBAC, masking, or UI. The methodology names this precisely: a metadata-governance catalog layered beyond the open format is Ceded. Data portability is a decoy here; the dependence accumulates in the catalog and compounds with every governed asset, policy, and lineage edge. The calibration is Palantir's Layer 1A almost line for line — open storage Delegated, governance Ceded. Both are software-platform peers with the same decoupled-capture shape, and Databricks is, if anything, more open on the data side (real open-source Delta in the customer's own bucket versus Palantir's Ontology-mediated virtual tables) — which sharpens rather than softens the finding. Peer also to VAST and Dell at strong, but those are coupled (your data lives in their storage) where Databricks and Palantir are decoupled. ### Borrowed Judgment Low for the governance logic — Unity Catalog, lineage, and ABAC are Databricks IP (Ceded to Databricks). The open-format data layer is Delegated (open-source formats in the customer's own cloud). The capture is decoupled: data open, governance captive. ### Working Notes The open-source vs. managed Unity Catalog gap is the capture proof point: OSS UC provides core APIs (Hive-compatible + Iceberg REST) but lacks automatic column-level lineage, audit logging, RBAC/row-filtering/column-masking, and the Catalog Explorer UI — the governance enterprises actually depend on is managed-only. Watch-list (not scored): Lakebase Azure (Beta) and GCP (not yet); external write to managed Delta tables (Public Preview); Governance Hub (Private Preview); RBAC in OSS UC (coming). Managed Iceberg, Iceberg REST, Iceberg v3, and external lineage reached GA at Data + AI Summit 2026. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Lakehouse-Native Governed Retrieval ### Vendor-Provided Components **Mosaic AI Vector Search (Native + Hybrid)** [DAPM: Ceded] GA. Vector and hybrid (reciprocal-rank-fusion) search over Delta-synced indexes; managed or BYO embeddings; Unity Catalog-governed permission-aware retrieval. A proprietary managed retrieval engine — the index, API, and sync pipeline do not lift. Proprietary Databricks platform — opinions captive, no open exit. **Feature & Function Serving (Online Feature Store)** [DAPM: Ceded] Low-latency feature and function lookups for RAG and ML. The Online Feature Store is Lakebase-backed (online tables deprecated). GA path; Lakebase gating applies (AWS GA, Azure Beta, GCP none). Proprietary feature-store opinions — feature tables, point-in-time joins, online serving — are captive even though Lakebase exposes PostgreSQL underneath. Proprietary Databricks platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No Structural NVIDIA Dependency** Embeddings and search run on cloud compute; GPU acceleration is incidental and hyperscaler-mediated. Vector Search and hybrid retrieval are Databricks IP. ### Gap Analysis RAG without a separate vector database. The buyer gets Mosaic AI Vector Search (GA) — index Delta tables directly, governed by Unity Catalog, with hybrid keyword + vector retrieval (GA) and managed-or-BYO embeddings, all serverless. Because the index is synced from governed Delta tables and access flows through Unity Catalog, retrieval is permission-aware by construction. A Feature & Function Serving path (Online Feature Store, Lakebase-backed) covers low-latency feature lookups. The decoupled split repeats one layer up: the source data (Delta) is open, but the vector index and retrieval engine are proprietary and captive. Mosaic AI Vector Search's index format, retrieval API, and sync-from-Delta pipeline do not lift — leaving means rebuilding RAG on Pinecone, Weaviate, or pgvector and re-wiring permission propagation. The score calibrates to VAST and Palantir at Layer 1B, both strong for native, governed, permission-aware retrieval owned by the platform. Databricks matches: GA native vector search, hybrid retrieval, UC-governed permission inheritance, and MLflow 3 GenAI evaluation/tracing for retrieval-quality observability. It sits above Dell, VMware, and Nutanix at 1B (all moderate), whose retrieval is Delegated (Elastic, pgvector) — Databricks' is proprietary, native, and GA. ### Borrowed Judgment Low — Vector Search, hybrid retrieval, and feature serving are Databricks IP, governed by Unity Catalog so retrieval inherits the Layer 1A permission model. Embeddings are optionally Databricks-served or BYO; NVIDIA is not required. ### Working Notes Online tables are deprecated; the replacement Online Feature Store is Lakebase-backed and therefore inherits Lakebase's cloud-gating (AWS GA, Azure Beta, GCP not yet). Mosaic AI Vector Search GA May 2024; hybrid search (RRF) GA Aug 2024. ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Data Engineering Heartland ### Vendor-Provided Components **Lakeflow Declarative Pipelines (DLT)** [DAPM: Ceded] GA. Declarative ETL/ELT with built-in data-quality expectations, incremental processing, and a managed runtime. A proprietary declarative framework — pipelines do not lift; the open-source Spark Declarative Pipelines donation is not yet GA. Proprietary Databricks platform — opinions captive, no open exit. **Lakeflow Connect (Managed Ingestion)** [DAPM: Ceded] GA. Managed ingestion connectors (SQL Server, Salesforce, Workday, ServiceNow GA; SharePoint, PostgreSQL, Oracle in Preview). Proprietary managed connectors — ingestion opinions are captive. Proprietary Databricks platform — opinions captive, no open exit. **Apache Spark + Structured Streaming** [DAPM: Delegated] The open-source distributed processing engine and streaming framework (DBR 17.x = Spark 4.0). The Spark API is open source with real alternatives — jobs lift to any Spark distribution. The Databricks Runtime packaging and Photon acceleration are the captive add-ons; the engine itself is Delegated. **Auto Loader (cloudFiles)** [DAPM: Ceded] Incremental and streaming file ingestion from cloud storage. A proprietary Databricks Runtime feature — not part of open-source Spark, so it does not lift. Proprietary Databricks platform — opinions captive, no open exit. **Photon (Vectorized Execution Engine)** [DAPM: Ceded] Proprietary closed-source C++ engine delivering 3-10x over JVM Spark SQL; the execution engine beneath pipelines and Databricks SQL. Off Databricks the customer falls back to open-source Spark on the JVM with no Photon. Proprietary Databricks platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 1C Dependency** Pipelines run on CPU compute; Photon is a CPU-vectorized engine. No structural NVIDIA dependency at the pipeline layer. ### Gap Analysis This is Databricks' birthplace and arguably its single strongest layer. The buyer gets Lakeflow, the unified pipeline platform, all GA since June 2025: Lakeflow Connect (managed ingestion connectors), Lakeflow Declarative Pipelines (formerly Delta Live Tables — declarative ETL/ELT with built-in data-quality expectations and incremental processing), and Lakeflow Jobs (orchestration). Plus Auto Loader for incremental file ingestion, Apache Spark and Structured Streaming as the engine, Photon accelerating it, dbt integration, and Unity Catalog column-level lineage threaded through all of it. For data engineering, this is best-in-class. The decoupled split is sharp here. Spark is open source and portable — you can run it anywhere. But the value-bearing pipeline opinions are captive: Lakeflow Declarative Pipelines is a proprietary declarative framework and managed runtime (its internals are being donated to Apache Spark as Spark Declarative Pipelines, but that is not yet GA in open-source Spark — the product you run today is proprietary); Auto Loader is a proprietary Databricks Runtime feature, not in open-source Spark; Photon is a closed-source C++ engine; and Lakeflow Connect connectors and Unity Catalog lineage are proprietary. You can lift raw Spark jobs, but your declarative pipelines, ingestion connectors, data-quality rules, incremental logic, and lineage graph do not lift without rebuilding. The score is peer to VAST at strong (DataEngine) and above Dell at moderate (Dataloop). Databricks essentially defines this category — the most mature data-pipeline platform in the instrument, and the most unambiguous strong in the row. ### Borrowed Judgment Low — Lakeflow, Auto Loader, Photon, and Unity Catalog lineage are Databricks IP. Spark is the open-source substrate (Delegated), but the Databricks Runtime packaging, Photon acceleration, and declarative-pipeline layer are proprietary. Decoupled pattern: open Spark API, captive pipeline and runtime opinions. ### Working Notes Lakeflow (Connect + Declarative Pipelines + Jobs) reached GA June 2025. Watch-list (not scored): Spark Declarative Pipelines (the DLT core donated to Apache Spark) is not yet GA in OSS Spark; Lakeflow Connect connectors vary by GA (SQL Server, Salesforce, Workday, ServiceNow GA; SharePoint, PostgreSQL, Oracle in Preview). DBR 17.x ships Spark 4.0. ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Managed Workload Orchestration; GPU Scheduling Pre-GA ### Vendor-Provided Components **Databricks Compute (Classic + Serverless Cluster Management)** [DAPM: Ceded] Autoscaling clusters in the customer's VPC (classic) or in Databricks' account (serverless, GA on all clouds); proprietary provisioning and lifecycle management on rented cloud VMs. Cluster and serverless orchestration opinions do not lift. Proprietary Databricks platform — opinions captive, no open exit. **Lakeflow Jobs (Workflow Orchestration)** [DAPM: Ceded] GA. Multi-task workflow DAGs with scheduling, retries, and dependencies across notebooks, pipelines, and SQL. Job and workflow definitions are Databricks-specific. Proprietary Databricks platform — opinions captive, no open exit. **Asset Bundles (Declarative Automation Bundles)** [DAPM: Ceded] GA. Infrastructure-as-code (YAML) for jobs, pipelines, notebooks, dashboards, and serving endpoints. The CLI is open-sourced, but the bundle opinions deploy to Databricks — the deployment target does not lift. Proprietary Databricks platform — opinions captive, no open exit. ### NVIDIA-Provided Components **GPU Scheduling Is Hyperscaler-Mediated** GPU scheduling depends on the underlying cloud; Databricks' own serverless GPU ('AI Runtime') is Public Preview, not GA. No structural NVIDIA orchestration dependency. ### Gap Analysis Fully managed compute orchestration the buyer never operates. They spin up autoscaling clusters (into their own VPC, or fully serverless — GA across all three clouds), run Lakeflow Jobs/Workflows with complex dependency DAGs, and codify it as infrastructure-as-code via Asset Bundles (now Declarative Automation Bundles). No VM lifecycle, no scheduler to run — Databricks handles provisioning, autoscaling, and job orchestration on top of the cloud. But this is proprietary orchestration on rented cloud IaaS. The cluster manager, Jobs scheduler, serverless layer, and bundle definitions are Databricks IP — job graphs, cluster configs, and orchestration opinions do not lift; leaving means rebuilding on the cloud's native orchestration or Kubernetes. And the AI-relevant part of this layer, GPU scheduling and fair-share, Databricks does not own: its serverless GPU is Public Preview, not GA, so dedicated GPU scheduling is still the hyperscaler's. The score calibrates to VAST at 2A (moderate): VAST owns infrastructure orchestration (Polaris) but GPU scheduling is partial; Databricks owns workload orchestration (Jobs, clusters, serverless) but GPU scheduling is pre-GA and it runs on the cloud's compute. It sits below VMware and Nutanix at 2A (strong), which are mature general-purpose orchestration platforms managing all infrastructure — Databricks orchestrates only its own data and AI workloads. This is where the data-plane strong streak correctly stops: 2A is not Databricks' to own. ### Borrowed Judgment Moderate. Databricks owns the workload orchestration (clusters, Jobs, serverless, Asset Bundles — Databricks IP, Ceded to Databricks), but the underlying compute capacity and GPU fair-share scheduling are the hyperscaler's — Databricks provisions cloud VMs and its own GPU scheduler is pre-GA. Orchestration authority is Databricks'; capacity and GPU-scheduling authority are the cloud's. ### Working Notes Watch-list (Preview, not scored): serverless GPU / 'AI Runtime'. Lakeflow Jobs (formerly Workflows) and serverless compute on all three clouds are GA. Asset Bundles GA since 2024; CLI is open-sourced but bundle definitions deploy to Databricks. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Mosaic AI Runtime — Serving + Agents (GA) ### Vendor-Provided Components **Mosaic AI Model Serving + Foundation Model APIs** [DAPM: Ceded] GA. Real-time and batch model serving; unified OpenAI-compatible endpoints for hosted (Llama, Gemma, GPT-OSS) and external (Claude, GPT, Gemini) models. Serves open and external models, but the serving infrastructure and unified gateway are proprietary managed services that do not lift. Proprietary Databricks platform — opinions captive, no open exit. **Mosaic AI Agent Framework** [DAPM: Ceded] GA (~Mar 2025). Author and deploy agents with tool-calling and evaluation, governed by Unity Catalog. Agent authoring and runtime opinions are Databricks-specific. Proprietary Databricks platform — opinions captive, no open exit. **Agent Bricks (Knowledge Assistant / Document Intelligence / Supervisor Agent)** [DAPM: Ceded] GA (Jan-Feb 2026). Turnkey agents plus multi-agent orchestration (Supervisor Agent). Proprietary turnkey agents — captive. Proprietary Databricks platform — opinions captive, no open exit. **Mosaic AI Model Training & Fine-Tuning** [DAPM: Ceded] GA. Managed fine-tuning and pretraining on H100; trained weights registered to Unity Catalog. Proprietary managed training stack. Proprietary Databricks platform — opinions captive, no open exit. **MLflow (Open Source)** [DAPM: Delegated] Open-source (Apache 2.0) experiment tracking, model registry, and GenAI tracing (MLflow 3). Core MLflow APIs are open and portable; the managed-MLflow governance and monitoring are the captive add-on. The open core is Delegated. ### NVIDIA-Provided Components **No Structural NVIDIA Runtime Dependency** Serving and training run on cloud NVIDIA GPUs (e.g. managed H100), but there is no NIM/NemoClaw-equivalent software dependency. Unlike Dell — where NVIDIA owns the 2B runtime — Databricks owns its runtime and NVIDIA is just the underlying GPU. ### Gap Analysis A complete, GA model-serving and agent runtime. The buyer serves any model — their own, open source (Llama, Gemma, GPT-OSS), or external (Claude, GPT, Gemini) — behind unified OpenAI-compatible endpoints via Mosaic AI Model Serving and Foundation Model APIs (both GA); builds agents with the Mosaic AI Agent Framework (GA); deploys turnkey agents via Agent Bricks (Knowledge Assistant, Document Intelligence, and the multi-agent Supervisor Agent — all GA); fine-tunes and trains with Mosaic AI Model Training (GA, managed H100); and tracks everything in MLflow (including MLflow 3 GenAI tracing). All governed by Unity Catalog. The decoupled pattern reaches the execution layer: open models, captive runtime. Model choice is genuinely open (any OSS or external model), and MLflow is open source (experiment and model tracking lift). But the serving infrastructure, the Agent Framework, the Agent Bricks products, and Mosaic training are proprietary Databricks managed services — agents and endpoints built here do not lift. The score calibrates to the hyperscalers at 2B (strong — comprehensive managed runtimes), not to the on-prem platform vendors. The deciding factor under the GA-gate: Databricks' agent runtime is GA (Agent Framework plus several Agent Bricks), whereas Nutanix's agents are preview, Dell's runtime is NVIDIA-owned, VAST's AgentEngine is newer and less proven, and VMware's is foundational. Databricks owns a GA-complete serving, agent, and training stack — the strong bar the clouds set. ### Borrowed Judgment Low. Serving, Agent Framework, Agent Bricks, and Mosaic training are Databricks IP (Ceded to Databricks); MLflow is open source (Delegated). Models are open or external (portable choice); the runtime is captive — and unlike the on-prem peers, the runtime is Databricks' own (not NVIDIA's) and is GA, not preview. ### Working Notes Watch-list (Beta/Preview, not scored): Agent Bricks Data + AI Summit 2026 additions (Agent Memory, Sandbox, MCP Catalog, Contextual Policies), Omnigent; serverless-GPU training path (Preview). The old Foundation Model Fine-tuning is deprecated (removal Aug 2026), folded into Mosaic AI Model Training; DBRX dropped from the Foundation Model APIs list. ## ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Agent Governance via Unity Catalog — Not Placement Reasoning ### Vendor-Provided Components **Unity Catalog Agent Governance (On-Behalf-Of Auth)** [DAPM: Ceded] GA. Agents, tools, and models governed as Unity Catalog assets; every tool and data call is permission-checked against the requesting user's UC entitlements — the agent is governed like the user. Proprietary Databricks platform — opinions captive, no open exit. **Mosaic AI Gateway (Serving Endpoints)** [DAPM: Ceded] GA. Rate-limiting, usage tracking, payload and inference-table audit logging, fallbacks, and traffic-splitting in front of serving endpoints. This is static routing configuration, not per-request placement reasoning. Proprietary Databricks platform — opinions captive, no open exit. **Agent Bricks Supervisor Agent (Multi-Agent Orchestration)** [DAPM: Ceded] GA (Feb 2026). Coordinates Genie spaces, Knowledge Assistant agents, MCP servers, Unity Catalog functions, and custom agents under UC governance. Multi-agent orchestration, not infrastructure placement. Proprietary Databricks platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** Agent governance and orchestration are Databricks IP. NVIDIA controls no identity, governance, or routing here — the same pattern as the clouds' 2C. ### Gap Analysis Databricks governs agents with the same authority it governs data. Unity Catalog agent governance (GA) registers agents, tools, and models as governed assets, and On-Behalf-Of auth validates every tool and data call against the requesting user's Unity Catalog permissions — so an agent is governed exactly like the employee it acts for. The Mosaic AI Gateway for serving endpoints (GA) adds rate-limiting, usage tracking, payload and inference-table audit logging, fallbacks, and traffic-splitting. And the Agent Bricks Supervisor Agent (GA, Feb 2026) coordinates multiple agents (Genie, Knowledge Assistants, MCP servers, Unity Catalog functions, custom agents) under Unity Catalog governance. Applying the 'Routing Is Not Reasoning' test: what ships GA is Intelligence-2C (governance of which agent may act, on what, under which policy) plus multi-agent orchestration. What does not ship is Infrastructure-2C — live, per-request placement reasoning. The AI Gateway's routing is static configuration (fallback order and traffic-split weights), not per-request cost/quality/compliance arbitration. The more control-plane-like Unity AI Gateway for agents and LLMs (service policies, spend caps, smart-routing) is Beta; AI Guardrails (safety and PII) is Preview; and the 'smart routing by task complexity, quality, and cost' language is announcement-only. There is no GA engine that reasons per-request and places inference. The score lands in the moderate cohort — AWS (AgentCore Policy/Guardrails), IBM (watsonx.governance + Bob), OCI, Cisco, HPE — all of which have productized, multi-component Intelligence-2C governance without live placement. Governance here is unusually deep because it inherits the full Unity Catalog authority (On-Behalf-Of), the same structural-governance property that earned Palantir and VAST credit. It sits above the Nutanix/VMware/Dell/VAST gap cohort because it adds GA multi-agent orchestration and UC-native agent identity rather than a lone gateway, and below Azure/Google/Palantir (strong) because the comprehensive control-plane gateway is still Beta and there is no Infrastructure-2C placement reasoning. ### Borrowed Judgment Low for what is provided — Unity Catalog agent governance, the AI Gateway, and the Supervisor Agent are Databricks IP (Ceded to Databricks), with no NVIDIA dependency. But this is low borrowed judgment for partial 2C: Intelligence-2C (governance and orchestration) is productized and GA; Infrastructure-2C (live placement reasoning) is Absent, not borrowed. The live per-inference placement gap is universal across the instrument, noted rather than penalized. ### Working Notes Watch-list (Beta/Preview, not scored): Unity AI Gateway for agents and LLMs (service policies, spend caps, smart-routing) — Beta; AI Guardrails (safety/PII) — Preview; 'smart routing by task complexity/quality/cost' — announcement-only, mechanism unspecified. ## ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** First-Party Data Apps + Marketplace Ecosystem ### Vendor-Provided Components **AI/BI (Genie + Dashboards)** [DAPM: Ceded] GA. Natural-language analytics over governed data (Genie) and the default BI dashboards layer. Proprietary first-party applications on Unity Catalog and the lakehouse — they do not lift. Proprietary Databricks platform — opinions captive, no open exit. **Databricks Apps** [DAPM: Ceded] GA. Serverless hosting for custom data apps (Streamlit, Dash, Flask, React) in the workspace, with Unity Catalog-governed data access. The app code uses open-source frameworks (portable), but the hosting runtime and UC integration are captive. Proprietary Databricks platform — opinions captive, no open exit. **Databricks One** [DAPM: Ceded] Workspace-GA. A business-user experience surface over the platform's apps and data. Proprietary UX layer. Proprietary Databricks platform — opinions captive, no open exit. **Databricks Marketplace (Delta Sharing)** [DAPM: Delegated] GA. Datasets, models, notebooks, apps, and MCP servers, built on the open Delta Sharing protocol. An ecosystem of substitutable third-party providers over an open sharing protocol — Delegated. ### NVIDIA-Provided Components **No NVIDIA Layer 3 Dependency** The value-plane applications are Databricks IP or ecosystem-provided. NVIDIA contributes nothing here. ### Gap Analysis Unlike the platform-enabler vendors, Databricks ships genuine first-party value-plane applications, not just tooling. The buyer gets AI/BI Genie (GA — ask questions of governed lakehouse data in natural language), AI/BI Dashboards (GA — the default BI layer), Databricks Apps (GA — build and host custom data apps serverlessly in the workspace), Databricks One (workspace-GA — a business-user experience), Databricks Assistant (GA), and the Databricks Marketplace (GA — datasets, models, notebooks, apps, and MCP servers, built on Delta Sharing). Real business capability, GA today. The first-party apps (Genie, AI/BI, One, Apps) are proprietary and built on Unity Catalog and the lakehouse — they do not lift; the value plane is captive even though the data beneath is open. The Marketplace is the open exception (Delta Sharing protocol, substitutable third-party providers). The value plane is also domain-bounded — data, analytics, and BI-centric (Genie, AI/BI) rather than the broad operational business applications of Palantir's Foundry or the hyperscalers' breadth. The score is above VMware and Nutanix at Layer 3 (moderate — platform-enabled, not platform-provided): those provide tools to build plus an emerging ISV ecosystem but do not ship the apps, whereas Genie is an application, not a platform service. It is not partner (Dell, Cisco, HPE, IBM), because Databricks' Layer 3 is not addressed entirely through an ISV ecosystem. Closest to Palantir and the clouds at strong (first-party apps), with the honest caveat that Databricks' value plane is analytics and data-app-centric. ### Borrowed Judgment Mixed, weighted to Databricks. The first-party apps (Genie, AI/BI, One, Apps) are Databricks IP (Ceded to Databricks) — the value plane is platform-provided, not merely enabled, which separates Databricks from the platform-enabler moderates. The Marketplace ecosystem is Delegated (third-party providers over the open Delta Sharing protocol). ### Working Notes Watch-list (Preview/announced, not scored): Genie Deep Research (Preview); Genie One / Agents / Ontology (announced); Databricks One account-level (Beta); Databricks Assistant Agent mode (Preview). Genie and Apps reached GA June 2025; AI/BI Dashboards GA June 2024; Marketplace GA June 2023; Databricks One workspace-GA Mar 2026. ════════════════════════════════════════════════════════════════════════════════ # Dell AI Factory with NVIDIA Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v2.6 - Codified-Criteria Audit & Watch-List Roll **Date:** July 12, 2026 **Source:** DTW 2026, GTC 2026, Dell press releases (March 16, May 18, June 22, 2026), Dell InfoHub documentation, Dell AI infrastructure product/marketing briefing (July 2026), published 4+1 model. v2.5 (briefing + DAPM axis reconciliation): DAPM re-reasoned from the customer's seat throughout — the rating measures the enterprise-to-vendor authority relationship; vendor-to-vendor framing (Dell-cedes-to-NVIDIA) retired to the nvidia[] dependency column and borrowedJudgment narrative. 2A Dell CSI Operator Ceded→Delegated (consumed interface is the multi-vendor CSI/K8s standard). GA-gate descores to dated watch-lists: 1A Exascale (targeted early 2H CY26, unconfirmed), 1C KV cache offload (NVIDIA CMX pre-GA), 2B Deskside Agentic AI (NemoClaw/OpenShell alpha). 1B GPU-accelerated DSE watch-listed: upstream Elasticsearch 9.4 took cuVS to GA April 2026, but upstream GA never moves an OEM cell — Dell-side GA unconfirmed in Dell docs. statusLabel convention ratified: capability vocabulary only; authority lives in the per-component DAPM chips. Summary rewritten: capture mechanism named (decoupled/invisible); 2C gap upgraded from instrument finding to vendor-confirmed deliberate ecosystem strategy. No layer status changed. v2.6 (July 12, 2026, codified-criteria audit): the canonical row re-read under the methodology's new capability grading criteria and channel-substitution rule. Trust3 Delegated→Ceded (proprietary governance platform; policies captive to the partner through any paper). 2B blueprint component re-reasoned to menu-altitude Delegated (deployed platforms capture per-choice, to the ISV). 1B held at moderate with the frontier check written in (rule 6: DSE is general but below the integrated-retrieval frontier, GPU path unconfirmed). Watch-lists rolled: Exascale still 2H 2026 (now expanding to 4-in-1 with PowerFlex, 1H 2027); GPU-accelerated DSE flagged as slipped — Dell's '1H 2026' passed without doc confirmation. No layer status changed. ## Summary Finding Dell has one of the most credible on-prem AI Factory infrastructure stacks in the market. Its credibility comes from physical infrastructure (Layer 0), storage and data lifecycle integration (Layers 1A/1B/1C, where the Dataloop acquisition gives Dell its first proprietary software asset in the data lifecycle), and ecosystem packaging (Layer 3). From the buyer's seat, the authority profile is just as legible: the enterprise retains commodity compute, cedes its storage and pipeline opinions to Dell, and inherits NVIDIA's judgment for everything GPU-aware above the rack. Dell's capture is decoupled and therefore mostly invisible. The data and table formats are genuinely open: Iceberg, Trino, S3-interface object storage. Schemas lift and catalogs swap. That openness is true and reassuring, and it's the layer Dell's own teams point to when they push back on authority scores. The captive layers sit above the data: the MetadataIQ catalog structure, Data Orchestration Engine pipeline logic, PowerRack integration, and Spectrum-X fabric opinions. The buyer feels portable because the bytes are portable. The opinions aren't. Dell's owned capability above the rack is thin by design: validated ISV blueprints and human services at Layer 2B, chassis-level management at Layer 2A. GPU scheduling judgment is NVIDIA's regardless of whose paper the license rides on. The deskside agentic runtime is pre-GA and watch-listed, not scored. The reasoning-plane gap is no longer just this instrument's finding; it's Dell's stated strategy. In a July 2026 briefing, Dell confirmed both halves of the boundary: the telemetry exists from Layer 0 up, and the decision layer is deliberately left to the ecosystem. 'I can give you the telemetry, but the ultimate decision comes down to somebody making the decision.' Security is not governance. Security constrains who can access the platform. Governance constrains what the platform does. The Dell AI Factory has security. It does not yet have governance at the infrastructure level. That does not make the AI Factory weak. It exposes where the next control-plane battle will be fought. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Dell Strength ### Vendor-Provided Components **PowerRack** [DAPM: Ceded] Turnkey rack-scale: compute, networking, storage integrated with thermal/power management as one unit. Integrated-system rule: the buyer's rack-integration opinions cannot be lifted to another vendor as deployed. **PowerEdge XE9812 (Vera Rubin NVL72)** [DAPM: Retained] 10x lower cost-per-token than Blackwell for agentic inference. Commodity-substitutable server: NVL72 systems ship from multiple OEMs, so workloads move without rebuilding. **Pro Max GB300 (Deskside)** [DAPM: Retained] 120B–1T parameter models. MaxCool liquid cooling. ~3 month break-even vs cloud. **PowerSwitch SN6000-series** [DAPM: Ceded] NVIDIA Spectrum-6 Ethernet. 800+ Tb/sec east-west. Networking opinions are embedded in the Dell/Spectrum-X stack with no abstraction layer offered that would make them portable to alternative switching infrastructure such as Juniper or Aruba. **PowerCool CDU C7000** [DAPM: Retained] First rack-mount CDU for Vera Rubin NVL72 density. 4U, 19", up to 40°C facility water. Physical plant accumulates no portable opinions — no lock-in surface. ### NVIDIA-Provided Components **GPU/Accelerator Silicon** Blackwell, Vera Rubin — the compute engines Dell builds around. **NVLink / NVSwitch** Intra-node high-bandwidth interconnect defining memory and compute topology. **Spectrum Ethernet Silicon** Dell brands and rack-integrates NVIDIA switching silicon. ### Gap Analysis Dell retains platform packaging authority at Layer 0, but the accelerator fabric and high-performance AI networking roadmap are structurally tied to NVIDIA. Dell provides genuine engineering differentiators in thermal design, rack integration, and mechanical authority. The networking silicon dependency is worth tracking. ### Borrowed Judgment Structural co-dependency: Dell retains mechanical authority, NVIDIA retains silicon authority. If NVIDIA changes the Spectrum roadmap, Dell's PowerRack networking story changes with it. From the buyer's seat, the ratings are unaffected by which vendor holds the captive layer: PowerRack and PowerSwitch opinions can't leave regardless of whose silicon sits under the Dell brand. ### Working Notes AMD alternative exists under 'Dell AI Platform with AMD' (separate SKUs). MI350P PCIe, air-cooled, ROCm/vLLM stack. Different Layer 2B story entirely. Watch-list (ISC, June 22, 2026, not scored): PowerEdge XE8812 — Vera Rubin NVL4, up to 144 GPUs per rack, fanless 100% direct liquid cooling, ORv3-style rack, 300kW+, Quantum-X800 InfiniBand. 'Globally available early next year' (early 2027). Scores as a component at GA. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Dell Strength ### Vendor-Provided Components **PowerScale (File Engine)** [DAPM: Ceded] MetadataIQ integration. NeMo Retriever connector. pNFS 25% throughput improvement. PowerScale storage opinions (configs, tiering policies, performance tuning) are captive to the Dell platform — the enterprise cannot lift them to Alletra or VAST without rebuilding. **ObjectScale (Object Storage)** [DAPM: Delegated] S3-compatible. S3 over RDMA. NVIDIA Omniverse integration. Palantir Ontology deploys here. S3-compatible interface; object-storage opinions (buckets, lifecycle, policies) lift to any S3 platform. ObjectScale is a proprietary implementation behind the multi-vendor S3 standard — the consumed interface keeps object opinions portable, so Delegated (not an open substrate the enterprise operates, which would be Retained). The captive surface is MetadataIQ, not ObjectScale itself. **MetadataIQ** [DAPM: Ceded] Indexes billions of files across PowerScale/ObjectScale. Foundation of the governance catalog. Metadata indexing opinions and catalog structure are captive to Dell. **Trust3 AI Integration** [DAPM: Ceded] Storage-layer governance: sensitive data discovery, 'write once, apply everywhere' policy, AI auditing. EU AI Act/GDPR/HIPAA. Proprietary single-vendor governance platform on Dell paper — the 'write once, apply everywhere' policies are authored in Trust3's engine and do not lift to another governance ISV without rebuilding. Channel-substitution rule: Ceded to the chosen partner through any paper. **Cyber Resilience (Built-in)** [DAPM: Ceded] Zero Trust, encryption, RBAC, immutable snapshots, XDR, data masking, air-gapped backup. Security implementation captive to Dell storage platforms — policies and configurations do not lift to another vendor's storage. ### NVIDIA-Provided Components **cuVS (Vector Search)** GPU-accelerated vector indexing (12x claimed). The GPU-accelerated path is not yet GA on Dell paper (watch-listed at 1B) — dependency finding, not shipping capability. **CX-8/CX-9 SuperNICs** Storage-side RDMA for GPU-direct access. **NeMo Retriever Connector** PowerScale integration for GPU-accelerated retrieval. ### Gap Analysis Dell's strongest layer after Layer 0. Trust3 AI provides agentic-AI-aware governance. The Exascale 3-in-1 architecture signals where the platform is heading on data locality, but it is watch-listed pending GA confirmation and not scored. The strategic question: is MetadataIQ metadata rich enough to drive Layer 2C placement decisions? Dell's marketing says yes. The proof is whether any Layer 2C can query it programmatically. The July 2026 briefing confirmed the telemetry and access-control surfaces exist; the reasoning consumer does not. ### Borrowed Judgment Low. Dell owns the storage platforms, metadata layer, and cyber resilience stack. NVIDIA provides acceleration, not governance logic. From the buyer's seat, the captive value is MetadataIQ: catalog structure and indexing opinions cannot leave the Dell platform. ObjectScale's S3 interface keeps object opinions portable to any S3 platform (Delegated — a proprietary implementation behind a multi-vendor standard, not an open substrate the enterprise operates); Trust3's governance policies are the partner's captive surface — Ceded to Trust3, not to Dell. ### Working Notes Data Analytics Engine Agentic Layer + MCP Server (Feb 2026) blur 1A/1B boundary — search, analytics, and orchestration surfaced as a single queryable service. Briefing (July 2026): AIDP access control is agent-aware — a user, an agent, or an agent acting on behalf of a user, with governance on where data lives and is processed through the pipeline; recorded as confirmation of this cell's framing. The briefing also named agent memory (store/retrieve of distilled agent memory) — not documented as GA anywhere the instrument can cite; unscored, fact question pending in Dell's written feedback. Watch-list (announced March 16, 2026; rolled July 12, 2026, not scored): Dell Exascale Storage (3-in-1: PowerScale + ObjectScale + Lightning FS on one platform, 10+ PB/rack, 6 TB/s reads) — current Dell materials still say 'available 2H 2026'; not doc-confirmed. Lightning FS shipped April 2026. New at the May 2026 announcements: Exascale expands to 4-in-1 with PowerFlex block (unified PowerRack architecture), itself dated 1H 2027. Scores when Dell documentation confirms GA. ## ◑ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Elastic-Powered Retrieval ### Vendor-Provided Components **Dell Data Search Engine (Elastic)** [DAPM: Delegated] Elasticsearch 9.4. Hybrid keyword+vector search. MetadataIQ integration. LangChain support. Available now on Dell paper; the GPU-accelerated path is watch-listed pending Dell doc confirmation. Elasticsearch opinions (indices, queries, pipelines) lift to any Elastic deployment. **Incremental Indexing** [DAPM: Delegated] Only updated files re-indexed. Keeps retrieval synchronized with governance catalog. **Analytics Engine Agentic Layer + MCP Server** [DAPM: Delegated] Unifies vector stores across Iceberg, Data Search Engine, PostgreSQL+PGVector. Agent-queryable through the open MCP interface. ### NVIDIA-Provided Components **NVIDIA cuVS** GPU-accelerated hybrid search (12x indexing claimed). Upstream Elasticsearch 9.4 took cuVS to GA April 2026; not yet GA on Dell paper — watch-listed, dependency finding. **NVIDIA STX Architecture** BlueField-4 + ConnectX-9 + Spectrum-X + DOCA. Storage-side acceleration — available to ALL storage vendors. **NeMo Retriever** PowerScale connector for GPU-accelerated retrieval. ### Gap Analysis Three-party dependency: Dell (storage + metadata), Elastic (search intelligence), NVIDIA (acceleration). STX is non-differentiating — every storage vendor has it. Dell's differentiation is MetadataIQ integration and the Elastic partnership, not NVIDIA acceleration. Frontier check (capability rule 6, July 2026): the 1B frontier is now integrated AI-retrieval platforms delivered on vendor paper — VAST InsightEngine natively and through the Cisco and Supermicro channels (strong). Dell's DSE is a general-purpose search engine (rule 4 favors its generality — Elasticsearch runs retrieval workloads Dell never anticipated) that the customer assembles into RAG, with the GPU-accelerated path still unconfirmed on Dell paper. Real, deployable, general, but below the frontier and gated — moderate holds on rules 5 and 6. ### Borrowed Judgment Moderate, distributed across two partners. Search intelligence is Elastic's. Acceleration is NVIDIA's. Dell's durable value is the data substrate — if you swap search engines, PowerScale data doesn't move. From the buyer's seat the retrieval opinions themselves (Elasticsearch indices, queries, pipelines) lift to any Elastic deployment, which is what keeps the layer Delegated rather than Ceded. ### Working Notes No retrieval quality observability (recall@k, latency percentiles) that a Layer 2C could use for placement decisions — the July 2026 briefing's telemetry discussion was token/usage metrics, not retrieval quality. Watch-list (rolled July 12, 2026, flag sharpened — not scored): GPU-accelerated vector indexing for the Data Search Engine — Dell's own materials dated it 'available 1H 2026'; 1H 2026 has ended, and the Q2 2026 AI Data Platform administrators guide (April 13, 2026) does not confirm it. Upstream Elasticsearch cuVS went GA April 2026, but upstream GA never moves an OEM cell. This is now a slipped date, not a pending one — the top fact question for the next Dell briefing; scores when Dell docs confirm. Briefing (July 2026): knowledge graphs, graph databases, and unified streaming ingest named as AIDP capabilities — none documented GA; fact questions pending written feedback. Placement note: the instrument scores retrieval (DSE) at 1B and movement (DOE/DAE) at 1C; Dell's AIDP framing bundles them. Vendor-feedback posture: portability claims on pre-GA features are untestable and carry no DAPM weight — written feedback should arrive with GA documentation. ## ◑ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Dell + Dataloop ### Vendor-Provided Components **Data Orchestration Engine (Dataloop)** [DAPM: Ceded] No-code/low-code AI data lifecycle. Dell's most meaningful software acquisition (~$120M, Dec 2025). GA confirmed: deployment documentation on Dell InfoHub and an orderable product page (verified July 2026). Proprietary pipeline orchestration — pipeline definitions built in the engine are captive to Dell. **Orchestration Engine Marketplace** [DAPM: Delegated] 200+ models, NVIDIA NIMs, Blueprints, AI-Q templates. Substitutable catalog content. **Data Analytics Engine (Starburst)** [DAPM: Delegated] GPU-accelerated SQL. Agentic Layer + MCP Server for agent access. Trino/Iceberg substrate: schemas and materialized views lift to another engine, catalogs swap — the briefing's own portability argument is the Delegated rationale. ### NVIDIA-Provided Components **NVIDIA CMX** BlueField-4 powered context memory tier (G3.5). Pre-GA — watch-listed here and on the NVIDIA row; dependency finding, not shipping capability. **NVIDIA STX Reference Architecture** Storage-side infrastructure reference. Non-differentiating for Dell. **Blueprints, NIMs, AI-Q Blueprint** Pre-built pipeline components through the Marketplace. ### Gap Analysis Dell's most significant strategic software move. Dataloop gives Dell proprietary orchestration logic — its first owned software asset in the data lifecycle. From the buyer's seat that ownership reads as decoupled capture: the data and table formats are open (Iceberg, Trino), so schemas lift and engines swap, while the pipeline opinions built in the no-code tooling are captive to Dell. The July 2026 briefing pushback made the instrument's case: open formats prove data portability, and data portability is the decoy — it says nothing about where orchestration authority accumulates. ### Borrowed Judgment From the buyer's seat: the enterprise inherits Dell's orchestration opinions outright at the engine (Ceded, first-party — Dell borrows little from partners there). Starburst-layer analytics judgment is swappable (Delegated). KV-cache judgment moved to the watch-list with the capability. ### Working Notes NAND Research flagged maturity concern: young acquisition as enterprise orchestration engine vs. established Databricks/Snowflake. HyperFRAME: only 14% of orgs have AI-ready data architecture. Watch-list (announced GTC, March 2026; checked July 7, 2026, not scored): KV cache offload to PowerScale/ObjectScale/Lightning FS via NVIDIA CMX — 19x TTFT and 5.3x QPS are demo claims; CMX is pre-GA (watch-listed on the NVIDIA row) and Dell's language is staged availability 'throughout the year.' Scores when Dell docs confirm; the 'Context Moves to Storage' framing travels with it. Briefing (July 2026): unified streaming ingest named as an AIDP capability — not documented GA; fact question pending. ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Rack-Level Management ### Vendor-Provided Components **Integrated Rack Controller** [DAPM: Ceded] Physical rack management — power, thermal, firmware, device inventory. Operates below Layer 2A. The buyer's rack-management opinions (power/thermal/firmware baselines) are captive to Dell PowerRack. **OpenManage Enterprise** [DAPM: Ceded] Infrastructure lifecycle management. Manages the chassis, not GPU workloads. Fleet templates, policies, and baselines do not lift to another OEM's management plane. **Dell CSI Operator** [DAPM: Delegated] Dell's one K8s operator — storage provisioning, not compute orchestration. The consumed interface is CSI/Kubernetes, a genuine multi-vendor standard: provisioning opinions live in Kubernetes-standard objects (StorageClasses, PVCs) and survive a swap to another vendor's CSI driver without rebuilding the layer. ### NVIDIA-Provided Components **GPU Operator + Network Operator + NIM Operator** Three of four K8s operators in the reference architecture are NVIDIA's. **NVIDIA Run:ai** GPU scheduling, quotas, fair-share. THIS IS the Layer 2A function. NVIDIA-acquired. **NVIDIA AI Enterprise** Commercial platform wrapping the full GPU orchestration and management stack. ### Gap Analysis Dell manages the rack. NVIDIA manages the GPU-aware substrate. That distinction matters because AI Factory differentiation depends less on whether the rack can be deployed and more on how scarce accelerated capacity is scheduled, partitioned, licensed, and governed at runtime. ClearML provides floating NVAIE license management — three authorities for one optimization function. ### Borrowed Judgment High. GPU-aware orchestration primitives are NVIDIA-controlled. Dell's authority is limited to physical chassis management (Layer 0), storage provisioning (Layer 1A), and deployment automation (Day 0/1). No alternative GPU scheduler exists within the Dell AI Factory. The July 2026 briefing confirmed the posture: Dell positions everything above the rack as ecosystem, not owned capability. ### Working Notes ClearML is the most interesting independent Layer 2A play. If Dell wanted proprietary 2A capability, acquiring or deep-partnering ClearML would be the most direct path. Pending (July 2026, not scored): Dell Automation Platform / private cloud briefing — the lifecycle-automation story (including the Nutanix-on-PowerStore private cloud landing July 2026) has not been briefed in over a year and may bear on this cell. Fact questions for that briefing: does the automation platform do anything GPU-aware itself (scheduling, partitioning, quota) or stop at deployment automation; who operates it day-2; can a customer run it without Dell services. ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Blueprints + Services ### Vendor-Provided Components **Agentic AI Platform (Blueprints)** [DAPM: Delegated] Cohere North, DataRobot, ClearML blueprints. Dell provides hardware substrate and services. Menu-altitude Delegated: the choice among validated ISV platforms is real before purchase. Once deployed, each chosen platform captures per its own terms — a production Cohere North estate's orchestration opinions are Ceded to Cohere through any paper, per the channel-substitution rule. Dell holds the menu, not the capture. **Accelerator Services for Agentic AI** [DAPM: Retained] Dell's human-delivered value: strategy, deployment, optimization. Services, not software — expertise transfers to the customer, no captive opinion layer. ### NVIDIA-Provided Components **NemoClaw (OpenClaw Stack)** Open-source agent runtime (Apache 2.0). Alpha — not GA, watch-listed. Jensen: 'the operating system for personal AI.' **OpenShell** Sandboxed agent runtime with security/privacy controls. Spans deskside to data center. Alpha — not GA, watch-listed. **NeMo Guardrails** Runtime safety boundaries — what agents are NOT allowed to do. Constraint enforcement, not placement. **Dynamo** Distributed inference framework. KV-aware routing to cache-warm nodes. Closest thing to a placement decision in the stack — but single-variable optimization. **NIMs + AI Enterprise** Containerized model serving + commercial platform. ### Gap Analysis What ships on Dell paper at 2B is exactly what a customer can buy and keep or swap today: validated ISV agent platforms delivered as Dell blueprints, and Dell's human expertise to deploy them. Dell owns no runtime software — model serving, agent execution, guardrails, and distributed inference are NVIDIA or ISV capabilities, carried in the dependency column. The NVIDIA deskside agent runtime (NemoClaw/OpenShell) is alpha and watch-listed, consistent with its treatment on the NVIDIA row. Dynamo's KV-aware routing remains the closest thing to placement reasoning in the stack — single-variable cache-locality optimization, not multi-variable policy. ### Borrowed Judgment Total for the runtime the enterprise actually consumes: the NVIDIA path arrives with NVIDIA's judgment end to end — closed NIM serving decisions the enterprise can't inspect, Dynamo routing behavior, guardrail defaults. That judgment is inherited through the dependency column, not through a scored Dell component. Dell's owned contribution is curation and services; the blueprint menu keeps pre-purchase choice open (deployed platforms capture per-choice, to the ISV rather than to Dell), and Accelerator Services expertise transfers to the customer. ### Working Notes Jensen's 'OS for personal AI' is a Layer 2B claim. An OS manages execution. A control plane manages placement and policy. Watch-list (checked July 7, 2026, not scored): Deskside Agentic AI (Pro Max workstations + NVIDIA NemoClaw/OpenShell + Dell Services) — the runtime is alpha (Apache 2.0), watch-listed on the NVIDIA row; scores when GA on Dell paper. Likely lands Delegated at GA (open substrate, self-hostable) — noted, not pre-scored. The deskside hardware is scored at Layer 0 (Pro Max GB300, Retained); the services are scored here (Retained). Dell's written feedback on the 2B ratings is pending (July 2026 briefing). ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Enterprise Responsibility ### NVIDIA-Provided Components **AI-Q 2.0 Reference Architecture** Multi-agent workflow scaffolding. Does NOT make placement decisions. **OpenShell Governance** Runtime security sandboxing. Layer 2B constraint enforcement, not 2C placement reasoning. Alpha — not GA. **Dynamo KV-Aware Routing** Performance-aware routing (single variable). Not multi-variable policy optimization. ### Gap Analysis The 4+1 model defines Layer 2C as a required function of AI infrastructure — policy-driven decisions about where compute runs relative to data, which model serves which request, and how cost/compliance/latency are arbitrated in real time. Dell does not provide this function. The enterprise retains responsibility for it. Applying the 'Routing Is Not Reasoning' test: AI-Q 2.0 = workflow scaffolding. OpenShell/NeMo Guardrails = constraint enforcement. Dynamo = performance routing. None provides multi-variable policy-driven placement. The July 2026 briefing confirmed the gap in Dell's own words, twice. First, the telemetry-versus-decision boundary: Dell has telemetry from Layer 0 'all the way to 2B or even 2C,' token metrics per query, tools and guardrails — 'but the ultimate decision of what is economical or not comes down to somebody making the decision. I can give you the telemetry.' Dell provides the inputs to a reasoning plane and explicitly disclaims the decision layer. Second, the strategy: 2C and above is a deliberate ecosystem play — 'we designed the ecosystem program for 2C and 3 to give the flexibility and not try to be the winner of all, because we don't know who's going to win those layers.' The gap is a decision, not an omission. For the buyer that matters: bringing your own 2C (Kamiwaza, ClearML, potentially Palantir Ontology) collides with no Dell roadmap. The enterprise must build custom 2C logic (6-12 months), bring a partner, or operate without it. Most will choose option 3 — the gap isn't visible until production agentic workloads expose it. Dell + Intel 'actively addressing' an AI-factory control plane (SiliconANGLE, May 2026) remains security/governance framing with no product announced — still not scoreable. ### Borrowed Judgment There is no judgment to borrow — the enterprise retains full responsibility for this function. That responsibility is implicit: most enterprises operating Dell AI Factory infrastructure do not yet recognize Layer 2C as a distinct function they need to provide. ### Working Notes Dave Vellante (theCUBE): 'The AI factory requires a new control plane — one that governs data, models and agents in real time.' That control plane is Layer 2C. Briefing (July 2026), vendor position recorded: 2C can't be productized because the intelligence is subjective — 'I can give you the tools, but I can't create the intelligence for you.' That argues the gap is universal, not that Dell fills it; it moves no cell. Scope clarification from the same exchange: the instrument rates infrastructure 2C (policy-driven placement — which model, which region, which cost/compliance tier), not domain reasoning; the GDPR-placement example is deterministic given telemetry Dell confirms it has. The capability spans 1B/1C (data boundaries) and 2C (agent/model placement) — confirmed as the instrument's split. ## ◇ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Partner Ecosystem ### Vendor-Provided Components **Dell AI Ecosystem Program** [DAPM: Delegated] Structured ISV validation path. Partners: Google, Hugging Face, OpenAI, Palantir, Reflection, ServiceNow, SpaceXAI. Substitutable partners on Dell paper. **Dell Enterprise Hub (Hugging Face)** [DAPM: Delegated] Curated open-weight models on PowerEdge. DeepSeek, GLM, Kimi, Gemma, Nemotron, Mistral, Arcee. Model choices and fine-tunes are portable — open-weight substrate. **Security Stack** [DAPM: Delegated] CrowdStrike + Fortanix + F5 + Intel confidential computing. Infrastructure security, not agent governance. Substitutable security ISVs — distinct from the Ceded cyber-resilience implementation built into Dell storage at 1A. ### NVIDIA-Provided Components **NemoClaw / OpenClaw Runtime** Execution surface for Layer 3 applications. NVIDIA provides substrate; ISVs provide business logic. Alpha — not GA, watch-listed at 2B. ### Gap Analysis One of the strongest on-prem AI ecosystem stories in market. Each partner maps to a coherent use case. But each brings its own governance domain — Palantir Ontology governs within Palantir's domain, ServiceNow Otto within ServiceNow's. Nobody governs ACROSS domains on shared infrastructure. Security protects the platform from threats. Governance constrains what the platform does. Both are necessary. Only security is present. ### Borrowed Judgment Distributed across partners, which is architecturally correct at Layer 3. The structural problem: no cross-domain infrastructure judgment (Layer 2C) constrains all agents regardless of which ISV built them. ### Working Notes 5,000+ AI Factory customers (Dell, May 18, 2026 release; up from 3,000 at GTC). As they move to production agentic workloads, the multi-agent governance problem becomes visible. More ISV partners = more independent agent populations = more urgent need for Layer 2C — now also Dell's stated design ('not try to be the winner of all,' July 2026 briefing). ════════════════════════════════════════════════════════════════════════════════ # Google Cloud AI Infrastructure Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.4 - Label Vocabulary & Chip Enum Reconciliation **Date:** July 12, 2026 **Source:** Google Cloud Next 2026 (Apr 22–24), GTC 2026, NVIDIA partnership, Forrester, SiliconANGLE, The New Stack, analyst coverage. v1.2 (instrument reconciliation): 2A GKE Retained→Delegated — managed K8s behind a standard interface is Delegated, not Retained. Cloud Storage remains Ceded (GCS-native API). v1.3 (lab-validated against the TFD corpus / vCTOA RAG build): added the retrieval surface to Layer 1B - Vertex AI Search (Discovery Engine, the Bedrock Knowledge Bases analog), BigQuery Vector Search, and Firestore Vector Search, all Ceded (single-vendor APIs); demoted the preview LookML Agent from a scored component to a notes watch-line per the GA-gate; and added the self-orchestrate-versus-managed Retain/Delegate fork to 1B/1C, mirroring the AWS row. ## Summary Finding Google Cloud is the only vendor in this assessment series that owns a frontier foundation model — and that single fact restructures the entire 4+1 analysis. Google built Gemini, trains Gemini on its own TPUs, optimizes its silicon for Gemini’s training requirements, and weaves Gemini’s intelligence into every layer of its cloud platform. This creates a model-integrated stack: an architecture where the frontier model is not a component plugged into infrastructure but the intelligence that pervades the infrastructure. No other vendor assessed possesses this vertical integration. Google owns every layer of the 4+1 model with proprietary IP: custom silicon (TPUs), custom networking (Virgo), proprietary storage (Colossus/Spanner/BigQuery), its own runtime and frameworks (JAX/Pathways), its own frontier models (Gemini), and a unified orchestration surface (Gemini Enterprise Agent Platform). The DAPM implication is not merely that every layer is Ceded — it is that every layer is ceded to a unified intelligence. With AWS, authority is distributed across multiple vendors’ judgment (AWS infrastructure, Anthropic model reasoning, ISV application logic). That distribution creates complexity but also structural checks. With Google Cloud + Gemini, the enterprise concentrates authority in one vendor’s judgment across every layer — from silicon to application. This is the deepest expression of vertical integration in enterprise technology since the mainframe era. The enterprise gains end-to-end optimization that no multi-vendor assembly can match. But the 4+1 framework makes visible what the integration obscures: the enterprise has no fallback position at any layer. Google Distributed Cloud (GDC) addresses data sovereignty without addressing judgment sovereignty — GDC still runs Google’s software stack and Google’s models. The structural question: does concentrating all layers of authority and all layers of model judgment in a single vendor deliver enough value to justify the governance position — and has the enterprise made that concentration explicit rather than inheriting it by default? ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** TPU + GPU Full Stack ### Vendor-Provided Components **Google Custom Silicon (TPUs)** [DAPM: Ceded] TPU 8t (training): 9,600-chip superpods, ~3x processing power vs Ironwood. TPU 8i (inference): 288GB HBM, 384MB on-chip SRAM, ~80% better perf/dollar. Designed for agentic AI, MoE models, large-scale RL. Google designs TPUs to train Gemini — the enterprise benefits from silicon optimized by one of the world’s most demanding ML workloads but does not direct the optimization priorities. **NVIDIA GPUs on Google Cloud** [DAPM: Ceded] A5X bare-metal on Vera Rubin NVL72 (among first cloud providers to deploy). A4 Ultra (NVL72 preview Q2 2026), A3/A3 Mega/A3 Ultra. Fractional GPUs (G4 VMs, industry-first RTX PRO 6000 Blackwell vGPU). Scale: 80,000 Rubin GPUs single-site, 960,000 across multisite. Third-party models (Claude, Llama, Mistral) run on NVIDIA, not TPU — the playing field is structurally uneven. **Virgo Network Fabric** [DAPM: Ceded] Purpose-built AI-optimized DC fabric. 134,000 TPU 8t chips connected at 47 Pb/s non-blocking bi-section bandwidth per DC. 4x bandwidth per accelerator, 40% lower unloaded latency vs prior gen. Also available for A5X. Designed for Gemini’s training topology. **Google Distributed Cloud (GDC)** [DAPM: Delegated] On-prem deployment of Google Cloud services. Connected and air-gapped configs. 4 racks to hundreds. NVIDIA Blackwell GPUs + Gemini Flash models on-prem. Managed GDC Provider initiative (Clarence, Gulf Energy, T-Systems, WWT). NATO deployment. Customer provides facility; Google provides and operates HW+SW. Addresses data sovereignty but not judgment sovereignty. ### NVIDIA-Provided Components **NVIDIA GPU Silicon** Vera Rubin NVL72, Blackwell B200/B300, H100/H200. 1M+ NVIDIA GPUs. NVIDIA instances serve third-party models that can’t run on TPU. **NVIDIA NIXL + Networking** NIXL for disaggregated inference. ConnectX/BlueField for GPU networking. Google manages the NVIDIA integration layer. ### Gap Analysis No Layer 0 capability gap — Google’s portfolio is the broadest of any single cloud provider. The gap is governance: the enterprise has no authority over any Layer 0 component beyond selecting instance types. The silicon-model feedback loop is structurally unique: Google designs TPUs to train Gemini, not primarily to sell cloud compute. TPU roadmap decisions reflect Gemini’s training topology, not enterprise customer workload requirements. The enterprise inherits optimization it didn’t direct. AWS’s Trainium is designed for customer workloads. NVIDIA designs for the broadest market. Google designs for Gemini and makes TPUs available to customers. The multi-accelerator matching problem (TPU vs NVIDIA vs Axion CPU) creates a workload-to-silicon decision that recurs per-workload in cloud vs once at procurement on-prem. No productized policy engine automates that matching. Fluid Compute (Layer 2A) begins to address it but doesn’t consult governance metadata. GDC follows the same inverted operating model as AWS AI Factories: Google operates infrastructure the customer houses. Unlike Dell PowerRack or HPE ProLiant (enterprise-owned hardware), GDC is Google-operated even when customer-hosted. ### Borrowed Judgment The silicon-model feedback loop: Google’s TPU roadmap is driven by Gemini’s training requirements. If Google decides TPU 9 should optimize for MoE architectures because that’s where Gemini is heading, every enterprise TPU workload inherits that architectural bet. Borrowed judgment at the silicon layer — a concept with no parallel in the Dell or HPE assessments. Virgo as borrowed network judgment: the enterprise inherits Google’s network optimization decisions without visibility or control. Cannot audit bandwidth sharing across tenants or prioritization of Google’s own Gemini training traffic. ### Working Notes The dual-architecture hedge (TPU + NVIDIA) gives Google pricing leverage and architectural independence. The enterprise benefits indirectly but does not control whether NVIDIA GPU instances remain first-class citizens as Google optimizes for its own silicon. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Model-Powered Governance ### Vendor-Provided Components **Cloud Storage + Rapid Tier (Colossus)** [DAPM: Ceded] Standard object storage at planetary scale. Rapid tier uses Colossus — Google’s internal distributed storage platform (previously powering Search, Gmail, YouTube, Gemini training). Sub-millisecond read/write. The enterprise gets the same storage engine that holds Google’s training data. **Smart Storage** [DAPM: Ceded] Automatically analyzes unstructured data and generates metadata/context on ingest using Gemini. Auto-tags images, PDFs. First recursive point: Smart Storage uses Gemini to enrich the data that will eventually be served to Gemini-powered agents. The model enriches the data that feeds the model. **Knowledge Catalog (Gemini-Powered)** [DAPM: Ceded] Universal business context and governance. Gemini-powered semantic extraction, entity relationship mapping, dynamic context graph construction. Sub-second semantic search for agent retrieval. Aggregates metadata across BigQuery, AlloyDB, Spanner, Cloud SQL, Firestore, Looker. Third-party integrations (Atlan, Collibra, Datahub). Enterprise Connectivity federates context from Salesforce, Palantir, Workday, SAP, ServiceNow. Column-level lineage (GA). **BigQuery Storage** [DAPM: Ceded] Serverless columnar storage for structured/semi-structured. Managed Iceberg tables. Separates storage and compute. BigQuery spans storage, analytics, ML, and governance in a single service. ### Gap Analysis Knowledge Catalog with Smart Storage represents the most ambitious attempt to solve the metadata-to-agent grounding problem. No other vendor has a production system that automatically extracts business semantics from raw data, builds a context graph, and serves that context to agents in real-time with governance enforcement. The recursive dependency is the central finding: Knowledge Catalog uses Gemini to perform semantic extraction and build the context graph. When Knowledge Catalog determines what context an agent receives, and that context graph was built by Gemini, the enterprise is consuming Gemini’s judgment about what its own data means — at the governance layer, before any application-level inference occurs. No other vendor has this pattern. Dell’s MetadataIQ indexes deterministically. AWS’s Glue doesn’t use Nova to enrich its catalog. HPE’s Ezmeral doesn’t use a foundation model for data semantics. VAST’s catalog is storage-native, not model-powered. The 1A→2C connection: Knowledge Catalog is not a passive registry. It makes decisions about what context agents receive, how data assets are ranked for retrieval, and which governance policies apply. The context graph determines agent grounding — an orchestration function, not a storage function. Governance gap: Knowledge Catalog is optimized for GCP. Third-party integrations federate INTO Knowledge Catalog, not out of it. An enterprise running PowerScale + S3 + GCS cannot use Knowledge Catalog as a federated surface across all three without making GCP the metadata authority. ### Borrowed Judgment The context graph as borrowed judgment: when Knowledge Catalog builds entity relationships and business meanings, every agent that queries the graph inherits its representation of reality. If Gemini’s semantic extraction misclassifies a data asset, every agent grounded in that context acts on the incorrect interpretation. Borrowed judgment at the governance layer — before application-level reasoning. Smart Storage as ingest-time judgment: a single Gemini classification at ingest (‘this document is about Project X’) becomes a persistent governance fact. Limited mechanisms to audit or correct model-generated metadata at scale. Comparison to AWS: AWS classifies 1A as Delegated because Lake Formation enforces customer-defined policies. Google’s 1A is Ceded because Knowledge Catalog generates governance intelligence using Gemini. The enterprise on AWS retains governance judgment. The enterprise on Google inherits it. ### Working Notes The Layer 1A / 2C boundary question: Knowledge Catalog’s context graph is architecturally Layer 1A (data catalog) but functionally Layer 2C (determines agent grounding and context routing). Model-powered governance layers that make orchestration decisions are functionally 2C even when architecturally 1A. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Model Prep & Managed Retrieval ### Vendor-Provided Components **BigQuery ML** [DAPM: Ceded] In-database ML training and inference. Supports linear/logistic regression, K-means, time series, XGBoost, DNNs, imported TF/PyTorch models. Collapses the boundary between data preparation (1B) and AI runtime (2B) by running ML directly in the warehouse. **Dataflow + Dataproc** [DAPM: Ceded] Managed Apache Beam (batch+stream) and managed Spark/Hadoop for large-scale processing. Open-source frameworks, Google-managed execution. **Data Agent Kit (Open-Source)** [DAPM: Delegated] MCP-based agents packaged as tools and skills. Supports Claude Code, Gemini CLI, Codex, VS Code. Enables intent-driven development: practitioners define goals, agents handle implementation. Creates governance recursion: the agent builds the pipeline that prepares the data that feeds the agent. **Vertex AI Feature Store** [DAPM: Ceded] Managed feature serving for online/offline ML models and agents. Consistent feature serving across training and inference. 1B→2B bridge. **Vertex AI Search (Discovery Engine)** [DAPM: Ceded] Turnkey managed RAG and enterprise search via the Discovery Engine API (Vertex AI Agent Builder). Create a data store over Cloud Storage, BigQuery, or websites and Google handles parsing, chunking, embedding, indexing, ranking, and grounded retrieval end to end. The direct analog to Bedrock Knowledge Bases, but on a proprietary managed datastore with no swappable open backend, so more captive than Bedrock KB: the retrieval opinions have no open exit. **BigQuery Vector Search** [DAPM: Ceded] VECTOR_SEARCH over an IVF vector index, GA, with hybrid keyword-plus-vector retrieval in-warehouse. (The ScaNN index type is still in preview.) A self-orchestrate primitive: the enterprise builds the retrieval pipeline, but the store is BigQuery, a single-vendor API with no open exit. **Firestore Vector Search** [DAPM: Ceded] K-nearest-neighbor vector search on transactional Firestore data, GA, with composite pre-filters and LangChain and LlamaIndex integrations. A self-orchestrate primitive on a single-vendor Firestore API: convenient for operational RAG, captive at the store layer. ### Gap Analysis No meaningful capability gap. Most mature Layer 1B in the assessment series. BigQuery ML eliminates data-to-model handoff. LookML Agent automates semantic model construction. Data Agent Kit enables agent-driven pipeline development. The gap is governance over model-generated data artifacts. When LookML Agent generates a semantic model, when Data Agent Kit writes a pipeline, when BigQuery ML trains a model — who reviews the output for correctness? These are AI-generated artifact governance problems. Google provides no productized capability for governing model-generated data artifacts at scale. Data Agent Kit’s explicit support for Claude Code and non-Google tooling is strategically significant — the one point in the stack where third-party model access is genuinely equal. Retrieval is the surface this layer is named for, and Google covers it comprehensively but captively. Vertex AI Search (the Discovery Engine API) is the turnkey managed RAG product, the direct analog to Bedrock Knowledge Bases: ingest, chunk, embed, index, and grounded retrieval handled end to end. BigQuery VECTOR_SEARCH and Firestore Vector Search are the self-orchestrate primitives. The governance fork mirrors AWS: consume Vertex AI Search and Google owns the retrieval decisions (Ceded), or self-orchestrate on BigQuery or Firestore vector and Retain the retrieval-strategy opinions in code. The twist against AWS: every native Google vector store is single-vendor (BigQuery, Firestore, and Vertex AI Vector Search are all Ceded), so self-orchestration Retains the pipeline logic but still lands on a captive store unless the enterprise deliberately chooses AlloyDB or Cloud SQL pgvector, the one portable escape. Even the turnkey surface is more captive than AWS: Bedrock Knowledge Bases is Delegated because it orchestrates over swappable open backends, while Vertex AI Search runs a proprietary managed datastore with no open exit. ### Borrowed Judgment The semantic model as borrowed judgment: when LookML Agent generates definitions, every analytics query and agent interaction using those definitions inherits Gemini’s interpretation of business logic. Powerful (automates weeks of manual semantic modeling) and risky (embeds model judgment in the analytical foundation). The pipeline-building agent as borrowed judgment: Data Agent Kit agents write Dataflow jobs and BigQuery transformations. The enterprise inherits the agent’s data engineering judgment — join strategies, filter logic, null handling. Previously human expertise, now model-generated. The retrieval surfaces deepen the capture: Vertex AI Search, BigQuery vector, and Firestore vector are all single-vendor APIs, so the enterprise inherits Google opinions on chunking, ranking, and grounding with no portable exit short of pgvector. ### Working Notes The comparison to AWS SageMaker Unified Studio: AWS provides a single governed environment across services (service-wide integration). Google achieves integration through BigQuery spanning storage, analytics, ML, and governance (service-deep integration). Google = tighter integration at cost of BigQuery lock-in. AWS = service diversity at cost of integration complexity. Watch-list (preview, not scored): LookML Agent (Gemini-powered semantic-model generation) - moved off the scored components per the GA-gate until it reaches GA; it remains a notable forward signal for model-generated semantic models. ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Managed Pipeline Portfolio ### Vendor-Provided Components **Cloud Storage Rapid Tier (Colossus)** [DAPM: Ceded] Sub-millisecond caching layer between persistent storage and compute. Previously internal-only, now customer-accessible. **Managed Lustre** [DAPM: Ceded] 10 TB/s bandwidth (10x YoY, claimed 20x faster than other hyperscalers), 80 PB capacity. RDMA-enabled. Training data movement layer between Cloud Storage and TPU/GPU clusters. **Cross-Cloud Lakehouse (Preview)** [DAPM: Ceded] Agentic AI workflows access data across AWS and Azure without egress — querying data in place rather than copying. Eliminates ETL and cross-platform data movement costs. Extends compute to wherever data sits. Google’s query engine becomes the universal data access layer regardless of physical data location. **BigLake** [DAPM: Ceded] Unified data fabric across data lake (Cloud Storage) and warehouse (BigQuery). Single schema, multiple processing engines. Row/column-level governance via BigQuery Storage API across all access paths including open-source engines. Multi-cloud via BigQuery Omni. **Smart Storage Ingest Enrichment** [DAPM: Ceded] Data entering Google Cloud is auto-enriched by Gemini with metadata and context tags. Data is not just moved — it is interpreted on arrival. No parallel in other vendors’ data movement layers. ### Gap Analysis Cross-Cloud Lakehouse is the most strategically important Layer 1C capability in the assessment series. It represents a fundamentally different approach to data gravity: extend compute to wherever data sits rather than moving data to compute. The DAPM implication: if Cross-Cloud Lakehouse delivers, the enterprise doesn’t need to move data to GCP. That reduces data lock-in at storage. But it increases lock-in at the query layer — analytical capabilities depend on Google’s query engine reaching across clouds. Data sovereignty improves. Analytical sovereignty does not. The data movement / enrichment coupling: Smart Storage’s Gemini enrichment at ingest means Layer 1C (movement) and Layer 1A (governance) are coupled through model inference. Moving files into Cloud Storage triggers model inference generating metadata that propagates into the context graph. No other vendor couples data movement with model-powered enrichment. Cross-Cloud Lakehouse is in preview. Performance, cost, and governance characteristics at enterprise scale are unproven. ### Borrowed Judgment Cross-Cloud Lakehouse as borrowed query optimization: when Google’s engine optimizes a cross-cloud query, the enterprise inherits Google’s optimization judgment. Opaque and unchallengeable — cannot tune the cross-cloud query plan or audit how Google’s engine accesses data in a competing cloud provider’s storage. The enrichment coupling: data engineers moving files into Cloud Storage unknowingly trigger model inference. Convenient (automatic enrichment) and opaque (the engineer may not know Gemini is interpreting their data on arrival). The same embed-and-index fork as 1B applies to the RAG ingest pipeline: consume Vertex AI Search and the ingest, chunk, and embed movement is managed for you (Ceded), or self-orchestrate it (for example Dataflow plus Vertex embeddings feeding BigQuery or Firestore vector) and Retain the pipeline logic, though the chosen store stays captive. ### Working Notes The asymmetry between data plane federation (Cross-Cloud Lakehouse) and control plane federation (absent at 2C) is a structural finding. Google invests in making data accessible across clouds but not in making agent governance portable across clouds. Data accessibility without orchestration portability draws workloads toward GCP as the governance center. ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** GKE + Sovereign Extensions ### Vendor-Provided Components **GKE + GKE Agent Sandbox** [DAPM: Delegated] Managed K8s with AI-era extensions. Agent Sandbox: gVisor-based secure isolation, 300 sandboxes/second/cluster with sub-second time to first instruction. Infrastructure built for the agentic era, not retrofitted. GKE is a managed service behind the standard Kubernetes API — manifests lift to another conformant cluster, so Delegated (operation delegated to Google; the standard interface keeps the opinions portable). **Fluid Compute** [DAPM: Ceded] GCE + GKE dynamically shifting workloads in real-time. CPUs for branchy agent logic, secure sandboxes, RL, SLM inference, RAG. GPU/TPU for training and large-model inference. Proto-Layer 2C: routes based on workload characteristics, not business context. **GDC (Sovereignty Analysis)** [DAPM: Delegated] On-prem Google Cloud services: GKE, Agent Platform, managed storage, Gemini Flash, Blackwell GPUs. Air-gapped for sensitive workloads. Addresses data sovereignty (where computation happens) but NOT judgment sovereignty (whose model drives computation). GDC runs Gemini on-prem — same recursive dependency, inside the enterprise perimeter. **Capacity Management** [DAPM: Ceded] CUDs (1yr/3yr), on-demand, preemptible/spot, Dynamic Workload Scheduler. Same pattern as AWS — capacity acquisition (2A), not workload placement (2C). ### NVIDIA-Provided Components **NVIDIA GPU Operator (on GKE)** Available for NVIDIA instances on GKE. Google manages the GPU integration layer. ### Gap Analysis GKE is the most mature managed K8s for AI workloads. GKE Agent Sandbox has no equivalent in Dell, HPE, or AWS portfolios — 300 sandboxes/second is built for agentic workload density. Fluid Compute sits at the 2A/2C boundary. Its dynamic workload shifting is more than capacity acquisition — runtime decisions about which compute type serves which workload. But less than full 2C — routes on workload characteristics, not business context (data residency, compliance tags, cost targets). The Fluid Compute → Knowledge Catalog connection does not exist: workload placement does not consult governance metadata. Same Infrastructure Layer 2C gap every vendor has. GDC: the full sovereignty analysis reveals that data sovereignty ≠ judgment sovereignty. Knowledge Catalog on GDC uses Gemini. Smart Storage on GDC uses Gemini. Agent Platform on GDC uses Google’s runtime. The enterprise gains physical sovereignty but retains the same judgment concentration. The ‘self-driving cloud’ narrative implies Gemini-powered infrastructure operations — autonomous root-cause analysis on infrastructure telemetry. If the Reasoning Plane is itself Gemini-powered, the operational intelligence and application intelligence are the same intelligence. ### Borrowed Judgment GKE consumed via the standard Kubernetes interface: the enterprise's manifests and operators lift to any Kubernetes (Delegated — managed K8s service; the enterprise could switch without rebuilding). Google's optional Gemini-driven scheduling optimization is the only borrowed-judgment layer, and only if the enterprise opts in. GDC as borrowed judgment in sovereign packaging: physical control over facility, Google’s judgment in software, model, governance, and operations. Data sovereignty with judgment concentration. Fluid Compute as proto-2C borrowed judgment: when Fluid Compute routes agent work to CPU vs GPU, that routing is Google’s judgment about optimal compute matching. Enterprise doesn’t configure the routing policy. ### Working Notes Data sovereignty vs judgment sovereignty: the 4+1 framework should distinguish between where data resides (which GDC addresses) and whose model’s reasoning shapes the AI system (which GDC does not address). An enterprise running GDC air-gapped has data sovereignty while fully ceding judgment sovereignty. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Model-Integrated Stack ### Vendor-Provided Components **Gemini Enterprise Agent Platform** [DAPM: Ceded] Unified platform for building, scaling, governing, optimizing agents. Subsumes all Vertex AI services. Agent Studio (low-code), ADK (code-first, Python/Go/Java/TypeScript, model-agnostic, open-source), Model Garden (200+ models incl Gemini, Claude, Llama, Gemma), Agent Runtime, Agent-to-Agent Orchestration, Agent Identity (GA), Agent Gateway, Agent Observability, Agent Registry, Memory Bank, Antigravity (desktop app + CLI). **The House Model Advantage** [DAPM: Ceded] Gemini on TPU: silicon designed for this model, networking designed for its topology, distributed runtime (Pathways) built for its coordination, inference optimized for its architecture, governance (Knowledge Catalog) powered by it, orchestration defaults to it. No other vendor achieves this degree of vertical optimization. Third-party models (Claude, Llama) run on NVIDIA GPUs — supported but not co-optimized. The playing field is structurally tilted. **Frameworks (JAX, Pathways, TorchTPU, vLLM)** [DAPM: Ceded] JAX: Google’s ML framework optimized for TPU. Pathways: distributed runtime for superpod-scale training. TorchTPU: full PyTorch support on TPUs (concession to ecosystem). vLLM: optimized across GPU + TPU. llm-d: open-source K8s-native inference serving (multi-vendor project). ### NVIDIA-Provided Components **NVIDIA GPU Instances** A3/A3 Mega/A3 Ultra/A5X for third-party model inference. CUDA ecosystem required for non-Gemini models. ### Gap Analysis Layer 2B is the center of gravity for the model-integrated stack. The model provider, runtime provider, and infrastructure provider are the same company. When the enterprise runs Gemini on Agent Platform on TPU, it borrows Google’s judgment at the model layer, runtime layer, framework layer, and silicon layer simultaneously. A single entity’s priorities shape the entire execution path. The Agent Platform collapses Layers 2B, 2C, and 3 into a single product surface: Agent Runtime (2B infrastructure), Agent Identity/Gateway/Registry/Orchestration/Observability (2C governance), Agent Studio/ADK/Antigravity (Layer 3 development). The product boundary does not align with the architectural boundary. AWS separates these: Bedrock (model access) is distinct from AgentCore Runtime (agent execution) is distinct from AgentCore Policy (governance). AWS’s separation preserves architectural boundaries the enterprise can independently govern. Google’s collapse optimizes integration but prevents swapping the governance layer (2C) while keeping the runtime (2B). The NVIDIA dependency at 2B is optional for Gemini (TPU-native) but required for third-party models. The enterprise using Claude on Google Cloud pays a structural performance tax — Claude runs on NVIDIA GPUs through a runtime designed for Gemini. Model Garden’s 200+ models are API-equal but not silicon-equal. ### Borrowed Judgment The model-integrated runtime: Gemini on TPU inherits Google’s judgment at model layer (training data, alignment, safety), runtime layer (scheduling, scaling, session management), framework layer (JAX/Pathways optimization), and silicon layer (TPU architecture). Most concentrated borrowed judgment in the assessment series. The open-source hedge: llm-d, TorchTPU, vLLM, ADK provide genuine open alternatives. Google opens components that reduce adoption friction (frameworks, SDKs) while keeping authority-concentrating components closed (Agent Runtime infrastructure, Agent Gateway, Pathways). Production deployment pulls open tools into Google’s managed surface where authority shifts from Retained to Ceded. ### Working Notes The 2B/2C collapse prevents the enterprise from independently governing the orchestration layer. An enterprise that wants Google’s Agent Registry and Agent Identity but AWS’s Bedrock for model access and its own governance engine for policy enforcement cannot compose that architecture. The components are bundled. ## ● Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Productized Placement ### Vendor-Provided Components **Agent Identity (GA)** [DAPM: Ceded] Agents as identity principals with authentication, authorization, audit. Control plane function: determines which agents exist as governed entities. **Agent Gateway** [DAPM: Ceded] Protocol-level governance for MCP and A2A communications. Security partner integrations (Broadcom, Check Point, Cisco, CrowdStrike, F5, Netskope, Okta, Palo Alto, Zscaler). Spans Layer 0 (networking), 2B (runtime), and 2C (orchestration). **Agent-to-Agent Orchestration** [DAPM: Ceded] Deterministic multi-agent workflow routing. Control plane function: determines which agent handles which subtask. **Agent Registry** [DAPM: Ceded] Catalog of agents with ownership, capabilities, protocols, invocation details. Administrator-controlled discoverability. Control plane function: determines which agents are available and who can use them. **Agent Observability** [DAPM: Ceded] Monitoring, tracing, debugging across production agent populations. Feedback loop for detecting faulty reasoning and intervening. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** All Layer 2C components are Google IP. NVIDIA does not control governance, policy, or reasoning in Google’s stack. ### Gap Analysis Google’s Intelligence Layer 2C is the most complete productized offering in the assessment series: Agent Identity + Gateway + Registry + Orchestration + Observability + Memory Bank. Together they constitute a genuine control plane for agent governance. Infrastructure Layer 2C — the autonomous placement engine — is NOT built as a customer-configurable product. The capacity primitives (Fluid Compute, CUDs, DWS) are building blocks, but they don’t compose into a policy-driven placement engine querying Knowledge Catalog governance metadata. Same gap as AWS and every other vendor. Google’s implicit Layer 2C is the most sophisticated in the assessment: managed services make autonomous placement, scaling, routing, and capacity decisions invisibly. The enterprise cannot see, configure, audit, or override these decisions. The model-integrated Reasoning Plane: if Google’s ‘self-driving cloud’ uses Gemini for infrastructure decisions, then the model powering the enterprise’s agents (Layer 3) is the same model governing agent orchestration (Intelligence 2C) is the same model deciding where agents run (Infrastructure 2C). One model’s judgment pervades every decision surface. Cross-cloud orchestration gap: Google federates the data plane (Cross-Cloud Lakehouse) but NOT the control plane. Agent Platform governs GCP agents only. Enterprise running agents across multiple clouds has no cross-platform agent governance surface — unless all agents route through Google’s Agent Gateway, which cedes cross-cloud governance to Google. The captive-but-best dilemma: this is evidence the control plane CAN be built as a coherent capability. The enterprise architect who wants it has one option: adopt Google Cloud. The federated alternative does not exist. ### Borrowed Judgment The captive control plane: enterprise inherits Google’s orchestration model — deterministic routing, Google-managed identity, Google-governed protocols. Well-engineered but unchallengeable — cannot substitute alternative orchestration logic within the Agent Platform boundary. The model-powered control plane: if the Reasoning Plane uses Gemini for infrastructure decisions, a model judgment error at the control plane layer is invisible to the enterprise, with no fallback to human decision-making. Intelligence 2C: Low borrowed judgment in the sense that the components are productized and configurable. High borrowed judgment in the sense that the governance logic itself (Agent Gateway protocol decisions, Agent Identity authentication model, Orchestration routing patterns) is Google’s, not the enterprise’s. ### Working Notes Google’s 2C proves the Control Plane Working Notes thesis: the control plane can be built. The question is whether it can be liberated from the vendor boundary — and whether the model-integrated dimension (control plane powered by the same model it governs) is a pattern to replicate or to avoid. The asymmetry: data plane federates (Cross-Cloud Lakehouse), control plane does not. This serves Google’s strategic interest — data accessibility without orchestration portability draws workloads toward GCP as governance center. ## ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Open Model Layer, Captive Platform ### Vendor-Provided Components **Gemini Model Family** [DAPM: Ceded] Gemini 3.1 Pro, Gemini 3.5, Gemini Flash. The model the entire stack was designed around. Gemma open-weight models for self-hosting (the one offering where enterprise can Retain model authority). **Model Garden (200+ Models)** [DAPM: Delegated] Gemini, Claude Opus/Sonnet/Haiku, Meta Llama, Gemma, open-source models. Model Evaluation service. Broadest model catalog of any cloud provider. Model-agnostic claim genuine at Layer 3 — more so than any other layer. **Application Surfaces** [DAPM: Retained] ADK (code-first, open-source, model-agnostic) — agents built on the open SDK lift out; the substrate is self-hostable. **Application Surfaces (Managed Studios)** [DAPM: Ceded] Agent Studio (low-code) and Agent Designer (no-code in the Gemini Enterprise app), plus Gemini Enterprise agent discovery and the Deep Research agent. Managed authoring surfaces — low/no-code agent definitions are captive to the platform; scored separately from the open ADK. **Consumer-Enterprise Feedback Loop** [DAPM: Ceded] Gemini powers Google Search, Gmail, Docs, Photos, Android, Chrome. Workspace Intelligence uses Gemini for agentic work. Model improvements from billions of consumer interactions directly benefit enterprise workloads. But: consumer-driven alignment and safety tuning may not align with enterprise needs. **Google Antigravity 2.0 (Agent-First Development Platform)** [DAPM: Ceded] Announced I/O 2026. Standalone desktop app + CLI + SDK — a full developer platform built around agent orchestration. Multi-agent parallel execution: orchestrate multiple agents and execute tasks simultaneously. Dynamic subagent workflows and scheduled background automation. Antigravity CLI (Go-based, replacing Gemini CLI — deprecated June 18, 2026) for terminal-native multi-agent workflows. Antigravity SDK for building custom agents with templates in AI Studio. Powered by Gemini 3.5 Flash (co-developed using Antigravity). Native voice command support. Ecosystem integrations: Google AI Studio, Android, Firebase. Export tool for AI Studio → local development. Search integration: real-time custom UI generation within Google Search answers. AI Ultra plan ($100/month, 5x usage limits). Google's most aggressive move in the agentic coding market — positioned as the hub for multi-agent development workflow orchestration, not just code assistance. ### NVIDIA-Provided Components **NVIDIA Models via Model Garden** NVIDIA Nemotron and other NVIDIA models available alongside all other providers. ### Gap Analysis No meaningful capability gap. Broadest model catalog. Most portable agent development framework (ADK). Application surfaces from no-code through code-first. Google deliberately keeps Layer 3 more open than any other layer — while ensuring every Layer 3 application is gravitationally pulled toward Agent Platform (2B/2C). By keeping Layer 3 open, Google maximizes platform adoption: enterprises wanting Claude on Google Cloud still consume Agent Platform’s runtime, identity, gateway, registry, observability. The model is portable; the platform is captive. Consistent with the 4+1 model’s prediction that vendor lock-in concentrates at Layer 2B/2C, not Layer 3. Google has understood this prediction and built strategy accordingly. Code portability vs operational portability: ADK is open-source and model-agnostic — agent code CAN run on AWS or on-prem K8s. But Agent Registry, Memory Bank, Agent Identity, Agent Gateway, Agent Observability are Google Cloud services with no portable equivalents. Agent code is an asset the enterprise owns. Agent operations are an asset it rents. Antigravity 2.0 deepens the Layer 3 gravitational pull toward Google's platform. The desktop app + CLI + SDK creates a development surface that integrates directly with Agent Platform (2B/2C): agents built in Antigravity inherit Agent Platform's identity, gateway, registry, and observability. The Gemini CLI deprecation (June 18, 2026) forces migration to Antigravity CLI — consolidating Google's developer AI surface into one opinionated platform. The SDK enabling custom agent templates in AI Studio means Antigravity is not just a coding tool but an agent construction platform that feeds directly into the Gemini Enterprise Agent Platform. Compare to AWS Kiro (spec-driven, methodology-opinionated, Bedrock-native) and GitHub Copilot (IDE-embedded, multi-model, GitHub-native). Google's differentiator is multi-agent parallel orchestration — Antigravity coordinates multiple agents simultaneously rather than single-agent sequential interaction. This maps to the 4+1 model's Layer 2C vision: orchestrating multiple agents is a control plane function that Antigravity surfaces through a developer tool. The consumer-enterprise feedback loop extends to Antigravity: Google is using Antigravity's capabilities in consumer Search to generate real-time custom UIs as part of search answers. Developer tool innovations flow to consumer products and back — a flywheel no other vendor in the assessment possesses. ### Borrowed Judgment Gemini as borrowed Layer 3 judgment: alignment changes affect agents (Layer 3), governance enrichment (Layer 1A via Knowledge Catalog), semantic models (Layer 1B via LookML Agent), and potentially infrastructure operations (Layer 2C via self-driving cloud). A single alignment decision propagates across the entire model-integrated stack. Platform defaults: Agent Studio and Agent Designer default to Gemini. Enterprise that adopts without explicitly selecting alternatives inherits Google’s model preference as a default rather than a decision. Strategic openness as borrowed judgment about lock-in location: Google’s decision to keep Layer 3 open and concentrate lock-in at 2B/2C is itself borrowed judgment the enterprise inherits. Evaluating Google on model diversity without evaluating platform captivity accepts Google’s framing of where portability matters. ### Working Notes The consumer-enterprise feedback loop has no parallel in the assessment. Model improvements from billions of consumer interactions benefit enterprise workloads — but consumer-driven alignment may constrain enterprise use cases. If Google tightens content policies for consumer safety, enterprise agents inherit that tightening. The Gemini CLI → Antigravity CLI forced migration is a significant authority move. Over 100,000 GitHub stars on Gemini CLI — all those developers must migrate to Antigravity by June 18, 2026. This concentrates Google's developer AI surface into one platform and one billing model (AI Ultra at $100/month). The deprecation timeline is aggressive but consistent with Google's pattern of consolidating developer tools around Gemini. Antigravity 2.0's scheduled tasks capability (agents running automatically in the background) converts the developer tool from a single-turn interaction to a persistent automation pipeline. This blurs the boundary between Layer 3 (application) and Layer 2C (orchestration) — when Antigravity schedules background agents to perform tasks autonomously, who governs those agents? The answer is Agent Platform — reinforcing the Layer 2C gravitational pull. ════════════════════════════════════════════════════════════════════════════════ # HPE Private Cloud AI Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.8 - 1A Frontier Promotion & Label Reconciliation **Date:** July 12, 2026 **Source:** GTC 2026, HPE GreenLake/storage May 2026 announcements, Discover 2025, Juniper acquisition, Town of Vail whitepaper, analyst coverage. v1.6: vendor name updated from 'HPE AI-Native Infrastructure' to 'HPE Private Cloud AI'. v1.7: Kamiwaza components at 1B/1C/2B/2C moved Delegated→Ceded under the instrument's retraction of the OEM-channel DAPM rule (a proprietary single-implementation platform is Ceded-to-partner through any paper; partnership substitutability claims fail the lift-to-leave litmus) — calibrated to Kamiwaza's own row, which scores the same surfaces Ceded. 2C statusLabel moved to capability vocabulary. Bracketing narrative corrected: the capability advantage over Dell survives; the authority reading is capability-with-capture versus no-capability-no-capture. No capability grades moved. v1.8 (July 12, 2026, /reconcile): 1A moderate→strong under the whose-paper and frontier rules (owned Alletra/Data Fabric/Polaris/Zerto depth; drift vs. Lenovo 1A strong resolved). 1B/2B statusLabels moved to capability vocabulary. 2C prose corrected: Google’s productized 2C is Ceded per the GCP row, not Retained. ## Summary Finding HPE presents the most structurally interesting comparison to Dell in the 4+1 model because it makes genuine software authority claims that Dell does not. Three capabilities differentiate HPE's architectural position: GreenLake Intelligence (agentic AI mesh using domain-specific LLMs via MCP — HPE-owned Layer 2A/2C for IT operations), the $14B Juniper Networks acquisition (full networking IP stack from silicon to software — owned networking depth that Dell entirely delegates to NVIDIA Spectrum), and the Unleash AI program with Kamiwaza as the chosen Layer 2C orchestration partner (validated in production at Town of Vail). The AI workload runtime (Layer 2B) remains structurally dependent on NVIDIA AI Enterprise, branded as ‘NVIDIA AI Computing by HPE.’ HPE co-engineers more deeply than Dell — Private Cloud AI is a jointly developed product — but the DAPM implication is the same: Layer 2B model execution authority is Ceded. However, HPE brackets the NVIDIA-controlled Layer 2B with governance surfaces above (GreenLake Intelligence at 2A/2C) and below (GreenLake platform at 2A) — governance coverage, not governance authority: the bracket is real, and it is a bracket of Ceded surfaces, HPE-owned above and partner-captive (Kamiwaza) beside it. Kamiwaza's capabilities span multiple layers — context orchestration (1B), governed data pipelines (1C), agent execution coordination (2B), and decision authority placement (2C) — making it a multi-layer platform, not a point solution. The Town of Vail deployment serves as a by-proxy assessment of Kamiwaza's capabilities across this full span. The Kamiwaza-provided functions are Ceded to the chosen partner: the platform's orchestration and governance opinions are Kamiwaza-captive, exactly as they score on Kamiwaza's own row. The structural advantage over Dell survives on the capability axis — the function exists, someone is accountable, and it is validated in production. The authority reading is capability-with-capture versus Dell's no-capability-no-capture. The Cray supercomputing heritage gives HPE a sovereign AI positioning that Dell cannot match — exascale systems for Argonne, HLRS, HammerHAI (EU AI Factory). This is a differentiated Layer 0 capability with implications for sovereign data governance at Layer 1A. HPE has one of the most credible on-prem AI infrastructure stacks in the market. Its credibility comes from genuine software authority (GreenLake Intelligence, Data Fabric, OpsRamp), depth of networking IP (Juniper/Aruba/Slingshot — three owned fabrics, each captive to its own stack but each representing real engineering investment), sovereign compute heritage (Cray), and a structured ecosystem model (Unleash AI) that deliberately addresses Layer 2C through a chosen partner rather than leaving it unaddressed. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** HPE Strength ### Vendor-Provided Components **HPE ProLiant Compute Gen12** [DAPM: Retained] Intel Xeon 6 / AMD EPYC. HPE iLO management silicon (HPE-owned). Improved perf/watt, security. Foundation for Private Cloud, Private Cloud AI, standalone. **HPE Cray EX4000/GX5000** [DAPM: Ceded] Exascale-class supercomputing. GX5000 unifies AI+HPC. Cray Slingshot interconnect. Liquid-cooled blade (GX240) with up to 16 NVIDIA Vera CPUs, 640 per rack. Deployed at Argonne, HLRS, HammerHAI. **HPE Cray Direct Liquid Cooling** [DAPM: Retained] Proprietary DLC supporting up to 400kW per rack with warm water operation. 100% DLC across GX5000 blades. As Blackwell/Vera Rubin density increases, cooling becomes the physical constraint — Cray heritage is genuine differentiator. **HPE Juniper Networking ($14B, July 2025)** [DAPM: Ceded] Full IP stack: Junos OS, MX routers, QFX switches, SRX firewalls, Mist AI-native ops, Apstra intent-based DC automation. Networking revenue 151.5% YoY to $2.7B Q1 FY2026. DC networking revenue up 380%+. Junos opinions are captive to the Juniper stack — there is no commodity substrate equivalent to x86 that would let the enterprise move these networking opinions to Aruba or Cisco without rebuilding. **HPE Cray Slingshot 400 Interconnect** [DAPM: Ceded] HPE-owned high-performance interconnect delivering 400 Gbps at scale with ultra-low tail latency for AI workloads. Distinct from NVIDIA InfiniBand — this is HPE networking IP for the supercomputing fabric. Slingshot opinions are captive to the Cray/HPE HPC stack — there is no commodity substrate that would let the enterprise move these fabric configurations to InfiniBand or Ethernet without rebuilding. **HPE Aruba Networking** [DAPM: Ceded] Campus and edge networking with AI-native Central platform. Being retooled on GreenLake Intelligence agentic mesh. Complementary to Juniper's DC focus and Slingshot's HPC fabric. Aruba opinions are captive to the Aruba/HPE stack — there is no commodity substrate that would let the enterprise move campus networking configurations to Juniper or Cisco without rebuilding. **HPE AI Factory (At-Scale + Sovereign)** [DAPM: Ceded] Full-stack AI infra: compute, GPUs, networking, liquid cooling, software, services. Blackwell (RTX PRO 6000 now) through Vera Rubin NVL72 (Dec 2026). Multi-tenancy via MIG with GPU passthrough (Spring 2026). Air-gapped configs for sovereign. NVIDIA Cloud Partner endorsed. STIG-hardened, FIPS-enabled. **Silicon Agnosticism (GX5000)** [DAPM: Retained] GX5000 supports NVIDIA AND AMD GPUs in the same rack architecture: GX440n blade (4 Vera CPUs + 8 Rubin GPUs), GX350a blade (1 AMD Venice CPU + 4 AMD MI430X GPUs), GX250 blade (8 AMD Venice CPUs, CPU-only). Up to 24 GPU blades per rack = 192 Rubin GPUs or 112 MI430X per rack. Neither Dell nor VAST offers multi-GPU-vendor blades in the same platform. **HPE Cray K3000 Storage System** [DAPM: Retained] First factory-built offering with embedded DAOS (Distributed Asynchronous Object Storage). Purpose-built I/O acceleration for AI/HPC workloads. Ships early 2026. Complements Alletra at the supercomputing tier. ### NVIDIA-Provided Components **NVIDIA GPU Silicon** RTX PRO 6000 Blackwell now. Vera Rubin NVL72 (72 Rubin GPUs, 36 Vera CPUs, NVLink, ConnectX-9, BlueField-4) Dec 2026. All AI acceleration depends on NVIDIA silicon. **NVIDIA Networking (InfiniBand, Spectrum-X)** Quantum-X800 InfiniBand for Cray GX5000 (144 ports, 800 Gb/s, 2027). ConnectX-9 SuperNICs, BlueField-4 DPUs, NVLink 6th-gen. Competes with HPE’s own Slingshot/Juniper/Aruba in AI fabric — structural tension. **NVIDIA Mission Control** AI Factory at-scale management planned for later 2026. GPU cluster operations, scheduling, resource allocation. HPE AI Factory will support Mission Control for large-scale deployments. ### Gap Analysis HPE's generic compute hardware is Retained: ProLiant servers and the underlying x86/NVIDIA silicon are a commodity substrate, so the enterprise can swap OEMs (Dell, HPE, Lenovo) without rebuilding workloads. But HPE's proprietary integrated systems are Ceded: the Cray EX/GX supercomputer and the full-stack AI Factory are turnkey, opinion-bearing platforms (proprietary blade form factor, Slingshot fabric, integrated software and services) that cannot be lifted to another vendor as deployed. The commodity-substrate test makes a generic server Retained; it does not make a proprietary supercomputer or a full-stack bundle Retained simply because commodity silicon sits inside it. HPE's networking is a different story. Juniper (Junos), Slingshot, and Aruba each represent deep proprietary opinion stacks with no commodity substrate equivalent. The enterprise cannot move Junos configs to Aruba gear, Slingshot fabric configs to InfiniBand, or Aruba campus configs to Juniper — each is captive to its own platform. All three score Ceded for the same reason Dell's PowerSwitch does. What HPE's networking depth does represent is capability breadth: three purpose-built fabrics (HPC, data center, campus/edge) vs. Dell's single NVIDIA Spectrum-X dependency vs. VAST's OEM networking reliance. The buyer gets more networking capability with HPE. They Cede it to three separate proprietary stacks rather than one. That is a different trade, not a better authority position. Two genuine differentiators remain: (1) Silicon agnosticism: GX5000 supports NVIDIA Rubin AND AMD MI430X in the same rack. Dell's AI Factory is NVIDIA-only under the primary SKU. VAST is NVIDIA-only. The enterprise retains GPU vendor optionality that peers don't offer. (2) Cray heritage for sovereign AI: national labs (Argonne), EU AI Factories (HammerHAI), government deployments where full stack traceability is required. This is a market position, not a DAPM distinction. The Cray K3000 with embedded DAOS adds HPC storage capability at the supercomputing tier that Dell Exascale and VAST DataStore do not provide in a factory-built form factor. ### Borrowed Judgment Generic compute: Low borrowed judgment. x86/NVIDIA is a commodity substrate, and ProLiant workloads are portable to Dell, Lenovo, or any other OEM (Retained). Proprietary integrated systems: the Cray EX/GX supercomputer and the AI Factory full-stack bundle are Ceded, their fabric, blade, and software opinions captive and not liftable as deployed. GPU silicon dependency is universal (NVIDIA or AMD); silicon agnosticism in GX5000 provides a hedge that Dell and VAST do not currently offer. Networking: High borrowed judgment across all three fabrics. Junos opinions are captive to Juniper. Slingshot fabric configs are captive to the Cray HPC stack. Aruba campus configs are captive to Aruba's platform. The enterprise Cedes networking authority to three separate HPE stacks. The depth of that networking capability is real; the authority position is Ceded across the board. ### Working Notes HPE Compute XD700 (OCP-inspired AI server on NVIDIA HGX Rubin NVL8, liquid-cooled, early 2027) targets neoclouds and service providers. Similar positioning to Dell's PowerRack but with OCP design philosophy. The three-tier networking portfolio (Slingshot for HPC, Juniper for DC, Aruba for campus/edge) is unique among the vendors assessed. Each fabric is purpose-built and deep — and each is captive to its own proprietary stack. The integration complexity (three platforms, three management tools) is the operational cost of that depth. Argonne, HLRS, HammerHAI (EU AI Factory), Hudson River Trading, and KISTI are named Cray GX5000 customers — reflecting a sovereign and hyperscale customer profile distinct from Dell's enterprise-focused AI Factory base. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Owned Storage + Governance Strength ### Vendor-Provided Components **HPE Alletra Storage MP X10000** [DAPM: Ceded] Disaggregated, all-flash, scale-out. Native file + object on single platform. 16 nodes, 23PB raw. 100% availability guarantee. RDMA-enabled for AI pipeline optimization across training, inference, KV cache. 2.5 PB/hr backup ingest. Storage opinions (configs, tiering, performance tuning) are captive to the HPE platform. **HPE Alletra Storage MP B10000** [DAPM: Ceded] Mission-critical block storage. 6-controller-node scaling (50% more perf vs 4-node). Dual-node fault tolerance. 5:1 data reduction guarantee. Real-time agentic support (v10.6.0): coordinated specialized AI agents for semantic understanding, adaptive reasoning, and prescriptive intelligence. Agents draw from system telemetry metadata, best practices, and accumulated product knowledge across installed base. Storage opinions captive to HPE platform. **HPE Data Fabric Software v8.1 (Ezmeral)** [DAPM: Ceded] Policy-based data placement and movement (tiering) across hybrid environments. Conversational interface and agentic AI assistant for natural language access to global namespace. Enhanced metadata integration for visibility, classification, lineage. Apache Polaris catalog support for Iceberg tables — consistent governance and compliance across platforms. Real-time S3-to-S3 object movement between any S3-compatible storage systems. Data Fabric governance opinions are captive to the HPE platform — Polaris provides open metadata portability but the orchestration layer is proprietary. **HPE Zerto Software** [DAPM: Ceded] Continuous data protection, AI-powered assistant, Microsoft Defender integration, live VMware-to-HPE VM migration. Near-zero RPO/RTO. Replication and recovery opinions are captive to the Zerto/HPE platform. ### NVIDIA-Provided Components **GPU-Accelerated Storage Integration** RDMA via CX-8/CX-9 SuperNICs for GPU-direct storage access. Same acceleration Dell and VAST also use. ### Gap Analysis HPE’s Layer 1A is a capable storage foundation with genuine HPE-owned governance intelligence. Three characteristics position it in the 4+1 model: First, the B10000’s agentic support architecture (v10.6.0) goes beyond predictive analytics into semantic understanding and adaptive reasoning — a coordinated set of specialized AI agents drawing from telemetry metadata and accumulated product knowledge. This is HPE-owned intelligence at the storage layer, architecturally aligned with GreenLake Intelligence’s domain-specific agent model. Dell’s storage management is infrastructure monitoring (CloudIQ, MetadataIQ indexing). VAST’s Element Store enriches metadata inline at write time. Three different approaches to storage intelligence. Second, Data Fabric v8.1 with Apache Polaris catalog for Iceberg tables provides cross-platform governance that participates in open-standard ecosystems. Dell’s MetadataIQ indexes within Dell storage boundaries. VAST’s Catalog indexes within the VAST namespace. HPE’s Polaris support means governance metadata is portable across platforms — a federated approach vs Dell’s and VAST’s platform-bounded approaches. Third, Data Fabric’s real-time S3-to-S3 object movement enables AI data ingestion from any S3-compatible source into the governed Data Fabric environment. This addresses the heterogeneous enterprise data ingestion problem — similar in function to VAST’s SyncEngine (which ingests from Google Drive, Jira, Confluence, S3) but operating at the storage protocol level rather than the application API level. X10000’s unified file+object on one platform reduces the number of storage engines vs Dell’s portfolio approach (PowerScale for file, ObjectScale for object, Exascale for combined). VAST’s Element Store goes further by collapsing file, object, table, and vector into a single data structure. HPE’s consolidation is at the platform level; VAST’s is at the data structure level. Calibration under the codified criteria (July 2026): HPE owns more opinion-bearing storage and governance software than any OEM peer — Alletra X10000/B10000, Data Fabric with the open Polaris catalog, Zerto — and the whose-paper rule credits it in full. With Lenovo’s 1A at strong on an owned high-end platform plus a productized partner menu, holding HPE at moderate made the grade mean different things across rows. Promoted to strong on owned-platform depth plus a genuine governance layer; the frontier comparison to Dell’s MetadataIQ-anchored strong now reads as different governance philosophies (federated/open-catalog vs. platform-bounded), not a grade difference. ### Borrowed Judgment Low to moderate. HPE owns storage platforms (Alletra X10000, B10000), Data Fabric software, and Zerto outright. GPU acceleration for storage I/O depends on NVIDIA networking silicon (CX-8/CX-9), but the storage intelligence — policy engine, metadata, agentic management agents — is HPE IP. Apache Polaris support is a deliberate governance strategy: by using an open standard for metadata catalog, HPE reduces governance vendor lock-in for its customers. Compare to VAST, where the governance catalog is proprietary (Ceded to VAST). The trade-off: HPE’s open-standard approach is more portable but less deeply integrated; VAST’s proprietary approach is tightly integrated but less portable. ### Working Notes Commvault and Veeam partnerships add data resilience capabilities (Delegated partners at Layer 1A). The agentic support in B10000 is distinct from GreenLake Intelligence: B10000 agents are storage-domain specialists drawing from storage telemetry and product knowledge. GreenLake Intelligence agents are cross-domain (networking + storage + compute). The two agent architectures are designed to complement each other — B10000 agents resolve storage-specific issues autonomously while GreenLake Intelligence correlates cross-domain patterns. Whether these agent systems actually interoperate via MCP or operate independently is an open question. The Data Fabric’s real-time S3 ingestion capability addresses a practical enterprise challenge: AI teams need to pull data from diverse S3-compatible sources (AWS, MinIO, other object stores) into a governed environment for AI pipeline consumption. This is not a differentiating capability on its own (any S3-compatible system can ingest from S3) but the governance integration — data lands in the Data Fabric namespace with policy-based placement and lineage tracking — is the value. Watch-list (HPE Discover 2026, not scored): Alletra MP X10000 auto metadata + governance policies and AI-ready pipelines - Q4 2026. ## ◑ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Reference Stack + Partner Orchestration ### Vendor-Provided Components **HPE Ezmeral Data Fabric (Retrieval Surface)** [DAPM: Ceded] Global namespace with conversational access for AI-driven retrieval. Federates data across hybrid environments. Natural language queries against namespace. The discovery layer that retrieval pipelines query — Data Fabric knows where data is and what policies govern it. Proprietary retrieval surface captive to HPE platform. **HPE Alletra X10000 RDMA Storage** [DAPM: Ceded] Low-latency file and object access for AI inference pipelines. RDMA via CX-8/CX-9 reduces retrieval latency for RAG. KV cache storage support for inference state persistence. The storage substrate that retrieval reads from. Proprietary storage platform captive to HPE. **Kamiwaza Context Orchestration (via Unleash AI)** [DAPM: Ceded] In Town of Vail: manages the full context pipeline for document-centric use cases. Identifies documents requiring processing (Section 508 compliance), ingests and extracts content (housing deeds), prepares contextual inputs for agent consumption. Determines what context each agent needs, from which sources, under what governance constraints. This is retrieval orchestration — above the storage layer, below the agent runtime. Governs cross-departmental context routing (legal, housing, admin). Lift-to-leave: the context-orchestration opinions are Kamiwaza-captive — Ceded to the chosen partner, as on Kamiwaza's own row. ### NVIDIA-Provided Components **NVIDIA NeMo Retriever** Embedding models (NV-EmbedQA-E5-v5, Mistral7B-v2, Arctic-Embed-L) and reranking in unified microservice. GPU-accelerated retrieval for RAG pipelines on Private Cloud AI. Provides the embedding intelligence that HPE’s storage does not. **NVIDIA AI-Q Blueprint** Research assistant and enterprise data agent blueprint. Connects enterprise data to AI agents via retrieval pipelines. Available on Private Cloud AI. **NVIDIA RAG Blueprint + Milvus** HPE’s reference RAG architecture uses NeMo Retriever for embedding, Milvus (open-source) for vector database, LangChain for chain serving. HPE does not own any retrieval intelligence component — it provides storage substrate and deployment platform. ### Gap Analysis Layer 1B is HPE’s thinnest proprietary layer. HPE provides storage infrastructure (Data Fabric namespace, Alletra RDMA) but does not own a vector database, an embedding engine, or a retrieval framework. The retrieval intelligence stack is entirely NVIDIA (NeMo Retriever) + open source (Milvus, LangChain). The three-vendor comparison at Layer 1B: • Dell: storage (PowerScale/ObjectScale) + Elastic (search intelligence, Delegated ISV) + NVIDIA (cuVS acceleration). Three authorities. • HPE: storage (Alletra/Data Fabric) + NVIDIA (NeMo Retriever, embedding) + open source (Milvus, LangChain). No proprietary retrieval intelligence. When Kamiwaza is added via Unleash AI, it provides governed retrieval orchestration above the storage and embedding layers. • VAST: storage + embedding + vector search + retrieval pipeline all in one platform (InsightEngine, native vector search, DataBase). One authority. HPE’s retrieval gap is structural: the company has no analog to Dell’s Elastic partnership or VAST’s native InsightEngine. This is a deliberate architectural choice — HPE provides infrastructure substrate and delegates retrieval intelligence to NVIDIA and open-source components. When Kamiwaza enters via Unleash AI, the retrieval story changes. Kamiwaza provides governed context orchestration that neither the storage layer nor the NVIDIA retrieval components provide independently: cross-departmental context routing, authority-constrained retrieval, and document-pipeline coordination. In the Town of Vail, this means the Section 508 compliance agent receives only the documents it’s authorized to process, with retrieval governed by department boundaries. This is a layer of retrieval intelligence that storage-native search (VAST) and embedding-accelerated search (Dell+Elastic) don’t address — the governance of who receives what context under what authority. The 4+1 model question: is governed context orchestration a Layer 1B function (retrieval) or a Layer 2C function (governance)? Kamiwaza’s context management spans both — it retrieves content (1B) according to governance policies (2C). The assessment classifies the retrieval function at 1B and the governance function at 2C. ### Borrowed Judgment Moderate to high. HPE’s own Layer 1B authority is limited to storage infrastructure. Embedding intelligence is NVIDIA (NeMo Retriever). Vector storage is open source (Milvus). Retrieval framework is open source (LangChain). Governed context orchestration is Kamiwaza (Ceded to the partner via Unleash AI). Compare to Dell: Dell delegates retrieval intelligence to Elastic (proprietary ISV partnership) and acceleration to NVIDIA. Dell’s borrowed judgment at 1B is Moderate — split between a proprietary ISV and NVIDIA. Compare to VAST: VAST’s borrowed judgment at 1B is Low — InsightEngine, vector search, and the retrieval pipeline are VAST IP. Only embedding model execution (NIM) is NVIDIA-provided, and InsightEngine is model-agnostic. HPE’s Layer 1B borrowed judgment is the highest of the three vendors because HPE owns the least retrieval IP. The mitigation: Kamiwaza’s governed orchestration adds a unique capability that pure retrieval engines don’t provide. ### Working Notes The HPE Developer Portal’s RAG reference architecture is instructive: NeMo Retriever embedding + Milvus vector DB + LangChain + Llama3-70B. This is a standard NVIDIA reference stack, not an HPE-differentiated architecture. Any NVIDIA partner (Dell, Lenovo, Supermicro) could deploy the identical stack. HPE’s Layer 1B differentiation comes not from the retrieval stack but from the storage substrate below it (Data Fabric governance, Alletra RDMA performance) and the orchestration layer above it (Kamiwaza context governance). The KV cache storage support in Alletra X10000 is worth noting as a Layer 1B/2B bridge: inference state persistence in storage allows agents to maintain context across sessions without holding GPU memory. Dell’s equivalent is the CMX KV cache offload (NVIDIA technology). HPE’s is storage-native. VAST’s CNode-X collocates cache and compute. ## ◑ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** HPE + Open Source ### Vendor-Provided Components **HPE Data Fabric Software (Pipeline Orchestration)** [DAPM: Ceded] Policy-based data placement considering performance, sovereignty, costs, compliance. Data lineage and compliance tagging. Agentic AI assistant for automated reporting and data placement decisions. Real-time S3-to-S3 object movement for AI data ingestion from external sources. Proprietary pipeline orchestration captive to HPE platform. **HPE Ezmeral Unified Analytics** [DAPM: Delegated] Enterprise-hardened packaging of the full open-source ML pipeline stack: Apache Airflow (workflow orchestration), Kubeflow (ML pipelines + model serving via KServe), Ray (distributed compute), Feast (feature store), MLflow (experiment tracking), Apache Spark (data engineering), Presto SQL (federated query), Apache Superset (visualization). Connectors to Snowflake, MySQL, Delta Lake, Teradata, Oracle. Built through acquisitions: BlueData (2018), MapR (2019), Ampool (2021), Arrikto/Kubeflow team (2023). Open-source substrate — enterprise could swap to alternative packaging of the same tools. **HPE Morpheus Software** [DAPM: Ceded] Hybrid and multicloud management, orchestration, migration, automation. VMware-to-HPE VM migration paths. Cloud-native workflow orchestration. **Kamiwaza Workflow Data Pipelines (via Unleash AI)** [DAPM: Ceded] In Town of Vail: orchestrates governed data flows across department boundaries — housing deeds from ingestion through verification to audit across legal, housing, and admin functions. Decision-driven data movement where pipeline logic is governed by authority constraints and compliance policy, not static ETL schedules. Lift-to-leave: the decision-flow and pipeline opinions are Kamiwaza-captive — Ceded to the chosen partner. ### NVIDIA-Provided Components **NVIDIA RAPIDS Accelerator for Apache Spark** GPU-accelerated data prep, model training, and visualization within Ezmeral Unified Analytics. Up to 29x faster development. Spark acceleration is the primary NVIDIA contribution at Layer 1C. **NVIDIA Blueprints** Pre-built AI application patterns deployed on Private Cloud AI. Pipeline templates, not pipeline infrastructure. ### Gap Analysis Three vendors, three distinct architectural strategies for Layer 1C: • Dell: acquired Dataloop (proprietary orchestration, no-code/low-code). Dell’s strongest software move, but the broader pipeline layer depends on ISV partners (ClearML, DataRobot, Starburst). Multiple authority boundaries. • HPE: packages the full open-source ML pipeline lifecycle (Airflow → Kubeflow → Ray → Feast → MLflow → Spark) under enterprise-grade guardrails. Four acquisitions (2018–2023) demonstrate deliberate investment. Value is in curation, hardening, support, integration — not proprietary technology. • VAST: built a proprietary DataEngine (event-driven serverless execution on CNodes). Entirely VAST IP. Tightly integrated with storage and retrieval layers. One authority. HPE’s open-source approach creates a specific DAPM trade-off: the enterprise avoids vendor lock-in (Airflow and Kubeflow are portable), but HPE’s authority is in packaging rather than core technology. If Apache Airflow’s community changes direction, HPE is affected. This is a different risk profile than Dell’s (proprietary Dataloop, partner-dependent beyond it) or VAST’s (proprietary DataEngine, VAST-dependent entirely). Data Fabric’s policy-based movement with Apache Polaris governance connects Layer 1C to Layer 2C: data moves according to explicit policies that consider performance, data locality, sovereignty, costs, and compliance. This governance-aware data movement feeds both GreenLake Intelligence (infrastructure decisions) and Kamiwaza (AI workload decisions). Dell’s Dataloop provides orchestration without integrated governance policy. VAST’s DataEngine has Event Broker for data-event-driven movement without an explicit policy engine. When Kamiwaza is selected via Unleash AI, it adds decision-driven pipeline capability: documents move across department boundaries based on decision logic (legal review required? accessibility compliance met? authority approval needed?). This connects infrastructure-level pipeline capabilities to business-level decision flows — where Layer 1C meets Layer 2C. ### Borrowed Judgment Low for pipeline packaging and integration (HPE owns Ezmeral, Data Fabric, Morpheus). Underlying components are open-source, limiting deep technical authority but also limiting NVIDIA dependency — RAPIDS for Spark is the only NVIDIA contribution at this layer. Open-source components are substitutable by the enterprise without HPE’s permission. Compare to Dell: Dell owns Dataloop (Retained) but depends on partners for everything else at Layer 1C. Four authority boundaries. Compare to VAST: VAST owns everything at Layer 1C (Retained by VAST, Ceded by the enterprise). One authority but total vendor dependency. HPE’s Layer 1C authority model is distinct: HPE curates and supports, the enterprise can substitute, NVIDIA acceleration is additive not required. ### Working Notes The acquisition history (BlueData 2018, MapR 2019, Ampool 2021, Arrikto 2023) shows deliberate multi-year investment in the data pipeline layer. HPE chose to build this capability rather than delegate entirely to partners. Dell’s Dataloop acquisition is a similar strategic move but more recent (2024) and narrower in scope. The NVIDIA RAPIDS Accelerator for Spark (up to 29x faster) is meaningful but optional — Ezmeral runs without GPU acceleration. Same pattern as VAST’s DataEngine (runs on standard CNodes, CNode-X adds GPU acceleration). Data Fabric’s real-time S3-to-S3 movement bridges Layer 1A and 1C: ingest from external S3-compatible sources into the governed namespace, where policy-based placement takes over. ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** HPE Strength ### Vendor-Provided Components **HPE GreenLake Cloud Platform** [DAPM: Ceded] Consumption-based hybrid cloud (4th gen). Unified VM + K8s management. Pay-per-use AI infrastructure. Dashboard for capacity, utilization, cost. Self-service cloud experience with full lifecycle management. Proprietary management platform — orchestration opinions captive to HPE. **HPE GreenLake Intelligence (Agentic AI Mesh)** [DAPM: Ceded] HPE-owned agentic AI framework infused across the entire hybrid stack (not a standalone product). Multiple domain-specific LLMs trained on HPE data, communicating via Model Context Protocol (MCP). Agents form an agentic mesh for inter-agent communication with secure, contextual data sharing. Domain agents for networking (Aruba/Juniper), storage (Alletra), compute (OpsRamp), orchestration, and FinOps. Cross-domain correlation: traces performance issues across application to storage to network chain. Can process real-time infrastructure metrics and execute actions across multiple vendor environments. Human-in-the-loop: agents take action subject to approval. Proprietary agentic framework captive to HPE platform. **HPE OpsRamp Software** [DAPM: Ceded] Multi-domain agentic system coordinating compute, network, storage, virtualization, and software layers. Use cases: root-cause analysis, explainability, capacity planning. AI-driven alerts, incident management, GPU monitoring, workload observability. Operations copilot with conversational product help and agentic command center. CrowdStrike integration for security monitoring. MCP support for connecting to GreenLake Intelligence and third-party tools. Proprietary management platform captive to HPE. **HPE Alletra X10000 MCP Servers (Native)** [DAPM: Ceded] Model Context Protocol servers built natively into X10000 storage. Enables GreenLake Intelligence agents to communicate directly with storage for data management orchestration. Connects storage operations to the broader agentic mesh via GreenLake Copilot and natural language interfaces. MCP is an open protocol but these servers are captive to Alletra/GreenLake. **HPE Compute Ops Management** [DAPM: Ceded] Cloud-native server lifecycle management for ProLiant fleet. Compute Copilot for AI-assisted infrastructure operations. Proprietary management platform captive to HPE. **HPE Private Cloud (4th Gen)** [DAPM: Ceded] K8s management with ProLiant Gen12. Unified cloud-native + virtualized workload management. Independent scaling for cloud-native workloads. Upgrade path to Morpheus for hybrid/multicloud. ### NVIDIA-Provided Components **NVIDIA MIG** GPU fractionalization for multi-tenancy in AI Factory portfolio. **NVIDIA Mission Control** AI Factory at-scale management. Planned later 2026. GPU cluster ops, scheduling, resource allocation. **NVIDIA AI Enterprise (Runtime Management)** Lifecycle management for AI software stack. Pre-integrated with ProLiant. ### Gap Analysis Layer 2A is where HPE makes its most substantial software authority claim. GreenLake Intelligence is not a rebranded monitoring tool — it is an agentic AI framework with domain-specific LLMs communicating via MCP, designed to be infused across the entire HPE hybrid stack. Four characteristics define HPE’s Layer 2A position: (1) Cross-domain agentic correlation. GreenLake Intelligence agents trace performance issues across application → storage → network chains, coordinating remediation across domains. Dell’s OpenManage and NVIDIA’s Run:ai operate within single domains (rack management and GPU scheduling respectively). VAST’s Polaris orchestrates VAST clusters but not the broader infrastructure around them. (2) MCP as the inter-agent communication standard. GreenLake Intelligence is compliant with MCP, enabling connection to third-party agents and devices. The X10000 has native MCP servers built in. This means the agentic mesh is architecturally open — ITSM systems can collaborate with GreenLake, and third-party infrastructure can be brought under GreenLake management. HPE positions this as ‘the mesh is open to more stitches.’ (3) Multi-vendor infrastructure support. NAND Research notes that GreenLake Intelligence agents can process real-time metrics and execute actions across multiple vendor environments, not exclusively HPE hardware. This extends Layer 2A authority beyond HPE’s own equipment — a broader orchestration scope than Dell’s OpenManage (Dell hardware only) or VAST’s Polaris (VAST clusters only). (4) FinOps agent for workload placement. Orchestration, networking, and FinOps agents collaborate to determine workload placement across private and public clouds. This is an economic placement decision — where should this workload run based on cost, performance, and policy? This function overlaps with Layer 2C territory. Dell’s Layer 2A is split between Dell-managed rack deployment (OpenManage) and NVIDIA-managed GPU scheduling (Run:ai). HPE’s Layer 2A is unified under GreenLake with OpsRamp providing a multi-domain agentic system that coordinates compute, network, storage, virtualization, and software layers. GreenLake’s consumption model (pay-per-use) creates a natural authority surface: HPE maintains an ongoing operational relationship with the infrastructure — metering, capacity management, utilization optimization — that traditional capex purchases don’t provide. The gap: GPU-specific scheduling is still Ceded to NVIDIA (MIG for fractionalization, Mission Control for at-scale management, planned later 2026). HPE orchestrates the infrastructure around the GPU cluster; NVIDIA orchestrates inside it. This is the same Layer 2A boundary as Dell, but HPE’s surrounding orchestration is a unified platform rather than separate point tools. ### Borrowed Judgment Low for infrastructure orchestration (GreenLake platform, GreenLake Intelligence, OpsRamp, Compute Ops Management are HPE-owned IP). Moderate for GPU-specific scheduling (MIG, Mission Control are NVIDIA-controlled). The NAND Research caveat is relevant for the DAPM assessment: tight coupling with GreenLake creates potential vendor lock-in for organizations with diverse infrastructure portfolios. However, MCP compliance and multi-vendor agent support partially mitigate this concern — the agentic mesh can extend beyond HPE hardware. Compare to Dell: Dell’s infrastructure orchestration is fragmented (OpenManage for servers, separate tools for storage and networking). GPU scheduling is fully NVIDIA-controlled (Run:ai). No agentic cross-domain correlation. Compare to VAST: Polaris provides fleet-level VAST cluster orchestration (Retained by VAST). DataEngine provides workload scheduling within the data platform. But Polaris doesn’t orchestrate non-VAST infrastructure. GreenLake Intelligence’s multi-vendor, multi-domain scope is broader. ### Working Notes GreenLake Intelligence’s cross-domain correlation and FinOps-aware placement push beyond traditional Layer 2A into Layer 2C territory for IT operations. The assessment classifies it as spanning 2A–2C for IT ops: infrastructure orchestration (2A) + cross-domain governance decisions and economic placement reasoning (2C). This dual classification is important — GreenLake Intelligence is both an orchestrator and a decision-maker. The MCP openness is architecturally significant: by using an open protocol for agent communication, HPE enables third-party integration without custom APIs. ITSM systems (ServiceNow, BMC) can collaborate with GreenLake agents. Third-party infrastructure can be managed. This is an open-ecosystem approach to infrastructure orchestration that Dell’s proprietary OpenManage and VAST’s proprietary Polaris don’t provide. The X10000 native MCP servers represent infrastructure-level agent communication — the storage array itself participates in the agentic mesh as a first-class agent endpoint, not just a managed resource. This is a specific implementation of the 4+1 model’s vision of infrastructure that is natively agent-aware. Watch-list (HPE Discover 2026, not scored): Morpheus Central + intent-based closed-loop network automation + GreenLake Intelligence agent orchestration/copilots - rolling out Q2-Q3 2026 through 2027. ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** NVIDIA Runtime + Framework Surface ### Vendor-Provided Components **HPE Private Cloud AI** [DAPM: Ceded] Co-engineered with NVIDIA. Pre-configured HW+SW stack with four right-sized configurations. Air-gapped capable. Scales to 128 GPUs with network expansion racks. OpsRamp integration for AI workload monitoring. Supports NVIDIA AI-Q, Omniverse, NeMo Retriever blueprints. Multi-tenancy via MIG with GPU passthrough (Spring 2026). The NVIDIA runtime is integral to the stack — the enterprise cannot substitute an alternative inference runtime without rebuilding. **HPE Ezmeral Unified Analytics (ML Runtime)** [DAPM: Delegated] Kubeflow for ML pipeline execution and model serving (KServe). Ray for distributed compute. MLflow for experiment tracking. Enterprise packaging of open-source ML runtime tools. Open-source substrate — enterprise could swap to alternative packaging. **Agentic Framework Ecosystem on Private Cloud AI** [DAPM: Delegated] CrewAI integration enables enterprises to build multi-agent solutions on Private Cloud AI. Deloitte Zora AI for Finance deploys on Private Cloud AI as an agentic platform for dynamic executive reporting (financial statement analysis, scenario modeling, competitive analysis). NIM Agent Blueprints provide pre-built agentic workflows. This is an emerging multi-framework agentic surface — not HPE-owned runtime IP but HPE-curated deployment options. **Kamiwaza Agent Execution Coordination (via Unleash AI)** [DAPM: Ceded] In Town of Vail: coordinates multiple specialized agents — accessibility identification, alt text generation, remediation guidance formatting. Determines which agents run, in what sequence, with what inputs, under what constraints. Enforces human-in-the-loop checkpoints. Manages agent lifecycle at the execution layer. ‘Coordination of multiple specialized AI agents and human workflows to execute multi-step decisions under explicit authority, governance, and audit constraints.’ Lift-to-leave: the execution-coordination opinions are Kamiwaza-captive — Ceded to the chosen partner. ### NVIDIA-Provided Components **NVIDIA AI Enterprise Software** The AI workload runtime: model serving (Triton), guardrails (NeMo Guardrails), distributed inference, training frameworks (NeMo). Pre-integrated on Private Cloud AI. STIG-hardened, FIPS-enabled for sovereign deployments. **NVIDIA NIM Microservices** Pre-optimized inference microservices for model deployment. Part of AI Enterprise platform. Available to all NVIDIA partners (Dell, Cisco, Lenovo) — not HPE-specific. **NVIDIA NIM Agent Blueprints** Pre-built agentic AI application patterns: Multimodal PDF Data Extraction, Digital Twins (Omniverse), AI-Q for enterprise data agents. Deployed on Private Cloud AI. Same blueprints available on Dell AI Factory, Cisco HyperFabric, Lenovo Hybrid AI. ### Gap Analysis Layer 2B reveals a more nuanced runtime architecture than Dell’s because authority is distributed across three actors rather than two. The three-actor model: • NVIDIA provides model execution — Triton (model serving), NeMo Guardrails (safety), NIM (optimized inference), NeMo (training). This is the compute execution layer. Identical across all NVIDIA partners (Dell, Cisco, Lenovo deploy the same stack). • HPE provides the deployment platform — Private Cloud AI (hardware, cooling, lifecycle), Ezmeral (ML runtime packaging), and increasingly an agentic framework surface (CrewAI, Deloitte Zora AI). This is infrastructure + curation. • Kamiwaza provides agent execution coordination (via Unleash AI) — determines which agents run, sequences execution, manages inputs/outputs, enforces execution-time constraints (authority boundaries, audit, human-in-the-loop). This is governance-aware agent coordination above model inference but below Layer 2C policy. This creates a layered runtime: NVIDIA executes individual model inference → Kamiwaza coordinates multi-agent workflows and enforces execution governance → Layer 2C (also Kamiwaza) makes policy decisions about what should run where. The 2B/2C boundary: 2B is execution coordination (how agents run), 2C is decision authority (why agents run, under what governance). The structural comparison across vendors: • Dell: NVIDIA at 2B (model execution + NemoClaw/OpenShell agent runtime). No agent coordination layer beyond NVIDIA. Dell provides packaging and services. • HPE: NVIDIA at model execution + Kamiwaza at agent coordination + CrewAI/ISV frameworks for agent building. Three layers of runtime capability from three sources. HPE provides infrastructure + curation. • VAST: AgentEngine provides a unified agent runtime (execution + coordination + lifecycle + observability) as VAST IP. NVIDIA provides GPU acceleration only. One authority. HPE’s ‘NVIDIA AI Computing by HPE’ branding signals co-engineering, but the DAPM question is precise: can HPE modify, extend, or replace NVIDIA runtime components independently? The answer appears to be no — ‘co-engineering’ means deeper integration and joint validation, not shared IP authority. NVIDIA controls the runtime; HPE controls the platform it runs on. The emerging agentic framework ecosystem (CrewAI, Deloitte Zora AI) on Private Cloud AI is worth noting: HPE is becoming a multi-framework agentic deployment surface, not locked to a single agent runtime. This is a platform strategy — provide the substrate that multiple agentic frameworks can run on — rather than a runtime strategy (build the definitive agent runtime, as VAST is attempting with AgentEngine). ### Borrowed Judgment High for AI workload runtime. NVIDIA controls model serving, inference optimization, guardrails, and training frameworks. The same NVIDIA AI Enterprise stack runs on Dell, Cisco, and Lenovo — this is not HPE-specific technology. The mitigating factor is the ‘bracketing’ architecture: HPE provides governance at Layer 2A (GreenLake Intelligence, HPE-owned) and sources Layer 2C governance from a chosen partner (Kamiwaza via Unleash AI). The NVIDIA-controlled Layer 2B runtime is sandwiched between two governance layers. The enterprise has governance coverage even where it doesn’t control execution — coverage, not authority: both bracket surfaces are Ceded from the customer's seat, one to HPE and one to Kamiwaza. Compare to Dell: Dell has NVIDIA at 2B with no governance brackets. No Layer 2C (Absent). Layer 2A is split between Dell (OpenManage) and NVIDIA (Run:ai). The enterprise has neither governance authority above nor unified governance authority below the NVIDIA runtime. Compare to VAST: VAST owns AgentEngine (2B) and is building PolicyEngine (2C). No bracketing needed because VAST controls both the runtime and the governance layer. The enterprise Cedes both to VAST. HPE’s borrowed judgment at 2B is the highest of any layer in the HPE assessment. The bracketing architecture is the mitigation, not the solution. ### Working Notes The CrewAI and Deloitte Zora AI integrations signal that Private Cloud AI is evolving from a single-stack NVIDIA deployment platform into a multi-framework agentic surface. This is architecturally different from both Dell’s approach (NVIDIA-only runtime) and VAST’s approach (proprietary-only runtime). HPE is positioning Private Cloud AI as the substrate that multiple agent frameworks deploy on. The NIM Agent Blueprints (PDF Extraction, Digital Twins, AI-Q) are available identically on Dell AI Factory, Cisco HyperFabric, and Lenovo Hybrid AI. These do not differentiate HPE at Layer 2B. HPE’s differentiation comes from the bracketing architecture (2A and 2C governance around the NVIDIA 2B runtime) and the emerging multi-framework agent deployment model. The bracketing architecture has a structural analog in HPE’s networking story: NVIDIA InfiniBand handles GPU-to-GPU interconnect (HPE doesn’t control it), but HPE’s Juniper/Aruba/Slingshot handles everything around the GPU fabric (HPE owns it). The pattern: cede the NVIDIA-specific function, retain authority over everything surrounding it. ## ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** IT-Ops Reasoning + Partner AI Orchestration ### Vendor-Provided Components **GreenLake Intelligence (IT Operations 2C)** [DAPM: Ceded] MCP-based agent communication across infrastructure domains. Domain-specific agents for networking, storage, compute, operations. Cross-domain correlation and autonomous remediation. Layer 2C for IT infrastructure operations — but not for AI workload placement and policy. Proprietary agentic framework captive to HPE platform. **Kamiwaza (AI Workload 2C, via Unleash AI)** [DAPM: Ceded] Agentic orchestration and decision routing for AI workloads. Policy-driven placement, mission decomposition, decision authority placement, cross-agent governance. Distributed Data Engines process data at the source without moving it or compromising security. Cross-environment evaluation: rather than treating anomalies as isolated alerts, Kamiwaza evaluates what else is happening across the environment to determine appropriate response. Town of Vail production agents: ARIA (accessibility auditing — independently audits websites, identifies Section 508 issues, provides developer fixes in days vs $1.5M and months for manual audits). Deed restriction processor (reviews documents spanning 60 years, extracts key data, answers compliance questions, generates Excel/PDF reports — work that previously required weeks of manual review). Fire detection coordinator (works with Vaidio/ProHawk video AI, evaluates cross-environment context, triggers workflows, supports operators as conditions change). HPE's chosen Layer 2C for AI workload orchestration, delivered under single-accountable-provider model. Lift-to-leave: the reasoning-plane opinions — orchestration flows, decision routing, ReBAC governance configurations — are Kamiwaza-captive and cannot be swapped without rebuilding. Deliberate choice and single accountability describe the procurement, not the exit: Ceded to the chosen partner, as on Kamiwaza's own row. **HPE Data Fabric Policy Engine** [DAPM: Ceded] Policy-based data placement considering performance, sovereignty, costs, compliance. Apache Polaris for cross-platform governance. Feeds governance signals into both GreenLake Intelligence (infrastructure) and Kamiwaza (AI workloads). Proprietary policy engine captive to HPE platform — Polaris provides open metadata portability but the policy orchestration layer is proprietary. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** GreenLake Intelligence and Kamiwaza are HPE-owned and partner-provided (Ceded to Kamiwaza) respectively. NVIDIA does not control the governance, placement, or policy reasoning layer in the HPE stack. ### Gap Analysis This is the most analytically interesting layer in the HPE assessment and where HPE’s approach diverges most from Dell’s. HPE has a two-part Layer 2C story: GreenLake Intelligence provides Layer 2C for IT infrastructure operations: correlates signals across networking, storage, compute to diagnose and resolve infrastructure issues. Routes decisions across domains. Takes autonomous action (subject to approval). Uses MCP for agent communication. HPE-owned IP. Kamiwaza provides Layer 2C for AI workload orchestration: agentic orchestration, decision routing, policy-driven placement, cross-agent governance. Distributed Data Engines process data at the source without data movement. Not an accidental ISV partnership — HPE’s deliberate architectural choice, curated, integrated, validated, and delivered under single-accountable-provider model. The Town of Vail as by-proxy Kamiwaza assessment — specific evidence: • ARIA accessibility agent: independently audits municipal websites, identifies Section 508 issues, provides developer fixes. Manual equivalent: $1.5M, months of work. ARIA delivers in days. This demonstrates decision automation with governance (the agent identifies what needs fixing, recommends how, but human developers implement). • Deed restriction processor: reviews housing documents spanning 60 years in disjointed legacy formats, extracts key data, answers compliance questions, generates reports. Previously required weeks of manual review and data entry in Excel. Single processing errors carry serious legal and financial consequences. This demonstrates governance-aware document intelligence with cross-departmental implications (legal, housing, administrative). • Fire detection coordinator: works with Vaidio video analytics and ProHawk enhanced vision. Rather than treating a video anomaly as an isolated alert, Kamiwaza evaluates what else is happening across the environment and determines the appropriate response. Agents surface relevant information, trigger correct workflows, and support operators as conditions change. This demonstrates cross-environment evaluation — the core Layer 2C pattern of reasoning across multiple data sources and agent outputs. • Deployment velocity: concept to first-phase production in three months. 20–30 additional use cases projected in first year. Additional use cases compose from existing primitives (decision flows, authority boundaries, governance constraints) rather than requiring new infrastructure. • Economic model: fixed-cost infrastructure on the town’s own solar/wind-powered data center. No cloud API token pricing. Billions of tokens without variable costs. • RBAC → REBAC governance: emerged from Kamiwaza’s production behavior. Traditional role-based access breaks when autonomous agents operate across department boundaries. Relationship-Based Access Control constrains agent permissions based on context, not just role. Structural comparison: • Dell: Layer 2C absent. No partner fills this role. No plan visible. • HPE: Layer 2C provided for IT ops by HPE-owned GreenLake Intelligence and for AI workloads by Kamiwaza (Ceded to the chosen partner) — validated in production with named agents and measurable outcomes. • Google: Layer 2C productized and shipping (Inference Gateway + DWS + Knowledge Catalog) — Ceded from the customer’s seat, per the GCP row. • VAST: Layer 2C Gap, emerging (Polaris ships as placement abstraction; PolicyEngine + TuningEngine GA end of 2026). Announced at VAST Forward 2026, GA end of 2026. ### Borrowed Judgment IT ops Layer 2C: GreenLake Intelligence is HPE-owned IP — low borrowed judgment for HPE, but the customer's ops-reasoning opinions are captive to the HPE platform. AI workload Layer 2C: High — Ceded to Kamiwaza. Deliberately chosen, integrated, and delivered under HPE's accountability, but the platform's opinions (orchestration flows, decision routing, ReBAC configurations) are Kamiwaza-captive and cannot be swapped without rebuilding. Structurally stronger than Dell's position on the capability axis because the function exists and someone is accountable; the dependency is captive. The trade is capability-with-capture versus Dell's no-capability-no-capture. ### Working Notes Strategic question: is partner-provided Layer 2C transitional (HPE eventually builds/acquires orchestration IP) or permanent (HPE’s value is ecosystem curation, not owning every layer)? Town of Vail evidence suggests HPE is comfortable with the ecosystem model — and that it works operationally. The RBAC → REBAC governance shift from Town of Vail validates the 4+1 model’s claim that Layer 2C requires governance architecturally distinct from Layer 2A infrastructure RBAC. Watch-list (HPE Discover 2026, not scored): native agent registry + governance via GreenLake Intelligence, and Private Cloud AI secure cross-framework agent registration - core GA July 2026, agentic observability/data intelligence Q4 2026. Directionally shifts HPE 2C from Kamiwaza-provided toward HPE-native, but it is Intelligence-2C (governance), not placement. Revisit at GA. ## ◇ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Unleash AI Ecosystem ### Vendor-Provided Components **HPE Unleash AI Program (26+ ISV Members)** [DAPM: Delegated] Curated (not open) ISV partner ecosystem. HPE is ‘highly selective’ with its ISV pool. 26+ members focused on different AI use cases from vision AI to agentic analytics. Validates interoperability, provides unified deployment and support. Positions HPE as ‘one accountable provider.’ Field motion targets decision friction, not infrastructure features. Training, certifications, and enablement support for partners. India-based partners taking locally developed AI use cases into global markets. **HPE Private Cloud AI (Deployment Platform)** [DAPM: Ceded] Pre-configured foundation for ISV solutions. Supports NVIDIA blueprints and partner applications. Air-gapped for regulated industries. Now includes dedicated turnkey development system for fast-tracking AI project validation. Evergreen and always current. **CrewAI (Pre-Installed on Private Cloud AI)** [DAPM: Delegated] Multi-agent automation framework pre-installed on HPE Private Cloud AI hardware. Enables enterprises to rapidly build and deploy tailored AI agents across industries: finance, healthcare, defense, retail, manufacturing, telecom, energy. On-premises deployment ensures data never leaves enterprise control. **Deloitte Zora AI for Finance** [DAPM: Delegated] Agentic AI solution reimagining executive reporting. Dynamic, on-demand, interactive experience driven by autonomous AI. Use cases: financial statement analysis, scenario modeling, competitive and market analysis. Deployed on Private Cloud AI. HPE is adopting internally first. Available worldwide. **Aible (Unleash AI Member, Discover 2025)** [DAPM: Delegated] AI agent platform for business users at enterprise scale. Completely autonomous specialized AI agents without requiring data science or ML engineering expertise. Auto-builds, coaches, and deploys AI agents across on-prem, hybrid cloud, and edge. ### NVIDIA-Provided Components **NVIDIA Blueprints + NIM** Pre-built AI application patterns (Multimodal PDF Extraction, Digital Twins, AI-Q) and inference microservices. Deployed on Private Cloud AI. Same blueprints available across all NVIDIA partners. ### Gap Analysis HPE correctly does not build Layer 3. The Unleash AI program is the most structured ecosystem curation approach among the infrastructure vendors assessed. Three characteristics define HPE’s Layer 3 approach: (1) Curated, not open. HPE is ‘highly selective’ with 26+ ISV members. Each partner is chosen for a specific AI use case domain and validated for interoperability. This is a deliberate contrast to Dell’s broader ISV partnership approach (more partners, less curation) and VAST’s smaller, focused ecosystem (CoreWeave, TwelveLabs, CrowdStrike). (2) Micro-focused agent model. Kamiwaza’s Luke Norris describes the approach: ‘tens if not hundreds of different agents that are micro-focused on particular jobs’ rather than one monolithic model. This is the operational philosophy behind Unleash AI — specialized agents from specialized partners, coordinated by Kamiwaza’s orchestration layer. (3) Pre-installed frameworks. CrewAI comes pre-installed on Private Cloud AI hardware. This is a different model than Dell’s (deploy NVIDIA NIM/NemoClaw as post-purchase software) or VAST’s (AgentEngine is the platform). HPE delivers the agent development framework as part of the infrastructure purchase. With Kamiwaza correctly positioned at Layer 2C (not Layer 3), the ecosystem layer map clarifies: • HPE provides infrastructure authority (Layers 0–2A) • Kamiwaza provides orchestration authority (Layer 2C, spanning 1B/1C/2B) • NVIDIA provides model execution runtime (Layer 2B) • ISVs provide domain applications (Layer 3): Deloitte Zora AI (finance), Aible (business users), ProHawk (video), Vaidio (vision AI), Blackshark.ai (geospatial), Gambit (citizen engagement) • Cross-cutting partners: CrowdStrike (security), Fortanix (confidential computing), Commvault/Veeam (data resilience), Red Hat (OS/K8s), SHI (integration services) The Town of Vail validates the ‘appliance-like operating model’ — unified deployment, lifecycle management, single escalation path. The coordination overhead that typically kills multi-vendor ecosystem solutions is addressed by HPE’s single-accountable-provider model and Kamiwaza’s orchestration layer. The SiliconANGLE analysis (May 2026) frames this as the emerging default for enterprise AI: ‘curated AI ecosystems’ where customers combine infrastructure, models, orchestration platforms, and ISV tooling without stitching every component together manually. HPE’s position is explicitly not a vertically integrated AI stack — it is a curated substrate model. ### Borrowed Judgment Distributed across partners, architecturally correct for Layer 3. Each partner maps to specific layers with identifiable authority boundaries. The structural comparison with Dell and VAST at Layer 3: • Dell’s ecosystem is load-bearing: ISV partners provide infrastructure-level functions (Cohere North for agent orchestration, DataRobot for lifecycle management) that Dell’s platform lacks. Remove Cohere North and Dell loses agent workflow orchestration. • HPE’s ecosystem is curated: ISV partners provide domain applications (Layer 3) while Kamiwaza provides orchestration (Layer 2C). Remove Deloitte Zora AI and HPE loses a finance use case, not a platform capability. • VAST’s ecosystem is additive: the platform is architecturally self-sufficient through Layer 2C. Partners add vertical use cases (TwelveLabs for video AI). Remove TwelveLabs and VAST loses a use case, not a platform function. HPE’s ecosystem structure is closer to VAST’s (additive) than Dell’s (load-bearing) at Layer 3, with the important distinction that HPE’s Layer 2C orchestration is sourced from an ecosystem partner (Ceded to Kamiwaza) rather than owned by the platform vendor itself (VAST’s PolicyEngine, Ceded to VAST). ### Working Notes The CrewAI pre-installation model is worth tracking: HPE hardware arrives with an agentic development framework already installed. This is a different go-to-market motion than selling infrastructure and then layering software. If this becomes the standard for Private Cloud AI, HPE is bundling Layer 3 development capability into the Layer 0 purchase. Deloitte and Aible represent enterprise-grade ISV deployments on Private Cloud AI — global SI (Deloitte) and AI platform vendor (Aible) choosing HPE’s infrastructure for agentic deployment. Dell’s equivalent is OpenAI Codex, SpaceXAI Grok, and ServiceNow. VAST’s equivalent is CoreWeave and TwelveLabs. Different ISV profiles reflect different customer bases. The SiliconANGLE ‘curated AI ecosystems’ framing from May 20, 2026 (one day ago) positions HPE’s Unleash AI approach as the emerging industry default. Whether this framing holds or whether vertically integrated stacks (VAST) or hyperscaler-controlled ecosystems (Google) prove more durable is an open question for the 4+1 assessment series. ════════════════════════════════════════════════════════════════════════════════ # IBM / Red Hat OpenShift AI Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.6 - ACP GA Promotion **Date:** July 12, 2026 **Source:** IBM Think 2026, Red Hat Summit 2026, Red Hat AI 3.4 GA, watsonx Orchestrate next-gen preview, IBM Sovereign Core GA, IBM Concert public preview, IBM Confluent acquisition, Granite 4.1 release, analyst coverage (SiliconANGLE, NAND Research, Futurum Group, ECI Research) v1.5 (July 12, 2026, /reconcile): OpenShell descored from 2B to a dated watch-list line — alpha on four peer rows; GA-gate applies uniformly. 1A Ceph+Fusion chip split: open-source Ceph Retained, proprietary Fusion Ceded (one chip, one substrate). v1.6 (July 12, 2026, /ga-check): watsonx Orchestrate Agentic Control Plane promoted from the watch-list on doc-confirmed GA ('available today on AWS and IBM Cloud,' IBM announcement + release notes) — scored as a Ceded 2C component; status holds moderate (Intelligence-2C governance, not placement). ## Summary Finding IBM / Red Hat is the only vendor in this assessment series attempting to build an enterprise AI operating model from middleware outward. Where Dell builds upward from hardware, VAST builds upward from storage, and Google builds downward from model intelligence, IBM builds from the platform layer — Red Hat OpenShift as the universal substrate — and extends authority in both directions: downward into infrastructure governance (Sovereign Core, Concert) and upward into agent orchestration (watsonx Orchestrate). The platform is the distribution vehicle, not the value. The value is the governance and orchestration intelligence that rides on top. The critical analytical lens for IBM is separating defensible proprietary IP from open-source packaging. The majority of IBM's AI platform capabilities — OpenShift (Kubernetes), vLLM (inference), KServe (model serving), Ray (distributed compute), Kubeflow (ML pipelines), MLflow (experiment tracking), Tekton (CI/CD), even InstructLab (model customization) — are open-source projects that run identically on VMware Tanzu, Amazon EKS, or bare Kubernetes. An enterprise could replicate most of IBM's Layer 2A/2B capabilities on any CNCF-compliant Kubernetes distribution. IBM's structural moats — capabilities that cannot be replicated without IBM — are concentrated in a narrow but strategically critical band: watsonx.governance (cross-platform AI assurance), watsonx Orchestrate (agentic control plane), Confluent integration with watsonx.data (governed real-time streaming), and Sovereign Core (runtime sovereignty). These are the components where IBM provides genuine authority above the Kubernetes baseline. IBM does not own compute silicon, does not own GPU scheduling, does not own networking fabric, does not own a high-performance AI-optimized storage platform, and does not own a frontier foundation model. Layer 0 is entirely Delegated or Absent — IBM provides no compute hardware, no server chassis, no networking switches, no cooling infrastructure, and no GPU fabric interconnect. IBM would be perfectly content for customers to run Layers 1A through 3 on a Dell AI Factory, HPE Private Cloud AI, or any OEM hardware. IBM's business model depends on someone else solving Layer 0. The consulting and services model reinforces the open-source strategy. IBM Consulting (~160,000 consultants) provides implementation expertise for the AI platform — but consulting is a competitive services market, not a platform dependency. Enterprises switch from IBM Consulting to Deloitte or Accenture for platform support the same way they switch SAP BASIS support providers: the structural moat is the platform IP (watsonx.governance, Orchestrate), not the services engagement. IBM Consulting is a competitive advantage in the services market, not a structural advantage in the platform architecture. The structural question for IBM is whether governance and orchestration authority — owning the narrow band of non-substitutable AI control plane software while everything else is open-source — is more durable than infrastructure authority (Dell, HPE), storage authority (VAST), or model authority (Google). The 4+1 model suggests this bet is architecturally sound — Layer 2C is where authority concentrates — but IBM must prove that watsonx Orchestrate's control plane is substantive, not just well-named, and that watsonx.governance's cross-platform assurance creates sufficient switching costs to justify the subscription when the rest of the stack is free. ## ○ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** No IBM Silicon (OEM-Provided) ### NVIDIA-Provided Components **NVIDIA GPU Silicon (via OEM Partners)** Blackwell, Vera Rubin available through Dell, HPE, Lenovo, Supermicro OEM partners running OpenShift. Red Hat Enterprise Linux for NVIDIA 26.01 provides Day 0 Blackwell support with Vera Rubin co-engineering underway. **NVIDIA AI Enterprise on OpenShift** GPU Operator, NVIDIA Run:ai (included in AI Enterprise), NIM microservices, DOCA runtime protection. Validated integration path through Red Hat AI Factory with NVIDIA. ### Gap Analysis IBM does not own compute silicon, server hardware, networking switches, cooling infrastructure, or AI-optimized network fabric. This is the emptiest Layer 0 of any vendor assessed. Dell owns PowerRack/PowerEdge/PowerSwitch/PowerCool. HPE owns ProLiant/Cray/Juniper/Slingshot/Aruba. VAST co-designs CNode-X storage hardware. Google owns TPUs and custom networking. AWS owns Trainium and EFA/SRD. IBM owns none of these. IBM's Layer 0 story is entirely indirect: Red Hat OpenShift runs on any x86/ARM hardware from any OEM. This is presented as a strength (hardware agnosticism, no vendor lock-in) but in 4+1 terms it means IBM has no Layer 0 authority whatsoever. Abstraction is not authority. Layer 0 defines what physical capabilities exist — compute density, thermal envelope, east-west bandwidth, accelerator topology. IBM has no opinion on any of these because IBM provides none of them. The networking gap is particularly significant for AI workloads. GPU-to-GPU bandwidth determines distributed training performance. Dell has PowerSwitch (NVIDIA Spectrum silicon). HPE owns Juniper + Slingshot for HPC fabric. AWS built EFA with SRD custom transport. IBM has no networking IP — not hardware, not software-defined networking for GPU fabrics. OpenShift's SDN handles container networking, not the GPU fabric networking that AI training requires. When an enterprise runs distributed training across 64 GPUs on OpenShift, east-west bandwidth depends entirely on whatever the OEM provided. IBM contributes nothing. The Red Hat AI Factory with NVIDIA co-engineering is significant: Day 0 Blackwell support, Vera Rubin co-engineering, NVIDIA Run:ai integration. But this is validation and integration work, not silicon or fabric authority. The same NVIDIA software stack runs on Dell, HPE, and Lenovo hardware. Note: IBM Z/Power with Telum on-chip AI inference is classified at Layer 3 (AI Application Layer), not Layer 0. Telum's value is transactional AI inference co-located with enterprise ledgers (banking, insurance) — this is an application-layer advantage where AI capability is adjacent to business data, not an infrastructure fabric capability. The same logic applies to SAP HANA on dedicated hardware: the value is in the application adjacency, not the compute fabric. ### Borrowed Judgment Total at Layer 0. IBM borrows all compute, networking, cooling, and fabric judgment from OEM partners and NVIDIA. The enterprise retains hardware vendor choice — a genuine governance benefit — but IBM adds no proprietary hardware value at any sub-layer: no silicon, no thermal engineering, no network fabric, no rack integration. Compare to Dell (retains mechanical/thermal/rack authority), HPE (retains networking end-to-end post-Juniper, cooling via Cray DLC, silicon agnosticism within owned chassis), VAST (retains storage hardware co-design with CNode-X). The abstraction-as-authority argument fails the 4+1 test. OpenShift abstracting hardware is a Layer 2A capability (infrastructure orchestration), not a Layer 0 capability. Layer 0 asks: what physical capabilities does the vendor provide? IBM's answer: none. ### Working Notes IBM Spyre AI Accelerator in Technology Preview is worth tracking. If IBM productizes custom AI silicon for OpenShift, the Layer 0 story changes fundamentally — IBM would join Google (TPU) and AWS (Trainium) as vendors with proprietary AI acceleration. But Technology Preview is not production. The multi-accelerator support story (NVIDIA, AMD ROCm, Intel Gaudi, IBM Spyre through different vLLM ServingRuntime variants) is the broadest of any platform assessed. But this is a Layer 2A/2B capability (platform support for multiple accelerators via Kubernetes operators), not Layer 0 authority. Supporting accelerators through software is fundamentally different from providing accelerators through hardware. Gemini's assessment frames Layer 0 abstraction as 'Silicon Decoupling' — a deliberate strategic choice. This framing is accurate as strategy but misleading as architecture. The enterprise architect choosing IBM accepts that Layer 0 is someone else's problem. The 4+1 model makes that acceptance visible. ## ◑ Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Governance Strength, Storage Gap ### Vendor-Provided Components **watsonx.governance** [DAPM: Ceded] Enterprise AI governance: Governance Graph (connected map of AI assets, policies, risks, regulations), model monitoring (bias, drift, fairness), agentic monitoring and security, regulatory library (EU AI Act, HIPAA, GDPR). Cross-platform: governs IBM, OpenAI, AWS, Meta models. IDC named IBM a Leader in AI governance. Proprietary IBM platform — opinions captive, no open exit. **watsonx.data (Data Lakehouse)** [DAPM: Ceded] Open data lakehouse with Presto and Spark engines, Iceberg table format. Shared metadata layer across clouds and on-premises. GPU-accelerated Presto (private preview). Context layer for AI-queryable metadata (private preview). watsonx.data intelligence provides data lineage, classification, quality, and Master Data Management. Proprietary IBM platform — opinions captive, no open exit. **IBM Confluent (Real-Time Streaming)** [DAPM: Ceded] Kafka + Flink + Tableflow integrated into watsonx.data. Zero-copy data sharing: query live Kafka streams as Iceberg tables without ETL. $11B acquisition positions IBM as the only infrastructure vendor with owned real-time streaming substrate. Proprietary IBM platform — opinions captive, no open exit. **IBM Storage Ceph** [DAPM: Retained] Ceph: distributed S3-compatible object storage for watsonx.data lakehouse — open-source substrate the enterprise can self-operate; the litmus's own Retained example. Competent but not AI-optimized at competitor level. **IBM Storage Fusion** [DAPM: Ceded] Proprietary storage services for OpenShift applications with data caching and acceleration; Storage Fusion HCI hosts watsonx on-premises. Fusion's data-services opinions are IBM-captive — split from the open-source Ceph substrate it sits beside, per the one-chip-one-substrate convention. **IBM Sovereign Core (Data Sovereignty)** [DAPM: Ceded] GA May 2026. Software platform enforcing data sovereignty across four pillars: operational, data, technology, AI sovereignty. Embeds policy at infrastructure runtime. Built on Red Hat OpenShift. Mistral AI as first certified model partner. Ensures data residency, model execution, and inference all operate within sovereign boundary. Proprietary IBM platform — opinions captive, no open exit. ### NVIDIA-Provided Components **GPU-Accelerated Presto (watsonx.data)** Private preview. Proof-of-concept with Nestlé showed 83% cost savings. GPU acceleration for analytical queries on the lakehouse. ### Gap Analysis IBM's Layer 1A is structurally split: strong governance, moderate storage. The governance stack is the strongest in this assessment. watsonx.governance provides end-to-end AI lifecycle governance — model monitoring, bias detection, drift monitoring, regulatory compliance (EU AI Act, HIPAA, GDPR), and a Governance Graph that maps relationships between AI assets, policies, risks, and regulatory requirements across platforms. Unlike Dell's Trust3 AI (partner overlay) or VAST's PolicyEngine (proprietary, platform-bounded), watsonx.governance operates across IBM and third-party platforms (OpenAI, AWS, Meta). This is the only cross-platform AI governance solution assessed. The watsonx.data lakehouse provides a governed data foundation with Presto and Spark engines, Iceberg table format, and shared metadata across cloud and on-premises. IBM Storage Ceph provides S3-compatible object storage. IBM Storage Fusion provides storage services for OpenShift applications. The Confluent acquisition adds real-time streaming (Kafka + Flink + Tableflow) with zero-copy data sharing — AI models can query live Kafka streams as Iceberg tables without ETL. But the storage infrastructure itself is not AI-optimized at the level of competitors. Compare: Dell Exascale provides 10+ PB/rack unified file+object+fast-file with MetadataIQ indexing billions of files. HPE Alletra X10000 provides unified file+object with Data Fabric policy-based placement. VAST Element Store collapses file, object, table, and vector into a single governed data structure. IBM Storage Ceph is competent distributed object storage but lacks the AI-specific metadata enrichment, inline embedding, or vector-native capabilities of Dell, HPE, or VAST storage platforms. The watsonx.data intelligence layer (data lineage, automated classification, Master Data Management, data quality) partially compensates by providing governance above storage — but the storage layer underneath is less performant and less AI-native than competitors' purpose-built solutions. ### Borrowed Judgment Low for governance (watsonx.governance, watsonx.data intelligence are IBM IP). Moderate for storage (IBM Storage Ceph and Fusion are IBM-owned but lack AI-specific optimizations). Low for streaming data (Confluent is now IBM-owned post-acquisition). The cross-platform governance capability is unique: watsonx.governance monitors models running on OpenAI, AWS, or Meta — not just IBM. This is the only assessed solution where governance authority is deliberately designed to extend beyond the vendor's own platform boundary. ### Working Notes The Confluent acquisition is the most strategically significant data-layer move from any vendor in this assessment. Real-time streaming + governed lakehouse + cross-platform governance creates a data foundation story that is architecturally distinct. Where Dell invested in Dataloop (pipeline orchestration) and VAST built DataEngine (serverless compute on data), IBM acquired the streaming substrate itself. The watsonx.data Context layer (private preview) adds contextual metadata directly into the lakehouse — making data AI-queryable without separate ETL. If this matures, it could address the metadata boundary problem identified in the Control Plane Working Notes: metadata that is both governed and real-time, not just batch-indexed. IBM's Governance Graph — mapping AI assets through policies, risks, and regulatory requirements as a connected graph — is the most sophisticated governance data model in this assessment. Whether it can serve as the queryable governance surface that a Layer 2C control plane needs is the open question. ## ◑ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Open-Source Assembled ### Vendor-Provided Components **OpenShift AI Model Serving (KServe + vLLM)** [DAPM: Retained] KServe for model serving orchestration. vLLM as primary inference runtime with support for NVIDIA, AMD, Intel Gaudi, and IBM Spyre accelerators. Serverless (Knative) and RawDeployment modes. Autoscaling based on request concurrency. **InstructLab (Open-Source Model Customization)** [DAPM: Retained] IBM-led open-source project for enterprise model customization. Structured taxonomy-based knowledge contribution without full retraining. Reduces dependency on model provider for domain-specific retrieval quality. Community-driven knowledge curation. **Vector Database Integration (External)** [DAPM: Delegated] OpenShift AI supports deployment of Milvus, Elasticsearch, pgvector, and other vector databases as containerized workloads. No IBM-owned vector database — enterprise selects and manages. Integration with watsonx.data for hybrid structured+unstructured queries. ### NVIDIA-Provided Components **NVIDIA NeMo Retriever** Embedding models and retrieval microservices available through Red Hat AI Factory with NVIDIA. Same components available across Dell and HPE deployments. ### Gap Analysis Applying the Kubernetes-baseline test: IBM has zero defensible IP at Layer 1B. Every component — KServe (CNCF), vLLM (open-source), Milvus (open-source), InstructLab (Apache 2.0), Granite Guardian (Apache 2.0) — runs identically on VMware Tanzu, Amazon EKS, or bare Kubernetes without IBM involvement. A competent ML engineering team can deploy the same retrieval stack on any CNCF-compliant cluster without IBM licensing, IBM consulting, or IBM support. This is IBM's weakest layer from a defensibility standpoint. The retrieval capability exists and works. The enterprise can build effective RAG pipelines on OpenShift AI. But nothing in the pipeline is IBM-specific. Compare: • VAST: InsightEngine provides end-to-end embedding + vector search + retrieval pipeline as a single integrated system on the same Element Store. Embeddings trigger the moment data lands — vectors are always current with source data. One authority, zero integration seams. This is proprietary IP that cannot be replicated outside VAST. • Dell: Data Search Engine (Elastic-powered) + MetadataIQ + NVIDIA cuVS. Three-party dependency, but MetadataIQ is Dell-proprietary metadata integration — a defensible asset. The Elastic partnership provides search intelligence Dell doesn't own but has engineered deep integration with. • HPE: Data Fabric namespace + NVIDIA NeMo Retriever + Milvus/LangChain. HPE has proprietary storage infrastructure underneath but no proprietary retrieval intelligence. When Kamiwaza enters via Unleash AI, it adds governed context orchestration — a defensible Layer 1B/2C capability. • IBM: Everything open-source. No proprietary vector database, no proprietary embedding pipeline, no proprietary search intelligence, no inline metadata enrichment at the storage layer. IBM provides the Kubernetes platform on which open-source retrieval components run. The platform is the value — but the platform is Layer 2A, not Layer 1B. The structural seam Gemini correctly identifies: context management exists as a distinct software layer on top of storage, not inline with data writes. IBM's architecture requires explicit pipeline configuration for embeddings. VAST's architecture makes embeddings structural. Dell's MetadataIQ makes metadata indexing structural. IBM has no structural retrieval integration — it's all application-layer assembly. InstructLab is IBM's most distinctive contribution — but it's a governance innovation (who controls what the model knows?) rather than a retrieval innovation. It's also Apache 2.0 and runs on any platform. The distinction matters: InstructLab is IBM-led community innovation, not IBM-owned defensible IP. watsonx.data's Context layer (private preview) may change this assessment. If contextual metadata becomes natively queryable for RAG within the governed lakehouse, IBM would have a proprietary retrieval integration point. But private preview is not production. ### Borrowed Judgment High — the highest borrowed judgment of any IBM layer. IBM borrows retrieval intelligence entirely from open-source communities (Milvus, Elasticsearch, pgvector), embedding models from NVIDIA (NeMo Retriever) or open-source (sentence-transformers), and search acceleration from NVIDIA (cuVS). IBM provides no proprietary retrieval logic, no proprietary vector indexing, and no proprietary search intelligence. This is not a criticism of the architecture — open-source retrieval works. It is a DAPM observation: the enterprise's retrieval pipeline at Layer 1B has no IBM authority in it. The judgment is borrowed entirely from open-source communities and NVIDIA. If Milvus changes its indexing heuristics, IBM inherits the change. If NVIDIA changes NeMo Retriever's embedding strategy, IBM inherits that too. IBM provides packaging, not judgment. Compare to VAST (low — owns InsightEngine, DataBase, Catalog, vector search) or Dell (moderate — owns MetadataIQ, delegates search to Elastic, depends on NVIDIA for acceleration). ### Working Notes InstructLab is IBM's most distinctive contribution at this layer — and it's deliberately positioned as an open-source community project rather than proprietary IP. InstructLab allows enterprises to contribute domain knowledge to model training through a structured taxonomy, reducing the enterprise's dependency on model providers for domain-specific retrieval quality. This is a governance innovation (who controls what the model knows?) rather than a retrieval innovation. The Granite Guardian models (safety and guardrails) operate at the boundary between Layer 1B and 2B — filtering retrieved content before it reaches the model. This is a retrieval governance function that Dell's Elastic and VAST's InsightEngine don't address at the retrieval layer. ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Confluent + Open-Source Pipeline ### Vendor-Provided Components **IBM Confluent (Kafka + Flink + Tableflow)** [DAPM: Ceded] IBM-defensible IP. Real-time data streaming platform. Kafka for event streaming, Flink for stream processing, Tableflow for zero-copy integration with watsonx.data Iceberg tables. $11B acquisition. Enables AI agents to reason over live event streams without ETL. No other infrastructure vendor owns a real-time streaming substrate. Proprietary IBM platform — opinions captive, no open exit. **IBM DataStage (watsonx Edition)** [DAPM: Ceded] IBM-defensible IP. Enterprise data integration engine with decades of maturity. Graphical and code-first pipeline building. Automated data cleansing, tokenization, formatting for LLM consumption. Comprehensive lineage tracking — trace which raw document fed into a specific fine-tuning or RAG dataset. Data sanitization (PII, hate speech, copyrighted content) reduces Layer 2C compliance burden. Proprietary IBM platform — opinions captive, no open exit. **ML Pipeline Stack (Kubeflow + Ray + MLflow + Spark + Tekton)** [DAPM: Delegated] Kubernetes-baseline capability. Open-source ML lifecycle on OpenShift AI. Kubeflow for pipeline orchestration, Ray for distributed training, MLflow for experiment tracking, Spark for data processing, Tekton for CI/CD. All run identically on VMware Tanzu, EKS, or bare Kubernetes. IBM provides enterprise packaging, not proprietary capability. **watsonx.data Query Federation** [DAPM: Ceded] IBM-defensible IP. Zero-copy query federation to external data platforms: Confluent Tableflow, Databricks Unity Catalog, Snowflake Open Catalog, Salesforce Data Cloud. Presto and Spark engines query data where it resides without copying. The federation logic is IBM-specific integration. Proprietary IBM platform — opinions captive, no open exit. ### NVIDIA-Provided Components **NVIDIA RAPIDS (via Spark Integration)** GPU-accelerated Spark on watsonx.data for pipeline processing. Same RAPIDS integration available across vendor platforms. ### Gap Analysis Applying the Kubernetes-baseline test, Layer 1C splits cleanly into IBM-defensible IP and open-source commodity — and the defensible half is genuinely strong. IBM-defensible IP (not replicable on Tanzu or bare Kubernetes): (1) Confluent (Kafka + Flink + Tableflow): IBM-owned post-$11B acquisition. Real-time streaming as infrastructure, not batch ETL. Tableflow zero-copy integration with watsonx.data Iceberg tables is IBM-specific. Kafka itself is open-source but the governed integration with the IBM lakehouse is proprietary. This is the most significant data-layer acquisition from any vendor in this assessment — no other infrastructure vendor owns a real-time streaming substrate. (2) DataStage (watsonx Edition): IBM proprietary enterprise data integration engine. Decades of maturity. Graphical and code-first pipeline building with automated data cleansing, tokenization, formatting for LLM consumption, and comprehensive lineage tracking. Enterprises can trace which raw document fed into a specific fine-tuning or RAG dataset. Compare to Dell's Dataloop (~$120M acquisition, less mature) or HPE's Airflow packaging (open-source, no proprietary lineage). (3) watsonx.data query federation: Zero-copy federation to Confluent Tableflow, Databricks Unity Catalog, Snowflake Open Catalog, Salesforce Data Cloud. IBM-specific integration logic. Kubernetes-baseline (replicable on any CNCF-compliant cluster): • Kubeflow — CNCF, pipeline orchestration, runs on any Kubernetes • Ray — open-source, distributed compute, runs anywhere • MLflow — open-source, experiment tracking, runs anywhere • Spark — open-source, data processing, runs anywhere • Tekton — CNCF, CI/CD pipelines, runs on any Kubernetes The 'Strong' classification is earned by the defensible half: Confluent + DataStage + watsonx.data federation. The open-source pipeline tools are commodity packaging — identical to HPE Ezmeral's approach, and replicable by any competent platform engineering team on any Kubernetes distribution. The architectural gap: unlike VAST's DataEngine (where pipeline functions execute directly on storage with CRDs — compute moves to data), IBM's pipelines run on OpenShift as separate containerized workloads — data moves to compute. For large-scale AI training pipelines, this creates more data movement than VAST's architecture. For real-time inference pipelines consuming Confluent streams, the data velocity advantage compensates. The resource overhead concern (correctly identified by Gemini's assessment): DataStage and Tekton are heavy enterprise platforms designed for complex corporate data architectures. For agile AI teams accustomed to Python scripts and LlamaIndex, IBM's data movement layer can feel over-engineered. This is a real practitioner concern — IBM's Layer 1C is enterprise-grade but not lightweight. ### Borrowed Judgment Low for defensible components. IBM now owns the streaming substrate (Confluent), the enterprise data integration engine (DataStage), and the lakehouse federation (watsonx.data). These are IBM IP with no external authority dependency. Moderate for open-source components. Kubeflow, Ray, MLflow, Spark, and Tekton are community-governed. IBM packages and supports them but inherits community judgment on architecture, performance, and API design. This is the same pattern as HPE's Ezmeral — enterprise packaging of open-source pipelines. NVIDIA provides GPU acceleration for Spark (RAPIDS) but the pipeline orchestration, streaming, data integration, and lifecycle are IBM-owned or community-governed. NVIDIA's authority at Layer 1C is limited to acceleration, not orchestration. Compare to Dell: Dell owns Dataloop (Retained) but depends on Starburst, NVIDIA, and ISV partners for the broader pipeline. Four authority boundaries. Compare to VAST: VAST owns everything at Layer 1C. One authority, total vendor dependency. IBM's model is distinctive: own the streaming substrate and enterprise data integration (defensible IP), package the open-source pipeline tools (commodity), federate across external data platforms (defensible integration). The enterprise retains more substitutability than VAST offers (can swap Kubeflow for Airflow without touching Confluent) but less than pure open-source (Confluent streaming integration is IBM-specific). ### Working Notes The Confluent acquisition creates a unique data velocity advantage. Dell, HPE, and VAST focus on data at rest (storage) and data in batch motion (pipelines). IBM now owns data in continuous motion (streaming). For agentic AI where agents need to reason over current events, current transactions, current sensor data — not yesterday's batch export — this is an architectural differentiator that no other infrastructure vendor possesses. Whether IBM can integrate Confluent deeply enough with watsonx.data to deliver on the 'zero-copy' promise is the execution question. The technology exists; the integration maturity is early. DataStage's data sanitization capabilities (removing PII, hate speech, copyrighted content prior to model training) are a crucial data-plane capability that directly reduces the compliance burden on Layer 2C downstream. This is the pipeline-to-governance connection that the 4+1 model identifies as critical: clean data in the pipeline means fewer governance exceptions at the reasoning plane. ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** OpenShift Strength ### Vendor-Provided Components **Red Hat OpenShift (Container Platform)** [DAPM: Retained] Enterprise Kubernetes platform. Hybrid cloud consistency across on-prem, AWS (ROSA), Azure (ARO), GCP, IBM Cloud, edge. Lifecycle management, security (SELinux, FIPS), multi-tenant isolation. The universal substrate for IBM's AI stack. **Red Hat OpenShift AI 3.4** [DAPM: Retained] MLOps + GenAIOps + AgentOps on OpenShift. Models-as-a-Service with token quotas, rate limiting, API key self-service, showback dashboards. AI gateway via Connectivity Link (Envoy/Kuadrant/Istio). Multi-accelerator model serving (NVIDIA, AMD, Intel Gaudi, IBM Spyre). **Red Hat Ansible Automation Platform** [DAPM: Retained] Infrastructure automation across hybrid environments. Ansible Lightspeed with IBM watsonx for AI-assisted automation content creation. Established enterprise automation authority — the bridge between AI platform operations and existing IT operations. ### NVIDIA-Provided Components **NVIDIA GPU Operator + Run:ai** GPU scheduling, fractionalization (MIG), workload management on OpenShift clusters. Run:ai now included in NVIDIA AI Enterprise for OpenShift deployments. ### Gap Analysis (GA-gate) The strong score rests on OpenShift, OpenShift AI, and Ansible (GA). IBM Concert, referenced below, is public preview (see watch-list). Layer 2A requires careful separation of Kubernetes-baseline capabilities from IBM-defensible IP. Kubernetes-baseline (replicable on Tanzu, EKS, or bare K8s): Container orchestration, namespace isolation, multi-tenancy, RBAC, lifecycle management, GPU scheduling via NVIDIA GPU Operator, model serving via KServe, distributed compute via Ray, CI/CD via Tekton. These are CNCF-ecosystem capabilities that IBM packages with enterprise support but does not own. VMware Tanzu provides the same primitives through VCF 9.1. An enterprise could switch from OpenShift to Tanzu without losing these capabilities — the same way enterprises switch SAP BASIS support from IBM to Accenture without touching SAP. IBM-defensible IP above the Kubernetes baseline: • NVIDIA bracketing: IBM is the only vendor in this assessment that makes NVIDIA AI Enterprise optional at Layer 2A. OpenShift AI's Kueue integration provides Kubernetes-native multi-tenant GPU queue management, quota enforcement, and fair-share scheduling without Run:ai. Dell's Layer 2A is 'Gap — Ceded to NVIDIA Run:ai.' HPE has the same NVIDIA GPU scheduling dependency. VMware depends on NVIDIA GPU Operator. IBM's platform team can dictate exactly how GPUs are carved up, queued, and billed using native OpenShift primitives, rendering NVIDIA's commercial scheduling layer optional. The enterprise can still use Run:ai on OpenShift — but it doesn't have to. This is a structurally significant Layer 2A differentiator. • Concert platform (Instana + Turbonomic + security modules): Cross-domain observability spanning applications, infrastructure, networks, and security with a graph-driven operations model. No CNCF equivalent. Datadog and Dynatrace compete at the product level but Concert's six-module integrated architecture (Observe, Operate, Optimize, Protect, Secure, Resilience) is IBM-specific. • Ansible Automation Platform: Established enterprise automation authority with Ansible Lightspeed (AI-assisted automation). Red Hat-owned IP with no Kubernetes-native equivalent. • OpenShift AI 3.4 Models-as-a-Service: Token quotas, rate limiting, self-service API keys, showback dashboards built as Kubernetes CRDs on Envoy/Kuadrant/Istio. The underlying components are open-source, but the composition and integration is IBM-specific packaging. An SI could replicate this on Tanzu with sufficient engineering — but IBM provides it out of the box. • Hybrid consistency: Same OpenShift control plane across on-prem, AWS (ROSA), Azure (ARO), GCP, IBM Cloud, edge. No other assessed vendor provides the same AI platform management plane across all major clouds. This is genuine differentiation — but OpenShift-specific, not a capability the enterprise retains if they move to Tanzu. The 'Strong' classification is justified by two structural differentiators that no other on-prem vendor matches: NVIDIA bracketing (making Run:ai optional through native Kueue scheduling) and hybrid consistency (same control plane everywhere). These are competitive advantages in the Kubernetes platform market. Concert and Ansible add IBM-specific observability and automation above the Kubernetes baseline. But the core orchestration primitives remain CNCF-baseline — maturity in Kubernetes packaging is a competitive advantage, not a structural moat. ### Borrowed Judgment Requires disaggregation: For Kubernetes-baseline capabilities: the enterprise borrows Kubernetes community judgment (scheduling, networking, storage orchestration) and NVIDIA judgment (GPU Operator, Run:ai). This borrowed judgment is identical regardless of whether the enterprise runs OpenShift, Tanzu, or EKS. It is not IBM-specific. For IBM-defensible IP: Low. Concert (Instana, Turbonomic), Ansible, and the OpenShift AI MaaS packaging are IBM-owned. The enterprise Cedes observability and automation authority to IBM when it adopts these — but can substitute with Datadog (observability) or Terraform (automation) without re-architecting the AI platform. The structural comparison requires a new framing: Dell's 2A (OpenManage) manages Dell hardware only. HPE's 2A (GreenLake) manages HPE infrastructure. VAST's 2A (Polaris) manages VAST clusters. VMware's 2A (VCF) manages virtualized infrastructure. IBM's 2A (OpenShift) manages any hardware — but so does any Kubernetes distribution. The scope is broad; the defensibility is in the packaging maturity and hybrid consistency, not in the orchestration primitives themselves. ### Working Notes The Models-as-a-Service architecture in OpenShift AI 3.4 is worth close attention. Token quotas, rate limiting, self-service API keys, and showback dashboards are Layer 2A functions that border on Layer 2C territory. When the platform decides which team gets how many tokens from which model — that's a placement decision. IBM positions these as 2A (resource management) rather than 2C (policy-driven placement), but the line is thin. LLMD (referenced in Summit sessions for intelligent resource orchestration) suggests IBM is building model-aware scheduling capabilities within OpenShift AI. If LLMD makes placement decisions based on model characteristics, load patterns, and cost constraints, it's a Layer 2C signal from the platform layer. Concert's six-module architecture (Observe, Operate, Optimize, Protect, Secure, Resilience) is the most comprehensive infrastructure operations platform in this assessment. Whether it constitutes a Layer 2C decision surface or a sophisticated Layer 2A monitoring/management system depends on whether Concert makes autonomous placement decisions or surfaces recommendations for human action. Watch-list (pending GA, not scored): IBM Concert Platform — agentic operations platform, public preview (Think 2026). ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Open-Source Runtime + NVIDIA ### Vendor-Provided Components **Red Hat AI Inference (Model Serving)** [DAPM: Delegated] Kubernetes-baseline capability. vLLM-based inference serving with KServe orchestration. OpenAI-compatible APIs. Models-as-a-Service architecture with enterprise authentication, token management, and showback. Supports NVIDIA, AMD, Intel Gaudi, IBM Spyre accelerators. vLLM and KServe are open-source — run identically on Tanzu or bare Kubernetes. IBM provides packaging and integration, not proprietary runtime. **Agent Lifecycle Management (MLflow-based)** [DAPM: Delegated] Kubernetes-baseline capability. LLM call tracing, tool execution tracking, reasoning step auditability. MLflow is open-source — runs on any platform. IBM provides integration with OpenShift AI and the watsonx governance stack. The auditability is enterprise-critical; the tooling is commodity. **Granite Model Family** [DAPM: Retained] IBM-defensible legal wrapper on open-source models. Granite 4.1 (3B, 8B, 30B dense models, Apache 2.0). ISO 42001 certified. Cryptographic model signing. Uncapped IP indemnity on watsonx.ai. Optimized for agentic workflows: tool calling, instruction following, function calling. Granite Guardian for safety guardrails. Models run anywhere (Hugging Face, Ollama, Dell Enterprise Hub, NVIDIA NIM). IP indemnity is IBM-specific — the only defensible element. ### NVIDIA-Provided Components **NVIDIA NIM + NeMo + Triton (via AI Factory)** Model serving, training, and inference runtime on OpenShift. Same NVIDIA runtime available across Dell, HPE, and Lenovo deployments. **OpenShell (Sandboxed Agent Runtime)** NVIDIA open-source project for sandboxed autonomous agent execution. Red Hat is a key contributor to upstream. Joint engineering underway to integrate with Red Hat's full-stack AI platform for infrastructure-level policy and oversight. **NVIDIA Agent Toolkit** Integrated into Red Hat AI Factory with NVIDIA for building autonomous agents. NemoClaw/OpenClaw agent runtime available on OpenShift. ### Gap Analysis (GA-gate) Confidential Containers, referenced below, is technology preview (see watch-list); the score rests on the GA runtime and governance components. Applying the Kubernetes-baseline test at Layer 2B reveals the same pattern as Layer 1B: the execution runtime is entirely open-source commodity, and IBM's defensible value is the governance and security wrapper around it. Kubernetes-baseline (replicable on Tanzu, EKS, or bare K8s): • vLLM — open-source inference engine, runs on any Kubernetes with GPU Operator • KServe — CNCF model serving orchestration, runs on any Kubernetes • MLflow — open-source lifecycle management, runs anywhere • Ray — open-source distributed compute, runs anywhere • NVIDIA NIM — containerized model serving, runs on any Kubernetes with NVIDIA AI Enterprise IBM-defensible IP above the baseline: • Confidential containers + agent security: SELinux, FIPS, sandboxed containers with NVIDIA Confidential Computing (Technology Preview). Hardware-enforced agent isolation protecting against runtime compromise even if another agent is breached. Red Hat's security hardening is genuine IP that vanilla Kubernetes and Tanzu don't match out of the box. This is the most comprehensive agent security architecture in this assessment. • OpenShell co-engineering: Red Hat contributing to upstream agent governance standards — infrastructure-level policy enforcement for autonomous agents. Not proprietary IP but IBM-shaped open-source that feeds into the 4+1 model's Layer 2C vision. • Granite IP indemnity: The models are Apache 2.0 and run anywhere. The uncapped IP indemnity is IBM-specific legal protection available only through watsonx.ai. The model is portable; the legal wrapper is not. • Granite Guardian: Safety guardrails, Apache 2.0. IBM-led but runs on any platform. One observation Gemini makes correctly: IBM deliberately treats the model runtime as a portable commodity layer. By standardizing on vLLM + KServe, IBM eliminates the runtime fragmentation seen in Dell's stack (NemoClaw, OpenShell, NeMo Guardrails, Dynamo, NIMs — five NVIDIA components creating a tightly coupled runtime). IBM's runtime simplicity is the strategy: fewer components, fewer dependencies, more portability. The trade-off is that IBM cannot achieve the extreme custom-silicon optimization of Google's Pathways/TPU integration or AWS's Trainium-optimized stack. The 4+1 model distinction: IBM's Layer 2B borrowed judgment is in execution (how models run — entirely from open-source and NVIDIA). IBM's Layer 2B authority is in governance (how model execution is constrained, audited, and secured — from Red Hat security hardening and confidential containers). This is the inverse of Dell's profile: Dell borrows governance at 2B, IBM borrows execution at 2B. The vendor comparison at 2B: • Dell: NVIDIA runtime + ISV blueprints (Cohere North, DataRobot, ClearML). Dell adds services, not runtime IP. Runtime is Ceded to NVIDIA. • HPE: NVIDIA runtime + Kamiwaza agent coordination + CrewAI/ISV frameworks. Three sources, three authorities. • VAST: AgentEngine provides a unified proprietary agent runtime. One authority. The strongest 2B defensibility in the assessment. • IBM: Open-source runtime (vLLM/KServe) + NVIDIA acceleration + Red Hat security wrapper + IBM governance overlay (watsonx.governance). IBM's value is not the runtime itself but the governance and security wrapper around an open-source runtime. The runtime is commodity; the wrapper is defensible. ### Borrowed Judgment Requires disaggregation: For execution runtime: High. vLLM, KServe, NVIDIA NIM, Triton are open-source or NVIDIA-controlled. IBM borrows all execution judgment from open-source communities and NVIDIA. This borrowed judgment is identical regardless of whether the runtime runs on OpenShift, Tanzu, or EKS. For governance and security: Low. Confidential containers, SELinux/FIPS hardening, OpenShell co-engineering, agent lifecycle auditability, and the Granite IP indemnity are IBM/Red Hat IP or IBM-led open-source. The enterprise Cedes security architecture to Red Hat when adopting OpenShift's agent security stack — but Red Hat's security opinions are the product. For model alignment: Variable. Granite models carry IBM's alignment choices (ISO 42001, rigorous data filtering, enterprise-focused tuning). If the enterprise swaps Granite for Llama or Mistral, it inherits those providers' alignment choices — but the execution fabric remains under the enterprise's authority regardless. The model is a Layer 3 choice; the runtime is a Layer 2B choice. IBM correctly separates them. ### Working Notes The Red Hat Advanced Developer Suite — trusted software factory, Trusted Libraries (SLSA Level 3), AI-driven exploit intelligence — adds a software supply chain governance layer that no other assessed vendor addresses at this depth. This is not traditional Layer 2B (model serving) but it's a critical enterprise concern: is the code that builds and deploys AI models itself trustworthy? OpenShift Dev Spaces supporting AWS Kiro, Microsoft Copilot, Claude CLI, Cline, Continue, and Roo from a single governed runtime is an underappreciated capability. Multi-assistant coding from one governed workspace means the developer's AI tools inherit OpenShift's security and governance posture — a concrete example of infrastructure-level governance applied to AI-assisted development. The three control points from Red Hat Summit 2026 (execution sandboxing, artifact provenance, short-lived agent identity) represent a governance-first approach to agent execution that aligns directly with the 4+1 model's Layer 2C thesis. The question: are these control points sufficient to constitute a Layer 2C function, or are they Layer 2B governance primitives that a separate Layer 2C must consume? Watch-list (pending GA, not scored): Confidential Containers + Agent Security — NVIDIA confidential computing in OpenShift, technology preview. Watch-list (descored July 12, 2026, /reconcile GA-gate uniformity): OpenShell integration (Red Hat co-engineering with NVIDIA on the sandboxed agent runtime) — OpenShell is alpha and watch-listed on the NVIDIA, Dell, Lenovo, and Supermicro rows; the same runtime cannot be a scored component here. Scores at doc-confirmed GA; the Red Hat upstream contribution is real context for the 2C vision, not shipping capability. ## ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Most Explicit 2C Claim ### Vendor-Provided Components **watsonx Orchestrate Agentic Control Plane** [DAPM: Ceded] GA per IBM's announcement and release notes — 'available today on AWS and IBM Cloud.' A centralized operational layer to observe, govern, and optimize AI agents regardless of where they were built or run: end-to-end visibility, policy enforcement, a shared agent catalog, native scheduling. Cross-framework Intelligence-2C — which agent may act, under what policy — not request-time placement (rule 4's routing-is-not-reasoning special case). Governance policies and catalog opinions authored in the control plane are IBM-captive. **IBM Bob (Multi-Model Orchestration)** [DAPM: Ceded] GA. Agentic developer assistant with multi-model routing: dynamically routes tasks to Claude, Mistral, Granite based on accuracy, latency, cost. Pass-through pricing. 80,000 internal IBM users, 45% average productivity gain. Demonstrates practical multi-variable placement decisions. **IBM Sovereign Core (Runtime Sovereignty)** [DAPM: Ceded] GA May 2026. Four sovereignty pillars: operational, data, technology, AI. AI sovereignty enforced at runtime — governing where inference happens, who controls models, how decisions are logged/traced/reviewed. Built on OpenShift + Red Hat AI. Mistral AI first certified model partner. Proprietary IBM platform — opinions captive, no open exit. **watsonx.governance (Cross-Platform AI Assurance)** [DAPM: Ceded] Governance Graph mapping AI assets through policies, risks, and regulatory requirements. Agentic monitoring and security capabilities. Cross-platform: governs IBM, OpenAI, AWS, Meta. Continuous compliance monitoring, not periodic audits. The governance query surface that a 2C control plane needs. Proprietary IBM platform — opinions captive, no open exit. ### NVIDIA-Provided Components **OpenShell Policy Layer** Governs agent execution, tool access, inference routing. 2B constraint enforcement with 2C policy potential. Same OpenShell available across NVIDIA partners. ### Gap Analysis Deployable today: IBM's shipping 2C is Bob (multi-model routing), Sovereign Core (runtime sovereignty, GA), and watsonx.governance. The headline cross-framework control plane discussed below - watsonx Orchestrate next-gen - is private preview (see watch-list); treat its capabilities as an announced direction, not a deployable option, until GA. The Kubernetes-baseline test produces its most significant finding at Layer 2C: this is the only layer where IBM provides 100% defensible IP. There is no CNCF equivalent for cross-framework agent governance, no open-source multi-variable model routing, no community-driven runtime sovereignty enforcement, no Kubernetes-native cross-platform AI lifecycle assurance. Every component at Layer 2C is IBM proprietary. Nothing here runs on Tanzu without IBM licensing. This finding validates IBM's entire strategic architecture. Layers 0 through 2B are progressively commodity — open-source packaging on delegated hardware. Layer 3 is consulting and models in a competitive services market. Layer 2C is the narrow band where IBM provides capabilities that cannot be replicated without IBM. The moat is here. IBM is the only vendor in this assessment that explicitly names its agent orchestration layer a 'control plane.' watsonx Orchestrate's next generation (private preview, Think 2026) is positioned as an 'agentic control plane for scaling and governing your AI' — and the capabilities described map directly to what the 4+1 model defines as Layer 2C. Applying the Intelligence 2C vs. Infrastructure 2C split established in the AWS assessment: Intelligence 2C (productized and portable): watsonx.governance provides continuous model monitoring, bias detection, drift tracking, regulatory compliance enforcement, and the Governance Graph mapping AI assets through policies, risks, and regulatory requirements. watsonx Orchestrate provides cross-framework agent governance — managing agents from IBM native, LangFlow, LangGraph, and A2A protocol with centralized policy enforcement, identity/credential management, and audit logging. IBM Bob demonstrates practical multi-model routing: tasks dynamically routed to Claude, Mistral, or Granite based on accuracy, latency, and cost — a multi-variable placement decision, not single-variable optimization like NVIDIA Dynamo's KV-aware routing. Sovereign Core enforces sovereignty as a runtime requirement, governing where inference happens, who controls models, and how decisions are logged within sovereign boundaries. IBM's Intelligence 2C is the strongest in this assessment for on-premises deployments. Four capabilities that no other on-prem vendor matches: (1) Cross-framework agent governance (watsonx Orchestrate) — manages agents regardless of which framework built them. Dell has no equivalent. HPE delegates to Kamiwaza. VAST's AgentEngine governs VAST-native agents only. (2) Cross-platform AI assurance (watsonx.governance) — governs models running on IBM, OpenAI, AWS, or Meta. No other governance solution spans vendor boundaries. (3) Multi-variable model routing (IBM Bob) — 80,000 internal users, demonstrated accuracy/latency/cost optimization. Production evidence at scale. (4) Runtime sovereignty (Sovereign Core) — sovereignty enforced at infrastructure runtime, not as a policy checkbox. No equivalent from any assessed vendor. Infrastructure 2C (absent/manual): watsonx Orchestrate governs agent behavior but does not autonomously calculate: 'Based on real-time token cost, data residency tags in watsonx.data, and current GPU cluster queue times, route this inference to on-prem PowerEdge versus burst to Azure.' That infrastructure placement coordination remains a manual configuration task for the platform architect. No productized engine queries Layer 1A governance metadata to make multi-variable infrastructure placement decisions in real time. This is the same split AWS exhibits: Intelligence 2C is productized (AgentCore Policy, Guardrails), Infrastructure 2C is implicit inside managed services. IBM's Intelligence 2C is more portable than AWS's (runs on-prem, multi-cloud). IBM's Infrastructure 2C is equally absent. Six-vendor Layer 2C comparison: • Dell: Absent. No productized control plane. • HPE: Retained (IT infrastructure ops via GreenLake Intelligence) + Delegated (AI workloads via Kamiwaza). • VAST: Gap, emerging (PolicyEngine + Polaris — middle-out from data layer, GA end 2026). • AWS: Intelligence 2C Delegated (AgentCore/Guardrails) + Infrastructure 2C Ceded/Implicit within managed services. • Google: Full 2C — Agent Platform with Inference Gateway + DWS. Most production-proven. Entirely Ceded to Google. • IBM: Intelligence 2C Retained/Ceded (watsonx.governance + Orchestrate — highly portable, productized, multi-cloud) + Infrastructure 2C Absent/Manual. The production maturity gap: watsonx Orchestrate next-gen is in private preview. The capabilities are described, the architecture is sound, Bob provides production evidence of multi-model routing at scale (80,000 users). But the full agentic control plane is not GA. Compare to Google's Agent Platform (GA, production-deployed) and AWS's Bedrock AgentCore (GA). IBM's 2C is the most explicitly named, the most framework-agnostic, and the least production-proven as a complete system. ### Borrowed Judgment Low — the lowest borrowed judgment of any IBM layer, because every component is IBM proprietary IP. watsonx Orchestrate, watsonx.governance, Sovereign Core, and Bob are all IBM-owned. No open-source dependency, no NVIDIA dependency, no partner dependency at this layer. The framework-agnostic approach means IBM borrows less agent-level judgment from any single framework vendor — but it also means IBM's orchestration authority depends on integration quality with frameworks it doesn't control (LangFlow, LangGraph, A2A). If LangGraph changes its execution model, IBM must update the integration. This is a different kind of dependency than NVIDIA runtime dependency — it's integration maintenance, not architectural dependency. The DAPM classification requires the Retained/Ceded distinction established in the series: • watsonx Orchestrate: Ceded to IBM. The enterprise consumes IBM's control plane — configures policies within it, but cannot replace it without re-architecture. Portable across clouds but not substitutable. • watsonx.governance: Retained by the enterprise. The enterprise defines its own governance policies, ethical thresholds, and compliance constraints. IBM provides the framework; the enterprise provides the judgment. The closest to genuinely Retained authority in IBM's stack. • Sovereign Core: Retained by the enterprise. The enterprise defines sovereignty boundaries; IBM enforces them at runtime. • Bob: Ceded to IBM. Multi-model routing logic is IBM's — the enterprise configures preferences but IBM's software makes the placement decisions. Compare to Google: Agent Platform provides 2C as a deeply integrated platform capability. More production-proven but less framework-agnostic. Entirely Ceded to Google — no on-prem option. Compare to HPE/Kamiwaza: Kamiwaza provides similar agent coordination but as a partner (Delegated). IBM provides it as owned IP. Compare to VAST: PolicyEngine provides data-layer governance with 2C ambitions. VAST builds 2C from data up; IBM builds 2C from platform out. Different architectural vectors toward the same Layer 2C function. ### Working Notes The Kubernetes-baseline finding at Layer 2C is the structural justification for IBM's entire AI strategy. Every layer below 2C is progressively more commodity — open-source runtime, open-source pipelines, open-source retrieval, delegated hardware. IBM's bet is that the narrow band of defensible IP at Layer 2C (governance + orchestration + sovereignty) captures more strategic value than the broad commodity layers beneath it. The 4+1 model suggests this bet is architecturally sound — Layer 2C is where authority concentrates. The ECI Research finding — two-thirds of enterprise AI leaders have already implemented multi-agent collaboration — validates the urgency. The problem watsonx Orchestrate addresses (governing hundreds of agents from different frameworks with consistent policy) is real and growing. The A2A protocol support is strategically important. If A2A becomes the standard for agent-to-agent communication (analogous to MCP for tool use), IBM's early support positions watsonx Orchestrate as the governance layer above the protocol — the control plane that governs A2A interactions. This is the platform-layer bet: don't own the protocol, govern the protocol. The Intelligence 2C vs. Infrastructure 2C gap is the open question for IBM's roadmap. If Concert's Turbonomic module (GPU cost optimization, workload placement recommendations) evolves from recommendations to autonomous placement decisions informed by watsonx.data governance metadata, IBM would close the Infrastructure 2C gap from the observability layer. The data to make infrastructure placement decisions exists across Concert (infrastructure telemetry), watsonx.data (data governance metadata), and Confluent (real-time operational data). The placement engine that consumes all three does not yet exist as a productized capability. Promoted July 12, 2026 (/ga-check): the Agentic Control Plane reached GA — 'available today on AWS and IBM Cloud' per IBM's announcement and release notes — and is now the cell's lead component alongside Bob, Sovereign Core, and watsonx.governance. ## ◇ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Consulting-Led Ecosystem ### Vendor-Provided Components **IBM Consulting (Competitive Services Market)** [DAPM: Retained] ~160,000 consultants with AI practice. Vertical industry solutions: banking (Banco Bradesco on ARO), healthcare, government, manufacturing. Competitive advantage in implementation services — but substitutable. Enterprise retains authority to select any SI (Deloitte, Accenture, Wipro, boutique firms) for platform implementation without architectural impact. The SAP BASIS pattern: platform IP is the moat, implementation services are a market. **Granite Model Ecosystem** [DAPM: Retained] Granite 4.1 (3B/8B/30B, Apache 2.0), Granite Guardian (safety), Granite Code, Granite Time Series. Available on watsonx.ai, Hugging Face, Dell Enterprise Hub, NVIDIA NIM, Docker Hub, Ollama, LM Studio. Multi-platform model family with strongest governance posture: ISO 42001, cryptographic signing, uncapped IP indemnity. **Multi-Model / Multi-Framework Support** [DAPM: Delegated] watsonx.ai supports Granite, Llama, Mistral, GPT-OSS, Nemotron. OpenShift AI serves any model via vLLM/KServe. watsonx Orchestrate manages agents from IBM native, LangFlow, LangGraph, A2A. The platform is model-agnostic and framework-agnostic by design. **Red Hat Partner Ecosystem** [DAPM: Delegated] OpenShift ISV ecosystem across industries. Microsoft (Platform Modernization Partner of the Year 2026), AWS (ROSA + Kiro integration), Salesforce (Agentforce reference architecture). Multi-cloud deployment: build on OpenShift, deploy across AWS/Azure/GCP/on-prem. **IBM Z/Power Transactional AI (Application-Layer Advantage)** [DAPM: Ceded] IBM Z16 Telum on-chip AI accelerator for real-time transactional inference co-located with enterprise ledgers — banking fraud detection, insurance claims, government transactions. This is a Layer 3 application advantage (AI adjacent to business data on proprietary hardware) not a Layer 0 infrastructure capability. Analogous to SAP HANA on dedicated hardware: the value is application adjacency, not compute fabric. watsonx Code Assistant for Z (IBM Bob Premium for Z) provides AI-assisted COBOL-to-Java modernization — 10x productivity gains in early users. ### NVIDIA-Provided Components **NVIDIA NIM + Agent Blueprints on OpenShift** NVIDIA NIM microservices and Agent Blueprints deployable on Red Hat AI Factory with NVIDIA. Same blueprints available across NVIDIA partners. ### Gap Analysis IBM's Layer 3 requires the same defensible-IP-vs-substitutable-services analysis applied throughout this assessment. IBM Consulting (~160,000 consultants) is a competitive advantage in a substitutable services market, not a structural platform dependency. Enterprises switch AI platform implementation partners the same way they switch SAP BASIS support: IBM to Accenture, Accenture to Deloitte, Deloitte to Wipro. The platform doesn't notice. The open-source components (OpenShift, vLLM, KServe, Kubeflow) run identically regardless of which SI assembled them. IBM Consulting's advantage is scale and experience, not lock-in. This is structurally different from every other vendor's services model: • Dell's Accelerator Services are additive — the NVIDIA runtime works without Dell's humans. • HPE's Kamiwaza partnership is structural — remove Kamiwaza and HPE loses Layer 2C orchestration. • VAST requires minimal consulting because the platform makes architectural decisions for you. • IBM's consulting is the primary go-to-market motion for a platform built from open-source components. The components are free. The assembly expertise is what IBM sells. But any competent SI can provide the assembly expertise. IBM Consulting competes for the engagement; it doesn't own the engagement by virtue of platform architecture. The Granite model family occupies a unique position: open-source (Apache 2.0), enterprise-grade, with the strongest governance posture of any model family (ISO 42001, cryptographic signing, uncapped IP indemnity). Granite is not competing with Claude or GPT-4 on raw capability — it's competing on trustworthiness, efficiency, and enterprise deployability. Critically, Granite runs on any platform — Dell Enterprise Hub, NVIDIA NIM, Hugging Face, Ollama. The model is portable; the IP indemnity is IBM-specific. IBM Z/Power transactional AI (Telum on-chip inference for banking, insurance, government) is a genuine Layer 3 application advantage. AI inference co-located with enterprise ledgers is not replicable on Dell or HPE hardware because the value is in the application adjacency to mainframe data, not in the compute architecture. watsonx Code Assistant for Z (Bob Premium for Z) with 10x productivity for COBOL modernization reinforces this — it's an AI application that only makes sense on IBM Z hardware. The ISV ecosystem comparison: • Dell's ecosystem is broad and horizontal: 5,000+ customers, OpenAI, Palantir, Google, ServiceNow, SpaceXAI. • HPE's ecosystem is curated and vertical: 26+ ISVs through Unleash AI. • IBM's ecosystem is consulting-driven and multi-cloud: IBM Consulting partnerships with SAP, Salesforce, Adobe, ServiceNow, plus the OpenShift ISV ecosystem deployable across all major clouds. • The multi-cloud deployment model (build on OpenShift, deploy across AWS/Azure/GCP/on-prem) is genuine differentiation at Layer 3 — ISVs building on OpenShift get portability that Dell and HPE can't match. ### Borrowed Judgment Distributed across ISV partners and model providers, which is architecturally correct at Layer 3. The consulting question is critical for DAPM: IBM Consulting is NOT a borrowed-judgment dependency. The enterprise retains full authority to select any SI for platform implementation, customization, and ongoing support. Switching SIs does not change the platform architecture, does not require re-engineering, and does not break running workloads. This is the SAP BASIS pattern: the platform IP (watsonx.governance, Orchestrate) is the structural dependency; the implementation services are a competitive market. The structural comparison: • Dell's ecosystem is load-bearing: ISV partners provide infrastructure-level functions. Remove Cohere North and Dell loses agent orchestration. • HPE's ecosystem is curated: partners provide domain applications. Remove Deloitte Zora AI and HPE loses a finance use case, not a platform function. • IBM's consulting is competitive: remove IBM Consulting and the enterprise hires Accenture. The platform remains intact. The adoption motion may slow but the architecture doesn't change. • VAST's ecosystem is additive: platform is self-sufficient. Partners add vertical use cases. ### Working Notes IBM's Layer 3 strategy is tightly focused on high-value, unglamorous enterprise utility — code modernization, automated compliance, legacy IT orchestration, mainframe fraud detection — rather than broad consumer-facing generative applications. Dell's Layer 3 partners are flashy (OpenAI, Palantir, SpaceXAI). HPE's Unleash AI partners target emerging AI use cases (video AI, geospatial, vision). IBM's Layer 3 targets the work enterprises actually need done: converting COBOL to Java, generating Ansible playbooks, detecting fraud in real-time transaction streams, modernizing mainframe applications. Nobody puts COBOL modernization on a keynote slide, but it's where regulated enterprise budgets concentrate. The watsonx application surfaces reinforce this enterprise utility focus: watsonx Assistant (conversational AI for customer service, HR, operations), watsonx Code Assistant / IBM Bob (AI-assisted development across the software lifecycle), watsonx Orchestrate applications (workflow automation binding agents to enterprise processes). These are not general-purpose AI platforms — they are purpose-built for specific enterprise operational domains. The 80,000 internal IBM Bob users represent the largest internal AI deployment from any vendor in this assessment. IBM is eating its own cooking at scale — and the 45% productivity gain claim, if sustained across production workloads, validates the agentic AI thesis more concretely than any vendor keynote. The EY tax technology partnership (Bob in private beta, described as 'closer to a collaborative agent than a simple coding tool') signals enterprise validation from a major professional services firm. Granite's Apache 2.0 licensing + uncapped IP indemnity is a distinctive governance posture. No other model family provides both open-source freedom AND vendor-backed legal protection. This addresses a specific enterprise concern: 'I want to run this model anywhere, and I don't want to worry about IP claims.' Neither OpenAI (closed-source, no indemnity) nor Meta (open-source, no indemnity) nor Google (Gemma open-weight but limited indemnity) matches this combination. The Kubernetes-baseline test at Layer 3: applications are inherently above the platform baseline, but IBM's distribution model matters. Granite models are Apache 2.0 — run on any platform. IBM Consulting is substitutable. The OpenShift ISV ecosystem targets any Kubernetes. IBM's defensible Layer 3 assets are narrow: Z/Power transactional AI (hardware adjacency), Granite IP indemnity (legal wrapper), watsonx Code Assistant for Z (mainframe-specific), and the watsonx application surfaces that integrate with Layer 2C governance (watsonx.governance integration creates application-level governance that doesn't port to Tanzu). The governance integration is the subtle lock-in: applications built to leverage watsonx.governance's Governance Graph inherit a dependency on IBM's governance architecture. ════════════════════════════════════════════════════════════════════════════════ # Intel AI Infrastructure Portfolio Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.0 — Initial Assessment **Date:** June 26, 2026 **Source:** Intel 2025/2026 earnings releases (SEC 8-K), Intel Newsroom (Gaudi 3 availability, Xeon 6 and Xeon 6+ 'Clearwater Forest' launch, agentic AI systems), IBM Cloud / Intel Gaudi 3 availability announcement, NVIDIA Investor Relations (Sep 2025 strategic agreement and $5B investment), OPEA (LF AI & Data) and oneAPI documentation, OpenVINO documentation, IT@Intel Agentic AI white paper (Oct 2025), DPDK and SPDK (Linux Foundation) project documentation, The Register / ServeTheHome Clearwater Forest launch coverage (Computex 2026), published 4+1 model. ## Summary Finding Intel is the only ingredient vendor in this instrument. Every other row — Dell, HPE, AWS, Google Cloud, Azure, VAST, VMware, Cisco, IBM/Red Hat, Oracle — is an integrator or operator that assembles Layers 0 through 3 into a system the enterprise buys whole. Intel is the component and open-standards company all of them buy from. Its authority sits entirely at Layer 0, where it is strong on CPU, present but thin on accelerator, and absent on fabric. Above Layer 0 it ships no procurable platform. But it is the open substrate the rest of the stack runs on, and that distinction is the assessment, not a deficiency to apologize for. For every other vendor this instrument names where the buyer is captured. Intel inverts the question. Its Layer 0 is unusually non-captive: Xeon is commodity x86, swappable across OEMs and to another x86 vendor without rebuilding (Retained), and Gaudi 3's open software stack — SynapseAI under Apache-2.0, community PyTorch and vLLM integration, OPEA in LF AI & Data — makes its opinions portable in a way CUDA is not (Delegated). The proprietary opinion layer most vendors use to hold a buyer is exactly what Intel leaves open. There is almost no decoupled capture to name. The real constraint is ecosystem thinness, few qualified Gaudi channels and no AI cluster fabric, and that is a capability limit, not an authority one. At Layers 1A through 3 Intel has no procurable offering, and those cells read as gaps. But gap means absent-as-product, not absent-as-presence. The open data-plane standards Intel originated, DPDK for networking and SPDK for storage, both now Linux Foundation projects, sit in the data path of the storage and pipeline stacks the enterprise actually buys (1A, 1C). Kubernetes operators let Gaudi and Xeon be scheduled inside Red Hat OpenShift AI (2A). At Layer 3, Intel silicon runs beneath nearly every enterprise AI application while Intel holds zero application authority. The one place deployable capability surfaces above Layer 0 is Layer 2B, scored moderate: OpenVINO, its model server, and the vLLM Gaudi backend are open, deployable inference runtimes. Developer-led rather than a managed service, but real, and depended on in production. The Reasoning Plane at Layer 2C is the universal gap, and Intel sits with everyone else in it. There is no Intel control plane and no policy-driven inference placement. OPEA's composable blueprints are workflow scaffolding; the internal IT@Intel agentic platform is Intel consuming agentic infrastructure, not selling it. Routing is not reasoning, and neither artifact crosses that line. The buyer's trade is straightforward. Intel is chosen as the open, lower-cost alternative to NVIDIA's pricing and CUDA lock-in, not as a full-stack AI infrastructure platform. Through Dell or IBM Cloud, a buyer gets a portable, open inference substrate whose accumulated opinions can leave when the workload outgrows it. What they do not get is an integrated AI infrastructure platform from Intel, because Intel does not sell one. The authority they keep is the point. The platform they must assemble around it is the cost. ## ◑ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** CPU Strong, Accelerator Present-but-Thin, No Fabric ### Vendor-Provided Components **Intel Xeon 6 (P-core and E-core Data Center CPUs)** [DAPM: Retained] AMX (Advanced Matrix Extensions) in every P-core for native per-core AI inference acceleration. Xeon 6776P hosts NVIDIA DGX B300; Xeon 6 hosts DGX Rubin NVL8. Google Cloud C4/N4 instances run on Xeon 6, and Intel co-develops custom ASIC IPUs with Google for networking, storage, and security offload. Fully standard x86 ISA: runs on any OEM platform and is substitutable across OEMs and to another x86 vendor without rebuilding workloads, which keeps the enterprise's opinions Retained. **Intel Xeon 6+ / Clearwater Forest (Intel 18A)** [DAPM: Retained] Up to 288 E-cores on Intel 18A (RibbonFET GAA plus PowerVia backside power), 576 MB L3, DDR5-8000. Launched at Computex 2026 and orderable immediately through Dell, HPE, Lenovo, and Supermicro; deployed in telecom and edge production. Tight 18A supply is a near-term caveat. Standard x86 across multiple OEMs, so the substrate is swappable and the cell is Retained. **Intel Gaudi 3 AI Accelerator** [DAPM: Delegated] Training and inference accelerator in PCIe and OAM form factors; rack-scale reference design with industry-standard Ethernet (not a proprietary fabric). IBM Cloud is the first CSP deployment; Dell AI Factory (PowerEdge XE9680, eight Gaudi 3 with 128 GB HBM each) is the primary on-premises channel; VMware VCF 9.0 support is confirmed. The software stack is open: SynapseAI is Apache-2.0, PyTorch and vLLM integration is community-maintained, OPEA recipes are LF-hosted. Because the accumulated opinions live in open tooling that lifts to other hardware, Gaudi is Delegated, not Ceded like CUDA-bound GPU silicon. The constraint is ecosystem thinness (limited qualified channels), which is a capability limit rather than an authority one. ### NVIDIA-Provided Components **Xeon 6 as NVIDIA DGX Host CPU** Intel Xeon 6776P is the host CPU for NVIDIA DGX B300, and Xeon 6 was selected as the host CPU for DGX Rubin NVL8. Intel's most visible current Layer 0 win is as a subcomponent supplier inside NVIDIA's AI system, not as a platform integrator of its own. **NVIDIA / Intel CPU-GPU Co-Engineering (Sep 2025)** Intel will build NVIDIA-custom x86 CPUs for NVIDIA's AI infrastructure platforms, coupled over NVLink, alongside x86 RTX SOCs for PCs. NVIDIA invested $5B in Intel common stock (~5% stake), completed Dec 2025. The agreement formalizes Intel as a CPU substrate supplier for NVIDIA's AI systems. It is supplier-to-platform, not platform-to-platform: Xeon becomes more strategic inside NVIDIA's AI factory; an Intel AI factory does not emerge from the deal. ### Gap Analysis Intel's Layer 0 is asymmetric in a way no other vendor in this assessment is: dominant at one sub-layer, absent at another. CPU silicon is Intel's strongest AI-era position. Xeon 6 with P-cores delivers AMX (Advanced Matrix Extensions) in every core, giving native AI inference acceleration without a discrete accelerator, and it is the host CPU inside NVIDIA's DGX B300 and Rubin NVL8 systems. Xeon 6+ 'Clearwater Forest' (Intel 18A, up to 288 E-cores) launched at Computex 2026 and is orderable through Dell, HPE, Lenovo, and Supermicro on day one, putting Intel's leading-edge node into production. CPU authority is the most durable Intel position across the entire 4+1 stack, and it rests on a genuine multi-vendor x86 ISA the enterprise can carry to another vendor. The AI accelerator is present but thin. Gaudi 3 is generally available, with IBM Cloud as the first cloud service provider to deploy it and the Dell AI Factory (PowerEdge XE9680, eight Gaudi 3 per node) as the primary on-premises channel. But qualified OEM channels are limited, and Intel's own executives have conceded the company is not yet participating in the cloud AI data-center market in a meaningful way. The accelerator's value is real where the workload fits; the ecosystem around it is the constraint. Networking is the clear absence. Intel exited AI networking in 2019 when it abandoned Omni-Path (spun out as Cornelis Networks), and it has no GPU-to-GPU scale-out fabric in market today, no competitor to NVIDIA Spectrum-X, HPE Slingshot, AWS EFA/SRD, or Cisco Silicon One. Intel's networking relevance is the open data-plane software standard it created, DPDK, rather than proprietary AI fabric silicon. The fabric absence is the most consequential Layer 0 gap for large-scale AI clusters. Calibration: NVIDIA (Layer 0 strong) owns the accelerator silicon every other vendor depends on; Dell (strong) integrates the full rack, thermal, and storage system. Intel cannot reach strong because it is dominant on CPU, thin on accelerator, and absent on fabric. The asymmetry caps the cell at moderate. The instrument has not assessed a pure silicon and open-standards vendor before; the cell reflects that Intel sells to OEMs and cloud providers who then sell systems to enterprises. ### Borrowed Judgment The judgment question inverts for Intel. Other vendors borrow Intel's judgment when they buy Xeon and Gaudi and build on top of it. Intel is the judgment source for the sub-layers it occupies, not a borrower. But the enterprise does not buy from Intel directly at Layer 0. Dell configures PowerEdge with Gaudi 3; IBM deploys Gaudi 3 on IBM Cloud; HPE and Lenovo integrate Xeon 6 and Xeon 6+. The enterprise chooses an OEM system and inherits Intel silicon inside it. Crucially, the opinions it accumulates stay portable: x86 binaries lift across OEMs and to another x86 vendor, and Gaudi's open software stack lifts to other hardware. Intel's Layer 0 authority is real, mediated through OEM systems, and deliberately non-captive. ### Working Notes Tight 18A supply is a near-term caveat for Xeon 6+: Intel is managing allocation closely and broader customer sampling continues through 2H 2026, though systems are orderable through major OEMs now. Watch-list (pending GA, not scored): Jaguar Shores, the Gaudi-branded successor to the cancelled Falcon Shores commercial product, with HBM4 and silicon-photonics interconnects, is confirmed on the roadmap but unlikely before 2027. It is not deployable today and is not scored. ## ○ Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Substrate Provider, Not a Storage Offering ### Gap Analysis Intel has no data storage product, no AI data-governance platform, and no lakehouse or catalog offering. The enterprise data foundation is built on Dell (PowerScale, ObjectScale), HPE (Alletra), VAST, or cloud-native services, and Intel does not own or operate any of them. But Intel is not absent from this layer in the structural sense, only as a procurable offering. SPDK (the Storage Performance Development Kit), an Intel-originated open-source userspace framework, sits in the storage data path of platforms across the industry, and Xeon silicon powers storage controllers nearly everywhere. Intel's confidential-computing primitives, SGX and TDX, are integrated by storage vendors and cloud providers into their own governance offerings. None of these is an Intel-procurable Layer 1A platform; they are the open substrate and security silicon the storage layer is built on. The enterprise that wants governed storage inherits the storage vendor's judgment, not Intel's, while running on Intel's substrate. Calibration: NVIDIA's Layer 1A is also a gap with no components, an accelerator that makes other vendors' storage faster without providing it. Intel is the same shape: critical to the layer as substrate, absent as an offering. The Absent score is an accurate reading of what Intel sells, not a criticism of its silicon business. ### Borrowed Judgment Not applicable as an authority cell. Intel provides no Layer 1A offering, so the enterprise owns the function by default and inherits its storage vendor's governance judgment (Dell, HPE, VAST, or a cloud provider). SGX/TDX are Layer 0 security primitives integrated by those vendors, not an Intel governance authority. ### Working Notes SGX and TDX are Layer 0 confidential-computing silicon, not Layer 1A governance. The instrument holds the security-versus-governance line: security constrains who can access data; governance constrains what the platform does with it. ## ○ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Open Reference Only, Not an Offering ### Gap Analysis Intel has no vector database, no embedding service, and no managed RAG pipeline. The closest artifact is OPEA (Open Platform for Enterprise AI), Intel's open-source initiative in LF AI & Data, which ships composable RAG reference architectures such as ChatQnA. But OPEA is a developer blueprint that OEM engineers and ISVs build from, not a product an enterprise deploys as its retrieval layer. Calibration: NVIDIA's Layer 1B is moderate because it ships real retrieval-acceleration components OEMs brand and deploy (cuVS, NeMo Retriever, NIM embeddings). Intel sits below that line: OPEA's partner count measures ecosystem engagement, not enterprise deployment, and nothing rises to a credited component. The instrument scores whether an enterprise can buy this from Intel at this layer. It cannot, so the cell is a gap, with OPEA named as a sub-threshold open-source reference rather than a scored component. ### Borrowed Judgment Not applicable as an authority cell. Intel provides no Layer 1B offering; the enterprise's retrieval architecture and its judgment come from the platform or search vendor it chooses. ### Working Notes OpenVINO appears in some Intel RAG discussions but is an inference-optimization toolkit; it is assessed at Layer 2B, not here. ## ○ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Substrate Provider, Not a Pipeline Offering ### Gap Analysis Intel has no ETL/ELT platform, no pipeline orchestration product, and no data-lineage offering. The enterprise that wants pipelines buys Databricks, Snowflake, AWS Glue, or GCP Dataflow; Intel silicon may accelerate those, but Intel does not operate them. As at Layer 1A, the absence is of an offering, not of presence. DPDK (the Intel-originated, now Linux Foundation Data Plane Development Kit) sits in the networking and data-movement path of the NFV, SDN, and storage stacks the enterprise actually runs, and the oneAPI Data Analytics Library (oneDAL) provides hardware-accelerated analytics primitives developers link into their own applications. The Intel/Google ASIC IPU work offloads storage and networking from host CPUs to improve pipeline throughput. All of this is open substrate and developer tooling beneath the pipeline layer; none is a procurable Intel pipeline or governance product. Calibration: this lands with NVIDIA's Layer 1C gap (its only candidate, CMX, was pre-GA) rather than with Dell's moderate (Dataloop gives Dell genuine proprietary orchestration IP). oneDAL is a build-time library, not orchestration, so it stays narrated in prose and the cell carries no scored components. ### Borrowed Judgment Not applicable as an authority cell. Intel provides no Layer 1C offering; the enterprise owns pipeline orchestration and lineage by default, running on Intel's open data-plane substrate. ### Working Notes oneDAL and DPDK are developer libraries and open standards, not procured or operated pipeline products; they are named here as substrate, not scored. ## ○ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Substrate Provider, No Orchestration Offering ### Gap Analysis Intel is critical to this layer without selling a product for it. The defacto x86 standard is the substrate every orchestration platform schedules over, and the open data-plane standards Intel originated, DPDK for networking and SPDK for storage, both now Linux Foundation projects, sit in the data path of the NFV, SDN, and storage stacks these orchestrators manage. The Intel Enterprise AI Foundation for OpenShift adds Kubernetes operators (device plugins, Gaudi health checks, RoCE provisioning) that let Gaudi and Xeon accelerators be scheduled inside Red Hat OpenShift AI. The internal IT@Intel '1AI' platform shows Intel's own engineers solving the orchestration problem on Intel silicon. None of this is a procurable orchestration platform. There is no Intel equivalent of NVIDIA Run:ai, HPE GreenLake Intelligence, or Red Hat OpenShift AI. The OpenShift operators enable Red Hat's orchestration; the scheduling authority is Red Hat's, not Intel's. The 1AI platform is Intel consuming agentic infrastructure, not selling it. Calibration: NVIDIA's Layer 2A is strong (Run:ai is the GPU-scheduling authority three OEMs rebrand) and Dell's is moderate (Dell at least owns rack orchestration and a CSI operator and rebrands Run:ai). Intel is below both: it brings a device plugin into someone else's scheduler, so the cell is a genuine gap. The orchestration authority belongs to Red Hat, Kubernetes, and the enterprise, and it is robustly Retained precisely because it runs on Intel's open, multi-vendor substrate rather than a captive one. ### Borrowed Judgment Not applicable as an authority cell, and the Retained-by-default here is the robust kind. The enterprise's orchestration opinions, Kubernetes and OpenShift manifests on x86, are portable because they rest on Intel-originated open standards (x86, DPDK, SPDK) that are genuinely multi-vendor. Intel is the reason the layer is Retained, not an absence within it. ### Working Notes The Intel Enterprise AI Foundation for OpenShift delivers accelerator-enablement primitives (Kubernetes device plugins, Gaudi health checks, RoCE provisioning), necessary for an OEM to offer a working Gaudi cluster but not sufficient to call it Intel orchestration. It is narrated here rather than scored. ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Open Runtime Tooling, No Managed Service ### Vendor-Provided Components **OpenVINO Toolkit and Model Server** [DAPM: Delegated] Model optimization (quantization, compression) and a KServe-compatible serving runtime across Intel targets (Xeon AMX, Gaudi 3, Arc GPUs, Core Ultra NPUs). Apache-2.0 and deployable today; widely used for edge and server inference. Serving and optimization opinions lift to other runtimes, so Delegated. **Gaudi 3 vLLM Backend (Open-Source Contribution)** [DAPM: Delegated] Intel-optimized vLLM backends for Gaudi 3, contributed to the open-source vLLM project and deployed via Dell AI Factory and IBM Cloud. Intel provides the optimization layer; the platform provides the lifecycle. The serving opinions live in vLLM (portable), consistent with the Gaudi-is-Delegated call at Layer 0. **Intel oneAPI Toolkit (2026.0)** [DAPM: Delegated] Open, SYCL-based cross-architecture programming model, UXL Foundation (Linux Foundation) governed, with the Base and HPC toolkits merged in the 2026.0 release. A build-time programming model for inference and HPC applications rather than a managed runtime, but open and swappable, so Delegated. ### NVIDIA-Provided Components **Gaudi 3 as NVIDIA NIM Alternative** Gaudi 3 runs vLLM and TGI with Intel-optimized backends, so enterprises can serve models on Gaudi without NIM. This is an inference-hardware alternative to NVIDIA on an open serving path, not an Intel-managed runtime product. ### Gap Analysis This is the one layer above Layer 0 where an architect has something from Intel to deploy to production today, and it is open rather than managed. OpenVINO is Intel's inference-optimization toolkit, and OpenVINO Model Server is a KServe-compatible, gRPC/REST serving runtime, one of the most widely deployed edge-inference runtimes in production. The vLLM Gaudi backend is the real serving path for Gaudi 3 deployments through Dell AI Factory and IBM Cloud. oneAPI is the open, SYCL-based cross-architecture programming model these build on, with millions of installations and UXL Foundation governance. All are open-source and deployable now. What Intel does not offer is a managed model-serving service. There is no Intel equivalent of NVIDIA NIM, AWS Bedrock, or GCP Vertex AI. Intel gives the enterprise the open runtime to stand up itself; it does not operate one. That is a 'you operate it' boundary, not an 'it does not exist' one. Calibration: NVIDIA's Layer 2B is strong (NIM closed and Ceded, plus Dynamo open and Retained, a complete runtime stack) and Dell's is moderate (Ceded to NVIDIA, owning no runtime IP). Intel earns moderate on a different basis than Dell: open, deployable runtime tooling it actually authored, not a rebrand. It is below NVIDIA's strong because there is no managed service and no proprietary differentiated runtime. The components are Delegated because the serving opinions live in open-source the enterprise can swap. ### Borrowed Judgment Moderate and delegated to open-source. An enterprise serving on OpenVINO Model Server or vLLM on Gaudi inherits Intel's optimization decisions (quantization, batching, kernel paths) but can swap to another open runtime without rebuilding application code, because the serving opinions sit in open-source tooling rather than a captive Intel layer. The judgment is real but portable. ### Working Notes Watch-list (active development, not scored): OPEA Enterprise Inference optimizes inference services on Intel hardware with Kubernetes orchestration and is positioned as production-grade, but it is an open-source project rather than an OEM-shipped managed runtime. If OEMs adopt it as the managed serving layer for Gaudi 3, Intel would gain an indirect Layer 2B presence through its open-source initiative. ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** No Control Plane ### Gap Analysis Intel has no Layer 2C capability and makes no Layer 2C claim in any reviewed product documentation. There is no policy-driven inference placement, no cross-agent governance control plane, no agentic resource-coordination product. Applying the 'routing is not reasoning' test: OPEA's composable blueprints can chain pipeline steps but make no policy-driven decision about where inference runs relative to data, which model serves a request, or how cost, latency, and compliance are arbitrated at request time. The internal IT@Intel '1AI' platform routes prompts to specialized agents on Intel's private cloud, but that is Intel consuming agentic infrastructure, not providing it. Neither crosses the line into a reasoning plane. Calibration: this is the universal 2C gap documented across the instrument. NVIDIA's 2C is a gap (runtime governance only), and Dell's is a gap (enterprise responsibility). Intel sits in parity with its peers; live Infrastructure-2C is a market-wide absence, not an Intel-specific deficit. ### Borrowed Judgment Not applicable. Intel makes no Layer 2C claim, so there is no judgment to borrow; the enterprise retains full responsibility for the function, as it does with every peer. ### Working Notes Watch-list (signal, not a finding): the Dell/Intel collaboration reported to be addressing the AI-factory governance gap (SiliconANGLE, May 2026). If a control-plane product ships through it, that would be Intel's first genuine claim above Layer 0. No product exists today. ## ○ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Silicon Enabler, Not Application Vendor ### Gap Analysis Intel has no AI application layer and no model. There is no Intel equivalent of Palantir Foundry/AIP, ServiceNow Now Assist, Salesforce Einstein, or Microsoft Copilot, and, unlike NVIDIA, no Nemotron-equivalent frontier model. The closest artifact is Intel Liftoff, a developer-relations program giving AI startups compute credits and go-to-market help; the applications those partners build are their products, not Intel's. This is where the assessment's thesis lands cleanly. Intel is the silent enabler: when an enterprise runs Copilot, Einstein, or Palantir AIP, some of that inference runs on Xeon, and in limited deployments on Gaudi 3. Intel captures Layer 3 value through silicon volume, not application authority. The substrate is present beneath nearly every enterprise AI application; the application authority is zero. That is the precise structural reading of an ingredient vendor at the value plane, and the gap is an offering absence, not an irrelevance. Calibration: NVIDIA's Layer 3 is moderate (it has Nemotron open models, a real if singular component) and Dell's is a partner ecosystem (a curated ISV program enterprises procure through). Intel matches neither: no model, and Liftoff is startup enablement rather than an application ecosystem the enterprise buys through. The enterprise gets its Layer 3 applications from Microsoft, Palantir, or ServiceNow directly, never with Intel as the procurement front, so the cell is a gap. ### Borrowed Judgment Not applicable as an authority cell. Intel provides no Layer 3 application or model, so it lends no application judgment; the enterprise's application authority comes entirely from the ISV or model provider it chooses, while running on Intel silicon. ### Working Notes Intel vPro and the Core Ultra AI PC are a client-compute story, outside the 4+1 instrument's enterprise-infrastructure scope (data center, cloud, edge server), and are not scored here. ════════════════════════════════════════════════════════════════════════════════ # Kamiwaza AI Orchestration Platform Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.2 - Reconcile: 2C Determinism Rule (held strong) **Date:** July 17, 2026 **Source:** Kamiwaza 1.0 launch (May 2026), Kamiwaza v0.9.3 docs, product pages, HPE Town of Vail whitepaper, Tracxn profile, SecurityBrief coverage, GitHub repos ## Summary Finding Kamiwaza is a software-only AI orchestration platform that enters the buyer conversation at Layer 2C and works downward. It is one of only two vendors in the instrument — alongside Articul8 — whose primary product includes the reasoning plane. Where Articul8's Intelligence-2C focuses on mission decomposition and domain-specific agent routing, Kamiwaza's Intelligence-2C focuses on governance: cross-agent authority, relationship-based access control, and policy enforcement at execution time. The enterprise gets production-validated cross-agent governance that no other on-prem vendor provides independently. The capture mechanism is governance-layer capture — a third pattern distinct from both coupled capture (data moves into the vendor's namespace) and decoupled capture (data stays open, proprietary opinion layer is captive). Kamiwaza's data stays in place, the infrastructure stays under enterprise control, and the enterprise feels free at the visible layers. What Kamiwaza captures is the authority to decide what the data means (the living ontology at Layer 1B) and what agents may do with it (the ReBAC governance at Layer 2C). These are arguably the most valuable layers to own — and the hardest to leave, because the ontology, relationship graph, and governance policies accumulate over time as proprietary Kamiwaza artifacts. Three gap layers — L0 (compute), L1C (data movement), L2A (infrastructure orchestration) — are by design. A software-only vendor should not be at Layer 0. Kamiwaza's explicit 'no data movement' thesis means L1C absence is architectural, not accidental. L2A absence reflects that Kamiwaza orchestrates AI workloads and agents, not the underlying infrastructure. These gaps define Kamiwaza's scope, not its weakness. The buyer's trade: production-grade cross-agent governance and distributed AI orchestration without moving data or committing infrastructure — in exchange for Ceding the governance and semantic layers to Kamiwaza's proprietary platform. The data is free. The understanding of the data is captive. A closed system is a closed system. Kamiwaza's position in the instrument is structurally inverse to Dell's: Dell is strong at the bottom of the stack (L0, L1A) and absent at the top (L2C). Kamiwaza is strong at the top (L2C, L1B) and absent at the bottom (L0, L1C, L2A). HPE's Unleash AI program bridges the two — Kamiwaza as the Delegated Layer 2C partner on HPE's infrastructure substrate. That pairing is the only assessed combination that covers all eight layers with identified authority at each. ## ○ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Enterprise Responsibility ### NVIDIA-Provided Components **NVIDIA DGX Spark Integration** Validated deployment target with dedicated .deb package. Kamiwaza runs on DGX Spark but does not provide or manage the hardware. **Intel Gaudi 3 / Ampere Validated** Whitepapers with Intel Gaudi 3 (DHS deployment, 85% analysis time reduction) and Ampere processors. Hardware partner validations, not hardware authority. ### Gap Analysis Kamiwaza provides no compute hardware, networking, or acceleration fabric. The enterprise brings its own infrastructure — on-prem servers, cloud instances, edge nodes, DGX Spark — and Kamiwaza runs on top of it. The platform is validated on NVIDIA DGX Spark, Intel Gaudi 3, and Ampere processors, but these are deployment targets, not owned hardware. The enterprise retains full responsibility for Layer 0. This is by design — a software-only platform vendor should not own the compute layer. ### Working Notes The DGX Spark integration includes a dedicated ARM64 .deb package with CUDA dependencies, suggesting meaningful optimization work for that platform. The Intel Gaudi 3 whitepaper (DHS deployment) and Ampere whitepaper demonstrate silicon-agnostic deployment — Kamiwaza runs on NVIDIA, Intel, and ARM compute without hardware lock-in. ## ◑ Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Derived Artifact Layer ### Vendor-Provided Components **Distributed Data Engine (DDE) Ingestion** [DAPM: Ceded] Connector-driven pipelines ingest from S3, Postgres, Kafka, SharePoint, Slack, file systems into Kamiwaza's vector stores. Scheduled or one-time runs. Credential management via Kamiwaza secrets (encrypted at rest). Job monitoring via observability dashboards. Proprietary ingestion pipeline — connector logic and scheduling captive to Kamiwaza platform. **Data Catalog** [DAPM: Ceded] Metadata catalog for ingested documents and data assets. Tracks connector provenance, security markings, ingestion status. Proprietary catalog — metadata schema and query interface captive to Kamiwaza. **Security Markings System** [DAPM: Ceded] Document-level security classification enforcement. system_high clearance validation, default_security_marking per connector, X-User-System-High header validation at retrieval time. Designed for regulated environments (federal, healthcare, financial services). Proprietary security model captive to Kamiwaza platform. **Vector Store Substrate (Milvus / Qdrant)** [DAPM: Delegated] Open-source vector databases for embedding storage and retrieval. Enterprise could extract vector data and operate Milvus/Qdrant independently. The vector data artifacts are portable; the pipeline that produced them is not. ### Gap Analysis Kamiwaza does not govern the enterprise's source data stores — they remain in place under whatever authority already manages them. What Kamiwaza does provide is a derived artifact layer: the DDE ingests from source systems (S3, Postgres, Kafka, SharePoint, Slack, file systems) into Kamiwaza's own vector stores (Milvus/Qdrant) and application database (CockroachDB). The Data Catalog indexes metadata. The security markings system (system_high, X-User-System-High header, default_security_marking) enforces classification at the artifact level. This is a genuine 1A function — Kamiwaza creates and governs its own data artifacts — but it operates alongside the enterprise's existing data governance, not in place of it. Comparable to how Palantir's Ontology creates a governed semantic layer over open data substrates without replacing the underlying storage. The vector store substrate is open-source (Milvus, Qdrant) — the data artifacts are technically portable. But the ingestion pipeline logic, connector configurations, chunking strategies, and security markings are Kamiwaza IP. ### Working Notes Supported sources as of v0.9.3: File, Amazon S3, Kafka, Postgres, Hive, Slack. The connector model is extensible — additional connectors available through support agreements. DDE connector APIs are mounted under /api/dde/ with full CRUD, trigger, and document management. Rate limiting per connector and requester (HTTP 429 with Retry-After). ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Kamiwaza Differentiator ### Vendor-Provided Components **Context Manager / Living Ontology** [DAPM: Ceded] Automatically builds and maintains a knowledge graph across all data sources — entities, relationships, insights spanning organizational boundaries. No manual mapping or data centralization. Grounds AI in up-to-date enterprise context, reduces hallucinations. Proprietary ontology engine — the semantic understanding of enterprise data is captive to Kamiwaza. **Inference Mesh** [DAPM: Ceded] Routes LLM reasoning to distributed data sources without moving the data. Decentralized inference adds intelligence to retrieval within the enterprise's security perimeter. Model-agnostic — supports multiple LLM providers via litellm integration. Proprietary routing and orchestration logic captive to Kamiwaza. **Retrieval Service** [DAPM: Ceded] RAG pipeline connecting DDE-ingested vector data to LLM inference. Integrates with Milvus/Qdrant vector stores and the Context Manager ontology. Proprietary retrieval orchestration captive to Kamiwaza. ### Gap Analysis Kamiwaza's primary product differentiator. The Context Manager automatically builds a living ontology across distributed data sources — a knowledge graph connecting entities, relationships, and insights that span organizational boundaries without manual mapping or data movement. The Inference Mesh routes LLM reasoning to data sources without moving the data, adding intelligence to retrieval without the security risk of sending internal data to public cloud APIs. The living ontology is the most architecturally significant 1B capability in the instrument alongside Palantir's Ontology. Both build a semantic layer over distributed data. The difference is architectural: Palantir pulls data into its namespace and operates the platform; Kamiwaza queries data in place and the enterprise operates the software on its own infrastructure. Both are Ceded — the semantic understanding is proprietary in both cases. The litellm fork on GitHub suggests open model routing for inference — Kamiwaza is model-agnostic at the inference layer, routing to whatever LLMs are available locally. This is a genuine differentiator vs. vendors locked to specific model providers. Open question: is the ontology exportable in an open format (RDF, OWL, JSON-LD)? If yes, the ontology data is portable even though the maintenance engine is captive. If no, both the engine and the accumulated knowledge graph are captive. This single fact would determine whether L1B capture is hard (nothing leaves) or soft (the snapshot leaves but the living maintenance doesn't). Either way, DAPM is Ceded — the ongoing ontology maintenance is proprietary regardless. ### Working Notes The Kaizen agent (v1.0) works across internal data sources through the Context Manager, drawing on information held across separate systems rather than a single repository. A skills library lets teams define what the agent can do and under which conditions. This is the L1B→L2B→L2C chain in action: retrieval (1B) feeds agent execution (2B) governed by policy (2C). ## ○ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Enterprise Responsibility ### Gap Analysis Kamiwaza explicitly does not move data — that is the product thesis. 'Stop moving data. Start running AI where your data lives.' ETL/ELT pipelines, lineage tracking, cost-aware data movement are not in scope. The enterprise's existing pipelines (Airflow, Spark, whatever they have) continue to operate. The DDE does ingest data into Kamiwaza's vector stores, but this serves Layer 1B retrieval, not Layer 1C data movement as the model defines it. There is no general-purpose pipeline orchestration, no lineage graph, no cross-system ETL. This absence is architectural, not accidental. A vendor whose core value proposition is 'we don't move your data' cannot also provide a data movement layer. The enterprise retains full responsibility for L1C. ### Working Notes The DDE ingestion from S3, Kafka, Postgres etc. creates vector artifacts for RAG — this is retrieval enablement (L1B), not data pipeline infrastructure (L1C). The boundary judgment: DDE serves retrieval context, not governed data movement. ## ○ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Enterprise Responsibility ### Gap Analysis Kamiwaza does not provision, schedule, or lifecycle-manage infrastructure. GPU scheduling, quota enforcement, VM lifecycle, cluster provisioning — none of this is Kamiwaza's domain. The enterprise's existing infrastructure orchestration (Kubernetes, VMware VCF, GreenLake, Run:ai) operates beneath Kamiwaza. The platform runs on Docker Swarm with Traefik, managed by Kamiwaza's own orchestration — but this manages Kamiwaza's internal services, not the enterprise's broader infrastructure estate. The enterprise retains full responsibility for infrastructure orchestration. Compare to VAST (Polaris manages the VAST fleet), VMware (VCF manages the entire estate), or HPE (GreenLake manages hybrid infrastructure). Kamiwaza has no equivalent — it is a tenant on the enterprise's infrastructure, not a manager of it. ### Working Notes Kamiwaza v0.5.0 addressed 'GPU waste, job starvation, dependency conflicts' — but this refers to managing Kamiwaza's own inference workloads on available compute, not infrastructure-level GPU scheduling across the enterprise. ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Kamiwaza Runtime ### Vendor-Provided Components **Workrooms (v1.0)** [DAPM: Ceded] Bounded collaboration spaces for teams and AI agents. Users and agents operate within existing permissions. Each Workroom contains its own data and tools available only to authorized users. Access boundaries enforced at the platform architecture level, not through manual policy changes. Proprietary runtime — execution model captive to Kamiwaza. **Kaizen Agent** [DAPM: Ceded] Configurable AI coworker that works across internal data sources through the Context Manager. Skills library defines what the agent can do and under which conditions. Multi-modal analysis and output. Expanded in v1.0. Proprietary agent framework captive to Kamiwaza. **Tool Shed** [DAPM: Ceded] Governed tooling platform enabling AI to access data and take action through controlled tools. Security enforced at the tool level — at execution time, AI can only perform actions the requesting user is authorized to take. ReBAC ensures permissions enforced when actions execute. Every action logged for audit. Proprietary governance model captive to Kamiwaza. **Infrastructure Substrate (Open-Source)** [DAPM: Delegated] Ray Serve (model serving), CockroachDB (application database), Milvus/Qdrant (vector stores), Keycloak (authentication), etcd (service discovery), Docker Swarm (orchestration), Traefik (reverse proxy). Open-source components the enterprise could operate independently — though the Kamiwaza services layer that composes them is proprietary. ### Gap Analysis Kamiwaza provides a genuine agent runtime: Workrooms (bounded execution environments with architecture-level access enforcement), Kaizen agent (configurable AI coworker with skills library), Tool Shed (governed tooling with ReBAC enforcement), and the Inference Mesh (local model serving). Chainguard-hardened containers provide attested infrastructure with SLSA Level 3 pipelines and verified SBOMs. What Kamiwaza does not provide is the broader model serving infrastructure — no NIM-equivalent optimized inference containers, no Triton, no KServe, no distributed training framework. The Inference Mesh serves models locally via Ray Serve, but the model serving story is thinner than a full 2B vendor like VAST AgentEngine or AWS Bedrock. The self-deploy model is notable: .deb packages for Ubuntu, .rpm for RHEL, MSI for Windows, macOS tarball. Enterprise Edition adds Terraform deployment. The enterprise operates the full platform on its own infrastructure with no Kamiwaza SaaS dependency at runtime. But self-deployable does not make it Retained — the runtime opinions are proprietary and captive. Authentication is built on Keycloak (open-source OIDC/JWT). RBAC policy is YAML-based. These substrates are open — the ReBAC governance layer above them is proprietary Kamiwaza IP. ### Working Notes Architecture stack: FastAPI gateway, Docker Swarm orchestration, CockroachDB (distributed SQL), Milvus/Qdrant (vector), etcd (service discovery), Ray Serve (model serving), Traefik (reverse proxy). The infrastructure substrate is heavily open-source — Python, React, CockroachDB, Milvus, Ray, etcd, Keycloak. The proprietary value sits in the Kamiwaza services layer that composes these into a governed platform. ## ● Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Kamiwaza Core ### Vendor-Provided Components **Living Ontology (Governance Layer)** [DAPM: Ceded] The ontology that determines what data means, how it relates across systems, and what policies apply — updated in real time across distributed sources. This is the governance foundation that the ReBAC layer and Tool Shed query to make authorization decisions. Proprietary semantic governance captive to Kamiwaza. **ReBAC Enforcement (Relationship-Based Access Control)** [DAPM: Ceded] Constrains agent permissions based on relationship context, not just role. Emerged from production behavior — traditional RBAC breaks when autonomous agents cross department boundaries. Enforced at execution time with full audit logging. Proprietary governance model captive to Kamiwaza. **Cross-Environment Evaluation** [DAPM: Ceded] Rather than treating anomalies or requests as isolated events, Kamiwaza evaluates what else is happening across the environment to determine appropriate response. Agents surface relevant information, trigger correct workflows, and support operators as conditions change. Validated in Town of Vail fire detection coordination. Proprietary orchestration logic captive to Kamiwaza. **Agent Lifecycle Governance** [DAPM: Ceded] Determines which agents run, in what sequence, with what inputs, under what constraints. Enforces human-in-the-loop checkpoints. Manages agent lifecycle at the execution layer. Mission decomposition and decision authority placement. Proprietary agent governance captive to Kamiwaza. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** Kamiwaza's governance and orchestration layer has no NVIDIA dependency. The platform is model-agnostic and silicon-agnostic at the reasoning plane. ### Gap Analysis Kamiwaza is one of two vendors in the instrument — alongside Articul8 — whose primary product includes Layer 2C. They separate on the determinism rule: Kamiwaza holds strong at Intelligence-2C on a deterministic mechanism, Articul8 is moderate because its core is a probabilistic planner. Kamiwaza's Intelligence-2C is governance-first, and the center of gravity is a deterministic, code-based mechanism: the living ontology determines context, the ReBAC enforcement constrains agent actions at execution time, the Tool Shed governs what tools agents can invoke, and every action is logged for audit. Kamiwaza asks: 'may this agent act, on this data, under this policy?' That is control executing code outside the model — the kind the determinism rule rewards — structurally in the class of Palantir's interaction-time security. Town of Vail production evidence validates it across multiple agent types — accessibility auditing, document processing, fire detection coordination — with cross-departmental authority boundaries and human-in-the-loop checkpoints. Articul8's Intelligence-2C is reasoning-first: mission decomposition and domain-specific agent selection. That core is LLM-driven planning (probabilistic), which is why Articul8 scores moderate at 2C under the determinism rule while Kamiwaza's deterministic ReBAC mechanism holds strong. The two are complementary in focus, but no longer at the same grade. Infrastructure-2C (where inference physically runs at request time): Partial for Kamiwaza. The Inference Mesh routes inference requests across available models and compute, which is adjacent to live placement. But the documentation frames this as locality-aware data access rather than multi-variable infrastructure placement reasoning (cost + compliance + latency + data residency simultaneously). Neither Kamiwaza nor Articul8 provides full Infrastructure-2C — Articul8 explicitly acknowledges this as the responsibility of hyperscalers and platform vendors. The RBAC-to-ReBAC evolution is architecturally significant. Traditional role-based access breaks when autonomous agents operate across department boundaries. Relationship-Based Access Control, which emerged from Kamiwaza's production behavior at Town of Vail, constrains agent permissions based on context, not just role. This validates the 4+1 model's claim that Layer 2C requires governance architecturally distinct from Layer 2A infrastructure RBAC. Compare to: • Dell: Layer 2C gap. Enterprise retains responsibility. • HPE: Delegated to Kamiwaza via Unleash AI. Production-validated. • Articul8: Ceded (Intelligence-2C moderate — mission decomposition on a probabilistic planner; moved strong→moderate under the determinism rule. Complementary in focus to Kamiwaza's deterministic governance). • Google: Ceded (Agent Platform — productized, comprehensive, captive). • VAST: Gap, emerging (PolicyEngine + Polaris — announced, GA end 2026). • Palantir: Ceded (Ontology + Apollo — Intelligence-2C strong, Infrastructure-2C adjacent). • Kamiwaza: Ceded (governance-first 2C; cross-agent authority and ReBAC capture). ### Borrowed Judgment Low, and Ceded to Kamiwaza. The governance mechanism the enterprise inherits — the living ontology, ReBAC enforcement, the Tool Shed, lifecycle governance — is Kamiwaza IP, and it is deterministic and code-based, not prompt-shaped: the authorization decisions are computed, not asked of a model. The relationship graph and governance policies accumulate over time as proprietary Kamiwaza artifacts that do not lift to another reasoning plane without rebuilding. The universal caveat still applies: the deterministic outcome-validator (did the decision achieve intent, under a graduated escalation policy) is absent here as everywhere, noted as universal rather than charged to Kamiwaza; live Infrastructure-2C placement is likewise only partial (the Inference Mesh's locality-aware routing). ### Working Notes Town of Vail deployment velocity: concept to first-phase production in three months. 20-30 additional use cases projected in first year. Additional use cases compose from existing primitives (decision flows, authority boundaries, governance constraints) rather than requiring new infrastructure. Economic model: fixed-cost infrastructure on the town's own solar/wind-powered data center — billions of tokens without variable cloud API costs. The Community Edition on GitHub provides a partial open-exit path, but it is not feature-equivalent to the Enterprise Edition (Workrooms, full ReBAC, Chainguard hardening). The Community Edition does not constitute a real escape hatch under the methodology's litmus test. ## ◑ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Platform + Kaizen ### Vendor-Provided Components **Kaizen Agent (v1.0)** [DAPM: Ceded] Configurable AI coworker with skills library. Multi-modal analysis and output. Works across internal data sources through Context Manager. The primary Layer 3 application Kamiwaza provides. Proprietary agent — captive to Kamiwaza platform. **App Garden** [DAPM: Ceded] Platform for discovering and deploying pre-packaged AI applications and services. Curated marketplace model. Deployment surface for enterprise and partner applications. Proprietary marketplace — captive to Kamiwaza platform. **Use Case Templates and Deployment Patterns** [DAPM: Delegated] Compliance automation, document processing, legal, HR, supply chain, sales acceleration, knowledge construction, smart cities. Templates and reference implementations, not packaged applications. Enterprise builds on these using Kamiwaza platform primitives. ### Gap Analysis Kamiwaza provides one first-party Layer 3 application — the Kaizen agent, a configurable AI coworker with a skills library, multi-modal analysis, and integration across internal data sources via the Context Manager. Kaizen is the primary user-facing surface for enterprise AI interaction. Beyond Kaizen, the App Garden provides a deployment surface for additional applications, and the Tool Shed enables governed tool composition. The use-case library is extensive — compliance automation, document processing, legal, HR, supply chain, sales acceleration, knowledge construction — but these are deployment patterns and templates, not packaged applications. Kamiwaza is primarily a platform vendor, not an application vendor. The analogy is closer to VMware (platform that enables enterprise-built applications) than to ServiceNow (application vendor with a platform). Layer 3 applications will be built by the enterprise's own teams or partners using Kamiwaza's platform primitives. The HPE Unleash AI partnership positions Kamiwaza-built applications (ARIA accessibility agent, deed restriction processor, fire detection coordinator) as reference implementations — proof points for what the platform enables, not the platform's Layer 3 offering. Partner integrations: Dell + Intel Gaudi 3 joint solution, HPE Unleash AI program, DHS deployment. These are delivery partnerships, not ISV ecosystem depth comparable to Dell's (5,000+ customers, OpenAI/Palantir/ServiceNow) or HPE's (26+ Unleash AI members). ### Working Notes $11M total funding over 2 rounds (seed, Jan 2025). Early-stage company. The partner ecosystem is nascent compared to established vendors — delivery partnerships with HPE, Dell, Intel rather than a broad ISV marketplace. ════════════════════════════════════════════════════════════════════════════════ # Lenovo Hybrid AI Advantage with NVIDIA Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.1 - DAPM Retraction & Capability Re-read **Date:** July 11, 2026 **Source:** v1.1 (same day): corrections from the Supermicro-session instrument rulings. The OEM-channel DAPM rule was retracted — a proprietary single-implementation platform is Ceded-to-partner through any paper; Delegated survives only via open/self-hostable substrates or multi-vendor standard interfaces. Chips moved: 1A DM/DG (ONTAP), DE (SANtricity), and WEKA Delegated→Ceded; 2B Nutanix Delegated→Ceded. Cloudian (S3), Red Hat, AI Library runtime, and TruScale survive on litmus grounds. 1A re-read under the whose-paper capability rule (the GPU parallel): moderate→strong on the productized menu plus owned InfiniBox. 1C re-examined under the exposure test and held at gap (no productized pipeline engine on Lenovo paper). Summary paragraph 2 rewritten accordingly. Original v1.0 source: CES 2026 (Lenovo Agentic AI + xIQ launch, AI Cloud Gigafactory, Jan 6), GTC 2026 (inference platforms, Vera Rubin support, March), Lenovo press releases (May 12 and June 24, 2026), Lenovo Press product documentation (Hybrid AI Software Platform lp2311, Hybrid AI 221/285/289 platform guides, Cloudian LVD lp2388, Centific LVD lp2180), Infinidat acquisition close (April 9, 2026), Cloudian/WEKA Lenovo SKU and reseller-agreement evidence, published 4+1 model. Assessment notes: InfiniBox components added in grading review (assessor initially missed the April 2026 Infinidat close; caught by human review). xIQ Agent Platform and xIQ Hybrid Cloud Platform descored to dated watch-lists under the GA-gate — announced January 6, 2026, no product documentation or part numbers as of July 11, 2026. Fabric DAPM reasoned from the customer's seat per the instrument's v2.5 Dell ruling: resale-vs-brand is vendor-to-vendor framing and does not move a customer-seat rating. ## Summary Finding Lenovo has one of the most credible on-prem AI infrastructure foundations in the market. Among the on-prem anchors on this instrument, each builds from what it owns: Dell from servers and storage, HPE from sovereign compute and three owned fabrics, Cisco from network silicon and security. Lenovo's anchor is manufacturing scale and thermal engineering — server breadth, Neptune sixth-generation liquid cooling, eight of the top ten public clouds as customers — and it is the only one of the four that owns neither a network fabric nor, until the April 2026 Infinidat close, any opinion-bearing storage software. Above the rack, capability arrives predominantly as other vendors' IP validated and SKU'd on Lenovo paper: ONTAP, Cloudian, WEKA, Red Hat, Nutanix, NVIDIA. The authority profile splits by lane. The open-substrate lanes genuinely delegate: Cloudian behind the S3 standard, Red Hat and the AI Library's runtime path on open substrates, TruScale on Slurm and Kubernetes. The proprietary partner platforms do not — ONTAP, WEKA, and Nutanix opinions cede to their owners through Lenovo's paper, because a proprietary platform's opinions have no exit regardless of the invoice. Validated solutions don't capture opinions for Lenovo; the proprietary ones capture for their owners. One contract with Lenovo is a procurement fact, not an authority fact — in both directions. Where Lenovo does capture, three surfaces matter, two of them new in 2026: InfiniBox and InfiniGuard (the Infinidat close gave Lenovo its first owned, opinion-bearing storage software), XClarity One fleet management, and — the inversion worth naming — the AI Library's first-party agents at Layer 3. Lenovo's strongest software capture surface is its applications, not its infrastructure software. The buyer who feels safely un-captured at the infrastructure layers is accumulating workflow opinions in Lenovo-built agents at the top. Everything GPU-aware above the rack is NVIDIA's judgment — Run:ai and NVIDIA AI Enterprise ride Lenovo paper as five-year subscription SKUs, which changes the invoice, not the authority. Same boundary as Dell and HPE. Two layers are unclaimed: no pipeline layer (1C) and no reasoning plane (2C) — no owned product, no packaged open-source alternative, and, unlike HPE's Kamiwaza designation or Dell's briefing-confirmed ecosystem strategy, no designated partner and no stated position. Whether those gaps are strategy or omission is this row's open briefing question. The buyer's trade: the broadest menu and the lowest OEM lock-in above the rack, in exchange for owning more of the assembly — and both missing planes — themselves. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Lenovo Strength ### Vendor-Provided Components **ThinkSystem GPU Servers (SR/SC lines, HGX B300, NVL-class)** [DAPM: Retained] Deskside to rack-scale on commodity x86/NVIDIA substrate — workloads move to Dell or HPE equivalents without rebuilding. 8x lower cost-per-token vs. cloud IaaS and sub-six-month ROI are Lenovo's claims for the inference platforms. **Neptune Direct-Water Cooling (6th Generation)** [DAPM: Retained] Lenovo-owned liquid-cooling IP; ~40-50% energy reduction claims; supports 100% liquid-cooled high-density Blackwell deployments. Physical plant accumulates no portable opinions — no lock-in surface. **AI Cloud Gigafactory with NVIDIA** [DAPM: Ceded] Turnkey rack-scale for AI cloud providers: pre-integrated Lenovo infrastructure, NVIDIA accelerated computing (B300, GB300 NVL72), Lenovo manufacturing, and Hybrid AI Factory lifecycle services. Integrated-system rule: the deployment's integration opinions cannot be lifted to another vendor as built. **Hybrid AI Factory Network Fabric (NVIDIA Spectrum-X)** [DAPM: Ceded] The factory's deployed fabric, sold and integrated by Lenovo as NVIDIA-branded product. Fabric opinions are embedded in the Spectrum-X stack with no abstraction layer that would make them portable to alternative switching — same customer-seat result as Dell's PowerSwitch. **ThinkStation PGX (Deskside AI)** [DAPM: Retained] ~1 petaflop personal AI, models to 200B parameters. Commodity deskside substrate, substitutable across OEMs. ### NVIDIA-Provided Components **GPU/Accelerator Silicon** Blackwell and Blackwell Ultra (B300, GB300 NVL72) now; Vera Rubin on the roadmap. The compute engines Lenovo builds around. **NVLink / NVSwitch** Intra-node high-bandwidth interconnect defining memory and compute topology. **Spectrum-X Ethernet Silicon** Consumed as NVIDIA-branded product — Lenovo holds no mechanical or brand authority over its fabric, a thinner position than Dell's branded PowerSwitch. ### Gap Analysis Lenovo's Layer 0 credibility is physical: ThinkSystem server breadth from deskside to NVL72-class racks, Neptune sixth-generation direct-water cooling (genuine Lenovo IP with a decade of heritage and ~40-50% energy-reduction claims), and manufacturing scale that powers eight of the top ten public clouds. The AI Cloud Gigafactory packages that manufacturing muscle for AI cloud providers — pre-integrated infrastructure, Lenovo manufacturing, and full-lifecycle Hybrid AI Factory services. What Lenovo does not own is networking. Dell at least brands NVIDIA silicon as PowerSwitch; HPE owns three fabrics. Lenovo resells NVIDIA Spectrum-X switches as NVIDIA-branded product, with a Cisco 800GbE Silicon One option in the Hybrid AI Factory's second phase — the one vendor whose networking aisle stocks two other vendors' captive fabrics. Either way the customer's fabric opinions are ceded to someone else's stack. ### Borrowed Judgment From the buyer's seat the fabric ratings are unaffected by whose logo is on the switch: fabric opinions are captive to the Spectrum-X (or Cisco Silicon One) stack regardless of whether the OEM brands, integrates, or merely resells it. What the resale posture does change is Lenovo's own leverage — Lenovo holds less mechanical authority over its networking story than Dell holds over PowerSwitch, and none of HPE's owned-fabric depth. Silicon roadmap authority is NVIDIA's throughout. ### Working Notes Watch-list (checked July 11, 2026, not scored): Vera Rubin NVL72 systems on Lenovo paper — NVIDIA states partner availability H2 2026; the AI Cloud Gigafactory names future Vera Rubin support. Scores when Lenovo documentation confirms GA. Confidence flag: GB300-class ThinkSystem GA is scored from CES/GTC 2026 launch language and Gigafactory support claims (B300/GB300), not from a per-model Lenovo Press product guide — doc inference, not confirmed fact. Fact question: is the Cisco Silicon One fabric option orderable in the Hybrid AI Factory today, or second-phase roadmap? Scored conservatively as menu context, not a component, pending confirmation. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Owned High-End + Partner-Delivered Strength ### Vendor-Provided Components **InfiniBox G4 / InfiniBox SSA G4 (InfuzeOS)** [DAPM: Ceded] High-end enterprise storage, owned Lenovo IP as of April 9, 2026. InfuzeOS opinions — Neural Cache behavior, provisioning, replication topologies — are captive to the platform, the same litmus result as Dell PowerScale and HPE Alletra. **InfiniGuard G4 + InfiniSafe Cyber Resilience** [DAPM: Ceded] Backup appliance with built-in cyber resilience. Recovery policies and immutability configurations do not lift to another vendor's platform — matches Dell's built-in cyber-resilience Ceded. **ThinkSystem DM/DG Arrays (NetApp ONTAP)** [DAPM: Ceded] Unified and capacity-flash arrays running NetApp ONTAP — a platform rated strong at this layer on NetApp's own row. The lift-to-leave litmus is indifferent to the paper: ONTAP opinions (SVMs, SnapMirror policies, tiering) are NetApp-captive with no exit to a non-ONTAP platform. Moving from Lenovo DM to NetApp AFF changes the invoice and keeps the captivity — channel substitution, not authority. Ceded-to-NetApp, matching NetApp's own row. **ThinkSystem DE Series (NetApp SANtricity)** [DAPM: Ceded] E-Series-based SAN arrays. Same litmus result: SANtricity opinions are NetApp-captive through any paper; portability within a single-owner ecosystem is channel substitution. **Cloudian HyperStore (Lenovo SKUs)** [DAPM: Delegated] S3-compatible object storage sold under Lenovo part numbers (7S0R-series license and support bundles). S3 multi-vendor standard keeps object opinions portable; the March 2026 Lenovo Validated Design adds GPUDirect/RDMA paths for AI workloads. **WEKA ThinkSystem SDS Ready Nodes + Certified AI Data Platforms** [DAPM: Ceded] WEKA productized under a global Lenovo reseller agreement (2023, 160+ markets), plus certified designs with DDN and VAST. WEKA is a proprietary single-implementation platform: filesystem and tiering opinions are WEKA-captive through any paper — Ceded-to-partner, matching WEKA's chips on the Supermicro row. The DDN/VAST certifications are not productized deliveries and are context, not components. ### NVIDIA-Provided Components **GPUDirect Storage / RDMA** GPU-direct data paths in the Cloudian and partner reference architectures (20GB/s per node claimed). The only NVIDIA surface at this layer — the near-empty column is itself a finding: Lenovo's storage acceleration story arrives through partner reference architectures, not first-party engineering. ### Gap Analysis Why strong: the capability axis credits what a customer can deploy on Lenovo paper today — the same rule that scores every OEM strong at Layer 0 on NVIDIA silicon — and Lenovo's deployable storage menu clears the bar twice over. The high-end tier is first-party: the Infinidat acquisition closed April 9, 2026, giving Lenovo InfiniBox G4 (hybrid), InfiniBox SSA G4 (all-flash), and InfiniGuard G4 (cyber-resilient backup) — owned, opinion-bearing storage software (InfuzeOS, Neural Cache, InfiniSafe) operating as a business unit in the Infrastructure Solutions Group. The productized partner tiers ride Lenovo paper with Lenovo support: ONTAP-based DM/DG arrays (a platform rated strong on NetApp's own row), SANtricity-based DE arrays, Cloudian HyperStore under Lenovo part numbers, WEKA as ThinkSystem SDS Ready Nodes under a global reseller agreement. That menu — an owned high-end platform plus four productized partner platforms — is comparable in kind to the seven-platform breadth that grades strong on the Supermicro row, with the owned exception Supermicro lacks. What remains missing at every tier is a Lenovo governance catalog — no MetadataIQ equivalent, no Data Fabric/Polaris equivalent. The governance surfaces that exist arrive with the platforms, so the function is deployable; the absence of a Lenovo-owned one is an authority fact the DAPM column carries, and a seam: each platform governs its own namespace, and nothing Lenovo-owned spans them. ### Borrowed Judgment Split by tier, and the buyer's authority position is decided at menu time. Buy InfiniBox and the enterprise cedes storage opinions to Lenovo (low borrowed judgment; Lenovo owns the IP). Buy DM/DG, DE, or WEKA and the opinions cede to NetApp or WEKA — proprietary platforms whose opinions have no exit through any paper. Buy Cloudian and the S3 standard keeps object opinions portable (Delegated, the genuine article). Nothing at this layer inherits NVIDIA judgment beyond partner-RA acceleration. ### Working Notes InfiniBox components added in grading review — the April 2026 Infinidat close was initially missed and caught by human review; it is Lenovo's Dataloop moment, the first owned software asset in the data layer. The v1.0 Delegated call on DM/DG rested on shared-ONTAP portability across paper; retracted in v1.1 — moving between sellers of one proprietary stack is channel substitution, and the chips now match NetApp's own row. Confidence flag: DE series currency scored from inference. Fact question for a briefing: is there any Lenovo-owned data-governance or metadata product not visible in documentation? ## ◑ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Validated Retrieval — Partner Stack ### Vendor-Provided Components **AI Library RAG Use Cases (Retrieval Infrastructure Facet)** [DAPM: Delegated] Validated retrieval pipelines — NeMo Retriever embedding, open-source vector store, LangChain-class chaining — deployed on the documented Hybrid AI Software Platform (Kubernetes/OpenShift + NIM). Indices and pipelines lift to any deployment of the same open stack. **Cloudian AI Data Platform (Lenovo Validated Design)** [DAPM: Delegated] Object-native data services for AI pipelines over GPUDirect/RDMA, per the March 2026 LVD. Partner data services behind the S3 standard — opinions move with Cloudian, not Lenovo. ### NVIDIA-Provided Components **NVIDIA NeMo Retriever** Embedding and reranking intelligence for the validated RAG pipelines. The retrieval judgment is NVIDIA's end to end — this is the load-bearing column at this layer. **NVIDIA RAG / AI-Q Blueprints** Pipeline patterns deployed through the AI Library on the documented Hybrid AI Software Platform stack. Identical blueprints ship on Dell AI Factory and HPE Private Cloud AI — non-differentiating. ### Gap Analysis Layer 1B is where the HPE row's observation lands on its named example: the NVIDIA RAG reference stack is one 'any NVIDIA partner (Dell, Lenovo, Supermicro) could deploy identically.' Lenovo owns no retrieval intelligence — no vector database, no embedding engine, no search partnership of Dell's Elastic kind, no owned namespace of HPE's Data Fabric kind. What Lenovo provides is a working retrieval capability in one procurement: AI Library validated RAG use cases on the documented Kubernetes/NIM platform, and the Cloudian AI Data Platform's object-native data services in the March 2026 LVD. This is the weakest moderate among the OEM rows — real deployable capability on Lenovo paper keeps it above gap; zero owned retrieval IP and the all-Delegated component mix carry the ownership finding. ### Borrowed Judgment High — structurally like HPE's 1B (the highest borrowed judgment in that row) but without even an owned storage substrate underneath it. Embedding intelligence is NVIDIA's, vector storage is open-source, data services are Cloudian's. The mitigation is the same as everywhere on this row: the opinions lift, so the borrowed judgment is at least not captive judgment. ### Working Notes One-contract procurement confirmed by SKU evidence (Cloudian part numbers, WEKA reseller agreement) — a procurement fact, not an authority fact. The AI Library appears at 1B, 2B, and Layer 3 by design; each cell scores a distinct facet (retrieval infrastructure here; deployment/runtime path at 2B; business capability at Layer 3). Confidence flag: the specific vector store in AI Library RAG deployments is inferred from the NVIDIA reference pattern (Milvus); Lenovo Press deployment guides would confirm. No retrieval-quality observability (recall@k, latency percentiles) that a Layer 2C could consume — same universal finding as the Dell row. ## ○ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Enterprise Responsibility ### NVIDIA-Provided Components **NVIDIA Blueprints / NIM Pipeline Templates** Pre-built pipeline patterns arriving through the AI Library — templates, not pipeline infrastructure. The thin column is the finding. ### Gap Analysis The 4+1 model defines Layer 1C as a required function — ETL/ELT, transformation, lineage, cost-aware orchestrated movement. Lenovo does not provide it. There is no owned pipeline product (no Dataloop equivalent), no packaged open-source pipeline stack (no Ezmeral equivalent), and no SKU'd marquee pipeline partner. What Lenovo's portfolio does move is data at the storage layer: ONTAP SnapMirror and FlexCache on DM/DG paper, InfiniBox replication, WEKA tiering, Cloudian AIDP ingest — all storage-platform features scored where the platforms are scored, at 1A. Storage replication is not ETL, and use-case-embedded data preparation inside AI Library agents is not pipeline infrastructure. Calibration: Dell's moderate rests on owned Dataloop orchestration; HPE's moderate rests on Data Fabric policy movement plus the Ezmeral packaging of Airflow/Kubeflow/Spark. Crediting Lenovo's nothing as moderate would make the score mean different things across rows. The enterprise brings its own pipeline tooling and owns the function. ### Borrowed Judgment There is no judgment to borrow — the enterprise retains full responsibility for pipeline logic on Lenovo infrastructure. Like the reasoning-plane gap, this responsibility is mostly implicit: it becomes visible when production data volumes expose it. ### Working Notes Sub-threshold signals, named in prose per the gap-cell rule: storage-native movement (SnapMirror/FlexCache, InfiniBox replication, WEKA tiering, Cloudian AIDP ingest) scored at 1A; NVIDIA blueprint pipeline templates via the AI Library; Veeam Kasten partnership (GTC 2026) for Kubernetes data protection. Fact questions for a briefing: is there any Lenovo-SKU'd pipeline or orchestration product not visible in documentation, and does the AI Library ship standalone data-prep/pipeline use cases (as opposed to prep embedded inside application agents)? Either fact would promote the cell. Re-examined July 11, 2026 under the corrected whose-paper and exposure rules (the tests that graded Supermicro's 1C strong): Lenovo productizes no pipeline-engine-bearing platform — the VAST/DDN relationships are certifications, not named products, and WEKA's movement features are filesystem properties scored at 1A. The gap held. ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Fleet Management + Consumption Orchestration ### Vendor-Provided Components **XClarity One** [DAPM: Ceded] Unified management platform — fleet lifecycle, zero-trust management, monitoring, automation; hybrid SaaS or on-prem local VM. Fleet templates, policies, and firmware baselines are captive to Lenovo's management plane — the OpenManage precedent. **TruScale GPUaaS (LiCO Orchestration)** [DAPM: Delegated] Lenovo-operated consumption model: metered NVIDIA GPU resources, workload scheduling and fair-share across tenant organizations. The scheduling substrate is Slurm/Kubernetes — genuine multi-vendor standards, so job definitions and manifests port to any Slurm/K8s environment (managed-service-behind-a-standard-interface rule). The captive surfaces (LiCO portal templates, TruScale metering) are thin convenience layers, not a proprietary placement engine — which is what separates this from Run:ai's Ceded. ### NVIDIA-Provided Components **NVIDIA Run:ai** GPU scheduling, quotas, fair-share — the Layer 2A function itself. Rides on Lenovo paper as a five-year subscription SKU (7S020050WW), which changes the invoice, not the authority: scheduling judgment is NVIDIA's regardless of whose paper the license rides on. **NVIDIA Mission Control** AI factory at-scale cluster management for the NVL-class deployments. **GPU / Network / NIM Operators + Base Command Manager** The documented Hybrid AI Software Platform's provisioning and orchestration substrate is NVIDIA's operator stack end to end. ### Gap Analysis Lenovo clears Dell's 2A bar and stays well under HPE's. The documented, orderable baseline: XClarity One (the go-forward unified management platform — fleet lifecycle, zero-trust management, hybrid SaaS or on-prem), TruScale GPUaaS (metered GPU consumption with workload scheduling and fair-share across tenant organizations via LiCO), and Run:ai/NVAIE riding Lenovo paper. TruScale GPUaaS is genuine workload-aware orchestration that Dell's rack-level-management cell explicitly lacks; nothing here approaches HPE's GreenLake Intelligence agentic mesh. The same OEM boundary applies throughout: Lenovo manages the fleet and the consumption model; NVIDIA manages everything GPU-aware inside the cluster. ### Borrowed Judgment High for GPU-aware orchestration — NVIDIA's judgment inside the cluster, Lenovo's around it, the same boundary as Dell and HPE. The distinctly Lenovo wrinkle: Lenovo has first-party scheduling judgment (LiCO), but its standalone SKUs are withdrawn, so the enterprise can now only consume that judgment by delegating operations to Lenovo through TruScale — an owned scheduler receding into a managed-service ingredient. ### Working Notes Watch-list (announced January 6, 2026; checked July 11, 2026, not scored): xIQ Hybrid Cloud Platform (AIOps + FinOps + DevOps across hybrid/multi-cloud) — production customers exist (DMEGC, possibly on the China-side xIQ Cloud predecessor) but no product documentation or part numbers confirm orderability; scores at doc-confirmed GA. Watch-list (June 2026, not scored): NVIDIA NemoClaw skills for AIOps — pre-GA, consistent with NemoClaw's alpha treatment on the Dell and NVIDIA rows. Fact questions: LiCO's status (standalone product guide and K8s part numbers withdrawn, yet lenovo.com still markets it and TruScale documents it as the orchestration engine); who operates TruScale GPUaaS day-2, and can a customer self-operate LiCO under a TruScale contract? ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Validated Runtimes + Services ### Vendor-Provided Components **AI Library Agent Deployment (Runtime Path Facet)** [DAPM: Delegated] Production-ready agents deployed onto the documented open substrate — Kubernetes/OpenShift with NIM microservices. The runtime path is standard and swappable; content is curated catalog, matching Dell's marketplace Delegated. **AI Services (AI Discover, AI Fast Start, Agentic Lifecycle Support)** [DAPM: Retained] Human-delivered strategy, deployment, and operation — proof-of-concept to production in as little as three months. Services, not software: expertise transfers to the customer, no captive opinion layer. Dell Accelerator Services precedent. **Hybrid AI Platform with Red Hat AI (CPU-Only Inference)** [DAPM: Delegated] Red Hat AI Enterprise on Xeon 6 (documented in the Hybrid AI 221 platform guide) — ~2x concurrent request claims for CPU inference. The Red Hat substrate is the instrument's portability benchmark: OpenShift AI/vLLM opinions lift to any OpenShift deployment. **Nutanix Enterprise AI on ThinkAgile HX650a** [DAPM: Ceded] Partner inference runtime on Lenovo paper. Nutanix is a proprietary single-vendor platform: its runtime opinions are Nutanix-captive through any channel — running identically on Dell or HPE gear is tin substitution, not an exit. Ceded-to-partner, matching the Nutanix chips on the Supermicro row. ### NVIDIA-Provided Components **NVIDIA AI Enterprise + NIMs** The model-serving runtime the primary path executes on, pre-validated on the Hybrid AI Platform configs and SKU'd on Lenovo paper (7S02001HWW). Inherited judgment, same as the Dell and HPE rows. **Dynamo** Distributed inference with KV-aware routing on the NVL-class racks — single-variable cache-locality optimization, not placement policy. **OpenShell / NemoClaw** Agent runtimes named in Lenovo's GTC 2026 materials. Alpha — not GA, watch-listed, consistent with their treatment on the Dell and NVIDIA rows. ### Gap Analysis Lenovo's 2B menu is broader than either peer's primary SKU — four runtime paths ship today. The NVIDIA path (AI Enterprise + NIMs on validated Hybrid AI Platform inference configs, RTX PRO 6000/4500 Blackwell). A CPU-only inference platform with Red Hat AI Enterprise on Xeon 6 (June 2026) — a non-NVIDIA lane Dell's AI Factory doesn't lead with. Nutanix Enterprise AI on ThinkAgile HX650a. And Lenovo's own agentic lane: the AI Library's production-ready agents (one-click deployment of autonomous and long-running agents, one-week-to-production claims, independently validated results). Lenovo owns no runtime software on any path — the AI Library's agents deploy onto the documented Kubernetes/NIM stack, and 'Lenovo Agentic AI' is a lifecycle program (services + library + platforms), not a discrete runtime product. Path diversity is the differentiator; ownership is not. ### Borrowed Judgment High on the primary path — model execution, serving optimization, and guardrail defaults are NVIDIA's, inherited through the dependency column. Genuinely lower than Dell's overall: the Red Hat lane is a real open-substrate alternative on Lenovo paper, the Nutanix lane is a captive alternative (path diversity without portability), and services expertise (AI Discover, AI Fast Start) transfers to the customer. ### Working Notes Watch-list (announced January 6, 2026; checked July 11, 2026, not scored): xIQ Agent Platform — no-code agent creation and deployment with built-in governance. No product documentation, part numbers, or non-China customers confirm orderability; scores at doc-confirmed GA. At GA the no-code capture question becomes live: agents authored in a proprietary no-code builder are opinions with nothing to export. Watch-list: 'limited-access co-development programs' (June 2026) are pre-GA by definition. The Knowledge Super Agent's runtime substrate is the documented platform stack (K8s pods on NIM/NVAIE per lp2311) — doc-supported inference. The AI Library facet scored here is the deployment/runtime path; its business capability is scored at Layer 3. ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Enterprise Responsibility ### NVIDIA-Provided Components **AI-Q Reference Architecture** Multi-agent workflow scaffolding. Does not make placement decisions. **Dynamo KV-Aware Routing** Performance-aware routing (single variable). Not multi-variable policy optimization. **OpenShell Governance** Runtime security sandboxing — Layer 2B constraint enforcement, not 2C placement reasoning. Alpha — not GA. ### Gap Analysis The 4+1 model defines Layer 2C as a required function — policy-driven decisions about where compute runs relative to data, which model serves which request, and how cost, compliance, and latency are arbitrated in real time. Lenovo does not provide this function, and applying the 'routing is not reasoning' test disposes of every candidate: TruScale GPUaaS schedules (a 2A function); Dynamo routes on cache locality (single variable); AI-Q is workflow scaffolding; the AI Library's governance is human-in-the-loop services configured around individual agents — governance-as-a-service, not a policy engine. Unlike HPE, Lenovo has designated no 2C partner: there is no Kamiwaza-equivalent in the Hybrid AI Advantage ecosystem. And unlike Dell — whose gap is briefing-confirmed deliberate ecosystem strategy — Lenovo has no stated position on the reasoning plane at all, so the gap reads as an omission rather than a decision. The enterprise must build custom 2C logic, bring a partner, or operate without it; most will choose the third and discover the gap when production agentic workloads expose it. ### Borrowed Judgment There is no judgment to borrow — the enterprise retains full responsibility for this function, mostly without recognizing Layer 2C as a distinct function it needs to provide. ### Working Notes The one first-party claimant — the xIQ Agent Platform's 'built-in governance' — is watch-listed with its platform (announced January 6, 2026, no product documentation as of July 11, 2026), and even at GA it would be Intelligence-2C (which agent may act, under what policy), not Infrastructure-2C placement. The FinOps placement language in the watch-listed xIQ Hybrid Cloud Platform is ops management, not request-time inference placement. The live-placement gap is universal across the OEM rows — noted as an instrument-wide finding, not a Lenovo-specific defect. Briefing question: does Lenovo see the reasoning plane as ecosystem territory, xIQ roadmap territory, or not at all? The answer moves the summary's read of the vendor, not the cell. ## ◑ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** First-Party Agents + ISV Ecosystem ### Vendor-Provided Components **AI Library Vertical Agents + Knowledge Super Agent (Business Capability Facet)** [DAPM: Ceded] Lenovo-built, production-validated agents for enterprise workflows: knowledge management (30% time reduction, independently validated), predictive maintenance, quality inspection, customer engagement. First-party application IP — workflow opinions accumulate in Lenovo's agent builds and do not lift to a competing framework. Owned-IP rule, same logic as InfiniBox at 1A. **Lenovo AI Innovators Ecosystem** [DAPM: Delegated] 50+ ISVs, 165+ validated solutions across vision AI, GenAI, and vertical use cases. Substitutable partners on Lenovo paper — matches Dell's AI Ecosystem Program and HPE's Unleash AI. **Centific Multimodal Agentic LVD** [DAPM: Delegated] Documented Lenovo Validated Design (lp2180) for multimodal agentic AI solutions. Partner IP, substitutable. ### NVIDIA-Provided Components **NIMs + Blueprints (Application Substrate)** The execution substrate under the AI Library agents. Identical blueprints are available on every NVIDIA partner — the differentiation is the Lenovo-built agent layer above them, not the substrate. ### Gap Analysis Two lanes, and that is the story of this cell. Lane one is the standard OEM play: the AI Innovators ecosystem (50+ ISVs, 165+ validated solutions across vision AI, GenAI, and verticals) plus documented validated designs like the Centific multimodal agentic LVD. Lane two is what Dell and HPE do not have: first-party application content. The AI Library's agents — the Knowledge Super Agent above all — are Lenovo-built business capabilities with independently validated production outcomes (30% reduction in knowledge-task time, up to 120 hours per employee annually), deployable in one to two weeks across manufacturing, retail, and healthcare workflows. That first-party lane is why this cell breaks from the peer rows' partner status: the methodology reserves partner for a layer addressed entirely through an ISV ecosystem, and Lenovo's is not. It is also the capture lane — a workflow built into a Lenovo agent keeps running on the open substrate without Lenovo, but only evolves with Lenovo, and its opinions do not lift to another vendor's agent framework. Self-deployable is not Retained. The ISV lane carries the same finding as every peer row: each partner governs within its own domain, and nothing governs across domains — which points back at the 2C gap. ### Borrowed Judgment Distributed across ISV partners, architecturally correct at Layer 3 — except in the first-party lane, where the enterprise inherits Lenovo's workflow judgment and cannot take it elsewhere. That inversion is the row's signature finding: the OEM whose infrastructure software captures least captures most through its own applications. ### Working Notes The AI Library facet scored here is the business capability; its runtime path is scored at 2B and its retrieval pipelines at 1B. Soft spot, low stakes: whether AI Library agents are cleanly products or services-configured deliverables per engagement — the May 2026 release and independent validation support product-grade repeatability; a briefing would firm the boundary. Instrument follow-up (logged, not this row's edit): re-test Dell's and HPE's Layer 3 partner status against the 'entirely ISV' boundary under this cell's lens — read here as both staying partner, since ISV frameworks and deployment services are not vendor-built business applications. ════════════════════════════════════════════════════════════════════════════════ # NetApp Intelligent Data Infrastructure Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.4 - AIPod Mini Whose-Paper Ruling **Date:** July 13, 2026 **Source:** NetApp INSIGHT 2025 (Oct 14-16, 2025), NVIDIA GTC 2026 (Mar 16-19, 2026), netapp.com / docs.netapp.com (ONTAP, AFF/ASA/AFX, StorageGRID, AIPod, AI Data Engine, NetApp Console/BlueXP, Trident), AWS/Azure/GCP first-party service docs, analyst/press coverage (TechTarget, Blocks & Files, StorageMath), published 4+1 model. v1.1 (instrument reconciliation): clarified that Layer 1C credits a data-movement decision/orchestration capability (policy-driven replication/tiering/caching), not transparent transport throughput — distinguishing NetApp's movement decisions from CoreWeave-style byte-transport acceleration. v1.2: Trident at 2A corrected Ceded→Delegated — the consumed interface is CSI/Kubernetes, a multi-vendor standard (Dell v2.5 CSI Operator precedent), and Trident is Apache 2.0; the v1.1 Ceded call confused what the driver connects to (NetApp storage, captured at 1A) with where provisioning opinions accumulate (Kubernetes-standard objects). 1A note added recording the Lenovo DM/DG channel pairing (capture is NetApp's through any paper, per the July 2026 OEM-channel retraction). v1.3 (same day, /ga-check): AI Data Engine promoted on doc-confirmed GA (docs.netapp.com/us-en/ai-data-engine — AIDE 9.18.1 U0/U1 and 1.0.0 release notes). 1B gap→moderate with AIDE vectorization/RAG retrieval endpoints scored Ceded; Metadata Engine added as a 1A component (Ceded); the watch-list's native-vector-DB question resolved native. Guardrails remain wrong-function at 2C (gap holds). 1C re-read resolved same day under the newly codified general-versus-fixed-function criterion (capability rule 4): AIDE's ingest pipeline scored as a Ceded 1C component, cell holds at moderate — the slice ships, the general half of the layer remains the enterprise's. v1.4 (July 13, 2026): NetApp AIPod Mini with Intel assessed against 2B and ruled out under the whose-paper rule (validated reference design on Supermicro compute, open-source OPEA runtime, SI-delivered — not a NetApp-paper product). No cell moved; recorded in 2B gap/notes as a second meet-in-the-middle path that confirms the storage-half finding. ## Summary Finding NetApp is a storage and data-management vendor whose authority concentrates in one of the most mature governed data foundations in the market (Layer 1A) and the storage-adjacent infrastructure around it — a software-defined storage fabric (Layer 0), best-in-class data movement (Layer 1C), and a hybrid-multicloud data control plane (Layer 2A). For the AI execution, reasoning, and application layers, NetApp is the data half of a meet-in-the-middle AI factory (AIPod): the compute, model-serving runtime, retrieval, and applications are NVIDIA's, sold alongside through the channel — not NetApp's. The capture is decoupled and Oracle-shaped: open at the surface, captive beneath. Data is reached through standard protocols (NFS, SMB, S3) and travels freely — uniquely, as first-party services on all three clouds (Amazon FSx for NetApp ONTAP, Azure NetApp Files, Google Cloud NetApp Volumes). But the value accumulates in ONTAP's opinions — SnapMirror relationships, clone hierarchies, efficiency and snapshot policies, the data-management workflows enterprises build on — and those do not lift. Leaving ONTAP is a multi-year rebuild even though the bytes move freely; running it on a hyperscaler does not loosen the hold, because it is still ONTAP. Eight of ten scored components are Ceded, all in the data plane where the ONTAP opinions live. NetApp's push up the stack — the AI Data Engine — reached GA and is now scored: the Metadata Engine gives Layer 1A its AI-metadata catalog, and AIDE's vectorization and retrieval endpoints give NetApp a first-party, storage-integrated retrieval surface at Layer 1B (moderate — first-release, license-gated, and hardware-gated to AFX plus Data Compute Node clusters; the software-only deployment is metadata-only). The architect can now deploy NetApp-native AI data governance and retrieval on an AFX estate; everywhere else, the mature data foundation plus data movement remains the deployable core. The meet-in-the-middle structure is the defining authority finding. Unlike Dell and Cisco — primes that sell, curate, and support the full NVIDIA stack as their own AI factory, and are therefore credited for the runtime and retrieval they deliver — NetApp brings the storage half of AIPod while NVIDIA brings the compute, runtime, retrieval, and application ecosystem, and the channel integrates. Each vendor is credited for its half: NetApp owns the data plane; NVIDIA owns the intelligence. The NVIDIA dependency is heavy exactly where NetApp partners (Layer 0 compute, Layer 2B runtime, Layer 1B retrieval, Layer 3 applications) and near-empty where NetApp owns the layer (1A, 1C, 2A). The buyer's trade: the buyer gets the most operationally proven, most portable, best-governed enterprise data foundation available — the safest home for data across on-prem and every major cloud — and in exchange accumulates ONTAP-specific opinions that are a multi-year lift to leave, even as the bytes stay open. For AI specifically, NetApp anchors the data; the intelligence is bought alongside from NVIDIA. NetApp's bet is that the governed data foundation is the durable position in the AI stack, and that when the AI Data Engine ships broadly it can convert that foundation into an AI-native data platform. The sharp contrast is VAST, its closest competitor: the same data-foundation strength, but where VAST verticalized into native retrieval and an agent runtime, NetApp partners for them while its own AI-native layer matures. ## ◑ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Software-Defined Storage Fabric ### Vendor-Provided Components **AFF / ASA / AFX (Disaggregated All-Flash Architecture)** [DAPM: Ceded] Proprietary all-flash storage systems running ONTAP. AFX disaggregates performance from capacity (AFX 1K controllers + NX224 NVMe enclosures, scale-out, DGX SuperPOD-certified); AFF/ASA cover unified and SAN all-flash. The architecture and its opinions cannot be lifted to another vendor as deployed. Proprietary NetApp platform — opinions captive, no open exit. **Storage Networking (NVMe-oF / NVMe-TCP / Ethernet)** [DAPM: Delegated] Standard fabric interfaces connecting compute to the storage estate — NVMe-over-Fabrics, NVMe/TCP, Ethernet, pNFS/RDMA. Substitutable switching against multi-vendor standards; the connectivity opinions lift. Mirrors how VAST's NVMe-oF fabric is scored. ### NVIDIA-Provided Components **NVIDIA DGX / DGX SuperPOD (AIPod Compute)** All AI compute in NetApp's AI factories is NVIDIA's: AIPod pairs ONTAP storage with NVIDIA DGX systems; AFF A90 and AFX are certified for NVIDIA DGX SuperPOD. NetApp supplies storage; NVIDIA supplies the accelerated compute. This is the compute half of the meet-in-the-middle. **NVIDIA-Certified Storage + AI Data Platform Reference Design** AIPod holds the NVIDIA-Certified Storage designation; AFX/AIDE align to the NVIDIA AI Data Platform reference design. The DX50 data-compute node ships with an NVIDIA L4 for storage-side data-engine acceleration (AIDE, now GA), not general AI compute. **Exception: AIPod Mini (Intel / OPEA)** AIPod Mini is the lone NVIDIA-free path — Intel Xeon 6 (AMX) compute with Intel's open-source OPEA RAG framework for departmental inference. Everywhere else, NetApp's AI compute is NVIDIA. ### Gap Analysis NetApp owns no silicon, GPUs, or networking fabric — but, like its closest competitor VAST, it owns a genuine software-defined storage architecture that belongs at Layer 0. AFX disaggregates ONTAP (separating performance from capacity: AFX 1K controllers + NX224 NVMe enclosures, scaling to 128 controllers and beyond an exabyte), connected over NVMe-oF. That fabric is NetApp's real Layer 0 capability, scored the way VAST's DASE/NVMe-oF architecture is scored at Layer 0. The AI compute is NVIDIA's. AIPod (NetApp storage + NVIDIA DGX), AIPod Mini (ONTAP + Intel/OPEA), and FlexPod AI (Cisco UCS + NetApp AFF + NVIDIA + OpenShift AI) are validated, channel-assembled designs — NetApp brings the storage half, partners bring the compute. So Layer 0 sits at moderate: a software-defined storage fabric (NetApp's own), with the acceleration and networking-for-GPU entirely NVIDIA's. Below Dell and Cisco (strong), which own compute silicon and networking; level with VAST, which scores its storage architecture at this layer the same way. ### Borrowed Judgment Multi-directional. NetApp retains its storage-fabric architecture (AFF/ASA/AFX, disaggregated ONTAP, NVMe-oF) — that is its Layer 0 IP. It borrows all accelerated-compute and GPU-networking judgment from NVIDIA (DGX, DGX SuperPOD, Spectrum/ConnectX in the AIPod designs) and server hardware from OEM partners (Lenovo, Cisco). The storage fabric is NetApp's; the compute fabric is NVIDIA's. ### Working Notes AIPod and FlexPod AI are reference/validated designs, channel-assembled ('meet in the middle') — NetApp does not manufacture or sell the GPU compute. AFX is GA and DGX SuperPOD-certified. Future support announced for NVIDIA RTX PRO Servers (Blackwell) and STX (BlueField-4 / Vera Rubin) is roadmap, not GA. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** NetApp Strength — Data Foundation ### Vendor-Provided Components **ONTAP Data Management (AFF/ASA/AFX, Cloud Volumes ONTAP, first-party FSx for ONTAP / Azure NetApp Files / Google Cloud NetApp Volumes)** [DAPM: Ceded] Multiprotocol data-management OS (NFS, SMB, S3, iSCSI, NVMe-oF) with SnapMirror, FlexClone, efficiency, and tiering — identical across on-prem and all three clouds. Standard protocols at the surface, but the ONTAP data-management opinions are captive: leaving ONTAP is a multi-year rebuild, and running it on a hyperscaler does not change that. Proprietary NetApp platform — opinions captive, no open exit. **StorageGRID (Object / S3)** [DAPM: Delegated] Geo-distributed S3-compatible object storage. The S3 interface is a multi-vendor standard — bucket, lifecycle, and policy opinions lift to any S3 platform without rebuilding. Delegated, mirroring the instrument's treatment of S3-interface object storage. **AI Data Engine — Metadata Engine (AI-Metadata Catalog)** [DAPM: Ceded] GA per product documentation: automated metadata extraction and cataloging of files and objects across local and peered ONTAP clusters; workspaces with RBAC and OIDC; Data Sync keeps catalogs current via policy-driven SnapMirror; centralized REST query and filtering APIs. Included with base ONTAP One licensing per the AIDE FAQ; full GPU-dependent capability gated to AFX + Data Compute Node clusters, with a metadata-only software deployment on customer RHEL servers (AIDE 1.0.0). The catalog structure and workspace opinions are captive to the NetApp platform. **Data Classification + Data Infrastructure Insights + Autonomous Ransomware Protection** [DAPM: Ceded] PII/sensitive-data classification and mapping (formerly Cloud Data Sense), observability with AI-driven anomaly detection (formerly Cloud Insights), and ML-based ransomware protection in ONTAP (on by default on recent AFF/ASA platforms). Compliance and resilience governance — proprietary opinions captive to NetApp's tooling, not an AI-metadata catalog. Proprietary NetApp platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No NVIDIA Dependency in the GA Data Foundation** ONTAP, StorageGRID, Data Classification, Insights, and Autonomous Ransomware Protection are NetApp IP over standard protocols — no NVIDIA. (The AI Data Engine's NIM-powered vectorization, now GA and scored at Layer 1B, adds an NVIDIA dependency inside the retrieval surface.) ### Gap Analysis This is NetApp's center of gravity and arguably the most mature data foundation on the instrument. ONTAP is a multiprotocol data-management OS (NFS, SMB, S3, iSCSI, NVMe-oF) with SnapMirror, FlexClone, efficiency, and tiering, running identically on-prem (AFF/ASA/AFX), as software (Cloud Volumes ONTAP), and as first-party services on all three clouds (Amazon FSx for NetApp ONTAP, Azure NetApp Files, Google Cloud NetApp Volumes). StorageGRID adds S3 object; Data Classification maps PII/sensitive data; Data Infrastructure Insights adds observability and anomaly detection; ONTAP Autonomous Ransomware Protection adds ML-based resilience, on by default in recent releases. It is the safest, most portable, most operationally proven data home in the market. The authority reading is Oracle-shaped. The protocols are standard and reassuring, but the value accumulates in ONTAP's opinions — replication topology, clone hierarchies, snapshot and efficiency policies, the feature-set workflows enterprises build on — and those are captive. The governance is real but compliance-flavored (PII classification, ransomware, observability), not the AI-metadata catalog a reasoning plane would query; that piece arrived: the AI Data Engine Metadata Engine reached GA (AIDE 9.18.1 U0; software-only on third-party RHEL servers as AIDE 1.0.0, April 30, 2026) — automated metadata extraction and cataloging across peered ONTAP clusters, workspaces with RBAC/OIDC, Data Sync via policy-driven SnapMirror, REST query APIs. The AI-metadata catalog a reasoning plane would query now exists on NetApp paper, scoped to ONTAP estates (full capability on AFX + Data Compute Nodes). NetApp earns strong here on data-foundation maturity and breadth — broader than any peer on pure enterprise data management, uniquely first-party across all three clouds — calibrated to VAST and Dell at this layer, now with the AI-native catalog shipping rather than forward-dated. ### Borrowed Judgment Low. ONTAP, StorageGRID, Data Classification, Insights, and ARP are NetApp IP — the data-management and governance opinions are Ceded to NetApp; StorageGRID's S3 surface keeps object opinions portable (Delegated). No partner or NVIDIA dependency in the GA foundation. ### Working Notes AI Data Engine Metadata Engine promoted from the watch-list July 12, 2026 on doc-confirmed GA (docs.netapp.com/us-en/ai-data-engine). Channel note: ONTAP also ships on Lenovo paper (ThinkSystem DM/DG) — the capture is NetApp's through any channel, and the Lenovo row scores those arrays Ceded-to-NetApp by design; the pairing is deliberate two-row structure, not drift. ## ◑ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** AIDE Retrieval — Storage-Integrated ### Vendor-Provided Components **AI Data Engine — Vectorization + RAG Retrieval Endpoints** [DAPM: Ceded] Data collections, embeddings, and retrieval endpoints created and served in AI Data Engine Console (AIDE 9.18.1 U1, license-gated), with guardrail policies defined in AIDE and bound to workspaces in System Manager. Storage-integrated: retrieval runs where the governed data lives, on AFX + Data Compute Node clusters. Proprietary NetApp surface — collection, embedding-pipeline, and guardrail opinions are captive to the platform. ### NVIDIA-Provided Components **NVIDIA NIM (Embedding Inside AIDE)** AIDE's vectorization pipeline is NVIDIA-NIM-powered — the embedding intelligence inside NetApp's retrieval surface is NVIDIA's, the same split as every storage vendor's GA retrieval. **Retrieval in AIPod (NeMo Retriever / NIM)** In the meet-in-the-middle AIPod, the co-sold retrieval path remains NVIDIA's NeMo Retriever / NIM within NVIDIA AI Enterprise. AIDE now gives NetApp a first-party alternative on AFX estates. ### Gap Analysis AI Data Engine reached GA and moved this cell. Per NetApp's product documentation (AIDE 9.18.1 U1; AIDE 1.0.0 followed April 30, 2026), AIDE creates data collections, embeddings, and retrieval endpoints in the AI Data Engine Console — a NetApp-owned, storage-integrated retrieval surface with guardrail-based governance, not embedding-prep for external vector databases. That settles the watch-list's open fact question in NetApp's favor. Why moderate and not more: the capability is first-release, license-gated (vectorization and RAG require the appropriate AIDE licenses), and hardware-gated — full capability requires ONTAP AI data platform clusters (AFX 1K storage nodes plus NetApp Data Compute Nodes), while the software-only third-party-server deployment (AIDE 1.0.0) is explicitly metadata-only with no vectorization, RAG, or GPU services. Calibration: at Dell's moderate (Elastic productized) and Nutanix's moderate (GA pgvector), still below VAST's strong (native vector search, mature, not gated to a single hardware line). Off-AFX estates and the AIPod channel path still consume NVIDIA's retrieval or bring their own. ### Borrowed Judgment Moderate. On AFX estates the enterprise can now inherit NetApp's retrieval judgment — collections, embeddings, retrieval endpoints, guardrails — with NVIDIA NIM providing the embedding intelligence inside it. The retrieval opinions accumulate in AIDE (workspaces, collections, guardrail policies) and are captive to the NetApp platform. Everywhere else the pre-AIDE reading holds: bring your own retrieval or consume NVIDIA's co-sold half of AIPod. ### Working Notes Promoted from gap July 12, 2026 on doc-confirmed GA (docs.netapp.com/us-en/ai-data-engine — release notes plus vectorization/data-collections and guardrails administration sections). The earlier watch-list question — native vector DB vs. embedding-prep to external DBs — resolved native: AIDE serves its own retrieval endpoints. Scope facts worth keeping: AFX 1K + Data Compute Node hardware gate for full capability; the AIDE 1.0.0 software-only path is metadata-only with a supported upgrade path when GPU services land there. ## ◑ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Best-in-Class Movement + AI-Ingest Pipeline ### Vendor-Provided Components **AI Data Engine — AI-Ingest Pipeline (Sync → Classify → Vectorize)** [DAPM: Ceded] Fixed-function AI data pipeline, GA per product documentation: Data Sync keeps sources current via policy-driven SnapMirror (AIDE 9.18.1 U0), guardrail classification screens sensitive data, NIM-powered vectorization transforms into embeddings and collections (U1) — terminating in AIDE's own retrieval surface. Fails the never-anticipated-workload test (no arbitrary stages, no external destinations, no lineage), so it is scored as the slice it is; the retrieval-endpoint facet is scored at 1B. Pipeline and policy opinions captive to the NetApp platform. **SnapMirror (Replication & Data Mobility)** [DAPM: Ceded] Async and sync replication across sites and clouds — the logistics layer for staging and distributing datasets, including AI training and inference data. ONTAP-to-ONTAP; the replication relationships and opinions do not lift to another platform. Proprietary NetApp platform — opinions captive, no open exit. **FlexCache + FabricPool (Caching & Cost-Aware Tiering)** [DAPM: Ceded] FlexCache caches hot data near compute; FabricPool auto-tiers cold data to object/cloud by temperature. A captive ONTAP caching and tiering engine — cost-aware data movement whose opinions are proprietary. Proprietary NetApp platform — opinions captive, no open exit. **BlueXP / NetApp Console Copy & Sync (Migration & Cloud Ingest)** [DAPM: Ceded] Moves and synchronizes data from external, file, and SaaS sources into the estate across hybrid multicloud. NetApp-proprietary sync and migration service; the movement opinions are captive. Proprietary NetApp platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No NVIDIA Dependency in Data Movement** SnapMirror, FlexCache, FabricPool, and Cloud Sync are NetApp IP — no NVIDIA. (The AI Data Engine's NIM-powered vectorization and Data Sync are now GA — see notes on the pending 1C re-read.) ### Gap Analysis NetApp's data movement is best-in-class for all data types, and it is a legitimate Layer 1C capability even though it is not marketed for AI: the layer's purpose explicitly includes movement, cost-aware movement, and cache tiering, and AI data logistics are exactly that. SnapMirror replicates and stages datasets across sites and clouds; FlexCache caches hot data near compute; FabricPool tiers cold data to object/cloud by temperature; BlueXP/Console Copy & Sync moves and synchronizes data from external and SaaS sources. An architect leverages these for AI today; the absence of an 'AI' label does not remove the capability. What earns the score is the movement-decision capability — configurable replication topologies, cost-aware tiering policies, and caching policy, where the enterprise accumulates opinions (the same opinions that make these Ceded to ONTAP) — not transparent byte-transport, which is a Layer 0/1A throughput property, not a Layer 1C decision. What keeps this cell at moderate is the general-versus-fixed-function rule. The AI Data Engine pieces of the transform story (Data Sync, vectorization, guardrail classification) reached GA in July 2026 and are scored below — but they are a fixed-function AI-ingest pipeline: fixed stages (sync, classify, vectorize) terminating in AIDE's own retrieval surface. Applying the decidable test — can it run a pipeline workload NetApp never anticipated? — the answer is no: there is no arbitrary transformation, no insertable stages, no choosable destination, and no lineage. The general half of the layer (ETL/ELT, transformation, lineage) remains the enterprise's, met with its own tooling (Airflow, Spark, Kubeflow). Best-in-class movement plus a shipping fixed-function slice is partial capability — moderate, calibrated against VAST's DataEngine (general, event-driven, arbitrary functions — strong) and Dell's Dataloop (general in kind, maturity-gated — the moderate anchor). This sits above Nutanix (gap — Time Machine is narrow copy-data-management, not a comprehensive movement platform) and below VAST (strong — DataEngine is a full event-driven AI pipeline platform). It is roughly level with Dell, for the inverse reason: Dell has pipeline IP (Dataloop) and thinner movement; NetApp has movement mastery and pre-GA pipelines. ### Borrowed Judgment Low for movement — SnapMirror, FlexCache, FabricPool, and Cloud Sync are NetApp IP, the data-movement opinions Ceded to NetApp. The AI-pipeline/transform/lineage half: the enterprise brings its own ETL/orchestration (Airflow, Spark, Kubeflow) today; NetApp's AIDE sync/vectorize/classify pieces are now GA and under re-read for whether they close that half. ### Working Notes Re-read resolved July 12, 2026 under the general-versus-fixed-function criterion (methodology, capability rule 4): AIDE's GA'd pipeline pieces are a purpose-built slice, so the cell holds at moderate with the slice scored as a component. What would move this cell toward strong: arbitrary/insertable transformation stages, destinations beyond AIDE's own surfaces, or lineage tracking — general-pipeline capability, documented. ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Data-Infrastructure Control Plane; No GPU Scheduling ### Vendor-Provided Components **NetApp Console (Hybrid Multicloud Data Control Plane)** [DAPM: Ceded] Unified provisioning, protection, governance, mobility, and health across the on-prem and multicloud storage estate (AFF/ASA/FAS/AFX, StorageGRID, FSx for ONTAP, Azure NetApp Files, Google Cloud NetApp Volumes). Proprietary control plane — orchestration opinions captive to NetApp. Proprietary NetApp platform — opinions captive, no open exit. **Trident (Kubernetes CSI Driver)** [DAPM: Delegated] Open-source (Apache 2.0), NetApp-maintained CSI driver provisioning NetApp storage to Kubernetes for AI and container workloads. The consumed interface is CSI/Kubernetes — a genuine multi-vendor standard: provisioning opinions live in Kubernetes-standard objects (StorageClasses, PVCs) and survive a swap to another vendor's CSI driver without rebuilding the layer. What the driver connects to is NetApp storage, and that capture is scored where it lives, at 1A. Matches the Dell CSI Operator precedent. ### NVIDIA-Provided Components **GPU Scheduling Is Not NetApp's** NetApp does no GPU scheduling, quotas, or fair-share. That function sits with NVIDIA (Run:ai / GPU Operator) and the enterprise's Kubernetes — the same universal Layer 2A gap that Dell and VAST carry. NetApp orchestrates the data infrastructure, not the AI compute. ### Gap Analysis NetApp orchestrates the data infrastructure, not the AI compute. NetApp Console (formerly BlueXP) is a mature, unified control plane for the entire storage and data estate across on-prem and all three clouds — provisioning, protection, governance, mobility, and health from one surface. Trident is NetApp's open-source Kubernetes CSI driver, dynamically provisioning ONTAP/StorageGRID/FSx/ANF storage to containerized and AI workloads. Together they are genuine infrastructure-orchestration capability for the data plane. The AI core of this layer — GPU scheduling, quotas, fair-share — NetApp does not provide; that belongs to the enterprise's Kubernetes and NVIDIA. So Layer 2A sits at moderate, calibrated to VAST and Dell: a storage/data control plane plus a vendor CSI driver, with GPU scheduling absent or borrowed from NVIDIA (universal across these peers). NetApp Console is a genuinely deep multicloud data control plane, comparable in kind to VAST Polaris; the GPU-scheduling gap is the same one Dell and VAST carry, so it does not dock NetApp below them. ### Borrowed Judgment NetApp owns the data-infrastructure control plane (Console) and the CSI driver (Trident) — Ceded to NetApp. AI-compute orchestration (GPU scheduling, quotas, fair-share) is not NetApp's at all; the enterprise's Kubernetes and NVIDIA own it. ### Working Notes Trident is open source (Apache 2.0, NetApp-maintained), but it is a NetApp-storage-specific CSI driver — the Kubernetes interface is portable, the driver is not. NetApp Console was renamed from BlueXP in 2025. ## ○ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Runtime is NVIDIA's (Meet-in-the-Middle) ### NVIDIA-Provided Components **NVIDIA NIM / AI Enterprise (AIPod Runtime)** The model-serving and inference runtime in a NetApp-anchored AI factory is NVIDIA NIM / AI Enterprise — the runtime half of the meet-in-the-middle AIPod, co-sold and channel-integrated, not a NetApp-delivered or NetApp-supported solution. NetApp supplies the governed data beneath. **Red Hat OpenShift AI (FlexPod AI)** In FlexPod AI (Cisco + NetApp), model serving comes from Red Hat OpenShift AI. Again a partner runtime; NetApp provides the storage. ### Gap Analysis NetApp has no model serving, no inference runtime, and no agent runtime — none. Where inference runs in a NetApp-anchored stack it is NVIDIA NIM / AI Enterprise (AIPod), Red Hat OpenShift AI (FlexPod AI), or open-source OPEA / Intel AI for Enterprise RAG on Intel Xeon 6 AMX (AIPod Mini with Intel) — three partner runtimes, none NetApp's. This is the runtime half of the meet-in-the-middle: NetApp brings the storage, NVIDIA brings the runtime, the channel integrates. Crucially, NetApp is not the prime that sells, curates, and supports the NVIDIA runtime as its own factory — unlike Dell (2B moderate, delivered NVIDIA runtime) or Cisco (2B moderate, own security IP over the runtime). So the capability is not credited to NetApp's row; it is NVIDIA's, scored there. The enterprise inherits NVIDIA's (NIM) or Red Hat's (OpenShift AI) execution layer entirely. NetApp is the data foundation beneath it. This sits below every storage and platform peer at Layer 2B — VAST (strong, built AgentEngine), Dell (moderate, delivered NVIDIA runtime), Nutanix (moderate, NAI serving) — because NetApp built no runtime IP and does not deliver the runtime as its own solution. ### Borrowed Judgment Total. NetApp provides no runtime. In a NetApp-anchored AI factory the runtime is NVIDIA's (NIM / AI Enterprise) or Red Hat's (OpenShift AI), co-sold via the channel rather than delivered by NetApp. NetApp is the storage beneath someone else's execution layer. ### Working Notes The meet-in-the-middle distinction is load-bearing: Dell and Cisco are primes that deliver and support the NVIDIA stack as their own factory (and are credited at 2B); NetApp is the storage half of a channel-assembled co-sell, so the NVIDIA runtime is credited on NVIDIA's row, not NetApp's. Assessed July 13, 2026 (NetApp AIPod Mini with Intel, solution brief SB-4332 + TR-5010): fails the whose-paper test and does not move this cell. Per NetApp's own documentation it is a 'validated reference design,' not a NetApp-paper product — compute is Supermicro (222HA-TN-OTO-37), switch is Arista, the runtime is open-source OPEA / Intel AI for Enterprise RAG cloned from GitHub, and delivery is through distributors and integration partners (Arrow, TD SYNNEX; Insight, CDW, Presidio, Long View). NetApp brings storage plus the reference architecture; it does not sell or support the runtime on its own paper. It is a second meet-in-the-middle path (Intel/OPEA-CPU alongside NVIDIA-GPU), which strengthens the storage-half finding rather than closing it. ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** No Reasoning Plane (Data Guardrails ≠ Agent Governance) ### NVIDIA-Provided Components **No NVIDIA Reasoning Plane Either** Neither NetApp nor NVIDIA provides a reasoning plane in this stack. As with Dell — which sells the full NVIDIA AI factory as prime and is still Layer 2C gap — NVIDIA's stack (AI-Q, Dynamo) is routing and scaffolding, not policy-driven placement. Selling or anchoring an NVIDIA factory does not create a 2C capability. ### Gap Analysis Applying the 'Routing Is Not Reasoning' test: NetApp has no agent governance, no model routing, no policy-driven inference placement, and no multi-agent orchestration. Its only adjacent capability is the AI Data Engine Data Guardrails, which scan, classify, and exclude sensitive data from AI/RAG pipelines — 'guardrails follow the data.' That is data-access governance (a Layer 1A-style claim), not agent-action governance (Intelligence-2C) and not request-time placement (Infrastructure-2C). It is the wrong function for this layer — a verdict GA does not change: the guardrails shipped (AIDE 9.18.1 U1) and the cell still reads gap, because what shipped governs data, not agents. This is a clean gap, consistent with the entire data and infrastructure cohort — Dell, Nutanix, VMware, VAST, NVIDIA, CoreWeave — none of which productizes a reasoning plane. The enterprise retains policy-driven placement and agent governance in full. The live-placement gap is universal across the instrument, noted rather than penalized. ### Borrowed Judgment Inverted — there is no Layer 2C to borrow. The enterprise retains policy-driven placement and agent governance entirely. NetApp's data guardrails (GA, scored at 1B) govern data, not agents; the meet-in-the-middle NVIDIA stack has no reasoning plane either. ### Working Notes AI Data Engine Data Guardrails reached GA (AIDE 9.18.1 U1) and are scored inside the 1B retrieval component, where they function. At 2C they remain the wrong function — data-access governance, not an agent or reasoning plane; the gap holds. ## ○ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** No Value Plane — Data Foundation Beneath Others' Apps ### NVIDIA-Provided Components **App Ecosystem in AIPod is NVIDIA's (AI Enterprise)** The application ecosystem delivered alongside AIPod is NVIDIA AI Enterprise — NIM microservices, blueprints, and models — NVIDIA's ecosystem, co-sold through the channel, not a NetApp-curated AI-application program. NetApp supplies the governed data beneath the apps. ### Gap Analysis NetApp ships no first-party AI applications and curates no AI-application ISV program of its own. Its ecosystem is infrastructure and compute (NVIDIA, Cisco, Lenovo, Intel, Red Hat) and channel (distribution and integration) — not value-plane ISVs. In the AIPod meet-in-the-middle, the application ecosystem is NVIDIA's (AI Enterprise) plus the customer's own apps; NetApp is the data foundation beneath the value plane, not a provider of it. This sits below the partner-ecosystem vendors at Layer 3 — Dell, HPE, and Cisco curate genuine AI-application ISV programs (Dell's OpenAI/Palantir/ServiceNow, HPE Unleash AI's 26+ ISVs) — and below VAST (moderate, Cosmos Community partner tracks). NetApp has neither a first-party app nor a curated AI-app ISV program, so it does not address the value plane even via partners: gap, not partner. The value plane is entirely the customer's, NVIDIA's, and ISVs'. ### Borrowed Judgment The value plane is the customer's plus NVIDIA's plus ISVs' — NetApp provides no AI applications and curates no AI-application ISV ecosystem. The enterprise brings its own apps; NetApp supplies the governed data beneath them. ### Working Notes AIPod and FlexPod AI position NetApp storage in AI solutions, but the application logic and ecosystem are NVIDIA AI Enterprise (NIM/blueprints/models) and the customer's — co-sold via the channel, not NetApp-curated. ════════════════════════════════════════════════════════════════════════════════ # Nutanix Cloud Platform with Enterprise AI Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.1 - Interface-Portability Reconciliation **Date:** June 20, 2026 **Source:** .NEXT 2026 (Apr 2026), .NEXT 2025, NAI 2.7 (May 2026), AOS 7.5 / AHV 11 / Prism Central pc.7.5, NUS 5.3, NDB v2.10, NKP 2.17, Nutanix Bible, NVIDIA vGPU product-support matrix for AHV, published 4+1 model. v1.1 (instrument reconciliation): 1A NUS Objects Retained→Delegated, resolving the internal split with NKP (2A) — proprietary implementation behind a multi-vendor standard is Delegated, not Retained. ## Summary Finding Nutanix occupies the same structural position as VMware — an HCI / private-cloud platform, an abstraction layer running on commodity x86 it does not manufacture — and against the 4+1 model it produces a strikingly similar profile: strongest at infrastructure orchestration (Layer 2A), competent across the data and runtime layers, and gapped at data pipelines (Layer 1C) and the reasoning plane (Layer 2C). Its deepest IP and authority sit in the orchestration control plane: Prism, AOS, and AHV. Two divergences lift Nutanix above a simple VMware echo. It owns its data foundation — Nutanix Unified Storage (file, object, block) plus Data Lens compliance governance — where VMware is bring-your-own-storage; this raises Layer 1A above VMware's, though it remains compliance governance, not the AI-metadata catalog Dell's MetadataIQ provides. And Nutanix Enterprise AI (NAI) is a genuinely good, less-NVIDIA-locked model-serving product: multi-source models (NVIDIA NIM, Hugging Face, or custom upload), a vLLM default engine, a CPU-only inference path, and portability across any CNCF-certified Kubernetes — arguably a cleaner standalone serving product than VMware's Model Runtime. The capture mechanism here is decoupled and invisible — the reassuring kind. The edges are open: commodity x86 hardware is Retained (swap Dell for HPE or Lenovo without rebuilding), NKP is unforked upstream Kubernetes (manifests lift to any conformant cluster), NAI runs on any cloud's Kubernetes and is not NIM-locked, and the models are open. The buyer feels unconstrained. Yet the value that actually accumulates — AOS/Prism orchestration opinions, Acropolis Dynamic Scheduling placement logic, NAI's serving and governance configuration, Data Lens policy — is captive to Nutanix. Eight of eighteen scored components are Ceded, concentrated exactly where operational opinions accumulate. Open hardware and open models do not make the control plane portable. Where the stack thins is identical to VMware. Layer 1C (data pipelines) is a gap: Time Machine and cross-cloud snapshots are storage mobility, not AI pipelines — there is no lineage and no cost-aware movement. Layer 2C (the reasoning plane) is a gap: Agent Gateway is a real, generally available agentic-governance control point (RBAC, rate-limiting, audit including MCP-call audit, health-based failover, capacity load-balancing), but governance and operational routing are not placement reasoning. There is no per-request, content/cost/quality model selection, and the data-side feeders a reasoning plane would query — an AI-metadata catalog at Layer 1A, lineage at Layer 1C — are absent, so data-relative placement is structurally out of reach. The forward 'Nutanix Agentic AI' full stack is Early Access (GA targeted second half of 2026) and does not move the cell. One Layer 0 caveat bounds the on-platform AI story: AHV virtualizes only vGPU-class accelerators (L40S, H100, L4, and a single Blackwell SKU — the RTX PRO 6000 Server Edition), not HGX training systems (B200/GB200), which remain NVIDIA bare-metal. Nutanix's on-platform AI is inference and fine-tune scale, not frontier training. The June 2026 'NVIDIA Certification' headline is a storage certification — Nutanix Unified Storage feeding external GPU servers over Spectrum-X and GPUDirect — not GPUs running on Nutanix. The buyer's trade is the VMware trade with a Nutanix accent: the lowest-friction on-ramp to private AI for the Nutanix (and VMware-refugee) installed base — same console, same operational model, no new vendor, plus real owned storage and a serving plane that is not captive to NVIDIA — in exchange for ceding the orchestration and serving control plane to Nutanix, and accepting that data pipelines and the reasoning plane remain the enterprise's own responsibility. Nutanix runs the private-AI platform. It does not yet govern the agents on top of it. ## ◑ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Hardware-Agnostic Abstraction ### Vendor-Provided Components **AOS + AHV (Distributed Storage OS + Hypervisor)** [DAPM: Ceded] AHV (KVM-derived) plus the proprietary AOS distributed storage OS, managed in Prism. GPU passthrough and GA NVIDIA vGPU with live migration. AHV/AOS storage policies, VM configuration, and Prism opinions are captive — they cannot be lifted to ESXi or Hyper-V without rebuilding. Proprietary Nutanix platform — opinions captive, no open exit. **Multi-Vendor Hardware Support** [DAPM: Retained] Runs on NX appliances and OEM platforms (Dell, HPE, Lenovo, Cisco UCS, Fujitsu) plus software-only on Supermicro/ProLiant/PowerEdge. Commodity x86 substrate — the enterprise swaps OEM without rebuilding workloads. The same hardware-agnostic position VMware occupies. **NVIDIA GPU Integration (vGPU on AHV)** [DAPM: Delegated] GA NVIDIA vGPU with live migration and GPU passthrough; L40S/H100/L4/RTX PRO 6000 Server Edition. The GPU runtime is NVIDIA's behind a Nutanix-managed integration — substitutable in principle, delegated in practice. **Flow Virtual Networking / Flow Network Security** [DAPM: Ceded] AHV software-defined networking: virtual routers and overlays (Flow Virtual Networking) plus distributed stateful microsegmentation (Flow Network Security, now with a Next-Gen policy model). Proprietary SDN — the network policy opinions do not port. Proprietary Nutanix platform — opinions captive, no open exit. ### NVIDIA-Provided Components **NVIDIA vGPU on AHV** GA NVIDIA vGPU with live migration on AHV. Supported accelerators are vGPU-class: L40S, H100, L4, and a single Blackwell SKU — the RTX PRO 6000 Server Edition (via vGPU 19.0). The runtime and drivers are NVIDIA's; Nutanix integrates and schedules them. **HGX Training GPUs Not Virtualizable on AHV** B200/GB200/GH200 HGX systems are NOT virtualizable on AHV — NVIDIA bare-metal only. This caps Nutanix's on-platform AI at inference and fine-tune scale, not frontier training. The single Blackwell SKU on AHV is server-graphics class, not an HGX training GPU. **NUS NVIDIA Certification (Storage Path)** The June 2026 NVIDIA Certification is a STORAGE certification — Nutanix Unified Storage feeding EXTERNAL HGX/x86 GPU servers via Spectrum-X / Spectrum-4 / BlueField-3 and GPUDirect over RDMA. It is not GPUs running on Nutanix, and does not extend the AHV GPU support matrix. ### Gap Analysis Like VMware, Nutanix is an abstraction layer, not physical infrastructure — it manufactures no silicon and no networking. The buyer runs AHV on the commodity x86 they already buy (NX appliances, or OEM platforms from Dell, HPE, Lenovo, Cisco UCS, Fujitsu, plus software-only on Supermicro/ProLiant/PowerEdge), with NC2 bare-metal on AWS/Azure. GPU passthrough and GA NVIDIA vGPU with live migration give them accelerated VMs under the Prism console they already operate. The calibration is identical to VMware Layer 0 (moderate): both Retain hardware vendor choice and Cede the hypervisor and software-defined networking. Nutanix's AI ceiling is, however, lower than VMware's — VCF 9.1 virtualizes Blackwell HGX with NVSwitch and GPUDirect RDMA for distributed inference, whereas AHV tops out at vGPU-class accelerators. That keeps Nutanix firmly at moderate, well below Dell and Cisco (strong), which own or specify compute and networking hardware. ### Borrowed Judgment Multi-directional, same shape as VMware. Nutanix borrows GPU silicon judgment from NVIDIA (same as everyone) and hardware engineering judgment from OEM partners (Dell, HPE, Lenovo, Cisco build the servers), but retains the abstraction layer — AOS, AHV, Prism, and Flow. The enterprise retains OEM choice (commodity x86), delegates the GPU runtime to NVIDIA, and cedes the AHV/AOS control plane and the Flow software-defined network. ### Working Notes AHV is KVM/QEMU/libvirt/OVS-derived, hardened by Nutanix; the control plane (AOS, Prism) is proprietary. AHV MIG-backed vGPU could not be confirmed from open sources (login-gated compatibility matrix) — does not affect the cell. AMD Instinct support is an announced roadmap item (strategic partnership Feb 2026; first platform targeted late 2026) — not GA, logged as a watch-list item. ## ◑ Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Platform Storage + Compliance Governance, Not AI-Native ### Vendor-Provided Components **NUS — Files & Volumes** [DAPM: Ceded] Unified file (SMB/NFS) and block (iSCSI, NVMe-oF/TCP) storage on AOS, managed in Prism. File and block platform opinions — tiering, performance policy, AOS integration — are captive and cannot be lifted to PowerScale or VAST without rebuilding. Proprietary Nutanix platform — opinions captive, no open exit. **NUS — Objects (S3 interface)** [DAPM: Delegated] S3-compatible object store on AOS. Bucket, lifecycle, and policy opinions lift to any S3 platform without rebuilding — the consumed interface is a genuine multi-vendor standard. NUS Objects is a proprietary implementation behind that standard, so the object opinions are portable and the call is Delegated (Retained is reserved for an open substrate the enterprise operates, e.g. Ceph); the same basis as NKP at Layer 2A. **Data Lens (Governance)** [DAPM: Ceded] Compliance and audit governance: access auditing, GDPR/HIPAA/SOX reporting, permission/risk analytics, ransomware detect-and-block. Data Lens 2.0 adds on-prem/air-gapped operation. Governs on metadata and access patterns, not AI/ML semantic content. Proprietary Nutanix platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No Direct NVIDIA Layer 1A Governance Dependency** NVIDIA provides nothing in the governance layer. Its only touch is the NUS storage data path (Spectrum-X / GPUDirect over RDMA) feeding external GPU servers — data-plane performance plumbing, not a Layer 1A governance catalog. ### Gap Analysis Unlike VMware (bring-your-own-storage), Nutanix actually owns the data foundation. Nutanix Unified Storage (NUS 5.3) provides file (SMB/NFS), S3-compatible object, and block (iSCSI, NVMe-oF/TCP) storage on AOS, all managed in Prism, and Data Lens adds genuine compliance governance — access auditing, GDPR/HIPAA/SOX reporting, permission and risk analytics, ransomware detect-and-block — that VMware lacks at the platform layer. But this is governance-as-audit, not AI-native governance. Data Lens classifies on metadata, permissions, and access patterns; there is no AI/ML semantic content classification, no AI metadata enrichment, and no data lineage of the kind Dell's MetadataIQ or HPE's Data Fabric provide. The 4+1 model defines Layer 1A as the governance catalog Layer 2C queries; Nutanix has a storage platform with compliance governance, not a queryable AI-metadata catalog. The result is stronger than VMware's vSAN-plus-external-arrays story but short of Dell's AI-specific metadata depth — moderate, at the top of the band. ### Borrowed Judgment Low. NUS and Data Lens are Nutanix IP — the storage platform and its governance are owned, not borrowed. The limitation is not authority but kind: the governance is compliance and audit, not the AI-metadata catalog a reasoning plane would query. NVIDIA contributes only the RDMA data path, not governance logic. ### Working Notes S3-over-RDMA is a roadmap item (later 2026); NetApp ONTAP integration is targeted second half of 2026 — both logged as watch-list, neither scored. The NUS object interface is split out as Delegated because S3 object opinions lift to any S3 platform — a proprietary implementation behind a multi-vendor standard — mirroring the Dell ObjectScale treatment. ## ◑ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Foundational RAG via Managed Postgres ### Vendor-Provided Components **Vector Search via NDB-Managed PostgreSQL (pgvector)** [DAPM: Delegated] GA. NDB v2.10 automates the lifecycle of PostgreSQL with the pgvector extension, framed by Nutanix as a vector DB for RAG. pgvector is OSS with real alternatives; the consumed interface is the standard Postgres wire protocol, so retrieval opinions lift to any Postgres platform without rebuilding. NDB operates the database — operation delegated, the open substrate keeps the opinions portable. **Milvus / NVIDIA AIDP on NUS (Layered Retrieval)** [DAPM: Delegated] The AI-native retrieval path layers third-party Milvus (OSS) and NVIDIA AI Enterprise on top of NUS. Nutanix provides the substrate, not the retrieval engine; Milvus is swappable. ### NVIDIA-Provided Components **NVIDIA AI Enterprise + NeMo Retriever (Accelerated Path)** The accelerated-retrieval path layers NVIDIA AI Enterprise and NeMo Retriever over NUS. The same non-differentiating stack is available on Dell, HPE, and VMware — NVIDIA is the accelerator, not the retrieval owner. ### Gap Analysis A Nutanix buyer who wants RAG has one genuinely shipping, owned path: NDB-managed PostgreSQL with pgvector (GA since NDB v2.7, Jan 2025; current v2.10), which Nutanix markets explicitly as turning Postgres into a vector database for RAG, with lifecycle automation (provision, patch, clone, HA). For moderate-scale enterprise RAG on data already kept in Postgres, that is a real, low-friction capability. The retrieval intelligence, however, is not Nutanix's. pgvector is OSS; the heavier AI-native retrieval path Nutanix points customers to is third-party Milvus plus NVIDIA AI Enterprise layered on NUS. Critically, NUS itself has no native vector database, no vector search, and no RAG retrieval in the storage layer — the AI-ready-data messaging is data-plane plumbing, not retrieval. There is no retrieval-quality observability (recall@k, latency percentiles) a Layer 2C could consume — the same universal gap noted in Dell and VMware. This calibrates almost exactly to VMware Layer 1B (moderate): VMware's owned vector path is also pgvector-on-Postgres, the same pragmatic, performance-ceilinged choice, short of Dell (Elasticsearch hybrid search + MetadataIQ) and VAST (native InsightEngine). ### Borrowed Judgment Moderate. The retrieval opinions live in stock OSS PostgreSQL + pgvector; NDB provides lifecycle automation around a portable open core. A pg_dump/restore moves the vector schema, embeddings, index definitions, and queries to any other Postgres — RDS, Azure, self-hosted — so the retrieval layer lifts out. Milvus is likewise OSS and swappable. Nutanix's contribution is operation, not retrieval intelligence. ### Working Notes pgvector index types (HNSW/IVFFlat) under NDB management are not confirmed from open sources — does not affect the cell. The KServe serving substrate is industry-attested but not officially named by Nutanix (vLLM is the documented engine). ## ○ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Gap ### NVIDIA-Provided Components **No NVIDIA Layer 1C Dependency** NVIDIA provides nothing at Nutanix Layer 1C. There is no GA CMX/KV-cache tiering on AHV; kvCache-aware routing in NAI is Tech Preview and is request-scheduling within an endpoint, not data-pipeline tiering. A near-empty NVIDIA column here is itself the finding — same as VMware Layer 1C. ### Gap Analysis Within the Nutanix world, data mobility is genuinely good — but it is storage mobility, not AI pipelines. NDB Time Machine provides copy-data management (point-in-time recovery, thin clones); snapshots replicate across clusters and clouds (NC2); native DB replication (Oracle Data Guard, Postgres pglogical) moves data between instances. None of that is a Layer 1C AI data pipeline. There is no dedicated ETL/ELT product, no CDC pipeline, no feature engineering, no lineage, no cost-aware movement, and no KV-cache tiering. The closest thing to data movement is OSS pglogical, which is not Nutanix IP. The enterprise running Nutanix for AI must bring its own pipeline orchestration (Airflow, Kubeflow, Spark, or a commercial alternative) and run it on NKP. This is the same call as VMware Layer 1C (gap): both lack an equivalent to Dell's Dataloop-powered orchestration, HPE's Ezmeral, or VAST's DataEngine. The gap has a downstream cost. A reasoning plane (Layer 2C) makes data-relative placement decisions — run compute where the data lives, weigh moving data versus moving compute — and to do so it needs the lineage and cost-aware-movement primitives this layer would provide. Their absence is one of the structural reasons a Nutanix Layer 2C cannot do data-relative placement. ### Borrowed Judgment The enterprise retains responsibility for the function by default — no vendor has claimed it. Any data-pipeline judgment is borrowed from whatever tooling the enterprise deploys on NKP (Airflow, Kubeflow, commercial), not from Nutanix. Storage and DB mobility (Time Machine, snapshot replication, pglogical) exist but are adjacent to, not the same as, the AI-pipeline function this layer scores. ### Working Notes Watch-list (dated, none scored): NetApp ONTAP integration targeted 2H 2026; S3-over-RDMA roadmap (later 2026); kvCache-aware routing in NAI 2.7 is Tech Preview (and is request-scheduling within an endpoint, not data-pipeline tiering). Time Machine and cross-cloud snapshots are described here as storage/DB mobility, not promoted to scored components — they do not clear the AI-pipeline bar. ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Nutanix Heritage Strength ### Vendor-Provided Components **Prism Central (Unified Fleet Management)** [DAPM: Ceded] Multi-cluster management, RBAC, quotas, self-service, and full-stack lifecycle across the Nutanix estate. The operational backbone — comparable to VMware's SDDC Manager and VCF Operations. Prism policies and management opinions do not port. Proprietary Nutanix platform — opinions captive, no open exit. **AOS + Acropolis Dynamic Scheduling (ADS)** [DAPM: Ceded] Automated VM placement, load-balancing, and resource management across the cluster. The placement opinions are captive to AOS — a proprietary scheduler. Proprietary Nutanix platform — opinions captive, no open exit. **Nutanix Kubernetes Platform (NKP)** [DAPM: Delegated] Turnkey Kubernetes built on pure upstream, CNCF-conformant Kubernetes (unforked, no proprietary API wrapper) via Cluster API. The consumed interface is the standard Kubernetes API, so manifests and workloads lift to EKS, AKS, or any conformant cluster without rebuilding. The Konvoy/Kommander management layer is Nutanix's; the consumed interface keeps the workloads portable. ### NVIDIA-Provided Components **NVIDIA GPU Operator** GPU lifecycle and scheduling on NKP via the NVIDIA GPU Operator (partner OSS) in the Kommander catalog. There is no Nutanix-owned GPU scheduler and no Run:ai equivalent — the GPU-aware layer is borrowed, the same dependency every on-prem vendor shares. **NVIDIA vGPU Manager** vGPU profiles and allocation on AHV. Multi-tenant GPU partitioning is managed through Prism, but the virtualization layer itself is NVIDIA's. ### Gap Analysis Layer 2A is Nutanix's home turf and its deepest IP. Prism Central provides mature, multi-cluster fleet management with RBAC, quotas, and self-service; AOS with Acropolis Dynamic Scheduling (ADS) automates VM placement and rebalancing; and NKP 2.17 delivers Kubernetes. One operational model — the console the buyer already runs — now extended to AI workloads. For a VMware-refugee shop (Nutanix's core go-to-market), this is the lowest-friction orchestration on-ramp to AI available. The calibration matches VMware Layer 2A (strong): both are HCI/private-cloud platforms whose deepest competency is unified infrastructure orchestration, and both carry the same NVIDIA GPU-scheduling dependency, which the layer treats as universal. Two honest qualifiers keep Nutanix a hair behind VMware: VMware's 2A remains the peak of the series on raw operational depth (100M+ cores, two decades, a single vSphere Supervisor managing VMs, containers, and AI from one plane), and Nutanix's unification is marginally less seamless — VMs via Prism/AHV and Kubernetes via a distinct NKP product rather than one supervisor. Still well ahead of Dell and Cisco 2A (moderate), where GPU orchestration is fully NVIDIA-owned. GPU-aware scheduling is the borrowed piece: NKP schedules GPUs via the NVIDIA GPU Operator, and there is no Nutanix-owned GPU scheduler. Policy-driven GPU scheduling (which workload gets which GPU on cost or compliance grounds) is NVIDIA's, not the platform's — the identical gap Dell and VMware carry. ### Borrowed Judgment Low — the lowest of any Nutanix layer. Prism, AOS, and ADS are Nutanix IP. The primary borrowed piece is GPU-aware scheduling (NVIDIA GPU Operator), the same dependency every on-prem peer shares. NKP's consumed interface — unforked upstream Kubernetes — keeps container workloads portable to any conformant cluster. ### Working Notes NKP is explicitly pure upstream CNCF Kubernetes with no proprietary API wrapper (D2iQ lineage: Konvoy/Kommander/NKP Insights are Nutanix IP; the runtime is unforked). NKP Metal (bare-metal Kubernetes) is Early Access with GA targeted 2H 2026 — logged as watch-list. ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Platform-Native Serving, Agents Pre-GA ### Vendor-Provided Components **Nutanix Enterprise AI (NAI) — Model Serving & Governance** [DAPM: Ceded] GA. RBAC-governed inference endpoints; 74-model validated catalog; multi-source models (NVIDIA NIM, Hugging Face, custom upload); CPU-only inference via Intel AMX; batch inference; vLLM default engine. Runs on any CNCF Kubernetes. NAI is proprietary Nutanix software — the serving and governance opinions (catalog, endpoint configuration, RBAC policy) do not lift to another serving stack without rebuilding, and running it on EKS does not change that. Proprietary Nutanix platform — opinions captive, no open exit. **Inference Runtime (vLLM / NVIDIA NIM / TensorRT-LLM)** [DAPM: Delegated] The model execution engines NAI orchestrates. vLLM is the documented OSS default; NIM (with TensorRT-LLM) is an optional path, not a lock. The consumed runtime is substitutable — applications call OpenAI-compatible endpoints — and is neither Nutanix-owned nor NVIDIA-mandatory. ### NVIDIA-Provided Components **NVIDIA NIM (One Serving Option, Not a Lock)** NIM is a model-serving option inside NAI — alongside Hugging Face and custom upload — not a requirement. NeMo/Nemotron are enabled and TensorRT-LLM is available via NIM. The finding is the de-emphasis: NAI is explicitly not NIM-locked. **Lighter NVIDIA Dependency Than Peers** A vLLM OSS default engine plus an Intel-AMX CPU-only inference path mean Nutanix's runtime is less NVIDIA-captive than Dell's all-NVIDIA NemoClaw/OpenShell 2B path. A genuine, scoreable differentiator. ### Gap Analysis Nutanix Enterprise AI (NAI 2.7, GA May 2026) is Nutanix's strongest owned AI asset. It deploys LLMs as secure, RBAC-governed inference endpoints from NVIDIA NIM, Hugging Face, or a custom upload; offers a 74-model validated catalog, batch inference, a CPU-only inference path (Intel AMX), and vLLM as the default engine; and runs on any CNCF Kubernetes (EKS/AKS/GKE/bare-metal), not just Nutanix. As a standalone model-serving product it is arguably cleaner and less NVIDIA-locked than VMware's Model Runtime. Two concerns. The inference runtime is borrowed — vLLM (OSS) and TensorRT-LLM via NIM (NVIDIA); NAI is the control and governance plane on top, and that plane's opinions are captive to NAI. And the agent-execution half of this layer is not GA: Layer 2B is model serving, agent execution, inference APIs, and distributed inference, but the native agent framework (NAI Labs – Agent), MCP servers, RAG pipeline, LoRA fine-tuning, and kvCache-aware routing are all Tech Preview in NAI 2.7. Per the GA-gate, none of those count toward the score. The result lands at VMware Layer 2B (moderate), with a different composition: VMware's Model Runtime and Agent Builder are both GA, whereas Nutanix's serving is GA and arguably more open, but its agent execution is preview where VMware's is GA. Nutanix trades VMware's GA agent-builder for a more open, less NVIDIA-captive serving runtime. Both sit well above Dell 2B (entirely NVIDIA NemoClaw/OpenShell-dependent). ### Borrowed Judgment Moderate. Nutanix owns the NAI serving and governance plane (its authority, captive to Nutanix), while the inference runtime is borrowed from vLLM (OSS) and NVIDIA (NIM) — the same structural runtime dependency Dell, HPE, and VMware carry, softened here by the OSS default engine and the CPU-only path. Agent execution is not yet GA, so there is no agent-runtime judgment to score. ### Working Notes GA-gate watch-list (dated, none scored): native agent framework / NAI Labs – Agent (Tech Preview), MCP servers (TP), RAG pipeline (TP), LoRA fine-tuning (TP), kvCache-aware routing (TP) — all in NAI 2.7; and the 'Nutanix Agentic AI' full stack (Early Access, GA targeted 2H 2026). Palo Alto Prisma AIRS model scanning is GA in NAI 2.7 (scored at Layer 3). ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Emerging Signals — Governance Without Placement ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** Agent Gateway is Nutanix IP. NVIDIA provides no agent governance, model routing, or placement reasoning in the Nutanix stack — the same pattern as Google's and Cisco's 2C, where the governance layer is vendor-owned, not NVIDIA-dependent. ### Gap Analysis With Agent Gateway (GA in NAI 2.7), the buyer gets a single, governed control point for agentic traffic: a unified API in front of both cloud and self-hosted models, RBAC, per-token rate-limiting, full audit including MCP-call audit, plus health-based endpoint failover and capacity load-balancing. For an enterprise worried about ungoverned agent and model sprawl, that is a real, shipping governance surface. Applying the 'Routing Is Not Reasoning' test: Agent Gateway governance (RBAC, rate-limiting, audit, MCP audit) is access control — genuine intelligence-layer governance, but not placement. Health-based failover and capacity load-balancing are single-variable operational routing (is the endpoint up, is it at capacity?), not multi-variable policy-driven placement. Per-request model selection on content, cost, or quality does not exist; kvCache-aware routing is Tech Preview, and even that schedules to GPU workers within a single endpoint, not among models. None of it is the reasoning plane Layer 2C defines. The gap is structural, not merely unbuilt. A reasoning plane makes data-relative placement decisions, and the data-side feeders it would query — an AI-metadata catalog at Layer 1A, lineage and cost-aware movement at Layer 1C — are absent. So even when the forward 'Nutanix Agentic AI' full stack reaches GA (Early Access today, targeted 2H 2026), its placement reasoning will be starved of data-side inputs. This calibrates to VMware Layer 2C (gap) — the closest decision-path peer — which likewise parks its agentic-governance signal short of placement reasoning. Nutanix has building blocks (governance, operational routing); it does not have a reasoning plane. ### Borrowed Judgment Inverted: there is no Layer 2C to borrow. Agent Gateway is Nutanix IP with no NVIDIA dependency, so the governance and operational-routing functions Nutanix does provide carry low borrowed judgment — but the placement, model-routing, and reasoning functions are Absent, not borrowed. The enterprise must build custom 2C logic, bring a partner, or operate without it. ### Working Notes Watch-list (dated, none scored): kvCache-aware routing (Tech Preview); the 'Nutanix Agentic AI' full stack (Early Access, GA targeted 2H 2026); native agent framework and MCP servers (Tech Preview). Agent Gateway is described here as a building block rather than scored — consistent with the gap-layer convention and with how the closest peer (VMware) treats its equivalent agentic-governance signal. ## ◑ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Platform-Enabled, Not Platform-Provided ### Vendor-Provided Components **Nutanix Enterprise AI (Platform Enablement)** [DAPM: Ceded] The integrated model-serving and governance platform enterprises build AI applications on — the tools to BUILD applications, not the applications themselves. Mirrors VMware's Private AI Services as a Layer 3 enabler. Proprietary Nutanix platform — opinions captive, no open exit. **ISV AI Partner Program** [DAPM: Delegated] DataRobot (flagship), Codeium, Instabase, Pryon, Lamini, UbiOps and others provide application logic, when-and-if-available. Substitutable partners — the enterprise can change ISVs without rebuilding the platform beneath them. **Hugging Face + Validated Model Catalog** [DAPM: Delegated] 74 validated models (Meta Llama, Google Gemma, NVIDIA Nemotron) deployed with the customer's Hugging Face token. Open and partner models, swappable — mirrors Dell's Enterprise Hub (Hugging Face) treatment. **Palo Alto Prisma AIRS (Model Security Scanning)** [DAPM: Delegated] GA in NAI 2.7. Partner-provided model and prompt scanning integrated into the serving path. A substitutable partner security capability. ### NVIDIA-Provided Components **NVIDIA Model Ecosystem (Nemotron + NIM)** Nemotron models and NIM containers are available through the NAI catalog. NVIDIA provides a slice of the model layer; Nutanix provides serving and governance. The same non-differentiating pattern as VMware Layer 3. ### Gap Analysis Nutanix does not sell AI applications — it sells the platform to build and run them, plus a curated on-ramp to models. The buyer gets Hugging Face integration (deploy validated Hub models with their own token, GA), a 74-model validated catalog (Meta Llama, Google Gemma, NVIDIA Nemotron), an ISV AI Partner Program (DataRobot as flagship, plus Codeium, Instabase, Pryon, Lamini, UbiOps), and Palo Alto Prisma AIRS model-security scanning (GA in NAI 2.7). Applications are built by the enterprise's own teams on NAI, or sourced from partners. This is the architecturally correct position for a platform vendor — the same one Dell and VMware occupy — so the concern is ecosystem depth, not capture. NAI is the platform-native enabler (Nutanix IP), but the application logic and models are partner, OSS, or enterprise-built. The ISV AI Partner Program is emerging and when-and-if-available — not at the curation depth of Dell's load-bearing ecosystem (OpenAI, Palantir, ServiceNow, 5,000+ deployments) or HPE's Unleash AI (26+ validated ISVs), and several models (Mistral, DeepSeek) are deployable but not officially validated. The calibration is VMware Layer 3 (moderate): both provide platform-native AI services (VMware's Private AI Services, Nutanix's NAI) that enable Layer 3, with an emerging AI-specific ISV ecosystem. That platform-native tooling is exactly what distinguishes moderate (VMware, Nutanix) from partner (Dell, Cisco — pure ISV ecosystem with no platform AI services of their own). ### Borrowed Judgment Distributed across partners and the enterprise's own development teams — the correct Layer 3 shape. NAI (the enabler) is Nutanix's; the applications and models are partner, OSS, or enterprise-built. The enterprise borrows application judgment from its chosen ISVs (DataRobot and the AI Partner Program) and model judgment from the open and partner catalog (Hugging Face, Llama, Gemma, Nemotron). ### Working Notes Hugging Face is an engineering-collaboration partner (since May 2024); models deploy with the customer's HF token. Mistral and DeepSeek are deployable but carry Nutanix's own not-officially-validated disclaimer. Prisma AIRS (Palo Alto) model scanning reached GA in NAI 2.7. ════════════════════════════════════════════════════════════════════════════════ # NVIDIA AI Platform A Components Company Becoming a Platform Vendor — Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.3 - Prose-Coherence (GA-Gate) **Date:** May 22, 2026 **Source:** GTC 2025, GTC 2026, Dynamo 1.0 GA, NemoClaw/OpenShell announcement, Run:ai acquisition, DGX Cloud, NVIDIA AI Enterprise, NIM GA, analyst coverage, SEC FY2026 annual report ## Summary Finding NVIDIA is the only vendor in this assessment series that appears inside every other vendor's assessment. Dell's Layer 2A GPU orchestration is NVIDIA Run:ai. HPE's Layer 2B runtime is NVIDIA AI Enterprise. VMware's GPU integration is NVIDIA vGPU Manager. AWS, Google Cloud, and Azure all run NVIDIA GPUs alongside their own silicon. VAST embeds NVIDIA GPUs, NICs, DPUs, and switches into its data platform. Mapping NVIDIA as a standalone vendor inverts the perspective: instead of asking 'where does Dell cede authority to NVIDIA,' the question becomes 'where does NVIDIA claim authority, and from whom?' The 4+1 mapping reveals NVIDIA as a vendor with deep authority at Layers 0, 2A, and 2B — and emerging ambitions at Layer 2C. NVIDIA designs the accelerator silicon that every other vendor depends on (Layer 0), provides the GPU orchestration platform that on-prem vendors brand as their own (Layer 2A via Run:ai), and controls the inference runtime and model lifecycle stack that sits between the enterprise's infrastructure and its AI applications (Layer 2B via NIM, NeMo, Dynamo, TensorRT-LLM). At Layers 1A through 1C, NVIDIA is an accelerator — it makes other vendors' storage and data pipelines faster without providing those capabilities directly. The structural tension is between NVIDIA as a silicon supplier and NVIDIA as a platform vendor. When NVIDIA was only selling GPUs, its interests aligned with every OEM and hyperscaler: more GPU adoption meant more revenue for everyone. As NVIDIA extends into GPU orchestration (Run:ai), inference runtime (NIM/Dynamo), agent governance (OpenShell/NemoClaw), and cloud infrastructure (DGX Cloud), it competes with the same customers who buy its silicon. Dell's AI Factory runs NVIDIA software. But DGX Cloud is NVIDIA competing with Dell for the same enterprise workload. The DAPM classification for NVIDIA is inverted from every other assessment: the enterprise consuming NVIDIA through an OEM (Dell, HPE, VMware) has already Ceded authority to the OEM, which has Ceded authority to NVIDIA. The enterprise consuming NVIDIA directly (DGX Cloud, NIM API) Cedes authority to NVIDIA without an intermediary. The enterprise self-hosting NVIDIA open-source software (Dynamo, OpenShell, Nemotron) Retains authority — but still runs on NVIDIA silicon. NVIDIA is the only vendor where every deployment path, at every layer, eventually depends on NVIDIA hardware. More than half of NVIDIA's engineers work on software. That statistic from NVIDIA's FY2026 annual report is the key to understanding the 4+1 mapping: NVIDIA is a software company that happens to sell the hardware its software requires. The assessment series has been documenting where NVIDIA's software authority appears inside other vendors' stacks. This assessment makes that authority explicit. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** NVIDIA Strength — Silicon Authority ### Vendor-Provided Components **GPU Accelerator Silicon (Blackwell, Vera Rubin)** [DAPM: Ceded] Blackwell B200/B300/GB200 (current generation). Vera Rubin NVL72 (next generation, deploying at hyperscalers). The accelerator silicon that every other vendor in this assessment depends on. Dell builds PowerEdge around it. HPE builds ProLiant and Cray around it. AWS offers it as P5/P6 instances. Azure offers it as ND-series. Google offers it alongside TPUs. VAST embeds it in CNode-X. No enterprise AI infrastructure exists without NVIDIA GPU silicon — or a deliberate decision to use an alternative (AWS Trainium, Google TPU, AMD Instinct). **Networking Silicon + Interconnect** [DAPM: Ceded] NVLink/NVSwitch (intra-node GPU interconnect). Spectrum-X Ethernet switches. ConnectX-7/8 SmartNICs. BlueField-3 DPUs. InfiniBand for GPU cluster fabric. NIXL for disaggregated inference data movement. Dell brands Spectrum switches as PowerSwitch. HPE integrates ConnectX into ProLiant. VAST uses ConnectX/BlueField for NVMe-over-Fabrics. The networking silicon is as structurally embedded as the GPU silicon. **DGX Platform (On-Prem Systems)** [DAPM: Ceded] DGX SuperPOD: leadership-class AI infrastructure for on-prem and hybrid. DGX Station: workgroup-scale AI compute. DGX Spark: desktop AI workstation. Pre-configured systems with NVIDIA software stack pre-installed. Competes directly with Dell PowerEdge, HPE ProLiant, and OEM AI server configurations — NVIDIA sells the assembled system, not just the components. **DGX Cloud (Hosted Infrastructure)** [DAPM: Ceded] GPU supercomputing as a service, hosted on AWS, Azure, GCP, and OCI. Includes NVIDIA AI Enterprise software and Base Command Platform. The enterprise accesses NVIDIA infrastructure through a hyperscaler substrate — Ceding to both NVIDIA (software/GPU) and the hyperscaler (facility/network). DGX Cloud competes with the hyperscalers' own GPU instance offerings while running on their infrastructure. ### Gap Analysis Layer 0 is NVIDIA's foundational authority. Every other vendor in this assessment depends on NVIDIA silicon at this layer — the only exceptions are AWS (Trainium/Inferentia), Google (TPU), Azure (Maia), and AMD Instinct instances on hyperscalers. The DGX Platform creates a structural tension with OEM partners. When NVIDIA sells DGX SuperPOD directly to an enterprise, that enterprise is NOT buying Dell PowerEdge or HPE ProLiant. NVIDIA is simultaneously its OEM partners' most critical supplier and their direct competitor. Dell's 'AI Factory with NVIDIA' branding and HPE's 'NVIDIA AI Computing by HPE' branding are attempts to keep the enterprise buying through the OEM rather than going to NVIDIA directly. DGX Cloud adds a second tension: NVIDIA competes with hyperscalers while running on their infrastructure. AWS, Azure, and GCP host DGX Cloud while simultaneously offering their own GPU instances. The enterprise choosing between Azure ND-series VMs and DGX Cloud on Azure is choosing between Microsoft-managed and NVIDIA-managed access to the same GPU hardware. The NVIDIA dependency at Layer 0 is the one dependency shared by every on-prem vendor assessed. Dell, HPE, VAST, and VMware all depend on NVIDIA GPU silicon. The difference is the scope of that dependency: at Layer 0, NVIDIA provides silicon. At Layer 2A, NVIDIA provides orchestration. At Layer 2B, NVIDIA provides runtime. The silicon dependency is structural and shared. The software dependency is where NVIDIA's authority claims create tension. ### Borrowed Judgment The enterprise consuming NVIDIA silicon inherits NVIDIA's GPU architecture decisions (memory bandwidth, interconnect topology, power/thermal profile), NVIDIA's driver and CUDA runtime decisions, and NVIDIA's product lifecycle and pricing decisions. This borrowed judgment is structural — it exists for every vendor in the assessment and for every enterprise running AI workloads on NVIDIA hardware. The DGX Platform adds system-level borrowed judgment: NVIDIA's hardware integration, thermal design, and rack architecture decisions. The DGX Cloud adds hyperscaler-layered borrowed judgment: NVIDIA software decisions on top of hyperscaler infrastructure decisions. ### Working Notes NVIDIA's FY2026 annual report segments its business into Compute & Networking (data center accelerated computing, networking, AI solutions, software, automotive) and Graphics (GeForce, Quadro/RTX). The Data Center platform — GPUs, DPUs, networking, DGX, software — is the revenue engine. The company's strategic direction is to expand from silicon supplier to platform vendor without alienating the OEM and hyperscaler customers who drive GPU volume. The NVIDIA dependency column in every other vendor's assessment can now be read as 'authority NVIDIA claims at this layer.' The standalone assessment makes that authority explicit and measurable. ## ○ Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Accelerator Only ### Gap Analysis NVIDIA provides no storage, no data governance, and no data platform. Zero components at this layer. NVIDIA accelerates other vendors' storage with GPU libraries (cuVS for vector search, RAPIDS for data processing) and networking hardware (BlueField DPUs, ConnectX NICs) — but those are acceleration functions assessed at their functional layers (1B for retrieval, 1C for data processing, Layer 0 for networking silicon). The storage platforms, governance catalogs, and data architectures are entirely owned by other vendors: Dell (PowerScale, ObjectScale, MetadataIQ), HPE (Alletra, Data Fabric), VAST (DataStore, DataBase, Catalog), AWS (S3, Glue, Lake Formation), Google (BigQuery, Knowledge Catalog), Azure (Blob, Fabric, Purview). The absence of NVIDIA-owned storage or governance is structurally significant for Layer 2C: a Reasoning Plane needs governance metadata — which data is sensitive, which models are approved, which compliance requirements apply. NVIDIA has no Layer 1A metadata to feed into a Layer 2C reasoning plane. Every other vendor's 2C ambition is anchored in governance metadata from 1A. NVIDIA's emerging Layer 2C (OpenShell/NemoClaw) operates without governance context because NVIDIA doesn't own the data layer. ### Borrowed Judgment None. NVIDIA has no data layer authority to lend or borrow. The enterprise's storage and governance judgment comes entirely from the storage vendor. ### Working Notes The STX Architecture observation from the Dell assessment applies here: STX is available to every storage vendor. It does not differentiate any OEM's storage offering — it raises the floor for all of them. NVIDIA's Layer 1A role is to make the data layer faster, not to provide or govern it. ## ◑ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Acceleration + Model Enablement ### Vendor-Provided Components **cuVS (GPU-Accelerated Vector Search)** [DAPM: Delegated] GPU-accelerated vector similarity search library. 12x faster vector indexing. Used by Dell (MetadataIQ integration), VAST (CNode-X vector search), and storage vendors for retrieval acceleration. NVIDIA provides the search acceleration; the platform vendor provides the retrieval infrastructure and index. **NeMo Retriever** [DAPM: Delegated] GPU-accelerated retrieval pipeline for RAG. Embedding models, reranking, and retrieval optimization. Integrated into Dell's Data Search Engine (PowerScale connector), HPE's retrieval stack, and VMware's AI Enterprise RAG Stack. Provides the retrieval intelligence that OEMs brand as part of their platforms. **NIM Embedding Models** [DAPM: Delegated] Pre-built inference microservices for text and multimodal embedding. Used by VAST's InsightEngine, Dell's retrieval pipeline, and hyperscaler RAG services. NVIDIA provides the embedding models; the platform vendor provides the retrieval infrastructure. ### Gap Analysis NVIDIA provides retrieval acceleration and embedding models but not retrieval infrastructure. The retrieval engines — Azure AI Search, OpenSearch, Elasticsearch, VAST InsightEngine, Google Vertex AI Search — are owned by other vendors. NVIDIA makes retrieval faster and provides the embedding models that make vector search work, but the enterprise's retrieval architecture is determined by the platform vendor. NeMo Retriever is a meaningful capability: it provides the GPU-accelerated RAG pipeline that multiple OEMs brand as part of their offerings. When Dell advertises 'GPU-accelerated hybrid search,' the GPU acceleration is NVIDIA's. The enterprise's retrieval quality depends on NVIDIA's embedding model quality — a borrowed judgment that is rarely made explicit. ### Borrowed Judgment Moderate. NeMo Retriever embedding models determine retrieval quality. The enterprise inherits NVIDIA's embedding model training decisions, architecture choices, and optimization priorities. This borrowed judgment is invisible — the enterprise interacts with Dell's search engine or VAST's InsightEngine, not with NVIDIA's embeddings directly. ### Working Notes The embedding model dependency is worth tracking across the assessment series. Multiple vendors (Dell, HPE, VAST, VMware) use NVIDIA NIM embedding models for their RAG pipelines. If NVIDIA changes embedding model architecture, quality, or licensing, it affects every vendor's Layer 1B simultaneously — a shared dependency that no single vendor controls. ## ○ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Gap — single pre-GA capability ### Gap Analysis Deployable today, this is a gap: NVIDIA's only Layer 1C capability, CMX (KV-cache offload via BlueField-4), is pre-GA (see watch-list). The analysis below describes CMX architecture; it is not an option an architect can deploy today. NVIDIA does not provide data pipeline orchestration (Data Factory, Dataloop, DataEngine, Airflow). Its Layer 1C presence is a single capability: KV cache management via CMX. CMX is architecturally significant because it addresses a data movement problem unique to AI inference — KV cache growing beyond GPU memory. This is a Layer 1C function (data movement) that directly affects Layer 2B performance (inference latency). Dell has validated it (19x TTFT improvement on PowerScale); HPE and VAST are expected integrations. The enterprise's KV cache strategy becomes a borrowed judgment from NVIDIA's CMX design decisions once adopted. NVIDIA also provides GPU-accelerated compute libraries (RAPIDS, cuDF) used within other vendors' data pipelines, but these are computation acceleration, not data movement — they make processing faster without providing pipeline orchestration, data lineage, or movement logic. They are assessed at the layers where they functionally operate (compute acceleration at Layer 0, retrieval acceleration at Layer 1B) rather than at Layer 1C. ### Borrowed Judgment Moderate for CMX — an architectural decision about KV cache management that affects inference performance and is harder to substitute once adopted. The enterprise inherits NVIDIA's decisions about cache eviction policy, offload thresholds, and storage tier targeting. ### Working Notes The KV cache tiering gap identified in the Azure assessment is relevant here: Azure has no CMX integration. Dell has validated it. HPE is expected. VAST's CNode-X architecture collocates cache and compute, potentially eliminating the need for CMX-style offload. The KV cache management approach varies by vendor — NVIDIA's CMX is one solution, not the only one. Watch-list (pending GA, not scored): NVIDIA CMX (Context Memory Extension, KV-cache offload, BlueField-4) — not yet GA. It was NVIDIA’s only Layer 1C component; with it pre-GA, NVIDIA ships no data-movement capability today, so the layer is a gap. ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** NVIDIA Authority via Run:ai ### Vendor-Provided Components **NVIDIA Run:ai (Acquired 2024)** [DAPM: Ceded] GPU orchestration and workload management platform. Kubernetes-native. Dynamic GPU pooling across hybrid environments. Fractional GPU sharing (no open-source equivalent). Fair-share scheduling with team-level quotas. Multi-cluster management from a unified control plane. Now part of NVIDIA AI Enterprise ($4,500/GPU/year standalone, included with DGX). Run:ai is the Layer 2A authority that Dell brands as part of AI Factory, HPE brands as part of Private Cloud AI, and VMware integrates through NVIDIA AI Enterprise. The OEM sells the relationship; NVIDIA controls the scheduling intelligence. GPU-level infrastructure (GPU Operator for driver/plugin lifecycle, MIG for hardware partitioning) provides the substrate Run:ai orchestrates — invisible plumbing, not standalone orchestration tools. ### Gap Analysis Layer 2A is where NVIDIA's platform ambition creates the most direct tension with its OEM partners. Run:ai is the most capable GPU-specific orchestration platform available — fractional GPU sharing, multi-cluster management, and fair-share scheduling are capabilities that no open-source alternative matches. But Run:ai is NVIDIA's product, not the OEM's. When Dell markets 'AI Factory with NVIDIA,' the GPU scheduling intelligence is Run:ai — NVIDIA's IP, NVIDIA's roadmap, NVIDIA's pricing. Dell provides the hardware, the rack integration, and the customer relationship. NVIDIA provides the scheduling brain. If NVIDIA changes Run:ai's architecture, licensing, or feature set, Dell's AI Factory Layer 2A changes with it — without Dell's input. The same dynamic applies to HPE (Private Cloud AI includes NVIDIA AI Enterprise with Run:ai) and VMware (VCF integrates NVIDIA AI Enterprise). Three OEMs, one scheduling authority. The hyperscalers avoid this dependency: AWS built Karpenter, Google built GKE Autopilot + Fluid Compute, Azure contributed DRA to upstream Kubernetes. Each hyperscaler owns its GPU scheduling intelligence. On-prem vendors do not — they consume NVIDIA's. The open-source alternatives (Kueue, KAI Scheduler, DRA) are catching up but lack Run:ai's fractional GPU sharing. The enterprise evaluating GPU orchestration choices is evaluating a NVIDIA proprietary vs. open-source trade-off — better capability (Run:ai) vs. more authority (open-source). ### Borrowed Judgment High. Run:ai's scheduling decisions — which team gets which GPU, how fractional sharing is allocated, when over-quota borrowing is permitted — are NVIDIA's judgment. The enterprise configures policies; NVIDIA's scheduler executes them. If Run:ai makes a scheduling decision that impacts training job completion time or inference latency, that's NVIDIA's borrowed judgment affecting business outcomes. NVIDIA AI Enterprise licensing adds commercial borrowed judgment: the enterprise's production deployment timeline depends on NVIDIA's licensing terms, pricing changes, and certification cycles. ### Working Notes The Run:ai acquisition (2024) is the most significant NVIDIA software acquisition for the 4+1 model. Before Run:ai, NVIDIA provided silicon and libraries. After Run:ai, NVIDIA provides the orchestration plane that sits between the enterprise and its own GPUs. The enterprise doesn't interact with GPUs directly — it interacts through Run:ai's scheduling layer. The open-source Kubernetes GPU scheduling landscape (DRA, Kueue, KAI Scheduler) is evolving rapidly. Microsoft contributed DRA to upstream Kubernetes at KubeCon 2026. If open-source GPU scheduling reaches feature parity with Run:ai's fractional GPU sharing, the enterprise case for Run:ai's licensing cost weakens. NVIDIA's response: integrate Run:ai deeper into AI Enterprise, making it harder to substitute. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** NVIDIA Authority — Inference + Agent Runtime ### Vendor-Provided Components **NVIDIA NIM (Inference Microservices)** [DAPM: Ceded] Pre-built, optimized inference containers for 100+ models. OpenAI-compatible API. Free for prototyping on DGX Cloud (build.nvidia.com). Production requires AI Enterprise license. Includes Nemotron, Llama, Mistral, and partner models. NIM is the inference runtime that multiple OEMs and hyperscalers brand as part of their platforms — AWS Bedrock offers NIM, Azure Foundry offers NIM, Dell deploys NIM on PowerEdge. **Dynamo 1.0 (Inference Operating System)** [DAPM: Retained] Open-source (Apache 2.0) distributed inference serving framework. GA March 2026. Disaggregated prefill and decode. KV-aware routing to GPUs with best cache match. KVBM for memory management. NIXL for GPU-to-GPU data movement. Grove for scaling. 7x performance boost on Blackwell. Adopted by AWS, Azure, GCP, OCI, CoreWeave, and dozens of inference providers. NVIDIA positions Dynamo as 'the operating system of AI factories.' Open-source but NVIDIA-optimized — runs best on NVIDIA hardware. **NeMo (Model Lifecycle)** [DAPM: Ceded] End-to-end model lifecycle management: data curation, model customization and evaluation, guardrailing and observability. NeMo Guardrails for content safety. NeMo Evaluator for model assessment. NeMo Data Designer for training data preparation (integrated into VAST's TuningEngine). The model lifecycle stack that operates above inference and below applications. **NeMo Guardrails (Runtime Content Safety)** [DAPM: Ceded] Programmable content safety framework inline with inference. Controls model output, topic boundaries, and factual grounding during model serving. Deployed as part of NIM containers or standalone. At Layer 2B, Guardrails functions as runtime content filtering — it controls what the model says during inference. The same capability serves a Layer 2C governance function when applied as policy enforcement for agent behavior. ### Gap Analysis (GA-gate) NemoClaw/OpenShell, referenced below as part of the stack, is alpha (see watch-list); the strong score rests on NIM and Dynamo (GA). Layer 2B is NVIDIA's deepest software authority and the layer where the platform ambition is most visible. NIM, Dynamo, NeMo, NeMo Guardrails, and NemoClaw/OpenShell constitute the complete inference, model lifecycle, and agent runtime stack. NeMo Guardrails and NemoClaw/OpenShell appear at both Layer 2B and Layer 2C because they serve dual architectural functions. At 2B they are runtime capabilities — content filtering inline with inference, agent execution environment. At 2C they are governance capabilities — policy enforcement for agent behavior, sandbox constraints on agent access. The same code, two architectural purposes. This dual-layer presence is itself evidence of NVIDIA's platform transition: a components company's software stays within one layer; a platform company's software spans layers. The open-source strategy is deliberate: Dynamo (Apache 2.0) and NemoClaw/OpenShell (Apache 2.0) are open-source, meaning the enterprise Retains the code. But both are optimized for NVIDIA hardware and NVIDIA's CUDA ecosystem. Running Dynamo on AMD or Intel GPUs is theoretically possible but practically disadvantaged. The open-source license provides code portability; the hardware optimization provides silicon lock-in. NIM is the more significant authority claim: it's closed-source, NVIDIA-only, and requires an AI Enterprise license for production. The enterprise using NIM to serve models has Ceded inference runtime authority to NVIDIA. The alternative — vLLM, SGLang, or other open-source serving frameworks — is slower but Retained. The Dell assessment's Layer 2B finding applies directly: Dell does not appear to own the core agent runtime, model-serving runtime, guardrail framework, or distributed inference framework. Those are NVIDIA's. This assessment confirms that observation from NVIDIA's perspective. ### Borrowed Judgment High for NIM (closed-source, NVIDIA-controlled inference optimization decisions). Low for Dynamo (open-source, enterprise can fork and modify). Moderate for NeMo (model lifecycle decisions — training data curation, evaluation metrics, guardrail policies — are NVIDIA's defaults that the enterprise inherits unless explicitly overridden). The inference optimization decisions in NIM and TensorRT-LLM directly affect model output quality, latency, and cost. Quantization choices, batching strategies, and KV cache management are NVIDIA's engineering decisions that the enterprise consumes without visibility. If an NIM container produces different outputs than a vLLM deployment of the same model, the enterprise may not know which is 'correct.' ### Working Notes Dynamo 1.0 GA (March 2026) is NVIDIA's strongest Layer 2B play. Positioning it as 'the operating system of AI factories' is explicitly a platform claim. Combined with Run:ai at Layer 2A and NIM at Layer 2B, NVIDIA controls the infrastructure orchestration, the inference optimization, and the model serving runtime — three layers of the enterprise's AI stack that sit between the hardware (which NVIDIA also provides) and the application (which the enterprise builds). The NemoClaw/OpenShell alpha status is important context: Futurum Research noted that NemoClaw addresses 'the deployment end of the agent trust chain well' but urged enterprises 'not to treat it as a complete governance solution.' Security and accountability need to be embedded throughout the development lifecycle, not just at runtime. This is the gap between NVIDIA's runtime governance (OpenShell) and Microsoft's lifecycle governance (Entra Agent ID + Agent Governance Toolkit). Watch-list (pending GA, not scored): NemoClaw + OpenShell agent runtime — open-source (Apache 2.0), alpha. Described in the gap narrative. ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Runtime Governance Only — Not a Reasoning Plane ### Gap Analysis Applying the 'Routing Is Not Reasoning' test from the VMware assessment: OpenShell provides runtime sandbox governance — it controls WHAT agents can access (filesystem, network, processes). NeMo Guardrails control WHAT models can say (content filtering, topic boundaries). Neither provides policy-driven decisions about WHERE compute runs relative to data, WHICH model serves WHICH request, or HOW cost/compliance/latency are arbitrated. OpenShell is agent runtime security. NeMo Guardrails is model output safety. Neither is a Reasoning Plane. NVIDIA's Layer 2C gap is structural: NVIDIA does not own storage (Layer 1A), data governance (Purview, Lake Formation, Knowledge Catalog), or enterprise identity (Entra, IAM). A Reasoning Plane needs governance metadata — which data is sensitive, which models are approved, which compliance requirements apply. NVIDIA has no data governance to query because it has no data layer. The consequence: NVIDIA's Layer 2C will always depend on another vendor's governance metadata. OpenShell can enforce sandbox policies, but it cannot make placement decisions informed by data classification, compliance status, or cost targets — because that information lives in Purview, Lake Formation, PolicyEngine, or MetadataIQ, none of which NVIDIA owns. This is the fundamental structural limitation of NVIDIA's platform ambition: NVIDIA can build runtime governance (2B/2C boundary) but cannot build a full Reasoning Plane (2C) because it lacks the data governance foundation (1A) that a Reasoning Plane queries. ### Borrowed Judgment Low for OpenShell (open-source, enterprise controls the policies). Low for NeMo Guardrails (configurable by the enterprise). The governance logic is transparent — the enterprise defines what agents can and cannot do. The missing borrowed judgment is the more significant finding: NVIDIA's Layer 2C cannot borrow data governance judgment from itself because it doesn't have a data governance layer. It must borrow from Dell (MetadataIQ), HPE (Data Fabric), VAST (Catalog), AWS (Lake Formation), Google (Knowledge Catalog), or Azure (Purview). NVIDIA's governance is runtime-only; other vendors' governance is data-informed. ### Working Notes The Futurum Research observation is the right framing: OpenShell addresses the 'deployment end of the agent trust chain' but enterprises should not treat it as a complete governance solution. Security and accountability need to be embedded throughout the development lifecycle. Compare to other vendors' Layer 2C: • Microsoft: identity + governance lifecycle (Entra Agent ID + AGT) — the broadest governance scope • Google: model-integrated orchestration (Agent Platform) — the deepest platform integration • AWS: policy + evaluation + registry (AgentCore) — the most modular approach • VAST: data platform governance (PolicyEngine + Polaris) — the most data-informed approach • NVIDIA: runtime sandbox (OpenShell) — the narrowest scope, addressing only execution-time security NVIDIA's Layer 2C is necessary but not sufficient. It complements other vendors' governance — it doesn't replace it. ## ◑ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Model + Blueprint Enablement ### Vendor-Provided Components **Nemotron Open Models** [DAPM: Retained] Post-trained on Llama, distilled from DeepSeek-R1. Deployment-ready for AI agents. Available through NIM API (build.nvidia.com) and as downloadable containers. Nemotron models are NVIDIA's answer to the model layer — open models optimized for NVIDIA hardware. Competes with OpenAI, Anthropic, Google, and Meta at the model layer while providing the hardware those competitors run on. ### Gap Analysis NVIDIA does not build enterprise AI applications. Its Layer 3 presence is a single component: Nemotron open models. NVIDIA also provides application enablement that falls below the Layer 3 threshold: Blueprints (pre-built reference patterns for PDF extraction, digital twins, RAG pipelines, AI-Q agent task decomposition — deployed through Dell, HPE, VMware, and hyperscaler marketplaces) and NIM API endpoints (build.nvidia.com — free API access to 100+ models, 1,000 free inference credits, GPU sandbox instances). Blueprints are reference architectures, not applications — the enterprise builds from them, not on them. NIM API is a developer on-ramp and go-to-market funnel, not an application platform. The Nemotron model strategy is the interesting Layer 3 finding: NVIDIA competes with the AI model providers (OpenAI, Anthropic, Google, Meta) whose models run on NVIDIA hardware. If Nemotron achieves quality parity with proprietary models, enterprises can run inference on NVIDIA hardware with NVIDIA models — a fully vertically integrated stack from silicon to model. No other silicon vendor has this: Intel doesn't have frontier models, AMD doesn't have frontier models, AWS Trainium serves other providers' models. The NIM API funnel is NVIDIA's developer moat: free prototyping creates adoption → adoption creates switching cost → production deployment requires AI Enterprise license on NVIDIA hardware. The funnel is silicon-to-model-to-lock-in. ### Borrowed Judgment Moderate. Nemotron model alignment, training data, and safety decisions are NVIDIA's. The model-to-silicon borrowed judgment is unique to NVIDIA: when the enterprise uses Nemotron on NVIDIA GPUs, both the model and the hardware are NVIDIA's. The enterprise borrows NVIDIA's judgment at every layer of the inference path. No other vendor has this — even Google (Gemini on TPU) separates the model team (DeepMind) from the silicon team. ### Working Notes The NVIDIA-as-model-provider dynamic creates an unusual competitive position: NVIDIA wants enterprises to adopt Nemotron (NVIDIA model revenue) AND wants enterprises to run OpenAI/Anthropic/Meta models on NVIDIA GPUs (NVIDIA hardware revenue). Both outcomes benefit NVIDIA, but they benefit NVIDIA in different ways. If Nemotron succeeds too well, it reduces the model diversity that drives GPU demand from multiple model providers. ════════════════════════════════════════════════════════════════════════════════ # Oracle Cloud Infrastructure (OCI) AI Infrastructure Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.5 - Compute Management GA Promotion **Date:** July 12, 2026 **Source:** Oracle AI World 2025, GTC 2026, OCI Enterprise AI GA (Mar 2026), Fusion Agentic Applications (Mar 2026), Oracle AI Database 26ai, Stargate/OpenAI partnership, NVIDIA/AMD partnerships, analyst coverage, Oracle Q3 FY2026 earnings. v1.3 (instrument reconciliation): 1A OCI Object Storage and 2A OKE Retained→Delegated — a managed service behind a multi-vendor standard interface is Delegated, not Retained. Also aligned the VAST Layer 2C cross-reference to VAST's current gap status. v1.4 (July 12, 2026, /reconcile): statusLabels moved to capability vocabulary per the instrument convention; no grades or chips changed. v1.5 (July 12, 2026, /ga-check): OCI Compute Management promoted from the watch-list on OCI product documentation (instance pools, cluster networks, capacity reservations) — scored as a Ceded 2A component; status holds moderate. ## Summary Finding OCI occupies a structurally unique position in this assessment series: it is the only hyperscaler whose AI infrastructure strategy is anchored by a database franchise. AWS builds down from managed services. Google builds out from a frontier model. OCI builds up from the enterprise data layer — Oracle AI Database 26ai, Autonomous AI Database, and the Fusion Applications estate that runs 97% of the Fortune 100. Every other hyperscaler treats the database as one service among many. Oracle treats the database as the gravitational center around which AI infrastructure orbits. The infrastructure story is more aggressive than the enterprise positioning suggests. OCI Zettascale10 connects up to 800,000 NVIDIA GPUs across multi-gigawatt clusters delivering 16 zettaFLOPS — the fabric underpinning the Stargate supercluster built with OpenAI in Abilene, Texas. Oracle Acceleron, a custom RoCE networking architecture with 2.5–9.1 microsecond latency, is genuine Layer 0 IP that positions OCI alongside AWS (Nitro/EFA/SRD) and Google (Virgo) as hyperscalers with proprietary networking stacks. The AMD partnership (50,000 MI450 GPUs, Q3 2026) makes OCI one of two hyperscalers with meaningful multi-vendor GPU strategy alongside AWS. The DAPM profile is heavily Ceded — structurally identical to AWS and Google Cloud in that the enterprise consumes managed services without controlling underlying architecture. But OCI adds a distinctive wrinkle: the database layer creates a gravitational pull that concentrates not just infrastructure authority but data authority. An enterprise running Fusion Applications on Autonomous AI Database on OCI Superclusters has Ceded compute, networking, database, application runtime, AND business logic to a single vendor. This is deeper vertical integration than AWS (which doesn't own the application layer) and comparable to Google's model-integrated stack — but achieved through the application and data layers rather than through a frontier model. OCI Enterprise AI (GA March 2026) is a credible but late entry to the agentic platform space. OpenAI Responses-compatible API, managed agent hosting, vector stores, MCP support, guardrails, and observability — the capabilities parallel AWS Bedrock AgentCore and Google's Gemini Enterprise Agent Platform. The differentiator is database-native AI: Select AI for natural language to SQL, AI Vector Search inside the database engine, and the Private Agent Factory pattern that keeps agent reasoning co-located with enterprise data. Whether 'AI at the database layer' is a structural advantage or an architectural constraint is the central DAPM question for OCI. The Stargate partnership and $553B RPO validate infrastructure demand. The Fusion Agentic Applications validate the application-layer strategy. The multicloud database deployments (Oracle Database@AWS, @Azure, @Google Cloud) validate the data-gravity thesis. But the 4+1 framework asks a different question: where does authority reside, and has the enterprise made that placement explicit? For OCI, the answer is that authority concentrates in the database — and the database concentrates in Oracle. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Full-Stack Cloud Compute ### Vendor-Provided Components **OCI Superclusters + Zettascale10** [DAPM: Ceded] Up to 800,000 NVIDIA GPUs across multi-gigawatt clusters. 16 zettaFLOPS peak performance. Underpins Stargate (OpenAI). Scales from 8 GPUs to 131,072 B200 GPUs, 100,000+ GB200 Superchips per cluster. Bare metal GPU instances with RDMA cluster networking. **Oracle Acceleron Networking** [DAPM: Ceded] Custom-designed RDMA over Converged Ethernet (RoCE v2). 2.5–9.1 microsecond GPU-to-GPU latency. Multiplanar network architecture with dedicated RoCE fabrics. Congestion-control-first (not PFC-dependent). Zero-Trust Packet Routing (ZPR) at the physical layer. Up to 3,200 Gb/s cluster network bandwidth. Oracle-owned networking IP. **NVIDIA GPU Fleet** [DAPM: Ceded] GB200 NVL72, B200/B300, H200, H100, A100, L40S bare metal instances. 1M+ GPUs. NIXL support for disaggregated inference. BlueField-4 integration for next-gen Superclusters (GTC 2026). DGX Cloud hosted on OCI. **AMD GPU Fleet (Q3 2026)** [DAPM: Ceded] 50,000 AMD Instinct MI450 Series GPUs. Helios rack architecture: 72 liquid-cooled GPUs per rack, AMD EPYC Venice CPUs, Pensando Vulcano DPUs. UALink/UALoE fabric. ROCm software stack. Up to 432 GB HBM4, 20 TB/s memory bandwidth per GPU. First hyperscaler with publicly available AMD AI supercluster at this scale. **OCI Dedicated Region25 + Oracle Alloy** [DAPM: Ceded] Full OCI stack (200+ services including SaaS) in customer data center, starting at 3 racks. 60+ Dedicated Region/Alloy regions live. EU Sovereign Cloud (Frankfurt, Madrid). Isolated Cloud Regions for classified workloads. Oracle Alloy enables partner-operated OCI. Fujitsu, SoftBank, Vodafone as anchor customers. ### NVIDIA-Provided Components **NVIDIA GPU Silicon + Rubin Roadmap** 1M+ NVIDIA GPUs deployed. Blackwell B200/B300, H200, H100, L40S, GB200 NVL72. Rubin roadmap committed. DGX Cloud hosted on OCI. NVIDIA BlueField-4 integration announced at GTC 2026 for OCI Superclusters. ### Gap Analysis OCI's Layer 0 is the most surprising story in this assessment series. A vendor perceived as a database company has built one of the largest GPU cloud fabrics in the world — the Stargate supercluster alone targets 800,000 GPUs. Oracle Acceleron is genuine networking IP: custom RoCE with congestion-control-first design (not PFC-dependent), multiplanar architecture, Zero-Trust Packet Routing at the physical layer. This is not leased NVIDIA networking — it's Oracle-designed fabric. The multi-accelerator strategy is more advanced than any hyperscaler except AWS. NVIDIA (Blackwell, Rubin), AMD (MI450 with Helios rack architecture, 50,000 GPUs Q3 2026), and Intel Xeon 6 processors. The AMD Helios rack — 72 liquid-cooled GPUs with Venice CPUs and Pensando Vulcano DPUs — is a fully integrated system comparable to NVIDIA DGX but from AMD's ecosystem. No other hyperscaler has committed to AMD at this scale. The Dedicated Region25 (full OCI in 3 racks, customer data center) and Oracle Alloy (partner-operated OCI) create a distributed cloud model with 60+ dedicated/Alloy regions live. This parallels AWS AI Factories, Azure Local, and Google Distributed Cloud — but Oracle's model delivers the full 200+ service stack, including SaaS, which none of the other hyperscalers match in dedicated form. The structural tension: OCI's GPU customers include OpenAI, xAI, Meta, and other frontier model trainers. These are not traditional Oracle enterprise customers. OCI is simultaneously serving the world's largest AI training workloads AND the world's most conservative enterprise database customers — two audiences with fundamentally different risk profiles and authority expectations. ### Borrowed Judgment The enterprise Cedes Layer 0 entirely to Oracle — GPU selection, networking topology, cluster design, physical infrastructure. Oracle Acceleron is Oracle IP, reducing NVIDIA networking dependency compared to hyperscalers using NVIDIA Spectrum-X. But the GPU silicon dependency on NVIDIA (and increasingly AMD) is structural and shared with every vendor in this assessment. The Stargate partnership creates a unique borrowed judgment dynamic: Oracle operates infrastructure for OpenAI's training workloads. The operational lessons from running the world's largest AI training cluster feed back into OCI's infrastructure decisions — but those decisions are made for frontier training requirements, not necessarily for enterprise inference workloads. ### Working Notes OCI's bare-metal GPU instances are architecturally distinctive. While AWS and Google abstract GPUs behind instance types, OCI exposes bare metal with RDMA cluster networking — giving customers direct hardware access without hypervisor overhead. This appeals to AI training customers (OpenAI, xAI) who need maximum GPU utilization. The trade-off: bare metal reduces Oracle's ability to multi-tenant and abstract, pushing more operational complexity to the customer. The $45–50B financing plan (Feb 2026) for OCI expansion and the $553B RPO (Q3 FY2026, up 325% YoY) demonstrate infrastructure investment at a scale that few anticipated from Oracle. Cloud infrastructure revenue grew 84% to $4.9B in the quarter. The CapEx intensity is comparable to AWS and Google — this is no longer a database company's side project. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Database-Anchored Data Foundation ### Vendor-Provided Components **Oracle Autonomous AI Database** [DAPM: Ceded] Self-managing, self-securing, self-patching database built on Oracle AI Database 26ai engine. AI Vector Search (native vector data type, HNSW/IVF indexes, SQL-based similarity search), Select AI (natural language to SQL), ONNX embedding models, Private AI Services Container integration, NVIDIA NIM container support. Autonomous AI Lakehouse with Apache Iceberg support for open multi-vendor data lakehouse. RAFT-based replication, JSON Relational Duality, quantum-resistant encryption, in-database SQL firewall. Platinum and Diamond-tier availability (Apr 2026). Anomaly detection, auto-indexing, auto-tuning. Autonomous AI Vector Database variant in Limited Availability (March 2026). Available on OCI, AWS, Azure, Google Cloud via multicloud deployments. **OCI Object Storage** [DAPM: Delegated] Standard object storage for unstructured data, embeddings, model artifacts. S3-compatible API. Regional and cross-region replication. Storage Classes for cost optimization. A managed service behind the multi-vendor S3 standard — object opinions lift to any S3-compatible platform, so Delegated (operation delegated to Oracle; the standard interface keeps the opinions portable). Reserve Retained for an open substrate the enterprise operates. **Oracle Database@AWS / @Azure / @Google Cloud** [DAPM: Ceded] Oracle AI Database running inside other hyperscalers' infrastructure with private interconnects. Multicloud Universal Credits for cross-cloud procurement. Teams use familiar AWS/Azure/Google tools and billing while running Oracle AI Database on OCI infrastructure within the hyperscaler. Oracle-AWS Interconnect expanded April 2026. ### NVIDIA-Provided Components **NVIDIA cuVS + CAGRA (Future)** GPU-accelerated vector indexing with NVIDIA CAGRA and cuVS designed for integration with Oracle AI Database. Not yet GA — future GPU acceleration for vector workloads. ### Gap Analysis Layer 1A is where OCI's structural differentiation is sharpest. Every other hyperscaler treats storage and governance as separate services composed by the customer. Oracle treats the database AS the governance layer — AI Vector Search, Select AI, data classification, audit, encryption, and access control are database-native capabilities, not services bolted on top. Oracle AI Database 26ai is the most significant Layer 1A product in this assessment because it collapses traditionally separate functions: relational storage + vector storage + semantic search + natural language querying + governance + encryption + lakehouse (Apache Iceberg) into a single authority boundary. AWS achieves comparable breadth by composing S3 + Glue + Lake Formation + OpenSearch — four services, four governance surfaces. Oracle delivers it in one. The Autonomous AI Lakehouse extends this to open formats: Apache Iceberg read/write in object store, enabling cross-cloud analytics without data movement. Oracle Database@AWS, @Azure, and @Google Cloud place Oracle's data authority inside other hyperscalers' infrastructure — a multicloud data-gravity strategy no other vendor in this series attempts. The DAPM implication: collapsing Layer 1A into a single database authority is simultaneously Oracle's greatest strength and greatest lock-in risk. The enterprise gains architectural simplicity and eliminates inter-service governance gaps. But substituting away from Oracle AI Database means losing vector search, semantic querying, governance, and the lakehouse capability simultaneously — a higher switching cost than any other hyperscaler's Layer 1A. ### Borrowed Judgment Low for database governance — Oracle AI Database 26ai governance (encryption, access control, audit, classification) is Oracle IP with 40+ years of enterprise hardening. The enterprise defines policies; Oracle enforces them inside the database. Moderate for AI-specific capabilities — Select AI and AI Vector Search embed model judgment (embedding quality, chunking strategy, SQL generation accuracy) inside the database layer. When Select AI generates SQL from natural language, the enterprise inherits the model's interpretation of business logic — the same borrowed judgment pattern identified in Google's LookML Agent, but at the database layer rather than the analytics layer. ### Working Notes Oracle's positioning of AI Database 26ai as 'the best memory core for enterprise agents' is architecturally significant for the 4+1 model. If agents store memory, context, and retrieval state in the database, then the database becomes the persistence layer for agentic intelligence — a Layer 1A function that directly enables Layer 2B/2C agent capabilities. This is the 'data-gravity for agents' thesis: agents that persist state in Oracle AI Database become progressively harder to move off Oracle's platform. The quantum-resistant encryption and in-database SQL firewall in 26ai address threats that other vendors' Layer 1A offerings don't yet productize. Security-forward positioning that aligns with sovereign AI requirements. 97% of Fortune 100 running on Oracle Database creates an installed base advantage no other vendor can replicate. The question is whether that installed base translates to AI workload adoption or whether enterprises run AI on one platform and databases on another. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Database-Native Retrieval ### Vendor-Provided Components **AI Vector Search (Oracle Autonomous AI Database)** [DAPM: Ceded] Native vector data type, HNSW and IVF vector indexes, SQL-based similarity search. Hybrid search combining vector similarity with relational predicates. Permission-aware retrieval through database access control. No separate vector database required. **OCI Enterprise AI — Managed Vector Store** [DAPM: Ceded] Managed vector storage with file ingestion, semantic search, and metadata filtering for RAG and NL2SQL use cases. Schema enrichment into semantic vector store for natural language queries that produce and execute SQL against customer databases with permission control. **Select AI + SQL Search (NL2SQL)** [DAPM: Ceded] Natural language to SQL generation and execution. Applications and analytics use LLMs to understand natural language questions and generate Oracle SQL. Bridges unstructured (vector) and structured (SQL) retrieval in a single query surface. **OCI Generative AI — Embeddings + Rerank** [DAPM: Ceded] Managed embedding and reranking services. Supports Cohere, NVIDIA Nemotron, and other embedding models. OpenAI-compatible APIs. Dedicated AI clusters for consistent latency. ### NVIDIA-Provided Components **NVIDIA NIM Containers for Embeddings** NVIDIA embedding models available through OCI Generative AI Model Import. cuVS acceleration planned for future vector indexing. ### Gap Analysis OCI's Layer 1B collapses into Layer 1A — and that's the architectural point. AI Vector Search lives inside Oracle AI Database, not as a separate vector database service. RAG queries execute as SQL against the same database that holds the enterprise's transactional data. Permission-aware retrieval inherits the database's existing access control model without a separate security overlay. This is the same architectural pattern as VAST (InsightEngine inside the data platform) but at the database level rather than the storage level. The comparison to AWS is instructive: Bedrock Knowledge Bases composes OpenSearch + S3 + embedding models across service boundaries. Oracle eliminates those boundaries by making vector search a database capability. The SQL Search (NL2SQL) capability adds a dimension no other vendor's Layer 1B provides: agents can retrieve structured enterprise data through natural language queries that generate and execute SQL. This bridges unstructured retrieval (vector search for documents) and structured retrieval (SQL for business data) in a single query surface — a genuine differentiation. The gap: Oracle's Layer 1B is database-bounded. Data outside Oracle AI Database isn't retrievable through AI Vector Search. AWS's Bedrock Knowledge Bases can index data from any S3-accessible source. Google's BigQuery can federate queries across storage boundaries. Oracle's retrieval requires data to be IN the database or accessible through database links. SyncEngine-like capabilities for ingesting external enterprise data (Google Drive, Jira, Confluence) are not evident in Oracle's published materials. ### Borrowed Judgment Low for vector search infrastructure — AI Vector Search, indexing, and SQL-based retrieval are Oracle IP inside the database engine. No external retrieval service dependency. Moderate for embedding quality — embedding models are either ONNX (customer-provided), NVIDIA NIM containers, or third-party models through OCI Generative AI. The quality of retrieval depends on embedding model choice, which the enterprise controls but doesn't build. The NL2SQL capability introduces a specific borrowed judgment: when Select AI generates SQL from natural language, the accuracy of retrieval depends on the model's understanding of the database schema. Schema enrichment into a semantic vector store (announced in OCI Enterprise AI) mitigates this — but the enterprise inherits the model's interpretation of table relationships and business logic. ### Working Notes The Private Agent Factory pattern (announced March 2026) keeps agent reasoning co-located with enterprise data inside Oracle AI Database — agents query the database directly rather than through external RAG pipelines. This reduces network hops and latency for retrieval but deepens the database dependency. Every agent built on Private Agent Factory inherits Oracle AI Database as a non-substitutable Layer 1B dependency. The Exadata for AI announcement (vector search offloaded to intelligent storage for dramatic speedups) bridges Layer 1A and Layer 1B at the hardware level — the storage system itself accelerates retrieval. This is comparable to VAST's CNode-X collocating cache and compute, but implemented at the database storage layer rather than the file system layer. ## ◑ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Database-Centric Pipelines ### Vendor-Provided Components **OCI GoldenGate** [DAPM: Ceded] Real-time data replication, transformation, and streaming across heterogeneous sources. Continuous data availability. Database migration, disaster recovery, real-time analytics. The deepest database connector ecosystem of any cloud data movement service. **OCI Data Integration + Data Flow** [DAPM: Ceded] Data Integration: managed ETL/ELT service with visual design. Data Flow: managed Apache Spark for large-scale data processing and ML data preparation. Integration with Autonomous AI Database and OCI Object Storage. **OCI Data Science** [DAPM: Ceded] Managed ML platform: JupyterLab notebooks, model training, model catalog, model deployment. AI Quick Actions for one-click model operations. GPU-enabled compute shapes. Separate service surface from OCI Enterprise AI. **Multicloud Database Connectivity** [DAPM: Ceded] Oracle Interconnect + AWS Interconnect for managed private high-performance connectivity (April 2026). Oracle Database@AWS, @Azure, @Google Cloud. Multicloud Universal Credits for cross-cloud procurement. Data stays in Oracle's governance model regardless of which cloud hosts the compute. ### Gap Analysis Layer 1C reveals the database-centric trade-off. Oracle's data movement capabilities are strong for database-to-database flows (GoldenGate) and analytics pipelines (OCI Data Integration, OCI Data Flow). But the ML-specific pipeline orchestration that Dell (Dataloop), HPE (Ezmeral Unified Analytics), VAST (DataEngine), and AWS (Glue + SageMaker Unified Studio) provide is less integrated. OCI Data Science provides notebooks, model training, and deployment but is a separate service from OCI Enterprise AI — creating a multi-surface problem similar to AWS's pre-Unified Studio fragmentation. There is no single governed environment that collapses data engineering, model training, and agent development the way SageMaker Unified Studio or Google's BigQuery ML attempt. GoldenGate is the most mature real-time data replication service in the assessment series — purpose-built for heterogeneous data movement across on-premises and cloud. No other hyperscaler has an equivalent with Oracle's depth of database connector support. The multicloud data movement story is strong: Oracle Database@AWS/@Azure/@Google Cloud moves the database layer to the customer's cloud of choice. But this is database replication, not general-purpose data pipeline orchestration. An enterprise needing Airflow-style DAG orchestration for ML workflows must deploy it on OKE — which is possible but not Oracle-managed. ### Borrowed Judgment Low to moderate. GoldenGate and OCI Data Integration are Oracle IP. Data Flow (managed Spark) and OCI Data Science (managed notebooks) are Oracle-managed wrappers around open-source technology (Apache Spark, JupyterLab). The enterprise retains pipeline logic but Cedes execution infrastructure. The pipeline gap means enterprises often bring third-party orchestration (Airflow, Kubeflow) to OCI — introducing borrowed judgment from those communities and creating governance boundaries that don't exist when using Oracle-native services. ### Working Notes The absence of a unified ML pipeline platform is the most significant gap relative to AWS and Google Cloud. Both competitors have invested heavily in collapsing the data engineering → model training → model serving → agent development pipeline into single governed surfaces. Oracle's approach is service-by-service composition — GoldenGate for replication, Data Integration for ETL, Data Flow for Spark, Data Science for ML, Enterprise AI for agents — without the horizontal integration layer. Oracle Fusion Applications data is already 'in Oracle' — the pipeline problem for Fusion customers is different than for greenfield AI. The data doesn't need to be moved; it needs to be made AI-ready. Select AI and AI Vector Search address this for Fusion customers in a way that no amount of pipeline tooling could match. The question is whether non-Fusion enterprises find the pipeline story compelling. ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** OKE + GPU Superclusters ### Vendor-Provided Components **OCI Compute Management (Instance Pools, Cluster Networks, Capacity Reservations)** [DAPM: Ceded] Documented in OCI product docs with CLI and SDK surfaces: instance configurations and pools for fleet management, cluster networks for RDMA-connected GPU/HPC groups, capacity reservations for predictable accelerator supply. Single-vendor management APIs — pool definitions and reservation opinions are OCI-captive, consistent with the instrument's treatment of hyperscaler-native management surfaces. **OCI Kubernetes Engine (OKE)** [DAPM: Delegated] Managed Kubernetes with GPU-aware node pools. Bare metal and virtual machine compute shapes. RDMA cluster networking for training workloads. MIG support for fractional GPU allocation. Autoscaling at pod level with GPU Device Plugin metrics. Karpenter support for node autoprovisioning. Free managed control plane. A managed service behind the standard Kubernetes API — manifests lift to another conformant cluster, so Delegated (operation delegated to Oracle; the standard interface keeps the opinions portable). **GPU Node Manager + Monitoring** [DAPM: Ceded] Kubernetes-native GPU, networking, and infrastructure monitoring for OKE clusters. NVIDIA DCGM integration. Health monitoring, utilization metrics, and alerting. Actively developing additional capabilities. ### NVIDIA-Provided Components **NVIDIA GPU Operator + Device Plugin on OKE** GPU discovery, health monitoring, scheduling within OKE clusters. MIG support for fractional GPU allocation. Node Manager for GPU/networking monitoring. NVIDIA DCGM integration. ### Gap Analysis OCI's Layer 2A follows the same pattern as AWS: Kubernetes-based GPU orchestration through a managed service (OKE) with NVIDIA GPU scheduling underneath. OKE provides managed Kubernetes with GPU-aware node pools, autoscaling, bare-metal RDMA networking, and MIG support for fractional GPU allocation. The distinction from AWS: OCI's bare-metal GPU instances give customers more direct hardware control than AWS's virtualized GPU instances. OKE autoscaling operates at the pod level using NVIDIA GPU Device Plugin metrics, and the Karpenter Provider for OCI (GA April 2026) brings flexible node autoprovisioning — matching AWS's Karpenter capability for just-in-time compute shape selection based on workload requirements. GPU scheduling authority follows the same pattern as Dell and HPE: NVIDIA controls GPU scheduling through the GPU Operator and Device Plugin. Oracle controls infrastructure orchestration through OKE. Policy-driven GPU scheduling (which workload gets which GPU based on cost, compliance, and performance) is not an OCI-native function — the same gap every vendor shares. The Dedicated Region model adds a Layer 2A dimension other hyperscalers don't match: OCI orchestrates infrastructure across public cloud, 60+ dedicated regions, isolated regions, and Alloy partner regions from a single control plane. The orchestration scope is broader than AWS (Outposts, AI Factories) or Google (GDC) in terms of the number and variety of deployment targets. ### Borrowed Judgment GPU scheduling: NVIDIA-controlled, same as every other vendor except AWS (Karpenter) and Google (TPU scheduling). Infrastructure orchestration: Oracle-controlled through OKE and the OCI control plane. Orchestration is consumed through the standard Kubernetes interface; the enterprise's manifests lift to another platform (Delegated — managed K8s service; the enterprise could switch without rebuilding). GPU scheduling intelligence is the Ceded part (NVIDIA-controlled). The bare-metal model creates a subtly different borrowed judgment profile: with bare metal, the enterprise has more direct GPU control (no hypervisor overhead, direct RDMA access) but also more operational responsibility. The judgment about GPU sharing, isolation, and scheduling is partially Retained by the customer — a more favorable DAPM position than fully managed GPU instances. ### Working Notes OCI's GPU monitoring through Node Manager is actively developing — the tool surfaces GPU, networking, and infrastructure metrics in a Kubernetes-native way. This is an operational foundation that could evolve toward Layer 2C if enriched with policy-driven decision-making. The multi-accelerator scheduling problem (NVIDIA vs AMD GPUs, different instance types) will become acute when AMD MI450 GPUs arrive Q3 2026. OCI will need workload-to-silicon matching capabilities — a Layer 2C function — to help customers choose between NVIDIA and AMD for each workload. No productized capability for this exists today. Watch-list (pending GA, not scored): OCI Compute Management (instance pools, cluster networks, capacity reservations) — Q3 2026. 2A still scored on OKE and GPU Node Manager. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** OCI Enterprise AI Platform ### Vendor-Provided Components **OCI Enterprise AI Platform** [DAPM: Ceded] End-to-end agentic AI platform (GA March 2026). OpenAI Responses-compatible API with multi-model routing. Enterprise AI agents with modular, composable primitives. Managed agent hosting for OSS frameworks and MCP servers. Vector stores, semantic search, NL2SQL, memory, tools. IAM integration, guardrails, observability, auditability. Dedicated AI clusters for isolated compute. **OCI Generative AI Service** [DAPM: Ceded] Foundation model access: xAI Grok (4.1 Fast, 4.3), Cohere Command A (Vision, Reasoning), NVIDIA Nemotron 3 Nano Omni, gpt-oss models. Chat, embeddings, rerank APIs. Model Import for custom models. OpenAI-compatible APIs. Sovereign AI options for data hosting. **Applications and Deployments (Hosted Runtime)** [DAPM: Ceded] Container-based hosting for custom agentic applications. Managed infrastructure, networking, storage integration, identity configuration. Public and private endpoint support. Build-in security. Supports OSS frameworks and custom runtimes. **Private Agent Factory (Oracle Autonomous AI Database)** [DAPM: Ceded] Agents run inside Oracle AI Database with direct data access. Co-locates agent reasoning with enterprise data. Eliminates external RAG pipeline latency. Private AI Services Container for on-premises inference without sending data to third-party services. ### NVIDIA-Provided Components **NVIDIA NIM + Nemotron on OCI** NVIDIA Nemotron models (including Nemotron 3 Nano Omni multimodal) available through OCI Enterprise AI. NIM containers deployable on OCI GPU instances. Model Import capability for custom NIM deployment. ### Gap Analysis OCI Enterprise AI (GA March 2026) is Oracle's unified agentic platform. The capabilities are comprehensive: OpenAI Responses-compatible API, managed agent hosting for OSS frameworks and MCP servers, vector stores, semantic search, NL2SQL, guardrails, observability, and auditability. The OpenAI API compatibility is strategically important — it reduces migration friction from OpenAI's platform and positions OCI as a drop-in alternative. The model catalog includes xAI Grok (4.1 Fast, 4.3), Cohere Command A (Vision, Reasoning), NVIDIA Nemotron 3 Nano Omni, and custom models through Model Import. The notable absence: no Anthropic Claude and no Meta Llama in published model availability. This is a narrower model selection than AWS Bedrock or Google's Model Garden. The agentic runtime distinguishes between three patterns: (1) OCI Generative AI APIs for direct model access, (2) Enterprise AI agents with managed orchestration, tools, memory, and retrieval, and (3) Applications and Deployments for container-based hosted agentic applications with custom runtimes. This three-tier model parallels AWS's Bedrock / AgentCore / SageMaker hierarchy. Dedicated AI clusters provide isolated compute for enterprise workloads — the opposite of shared-tenant model serving. This addresses a specific enterprise concern: inference latency predictability and data isolation. The trade-off is cost — dedicated clusters have fixed cost regardless of utilization. The Fusion Agentic Applications (22 agents across HR, finance, supply chain, CX — GA March 2026) represent Layer 2B and Layer 3 simultaneously: the runtime executes agents that are pre-built for Oracle's application estate. No other hyperscaler ships pre-built enterprise application agents at this scale. ### Borrowed Judgment Model providers (xAI, Cohere, NVIDIA) bring training data, alignment, and safety decisions as borrowed judgment. Oracle's guardrails constrain output but reasoning in model weights is not customer-configurable — same pattern as every other hyperscaler. The Fusion Agentic Applications introduce a unique borrowed judgment dynamic: these agents execute decisions within business processes by accessing unified enterprise data, workflows, policies, approval hierarchies, and permissions. The enterprise inherits Oracle's judgment about how HR, finance, and supply chain processes should be automated. This is not model-level borrowed judgment — it's business-process-level borrowed judgment. When a Fusion Agentic Application automates a talent review or maintenance troubleshooting, the enterprise inherits Oracle's encoding of what 'good' process execution looks like. ### Working Notes The Private Agent Factory pattern (Oracle AI Database 26ai) is architecturally significant: agents run inside the database, with direct access to enterprise data without external API calls. This eliminates the retrieval latency that plagues cloud-native RAG architectures but creates total database dependency for the agent runtime. OCI Enterprise AI's MCP support and OSS framework hosting (container-based deployment with managed infrastructure, networking, storage, and identity) position Oracle to benefit from the open agentic ecosystem without building a proprietary framework. This is a pragmatic strategy: let the frameworks proliferate, provide the managed hosting. The IBM partnership (watsonx Orchestrate agents on Red Hat OpenShift on OCI, IBM Granite models via OCI Data Science) adds an enterprise-focused model and agent ecosystem that differentiates from the consumer-AI-focused model catalogs of AWS and Google. ## ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Intelligence 2C: Emerging | Infra 2C: Implicit ### Vendor-Provided Components **OCI Enterprise AI Governance** [DAPM: Ceded] IAM-based access control for AI resources. Guardrails for content safety and policy enforcement. Observability for agent behavior, tool usage, and data flow monitoring. Auditability with audit logs capturing all AI interactions. Sovereign AI options. Project-level isolation for agent workloads. **Autonomous Database Self-Management** [DAPM: Ceded] Self-managing, self-securing, self-repairing database operations. Auto-indexing, auto-tuning, anomaly detection, autonomous performance optimization. Infrastructure 2C at the database layer — autonomous placement and resource decisions within the database boundary. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** All Layer 2C components are Oracle IP. NVIDIA does not control governance, policy, or reasoning in the OCI stack. ### Gap Analysis Intelligence Layer 2C (partially present): OCI Enterprise AI governance includes IAM-based access control, guardrails for content safety, observability for agent behavior monitoring, and auditability for compliance. Oracle's blog series on runtime governance (April–May 2026) articulates sophisticated 2C concepts: runtime budget guardrails, approval-aware execution, pre-execution veto, safe degradation, evidence-backed runtime control, and the WORM Evidence Vault for audit-grade trace preservation. But there is a gap between the conceptual architecture Oracle's engineering team has published and the productized capabilities in OCI Enterprise AI GA. The runtime governance blog describes a governed execution layer with budget guardrails, circuit breakers, and safe-mode execution. The GA product offers IAM, guardrails, and observability — necessary but not sufficient for the full 2C vision Oracle's own engineers have articulated. Infrastructure Layer 2C (not built): No OCI service answers 'given data residency, cost, latency, GPU availability across NVIDIA and AMD, and compliance requirements, should this workload run on dedicated clusters in us-ashburn-1 or on Dedicated Region infrastructure in Frankfurt?' The capacity management primitives exist. The policy-driven placement engine does not. The Autonomous AI Database is the closest OCI comes to Infrastructure 2C: it self-manages, self-secures, auto-indexes, auto-tunes, and detects anomalies autonomously. But this autonomy operates at the database layer, not at the infrastructure-wide placement layer that the 4+1 model defines. OCI's unique 2C opportunity: Oracle is the only hyperscaler with both the application layer (Fusion) and the data layer (AI Database) under one authority. A Layer 2C that queries Fusion application state, database governance metadata, GPU utilization, and compliance posture to make autonomous placement decisions would have richer context than any other hyperscaler's 2C — because Oracle sees from application logic through data governance to infrastructure. The data to build 2C exists. The product does not. Cross-vendor Layer 2C comparison: • Dell: Absent. • HPE: Retained (IT ops, GreenLake Intelligence) + Delegated (Kamiwaza). • VAST: Gap, emerging (Polaris ships as placement abstraction; PolicyEngine + TuningEngine GA end 2026). • AWS: Intelligence 2C Delegated (AgentCore Policy). Infra 2C implicit. • Google: Most complete productized Intelligence 2C (Agent Identity + Gateway + Registry + Orchestration + Observability). • OCI: Intelligence 2C emerging (guardrails + observability + auditability). Conceptual vision published but not yet fully productized. Infra 2C implicit. ### Borrowed Judgment Intelligence 2C: Low — guardrails, observability, and auditability are Oracle IP. Customer defines IAM policies; Oracle enforces. Infrastructure 2C: Ceded (implicit) — the Autonomous AI Database makes placement, scaling, and optimization decisions autonomously. These are 2C functions at the database layer that the enterprise has Ceded without explicit classification. The parallel to AWS's implicit 2C (managed service decisions) applies, but OCI's implicit 2C is concentrated in the database rather than spread across managed services. ### Working Notes Oracle's engineering blog series on runtime governance for agentic AI (April–May 2026) is the most sophisticated published thinking on Layer 2C from any hyperscaler's engineering team. The concepts — runtime budget guardrails as governed execution, evidence-backed observability with WORM vaults, approval-aware execution with pre-execution veto — map directly to what the 4+1 model defines as Intelligence Layer 2C. These capabilities are not yet shipping as GA products, but the engineering direction signals that Oracle understands the 2C problem and is building toward it. The question is execution velocity: can Oracle productize these concepts faster than AWS (which already has AgentCore Policy GA) or Google (which already has Agent Identity + Gateway + Registry GA)? Oracle's published vision is ahead of its shipped product — a familiar pattern for Oracle, which historically leads with database innovation and follows with cloud operationalization. The Fusion Agentic Applications contain implicit 2C: these agents make and execute decisions within business processes by accessing policies, approval hierarchies, and permissions. The agent governance is embedded in the application logic, not exposed as a configurable control plane. This is 2C — but it's Ceded 2C, where Oracle defines the governance model and the enterprise consumes it. ## ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Strongest Application-Layer Authority ### Vendor-Provided Components **Oracle Fusion Agentic Applications** [DAPM: Ceded] 22 specialized AI agents across HR, Finance, Supply Chain, and CX (GA March 2026). Outcome-driven, proactive, reasoning-based. Execute decisions within business processes by accessing unified enterprise data, workflows, policies, approval hierarchies, permissions, and transactional context. Built into Oracle Fusion Cloud Applications. **ISV + Model Ecosystem** [DAPM: Delegated] xAI Grok, Cohere Command A, NVIDIA Nemotron, IBM Granite models. IBM watsonx Orchestrate agents on Red Hat OpenShift on OCI. SoftBank sovereign cloud with custom AI models on OCI. Oracle Analytics with AI-powered assistants. ### NVIDIA-Provided Components **NVIDIA NIM + Nemotron Models** NVIDIA models via OCI Enterprise AI alongside xAI, Cohere, and customer models. ### Gap Analysis Layer 3 is OCI's most distinctive position in the assessment series. Oracle is the ONLY vendor that owns a comprehensive enterprise application suite AND the AI infrastructure to power it. AWS provides infrastructure and some first-party applications (Q, Connect). Google provides infrastructure and productivity applications (Workspace). Neither owns ERP, HCM, SCM, or CX at Oracle's scale. Fusion Agentic Applications (22 agents, GA March 2026) demonstrate what happens when the application vendor controls the AI infrastructure: agents that can access unified enterprise data, workflows, policies, approval hierarchies, permissions, and transactional context — without integration middleware, without API gateways, without cross-vendor authentication. This is not an ISV ecosystem (Dell's model) or a curated partner program (HPE's Unleash AI). This is first-party AI applications running on first-party infrastructure accessing first-party enterprise data. Specific agent domains: workforce scheduling, payroll issue resolution (HR), financial process automation (Finance), supply chain optimization (SCM), and customer experience enhancement (CX). Each agent is pre-trained on Oracle's understanding of enterprise processes — borrowed judgment at the business logic layer. The Oracle AI Data Platform for US Federal Government bundles OCI, Autonomous AI Database, and Enterprise AI with FedRAMP High and DISA IL4/IL5 authorization — extending Layer 3 into classified and sovereignty-constrained environments. The US Department of War agreement (May 2026) for AI on classified networks across 10 cloud regions at DISA IL2 through Top Secret demonstrates sovereign Layer 3 that no other hyperscaler matches in classification depth with equivalent application-layer integration. Custom agent development uses the OCI Responses API (OpenAI-compatible) with managed hosting for OSS frameworks and MCP servers — the same runtime surface described at Layer 2B, consumed here as an application development platform. The IBM partnership adds watsonx Orchestrate agents on OCI for HR use cases, extending the agentic ecosystem beyond Oracle's own applications. IBM Granite models via OCI Data Science provide additional model options. The SoftBank sovereign cloud platform on OCI (May 2026) demonstrates Layer 3 for sovereign AI: SoftBank's own generative AI models running on OCI Enterprise AI with full data control within Japanese data centers. The DAPM question: Fusion Agentic Applications are the deepest expression of Ceded authority in this assessment. The enterprise Cedes business logic, process automation, and decision-making to agents that Oracle built, Oracle trained, and Oracle hosts. The efficiency gains are real. The authority concentration is total. ### Borrowed Judgment The highest borrowed judgment of any Layer 3 in this assessment. Fusion Agentic Applications encode Oracle's interpretation of enterprise processes: what constitutes a complete talent review, how maintenance troubleshooting should proceed, when payroll exceptions should escalate. The enterprise inherits decades of Oracle's process design as embedded agent behavior. Model providers (xAI, Cohere, NVIDIA) add model-level borrowed judgment. Oracle's application logic adds process-level borrowed judgment. The combination is unique: no other vendor in this assessment embeds both model judgment AND business process judgment into a single managed AI application layer. DAPM Action 3 applies with maximum force: when you move off Oracle, what judgment doesn't move with you? Answer: the process logic, the data governance, the application state, the agent memory, and the transactional context — essentially everything above Layer 0. ### Working Notes The Oracle AI Data Platform for US Federal Government (March 2026) combines OCI, Autonomous AI Database, and Enterprise AI into a unified offering for government agencies with FedRAMP High and DISA IL5 authorization. This is Layer 3 for regulated environments — AI agents operating within classified and sovereignty-constrained boundaries. The US Department of War agreement (May 2026) for advanced AI capabilities on classified networks, leveraging 10 cloud regions at DISA IL2 through Top Secret and Special Access Program levels, demonstrates sovereign Layer 3 that no other hyperscaler matches in classification depth with equivalent application-layer integration. The 531% growth in multicloud database revenues suggests enterprises are increasingly running Oracle's data layer inside other hyperscalers — creating a cross-cloud data authority that could enable cross-cloud Layer 3 applications. Oracle's vision may be: own the data and application layers, be agnostic about the infrastructure layer. This is the inverse of AWS's strategy (own the infrastructure, be agnostic about the application layer). ════════════════════════════════════════════════════════════════════════════════ # Palantir AIP + Foundry + Apollo Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v3.2 - Label Vocabulary Reconciliation **Date:** July 12, 2026 **Source:** Palantir Architecture Center (Platforms, Ontology System, Multimodal Data Plane, Interoperability, Rubix, AIP Architecture), Apollo docs (How Apollo Works, Plans & Constraints), 'Securing Agents in Production' (Palantir blog, Jan 2026), AIPCon 9, Q1 2026 earnings (May 4, 2026), published 4+1 model v3.2 (July 12, 2026, /reconcile): 2A/2C statusLabels moved to capability vocabulary; no grades or chips changed. ## Summary Finding Palantir is the first vendor in this series that is not an infrastructure vendor at all, and reading its own documentation makes the inversion precise. Dell, HPE, VAST, Cisco, and the clouds build upward from silicon, storage, or a data center; Palantir builds downward from the decision. Its Architecture Center is explicit that the Ontology is designed to represent 'the complex, interconnected decisions of an enterprise, not simply the data.' Where Dell is strongest at Layer 0 and absent at Layer 2C, Palantir is structurally absent at Layer 0 and strongest at Layers 1A (governance), 2B (governed agent runtime), 2C (reasoning/orchestration), and 3 (value). It is the mirror image of every infrastructure vendor assessed so far. Palantir's data and compute architecture is genuinely open, and that openness is exactly what makes the capture hard to see. The Multimodal Data Plane uses Apache Iceberg as its primary table format, registers Databricks / Snowflake / BigQuery data through Virtual Tables 'without needless data duplication,' pushes compute down to those same engines, stores data at rest in open formats (Iceberg/Parquet) reachable over REST/JDBC/S3, and supports being 'one participant in a wider' data/AI mesh — Palantir's own 'unwalled garden.' Every one of those openness claims is true. But openness at the data and access layers is not portability of the thing the enterprise actually builds. The test that matters is whether a vendor's opinions — its proprietary way of modeling, governing, and orchestrating — can be lifted out and operated elsewhere. Palantir's opinions are the Ontology, and they run only on Palantir. The boundary, then, is not the data — it is the opinions plus the operating model. Even when data stays in Snowflake and compute pushes down to Databricks, the Ontology (objects, links, actions, and interaction-time security), the agent runtime, and the deployment control plane remain Palantir's, deployed on Palantir's hardened Kubernetes substrate (Rubix) and operated under Palantir's connection. The four-fold integration of data, logic, action, and security is the proprietary IP; the storage underneath is deliberately commoditized and open. The capture is Oracle-shaped: open at the access layer, captive at the layer of accumulated proprietary dependence — and like Oracle, the lift to leave compounds with every use case built, because every object model, action, agent, and workflow is Palantir-specific surface that would have to be rebuilt elsewhere. Apollo is the most important find for the 4+1 model, and the documentation makes both its power and its limit exact. Apollo is a genuine constraint-solving orchestration engine: a Hub continuously evaluates every possible Plan for each Spoke, evaluates all constraints attached to each Plan (maintenance windows, product-dependency version ranges, suppression windows, artifact availability), and issues only Plans whose constraints are satisfied — across connected, disconnected, and air-gapped estates under FedRAMP High, IL5, and IL6. Its explicit 'propose-a-Plan-then-execute' paradigm, with a dependency-graph invalidation model and break-glass overrides, is the closest thing in the entire series to the multi-variable policy reasoning the working notes describe. The limit: Apollo's constraints govern software deployment and day-2 operations — versions, dependencies, time windows, artifact presence — not live per-inference placement of which model, which region, which cost tier at request time. Palantir has built the control-plane MECHANISM the 4+1 model wants; it is pointed at the platform lifecycle and at agent actions, adjacent to the Infrastructure-2C placement function rather than identical with it. The DAPM crux cuts as sharply as it favors. Palantir hands the enterprise a real governance surface, a real governed agent runtime, and a real constraint-based control plane — but as a Ceded dependency operated under Palantir's judgment. The platform is not self-deployable; Palantir engineers deploy and manage it, Apollo holds a persistent connection back to Palantir for updates, monitoring, and orchestration, and Forward Deployed Engineering bridges the last mile. This places Palantir in the proprietary-captive cluster with Google, Oracle, and VAST — heavily Ceded — distinguished by a decoupled capture mechanism: where VAST couples (your data must enter its namespace, so the commitment is visible), Palantir decouples (your data stays open, so the commitment is invisible until you try to leave). The inverse-of-Dell shape holds end to end. Dell owns the floor and is Absent at the reasoning plane, so it says 'bring any Layer 2C' — and that 2C arrives Delegated and substitutable, leaving authority distributed. Palantir Cedes the floor and owns the reasoning plane, so it says 'bring any Layer 0' — and everything above the floor consolidates into one Ceded-and-operated authority. Same invitation, opposite consequence: a federated answer to 'who ties the control plane together' versus a unified one. Q1 2026 (revenue +85% YoY, 1,007 commercial customers, US commercial +133%, Rule of 40 at 145%) shows a growing number of enterprises accepting the unified trade. ## ○ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Not Palantir's Layer (By Design) ### NVIDIA-Provided Components **Sovereign AI OS (Palantir + NVIDIA)** Announced at AIPCon 9 (May 2026): a turnkey system combining NVIDIA Blackwell Ultra hardware with Palantir's full software suite, aimed at data-sovereignty and latency-sensitive deployments. This is the single place Palantir touches Layer 0 — and the silicon, interconnect, and acceleration are entirely NVIDIA's. It is the clearest market signal in the series of the control plane being assembled from both ends: the most application-down vendor (Palantir) paired with the most silicon-up vendor (NVIDIA) in one offering spanning Blackwell Ultra to the Ontology. **No Palantir Silicon, Switching, or Interconnect** Palantir designs no chips, switches, or fabric. By the Rubix documentation it deploys with identical operational characteristics across AWS, Azure, Google Cloud, Oracle Cloud, or on-premises — consuming whatever Layer 0 exists beneath it. ### Gap Analysis Layer 0 is simply not Palantir's layer, and the documentation treats this as a feature rather than a gap. Rubix (the hardened Kubernetes substrate) is explicitly designed to 'abstract away the peculiarities of different environments and providers,' giving identical operational characteristics on AWS, Azure, GCP, OCI, and on-prem. Palantir's value is structurally indifferent to the silicon underneath. The contrast with the infrastructure vendors is total and clean: Dell's strength is Layer 0 and its gap is Layer 2C; Palantir's gap is Layer 0 and its strength is Layer 2C. The Sovereign AI OS partnership with NVIDIA is the exception that proves the rule — when Palantir needs a Layer 0 story (sovereignty, latency, turnkey on-prem), it borrows NVIDIA's Blackwell Ultra stack wholesale rather than building its own. For the 4+1 model this is significant: because Palantir floats above Layer 0, its governance and reasoning surfaces are the only ones in the series that are not anchored to a particular infrastructure vendor's hardware. That is both the source of its federation claim (it can sit atop Dell, HPE, VAST, or a hyperscaler equally) and the reason its lock-in lives entirely in the upper layers rather than the silicon. ### Borrowed Judgment Total at Layer 0, and irrelevant to the value proposition by design. Palantir inherits all silicon, networking, and acceleration judgment from the host environment. The consequence worth tracking: the enterprise's Layer 0 choice (and its Layer 0 DAPM position) is made with a different vendor (Dell, a hyperscaler, or NVIDIA via the Sovereign AI OS), and Palantir simply rides on top — so a Palantir adoption decision does not, by itself, resolve any Layer 0 authority question. ### Working Notes The Sovereign AI OS is the entry worth tracking across the whole series, because it is the first offering that explicitly fuses the top and bottom of the 4+1 stack into one SKU. If the control plane ends up being co-built by an application-down vendor and a silicon-up vendor meeting in the middle (working notes, Pattern 4 and Open Question #2), this partnership is the prototype. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Palantir Strength — Governance Authority, Open Storage ### Vendor-Provided Components **The Ontology (Four-Fold: Data + Logic + Action + Security)** [DAPM: Ceded] Models the enterprise's DECISIONS, not just its data: semantic 'nouns' (objects, properties, links) paired with kinetic 'verbs' (actions, automations) and the logic behind them (business rules, ML models, LLM-driven functions, multi-step orchestrations), all woven through a security layer. A documented 'digital twin' enabling read-write loops between humans and agents. **Interaction-Time Security (Role / Marking / Purpose)** [DAPM: Ceded] Reconciles granular role-, marking-, and purpose-based policies at the moment of interaction across data, logic, actions, and LLM calls — down to row/column level. Agents inherit security scopes from a human user or a project's permission structure, so agent governance equals employee governance. Cataloged in expressive audit logging. **Open Storage & Virtual Tables (MMDP)** [DAPM: Delegated] Iceberg as primary table format; data at rest in open formats (Iceberg/Parquet, original CSV) over REST/JDBC/S3. Virtual Tables register Databricks/Snowflake/BigQuery data without duplication. The storage layer is deliberately open and commoditized — the 'unwalled garden.' **Metadata & Semantic Interoperability** [DAPM: Ceded] Metadata services expose mandatory (security/attribution/lineage) and discretionary (tags/enrichments) metadata across datasets, ontology elements, agents, models, and pipelines for connection to existing catalogs/MDM. Ontology elements are REST/JSON-accessible with bidirectional sync to external semantic tools; Palantir MCP enables agent-driven semantic interop. **Ontology SDK (OSDK) — the 'Operational Bus'** [DAPM: Ceded] Turns the Ontology into a programmatically queryable API gateway / operational bus across the enterprise — the queryable, action-bearing control surface the working notes call for, as opposed to a display-only catalog. **Lineage, Versioning & Global Branching** [DAPM: Ceded] Every data query tied to full version history and the transformation logic that produced it; object types, actions, logic, and policy rules are versioned. Global Branching applies software-engineering change management ('version control for reality') to operational data and logic for both humans and agents. ### NVIDIA-Provided Components **No NVIDIA Layer 1A Dependency** The Ontology, its security system, lineage, and the MMDP open-data architecture are Palantir IP. NVIDIA contributes nothing to the governance layer. ### Gap Analysis This layer splits cleanly into two findings: storage (open) and governance (Palantir's, and very strong). Storage is deliberately open. The Multimodal Data Plane commits to Apache Iceberg as the primary table format; data at rest is stored in original/open formats (CSV, Iceberg, Parquet) and reachable through REST, JDBC, and S3-compatible interfaces. The Virtual Tables framework registers data from Databricks, Snowflake, and BigQuery 'without needless data duplication,' and Palantir explicitly frames itself as able to be 'one participant in a wider, more heterogenous enterprise architecture' — an 'unwalled garden.' This is a documented openness at the data layer — it engages Pattern 1 (the metadata boundary problem) head-on: Palantir's Interoperability documentation describes metadata services that expose mandatory metadata (security, attribution, lineage) and discretionary metadata (tags, enrichments) across datasets, ontology elements, agents, models, and pipelines for connection to existing catalogs and MDM tools, plus bidirectional semantic synchronization with external semantic models. Governance is where Palantir is genuinely one of the strongest in the series. The Ontology binds data, logic, action, and security into one model, and the security system reconciles role-, marking-, and purpose-based controls 'at the time of interaction, across tens of thousands of humans and agents,' down to row/column-level restrictions. Agents take security scopes that inherit from a human user or a project's permission structure — so an agent is governed exactly like the employee it acts for. This is the 'control surface,' not the 'checkbox,' that Pattern 2 asks for: governance metadata that is programmatically queryable (via the Ontology SDK / OSDK as an 'operational bus'), real-time (evaluated at interaction time), and policy-aware (purpose-based controls). The honest residual boundary: the queryable, action-bearing governance lives in the Ontology, which is Palantir's. You can keep your data in Snowflake; the decision graph that makes it operational — and the authority that governs it — is Palantir's. The five working-notes criteria are largely met (queryable, real-time, policy-aware, observable), and 'cross-platform' is met via virtual tables and metadata/semantic interop rather than only via ingestion. ### Borrowed Judgment Low for the governance logic itself — the Ontology Language/Engine/Toolchain, the interaction-time security system, and lineage are Palantir IP. The structural dependency is subtle: data need NOT enter Palantir's storage (open formats, virtual tables, pushdown), but to be GOVERNED and made operational by Palantir, it must be modeled in Palantir's Ontology and mediated by Palantir's security system. The enterprise can retain its storage substrate while ceding the decision-and-governance layer. Compared to Dell's MetadataIQ (indexes Dell storage in place, stays at Layer 1A), Palantir's governance reaches up into Layer 2C — but the authority that does the reaching is Palantir's. ### Working Notes The load-bearing concepts are the 'four-fold integration' (data + logic + action + security) and the Language/Engine/Toolchain decomposition: together they are why this is a governance authority rather than a catalog. The enterprise can keep its data substrate open (virtual tables, pushdown) while the decision graph that makes that data operational remains Palantir's. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Palantir Strength — Governed Retrieval ### Vendor-Provided Components **Ontology-Object Retrieval** [DAPM: Ceded] Agents retrieve governed objects (with links, logic, and security) rather than raw chunks; context is continuously integrated into the Ontology. Retrieval inherits interaction-time security automatically. **Vector, Compute & Tool Services** [DAPM: Ceded] Integrated vectorization to produce/manage embeddings; extensible compute (multi-node Spark/Flink, single-node DuckDB/Polars, or BYO containerized engines); and a tool-services layer that functions as an evolving 'tool factory' for agents. Modular and model-agnostic. **AIP Assist / Context-Aware Surfaces** [DAPM: Ceded] Context-aware assistance across out-of-the-box applications, grounded in the Ontology to shorten time-to-value when exploring governed data. ### NVIDIA-Provided Components **No Hard NVIDIA Dependency** Vectorization and retrieval are Ontology/AIP-native services on the Rubix compute mesh. GPU acceleration is inherited from the host Layer 0 when present, not architecturally required; embeddings can be produced by any registered model. ### Gap Analysis Palantir's retrieval story is distinctive and well-documented: rather than RAG-over-raw-text producing 'a slightly better search engine,' agents retrieve and reason over Ontology objects — governed business entities carrying their links, logic, and permissions. The AIP architecture lists 'vector, compute, tool services' as a first-class capability: integrated vectorization to produce and manage embeddings, plus an extensible compute framework. Context is continuously integrated into the Ontology rather than assembled ad hoc at query time, which raises retrieval quality and governance simultaneously — every retrieved object carries its security markings, so retrieval respects the same interaction-time policy as everything else. The trade-off mirrors Layer 1A: retrieval is rich and governed within the Ontology's semantic frame. Compared to VAST's InsightEngine (purpose-built vector retrieval native to the data platform with permission-inheriting vector rows) or Dell's Elastic-based hybrid search, Palantir's retrieval value is semantic and governed rather than infrastructural — it is less a raw vector-DB play and more 'retrieval over a permissioned decision graph.' Because embeddings can be generated by any registered model and data can sit in virtual tables, retrieval does not force a storage migration. The score here is strong on the basis of governed retrieval — permission-aware object retrieval is what Palantir's buyers come for — not on raw vector-search performance or scale, where the purpose-built players lead. VAST and Palantir are both strong at 1B for opposite reasons: VAST on retrieval performance native to the data plane, Palantir on retrieval governance native to the Ontology. Against the working-notes criterion of retrieval-quality observability feeding placement: Palantir partially supplies it because AIP Evals suites are automatically tracked against the functions and sub-agents that consume context, giving a governed record of how retrieval-dependent behavior shifts over time — closer to the 'observable' criterion than most infrastructure vendors, though still oriented to agent quality rather than to a Layer 2C placement engine optimizing recall@k against cost. ### Borrowed Judgment Low. Retrieval semantics, vectorization services, and the context-integration model are Palantir's. The enterprise inherits Palantir's judgment about how context is modeled, embedded, and surfaced — which is precisely the value, but also means retrieval behavior is defined inside Palantir's framework rather than configured against an open, swappable vector store the enterprise operates independently. ### Working Notes Because retrieved units are governed objects rather than text chunks, retrieval inherits the row/column/marking/purpose security of Layer 1A automatically. This is the same structural-security property VAST achieves via permission-inheriting vector rows, reached here through the Ontology's unified security model rather than a shared Element Store. ## ◑ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Foundry Pipelines + Open Compute (MMDP) ### Vendor-Provided Components **Foundry Pipelines & Context Engineering** [DAPM: Ceded] Extensible multimodal connection/transformation across batch, streaming, and real-time replication (CDC) on any bundled runtime (Spark, Flink, DataFusion, Polars), with cohesive security, governance, and provenance tracking. Feeds the Ontology ('South of the Ontology'). **Open / Pushdown Compute & BYO Compute** [DAPM: Delegated] Pushdown to Databricks/Snowflake; orchestration with external inference infra, Spark clusters, and on-prem HPC; Compute Modules import any containerized runtime/model/executable, securely orchestrated by Rubix. Pipelines can run where existing compute lives. **Global Branching Change Governance** [DAPM: Ceded] Proposed changes (human or agent) live on a branch; reviewed and merged. Versioning spans object types, actions, logic, and policy rules — software-engineering governance applied to operational data and logic. ### NVIDIA-Provided Components **No NVIDIA Layer 1C Dependency** Pipeline authoring, transformation, lineage, and the open compute framework are Foundry/MMDP IP. Acceleration, if any, comes from the host Layer 0. ### Gap Analysis Foundry is a mature data-operations platform — connection, transformation, pipeline authoring, and lineage feed the Ontology — and the MMDP documentation makes the compute side notably open. The 'any compute' architecture supports pushdown to cloud-native runtimes like Databricks and Snowflake, 'Bring Your Own Compute' via Compute Modules (any containerized runtime securely orchestrated by Rubix), and orchestration with external inference infrastructure and on-prem HPC. Pipelines can run where the data and existing compute investment already live, not only inside Palantir. The contrast with Dell's Layer 1C is instructive. Dell's most distinctive 1C capability is KV-cache-to-storage offload — an infrastructure-physics optimization with direct inference economics. Palantir has no equivalent, because it does not operate at the storage-physics layer; its movement is logical and operational rather than infrastructural. Where Dell solves 'move bytes between PowerScale and the GPU cluster efficiently,' Palantir solves 'transform and govern data into operational objects, wherever the bytes physically sit.' The governance overlay is the differentiator: Global Branching means a pipeline or logic change — proposed by a human or an agent — lands on a branch, is reviewed, and is merged, with the change versioned across object types, actions, logic, and policy. That makes data movement auditable and reversible at the operational layer, not just the file layer. ### Borrowed Judgment Low for pipeline, transformation, lineage, and the open compute framework (Palantir IP). The dependency is that the governed terminus is the Ontology: movement and transformation, however open the runtime, exist to populate and update Palantir's governed model. Pushdown to Snowflake/Databricks genuinely reduces substrate lock-in; the decision layer that consumes the result does not move. ### Working Notes Palantir's interoperability posture is open in both directions — open formats at rest and governed egress to external systems — which is a meaningfully stronger stance than the on-prem storage vendors, whose openness is mostly about ingest. ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Platform-Scoped Orchestration ### Vendor-Provided Components **Rubix (Hardened, Autoscaling Kubernetes)** [DAPM: Ceded] Palantir's own hardened, autoscaling Kubernetes substrate that orchestrates the platform's workloads — secure-by-default networking, policy-driven node management, continuous cost optimization, FedRAMP High/IL5-6/CMMC. Real orchestration the enterprise consumes without configuring; platform-scoped, not a scheduler for the enterprise's separate fleet. **Rubix↔Apollo Execution Layer** [DAPM: Ceded] Apollo computes constraint-satisfied Plans; Rubix executes them via zero-downtime rollouts with automated rollback. Orchestration intelligence separated from execution mechanism — the bridge to the 2C control plane. **Compute Modules (Customer Workloads on Rubix)** [DAPM: Ceded] Lets a customer run a containerized workload on Rubix and inherit its placement, isolation, and cost optimization. A runtime convenience, not an exposed accelerator-fleet scheduler with quotas or fair-share the enterprise governs. ### NVIDIA-Provided Components **No NVIDIA Scheduler Dependency** Rubix is Palantir's own hardened, autoscaling Kubernetes implementation. There is no Run:ai dependency; GPU primitives, when needed, are inherited from the host environment. Rubix runs identically across AWS, Azure, GCP, OCI, and on-prem. ### Gap Analysis Layer 2A is Ceded, not absent — and the distinction from Layer 0 is the key to scoring it correctly. At Layer 0 Palantir genuinely provides nothing; you bring the floor. At 2A, orchestration absolutely happens: the agents and applications the enterprise buys are scheduled, autoscaled, isolated, and cost-optimized on Rubix, Palantir's hardened Kubernetes substrate. The capability exists and Palantir controls it. What the enterprise does not get is governance authority over it — no exposed GPU-scheduling, quota, or fair-share surface its infrastructure teams or AI application models can configure. Capability exists, vendor controls it, enterprise consumes without authority: that is the textbook definition of Ceded. This matches how the clouds score at 2A. AWS, Google, and Azure all orchestrate customer workloads through managed, largely invisible scheduling that the enterprise cannot configure or override — and that scores as present-and-Ceded, not absent. Rubix is architecturally the same fact: real orchestration the customer consumes without holding the keys. Scoring Palantir's managed orchestration as absent while the clouds' managed orchestration is present-and-Ceded would treat the same architectural pattern two different ways. The strength is moderate, not strong: Rubix is real and capable, but it is platform-scoped (it orchestrates Palantir's workloads, not the enterprise's separate GPU fleet) and it is not a differentiator the buyer chooses Palantir for. Compute Modules lets a customer run a containerized workload on Rubix and inherit its placement, but it is a runtime convenience, not an accelerator-fleet scheduler the enterprise operates. If the enterprise runs a separate GPU cluster for its own training/inference, that cluster still needs its own scheduler (the cloud's, Run:ai, or Kubernetes) — Palantir orchestrates what runs on Palantir, not the wider estate. Against the inverse-of-Dell spine: Dell's 2A is a gap (the capability is needed and NVIDIA's Run:ai holds it); Palantir's 2A is Ceded (the capability exists, is real, and Palantir holds it). Both thin at 2A in the sense that neither hands the enterprise authority — but for opposite structural reasons. ### Borrowed Judgment Moderate and Ceded. The enterprise inherits Palantir's judgment about how its Palantir-run workloads are scheduled, scaled, isolated, and cost-optimized — node ephemerality, placement heuristics, workload distribution — without the ability to configure or override it. This is the same borrowed-judgment posture as the clouds' managed orchestration: you get the benefit of the vendor's operational opinion and you do not get to change it. It is bounded, though, because it governs only the Palantir footprint; the enterprise's separate fleet, if any, runs on its own scheduler under its own authority. ### Working Notes Rubix is genuine engineering, productized beyond Palantir's own software (offered to other software vendors for regulated-environment deployment via Palantir FedStart, and underpinning the Mission Manager government onboarding offering). The Rubix↔Apollo split is the architecturally notable detail and the bridge to 2C: Apollo computes the constraint-satisfied Plans, Rubix executes them via zero-downtime rollouts with automated rollback — orchestration intelligence separated from execution mechanism. For the 4+1 buyer, the honest framing is that this orchestration is real and Ceded: you benefit from it, you don't govern it, and it covers the Palantir footprint rather than your whole accelerator estate. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Palantir Strength — Governed Agent Runtime ### Vendor-Provided Components **AIP Agent Runtime (Stateful Loop on Rubix)** [DAPM: Ceded] Stateful control loop over a stateless reasoning core, executing tools/memory under infrastructure- and platform-level guardrails plus developer-configured controls. Per-workload isolation; operational vs. application executions distinguished; every interaction authenticated, authorized, logged. Agents are governed identities, like employees. **Secure 'Any Model' Integration (Model Catalog)** [DAPM: Ceded] Governed access to commercial and open LLMs plus customer/fine-tuned models on a level playing field, via Palantir-managed infra with no provider retention or retraining; regional endpoints where available; token-limit governance across use-cases. Models are swappable and Evals-comparable. **Agent Lifecycle: AIP Logic, Chatbot Studio, Code Workspaces** [DAPM: Ceded] No-/low-/pro-code construction of LLM-driven functions and durable agent orchestrations on top of the Ontology, with tool access scoped by permission (data tools, logic tools, operational tools). **AIP Evals** [DAPM: Ceded] Evaluation suites operating directly against the Ontology: create test cases, debug/iterate agent definitions, compare performance across LLMs, examine execution variance. Automatically tracked against the functions and sub-agents they govern. **End-to-End Observability** [DAPM: Ceded] Monitoring of every Ontology-feeding data flow, every human/agent action, chained-execution traces, and token/resource consumption. Telemetry and log access are themselves governed by data markings. ### NVIDIA-Provided Components **Model-Agnostic 'Any Model' (MMDP)** AIP's Model Catalog offers commercial models (OpenAI, Anthropic, Google, xAI) and open models (Meta/Llama) on a level playing field with customer-registered, fine-tuned, and existing enterprise models, via Palantir-managed infrastructure that guarantees no provider data retention and no retraining on transmitted data. NVIDIA NIM can be one registered model source among many; none is privileged. Access can be governed with token limits across use-cases. ### Gap Analysis This is one of Palantir's two strongest layers and a direct contrast to Dell, which Cedes the entire runtime to NVIDIA (NemoClaw/OpenShell/Dynamo). The 'Securing Agents in Production' documentation defines an agent as 'a stateful control loop that repeatedly invokes a stateless reasoning core (a frontier language model), interprets its outputs, executes tools and memory options, and feeds the results back until a termination condition is met' — and then makes that loop the unit of governance. Agents run on Rubix with per-workload isolation; every interaction is authenticated, authorized, and logged; and crucially the runtime distinguishes operationally-privileged executions from application-driven executions operating under precisely governed permissions. The 'any model' philosophy makes the model a swappable component rather than the center of gravity — the Ontology is the center. This is the structural inverse of Google's model-integrated stack, where one model's (Gemini's) judgment pervades every layer; in Palantir the reasoning core is deliberately interchangeable and compared via Evals. The AIP agent lifecycle is a documented build/orchestrate/evaluate loop: no-/low-/pro-code construction (AIP Logic for low-code durable orchestrations, Code Workspaces for pro-code), with AIP Evals operating directly against the Ontology to create test cases, debug, compare performance ACROSS different LLMs, and examine variance across executions. Observability is end-to-end: every data flow into the Ontology, every action by a human or agent, the cascade of chained executions, and even token consumption are monitored — and log access is itself governed by data markings. The agent-as-governed-identity model is the decisive property: agents operate atop the same foundation as human users, abide by the same change management (Global Branching), and weave human-in-the-loop with autonomous operations. Specialized builder agents (AI FDE, AIP Analyst) can themselves construct pipelines, write logic, train models, and build ontologies — under the same governance. This is genuinely Palantir's IP, not a partner framework wrapped in packaging. ### Borrowed Judgment Low for the runtime, governance, isolation, and evaluation machinery — all Palantir IP. The reasoning core is explicitly borrowed but equally explicitly interchangeable (any model, no retention, compared via Evals). The enterprise inherits Palantir's judgment about HOW agents are isolated, permissioned, evaluated, observed, and audited; it retains choice over WHICH model reasons. That decoupling is the opposite of the hyperscaler model-integrated pattern and is, for many regulated buyers, the point. ### Working Notes The 2026 forward bet is 'Agentic AI Hives' — autonomous agent networks coordinating on complex problems (e.g., supply-chain disruptions) without human intervention: a shift from decision-support to decision-execution. Shared-Ontology memory plus uniform governance is what Palantir argues makes scaling from single agents to coordinated networks an engineering problem rather than an architectural rewrite. Karp's Q1 2026 framing — differentiating Palantir from model developers amid a 'thousandfold' token-cost decline — is precisely a claim that the durable value is this governed runtime/decision layer, not the model. ## ● Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Closest Productized Mechanism ### Vendor-Provided Components **Apollo Orchestration Engine (Constraint Solver)** [DAPM: Ceded] Hub/Spoke topology; continuously evaluates possible Plans against their attached constraints (deployment safety, dependencies, timing, environment health) and issues only satisfied Plans. Automated rollback that always respects human-set holds; connected/disconnected/air-gapped under FedRAMP High/IL5/IL6. Governs platform lifecycle and day-2 ops — not live inference placement. **Interaction-Time Policy Reasoning (Ontology)** [DAPM: Ceded] Reconciles granular role/marking/purpose security across data, logic, actions, and LLM calls for tens of thousands of humans and agents at the moment of interaction — the decision plane for what agents may do. Productized and shipping. **Agent Governance & Observability** [DAPM: Ceded] Agents as governed identities; tool calls and queries versioned and tied to Evals; chained-execution tracing; telemetry and log access governed by data markings. A feedback loop for detecting and constraining faulty agent reasoning. **Persistent, Vendor-Operated Connection (Not Self-Deployable)** [DAPM: Ceded] Palantir engineers deploy and manage; Apollo holds an ongoing connection for updates, monitoring, and orchestration ('shared security model'). The reasoning plane is operated under Palantir's authority — the defining DAPM caveat for this layer. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** All reasoning-plane logic — Ontology interaction-time security, Apollo's constraint solver, agent governance and observability — is Palantir IP. ### Gap Analysis Applying the working notes' 'Routing Is Not Reasoning' test, Palantir comes closer than any infrastructure vendor — and the answer has two halves plus a limit. (1) DECISION governance (Intelligence-2C): the Ontology's interaction-time security plus the AIP runtime constitute a genuine reasoning plane for WHICH agent may take WHICH action on WHICH object under WHICH policy, with role/marking/purpose controls reconciled at the moment of action across tens of thousands of humans and agents. Productized and shipping, and the most complete agent-action governance surface in the series alongside Google's. (2) ORCHESTRATION (the surprising part): Apollo is a real constraint-solving control plane. A Hub continuously evaluates every possible Plan for each Spoke, evaluates the constraints attached to each Plan, and issues only those whose constraints are satisfied — with automated rollback that always respects human-set holds, across connected, disconnected, and air-gapped estates under FedRAMP High / IL5 / IL6. Palantir's own framing — 'different from other control-loop systems' because it proposes a transparent Plan and executes only on satisfied constraints rather than acting silently — is almost verbatim the auditable, multi-variable policy reasoning the working notes say is missing. THE LIMIT (the crux for the 4+1 model): Apollo's constraints govern SOFTWARE DEPLOYMENT and day-2 OPERATIONS, not LIVE PER-INFERENCE PLACEMENT of which model serves which request, in which region, at which cost/latency/compliance tier. Palantir has built the control-plane MECHANISM the model wants, and the DECISION plane for agent actions — but the mechanism is pointed at the platform lifecycle and at agent governance, ADJACENT to the live inference-placement function rather than identical with it. No vendor in the series fully productizes that live-placement engine; Palantir is the closest on mechanism, Google the closest on the data→placement chain. ### Borrowed Judgment This is the heart of the assessment. Unlike Dell (no judgment to borrow — build custom 2C in 6–12 months, bring a partner, or operate without it), Palantir hands the enterprise a real reasoning plane — but as a fully Ceded, vendor-OPERATED dependency. The platform is not self-deployable: Palantir engineers deploy and manage it; Apollo maintains a persistent connection back to Palantir for updates, monitoring, and orchestration (Palantir's documented 'shared security model'); and Forward Deployed Engineering sustains it. So the judgment exercised by both the decision plane (Ontology security) and the orchestration plane (Apollo) is Palantir's, running inside Palantir's operating relationship. This is the same answer VAST and Google give — the vendor holds Layer 2C — but with a distinctive shape: open data substrate underneath, closed governance/decision authority above, vendor-operated throughout. Absent (Dell) is worse than Ceded (Palantir); but Ceded-AND-vendor-operated is a heavier governance position than Ceded-but-self-operated (e.g., software a customer runs themselves). ### Working Notes Maps directly onto the working notes. Open Question #1 (product or pattern?): Apollo is the series' strongest evidence that the control plane can be a PRODUCTIZED constraint engine, not merely a pattern. Open Question #2 (who builds it?): Palantir is the concrete realization of the 'governance platform / new category' candidate, and the option the Dell assessment named ('potentially Palantir Ontology'). Open Question #7 (does 2C unify Infrastructure-2C and Intelligence-2C?): Palantir unifies them organizationally — one vendor, one authority — but not functionally; agent-action governance and operational orchestration are strong, while live inference placement is the unaddressed third. The DAPM distinction the model surfaces is precise: Dell at 2C is ABSENT (no authority to cede); Palantir at 2C is CEDED-AND-OPERATED (authority exists, is productized, and is held and run by the vendor). ## ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Palantir Strength — Where It Starts, Not Where It Ends ### Vendor-Provided Components **Workshop & AIP Applications** [DAPM: Ceded] Operational apps and workflows built directly on governed Ontology objects — object-oriented analytics, real-time app building, multimodal governance workflows, persona-tailored out-of-the-box applications, with AI infusion controlled and transparently assessed. **Action-Bearing Agents (Decision Execution)** [DAPM: Ceded] Agents propose and execute real business actions on Ontology objects under human-in-the-loop branching governance — the 'decision-execution' shift toward Agentic AI Hives. **Enterprise-Automation Builder Agents (AI FDE, AIP Analyst)** [DAPM: Ceded] Specialized agents that construct pipelines, write business logic, train models, build ontologies, and develop applications — operating atop the same foundation and governance (Global Branching, interaction-time security) as human builders. **Forward Deployed Engineering** [DAPM: Ceded] Palantir's human delivery model bridges the last mile from data to operational reality. Value-accelerating, but it deepens the vendor-operated dependency rather than reducing it. ### NVIDIA-Provided Components **Model Source Only** NVIDIA's role at Layer 3 is as one possible model/runtime source (NIM) under 'any model.' Business logic, workflows, and applications are Palantir-and-customer-owned and built on the Ontology. ### Gap Analysis Layer 3 is where every infrastructure vendor is weakest (partner ecosystems) and where Palantir is native — indeed it is where Palantir STARTS and reaches down from. Because the Ontology binds data, logic, action, and security, applications and agents are built directly against governed business objects: object-oriented analytics, real-time application building (Workshop), multimodal governance workflows, and out-of-the-box applications tailored to operational users, compliance teams, engineers, and analysts, all with the 'infusion of AI carefully controlled and transparently assessed, ensuring a smooth journey from augmentation to automation.' Agents propose and execute real actions on Ontology objects (reroute shipments, trigger purchase orders) under human-in-the-loop branching governance. The contrast with Dell's Layer 3 is the most illuminating in the series. Dell assembles independently-governed ISV agent populations (Palantir is itself one of Dell's named ISVs) with no cross-domain governance binding them on shared infrastructure. Palantir is one such population — but INTERNALLY it has exactly the cross-domain governance the Dell stack lacks: every app and agent shares the same Ontology, the same interaction-time security, the same Evals, the same Global Branching change control. The 4+1 tension this exposes: Palantir solves the multi-agent governance problem WITHIN its boundary; it does not solve it ACROSS an enterprise running Palantir AND ServiceNow AND a homegrown stack — which is the federated reality the working notes (Pattern 3) insist on. Palantir's answer to heterogeneity is MMDP openness at the data layer, not governance federation at the agent layer. Validated traction is the strongest at this layer in the entire series: Q1 2026 revenue +85% YoY to $1.633B, 1,007 commercial customers (+31%), US commercial +133%, Rule of 40 at 145%, with named production deployments (SAP reporting >99% validation accuracy and large cloud-migration reductions; GE Aerospace; Airbus; Stellantis). Analysts characterized AIP as 'an operational system for deploying agents with governance, cost attribution, and auditability, not a model wrapper' — a Layer 3/2C claim, not a Layer 0/1 one. ### Borrowed Judgment Distributed between Palantir and the customer, which is architecturally correct at Layer 3 — but unlike a neutral application platform, the value layer is inseparable from Palantir's governance, runtime, and operating model beneath it. You do not adopt Palantir's Layer 3 without adopting the Ontology, Rubix, Apollo, and the vendor-operated relationship. The value is real and the governance is real; both are Ceded together. ### Working Notes Forward Deployed Engineering is the human mechanism that bridges the 'last mile' from data to operational reality and is central to Palantir's traction — but it deepens the operating dependency rather than reducing it, reinforcing the Layer 2C 'vendor-operated' finding. The AIP architecture's 'enterprise automation' category (specialized builder agents like AI FDE and AIP Analyst constructing pipelines, logic, models, and ontologies under the same governance as humans) is the clearest expression of the augmentation→automation trajectory and of why the value layer cannot be cleanly separated from the layers beneath it. ════════════════════════════════════════════════════════════════════════════════ # Qlik (Cloud Analytics + Talend Cloud + Open Lakehouse) Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.0 - Initial Assessment **Date:** July 7, 2026 **Source:** help.qlik.com (Open Lakehouse architecture, Qlik MCP server, Qlik Predict, client-managed Qlik Sense), Qlik press releases (Open Lakehouse GA Sept 2025; agentic analytics GA + MCP Server Feb 2026; Qlik Connect 2026; agentic data engineering GA June 2026), Qlik Community FAQ and support articles, published 4+1 model ## Summary Finding Qlik is a data integration and analytics software vendor that floats above Layer 0 (SaaS on AWS, and AWS-only) and is strongest at the two ends it has always owned: data movement (1C — the Attunity/Talend/Upsolver heartland) and the analytics value plane (Layer 3 — the application estate plus a GA detect-predict-act loop). The middle of the stack is uniformly moderate — storage/governance, retrieval, orchestration, and runtime are all real but product-bounded — and the reasoning plane is a gap by explicit posture. Authority sits in Qlik's opinion layer wherever Qlik ships one; the substrate never. This is the cleanest zero-NVIDIA row in the instrument: no GPU plane, no accelerator dependency, no NVIDIA software surface at any layer. Qlik's entire generative chain is instead double-stacked borrowed judgment — Qlik chooses the model, AWS serves it, Anthropic reasons — with no customer model choice anywhere. The enterprise that adopts Qlik's AI inherits two vendors' judgment above it and holds authority over neither. Two capture generations coexist in one platform, and the row's shape inverts Databricks': Qlik's newest surface is its most open and its oldest is its most captive. Open Lakehouse is the most Retained-friendly lakehouse posture in the series — Iceberg tables in the customer's own S3, cataloged in the customer's own AWS Glue, valid and readable if Qlik vanishes tomorrow. The legacy estate is classic coupled capture — QVD libraries and set-analysis logic that lift to nowhere. And the new decoupled ask sits between them: the Trust Score and data-product layer is Qlik asking the enterprise to cede trust itself to the platform, an ask made to feel free precisely because the bytes underneath stay visibly yours. Qlik's reasoning-plane answer is MCP-first: be a well-governed tool inside somebody else's agentic infrastructure rather than ship a thin gateway and call it a control plane. That scores as a capability gap and reads as architectural honesty — the 'bring any Layer 2C' invitation, made one layer up from where Dell makes it, with the same consequence: the reasoning plane stays the enterprise's to build, and the authority stays Retained by default. The DAPM profile makes the trade exact: of 20 scored components, 19 are Ceded and one is Delegated (the lakehouse storage interface — the single genuinely portable thing Qlik touches). The buyer gets best-in-class heterogeneous data movement and a complete, GA insight-to-action loop for the analytics domain, in exchange for ceding every opinion layer Qlik ships — pipelines, trust, retrieval, models, agents, apps — while retaining the most portable data substrate in the series and full default ownership of the reasoning plane Qlik declined to claim. ## ○ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Not Qlik's Layer (By Design) ### NVIDIA-Provided Components **No NVIDIA Surface Anywhere in the Stack** Effectively zero NVIDIA dependency across the entire Qlik estate — no GPU story at all. The associative engine is CPU-native, Qlik Predict is classical AutoML, and generative inference is mediated by Amazon Bedrock. The emptiest NVIDIA column in the series: Qlik's borrowed AI judgment flows to AWS and Anthropic, not NVIDIA. ### Gap Analysis Layer 0 is not Qlik's layer, by design. The buyer never thinks about it, and that is the pitch: Qlik Cloud is SaaS on AWS. The one place Qlik touches infrastructure is Open Lakehouse, where compute runs on EC2 Spot instances inside the customer's own AWS VPC — the customer pays the EC2 bill and keeps data sovereignty while Qlik provisions and manages the cluster. Client-managed Qlik Sense still exists for regulated shops (May 2026 release shipping, EOS 2028), running on the customer's own Windows servers. The consequence for the 4+1 model is the same as Databricks' and Palantir's: a Qlik adoption decision resolves no Layer 0 authority question — whatever capture exists at silicon and fabric belongs to AWS, a different row on this map. The asymmetry worth naming: Qlik Cloud is effectively AWS-only (Open Lakehouse requires EC2, S3, and Glue), so the Layer 0 question is not even multiple-choice. The VPC-resident lakehouse compute is a deployment topology, not a Layer 0 capability — the same reasoning that keeps Databricks' classic-clusters-in-your-VPC from moving its Layer 0 cell. ### Borrowed Judgment Total at Layer 0, and irrelevant to the value proposition by design. Qlik inherits all silicon, networking, and acceleration judgment from AWS. The enterprise's Layer 0 authority position is set by AWS's row, and adopting Qlik does not, by itself, resolve it — it does, however, quietly commit the AI-compute question to AWS, because there is no non-AWS path. ### Working Notes No documented non-AWS Qlik Cloud region was found (including the AWS European Sovereign Cloud collaboration), but the docs do not state AWS-only as policy — inference, flagged. Does not move the cell; it is a gap regardless. ## ◑ Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Open Lakehouse on Your Substrate; Captive Trust Layer ### Vendor-Provided Components **Qlik Open Lakehouse (Iceberg on Customer S3 + Glue)** [DAPM: Delegated] GA Sept 2025. Iceberg tables written to the customer's own S3, cataloged in the customer's own AWS Glue catalog; interoperable with Snowflake, Athena, Trino, SageMaker. The consumed interface is a genuine multi-vendor standard — the tables lift to any Iceberg writer without rebuilding. Delegated (a standard interface consumed through Qlik's managed service; the customer owns the account but AWS operates the substrate and Qlik operates the writing layer). **Adaptive Iceberg Optimization (Upsolver Engine)** [DAPM: Ceded] Continuous compaction, dynamic partitioning, and file cleanup. Proprietary optimizer with no open exit — but its output is standard Iceberg, so leaving means swapping optimizers, not rebuilding tables. Ceded, with a low blast radius the litmus itself concedes. **Data Products + Trust Score** [DAPM: Ceded] Curated, governed datasets scored on accuracy, timeliness, diversity, and completeness; Data Product Agent GA June 30, 2026. The curation opinions — product definitions, the trust-scoring model, the lineage graph — are Qlik-captive. Proprietary Qlik platform — opinions captive, no open exit. **Data Quality & Semantic Types (Talend Heritage)** [DAPM: Ceded] Validation rules, semantic types, and quality management threaded through Qlik Talend Cloud; Data Quality Agent GA June 30, 2026. Rules and validation logic do not lift to another quality platform. Proprietary Qlik platform — opinions captive, no open exit. **QVD/QVF Associative Data Layer** [DAPM: Ceded] Proprietary extract and application data format readable only by the Qlik engine. The accumulated QVD estate — often a decade deep — is the longest-lived lock-in surface in the row: coupled, visible, classic capture. Proprietary Qlik platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 1A Dependency** Storage, catalog, quality, and trust scoring are Qlik IP or AWS services. NVIDIA contributes nothing to the governance layer. ### Gap Analysis Two generations of storage story. The new one: Qlik Open Lakehouse (GA September 2025) writes Iceberg tables into the customer's own S3 buckets, cataloged in the customer's own AWS Glue catalog, in the customer's own AWS account. Adaptive Iceberg Optimization (the Upsolver engine) handles compaction, partitioning, and cleanup continuously; any engine — Snowflake, Athena, Trino, SageMaker — queries the same tables. Layered on top: data products with Trust Score (accuracy, timeliness, diversity, completeness), data quality rules and semantic types (the Talend heritage), end-to-end lineage, and, GA June 30, 2026, the Data Product Agent and Data Quality Agent. The openness claim here is stronger than Databricks' — and mostly true. Databricks keeps data in your bucket but governs through its captive managed Unity Catalog; Qlik did not build a captive lakehouse catalog at all — the catalog is your Glue. If Qlik walks away, valid Iceberg tables remain in your account, readable by anything. What is captive is the trust layer: quality rules, semantic types, Trust Scores, data product definitions, and the lineage graph are Qlik opinions that do not lift. The intent is the instrument's concern regardless of field adoption: Qlik presents the Trust Score and data products as the governance surface — the ask is that the enterprise cede trust to the platform, and the openness of the substrate underneath is what makes that ask feel free. And the old capture surface still runs: enterprises with a decade of QVD/QVF extract libraries hold a proprietary analytic data layer readable only by the Qlik engine — the classic, coupled kind of capture, sitting right next to the new open one. Calibration: Databricks and Palantir are both strong at 1A because their governance is an authorization authority — Unity Catalog does RBAC, masking, and row-filtering at query time; the Ontology reconciles security at interaction time. Qlik's governance is curation and quality metadata (trust scoring, semantic types, data products), not an access-control authority — the lakehouse delegates access enforcement to AWS IAM and Glue. Storage capability is real but one year old and AWS-only. Partial-but-genuine on both halves of the layer: moderate. The inverse of the Databricks trade — less capability, materially less capture. ### Borrowed Judgment Split by generation. The lakehouse substrate borrows AWS's judgment (S3, Glue, IAM) through a genuinely open interface the enterprise could repoint. The trust layer is where Qlik asks you to cede: quality rules, Trust Scores, data product definitions, and lineage are Ceded to Qlik and do not lift. The QVD estate is the oldest and heaviest cession in the row. ### Working Notes Vendor follow-up: whether Open Lakehouse fine-grained access control integrates AWS Lake Formation or stops at IAM/Glue permissions — does not move the DAPM (the authority is AWS's either way) but sharpens the governance-is-not-authorization reasoning. Watch-list: Open Lakehouse streaming ingestion/transformations were announced with planned GA Q1 2026; delivery not yet confirmed in docs (see Layer 1C). ## ◑ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Turnkey RAG, Closed Box ### Vendor-Provided Components **Qlik Answers Knowledge Bases** [DAPM: Ceded] GA Sept 2024. Managed indexing, retrieval, and cited answers over unstructured sources. Proprietary managed RAG: indexes, chunking, and retrieval opinions do not lift; the OpenSearch underneath is invisible and not customer-operable. Leaving means rebuilding RAG elsewhere from the source documents. Proprietary Qlik platform — opinions captive, no open exit. **Associative Engine as Agent Context (via MCP)** [DAPM: Ceded] The agentic experience and MCP Server retrieve from Qlik apps and data products; the exclusion primitive (non-association as a native engine state) is the singular capability no other row matches. MCP is a genuine multi-vendor interface, but the opinions that accumulate — Qlik app data models, associative logic — live behind the interface, not at it, and run only on the Qlik engine. Open access at the input layer, capture at the dependence layer. Standard plug, captive socket. **Embedded Model Layer (Bedrock/Claude, Qlik-Selected)** [DAPM: Ceded] Generation and embeddings are Amazon Bedrock-mediated (Anthropic Claude), selected by Qlik with no customer model choice and no BYO endpoint. The generative judgment is inherited two vendors deep. Proprietary integration — no customer authority at either link. ### NVIDIA-Provided Components **No NVIDIA Layer 1B Dependency** Retrieval runs on Qlik-operated OpenSearch; embeddings and generation are Amazon Bedrock-mediated. Whatever silicon serves them belongs to AWS's row. ### Gap Analysis Qlik Answers (GA September 2024) is a turnkey RAG product: point it at unstructured sources, Qlik indexes them into knowledge bases, and assistants answer with citations — embeddable in your own apps via qlik-embed, and queryable by third-party assistants through the MCP Server (GA February 2026). On the structured side, the associative engine itself is the context provider: the agentic experience and MCP tools retrieve from Qlik apps and data products, and the engine's associative model — every value linked to every related value — is genuinely distinctive context for multi-step questioning. Its singular contribution is the exclusion primitive: 'what is NOT associated' is a native engine state, a signal neither SQL joins nor vector similarity provide natively. Zero retrieval infrastructure to build. The grading line is not the managed abstraction — every managed offering on this map carries that cost. It is that Qlik's retrieval is a product feature, not an infrastructure service. Amazon Bedrock Knowledge Bases is equally managed, yet hands the customer the infrastructure interface: embedding model choice, index parameters, a retrieve API to build against. Qlik hands over none of that: no index management, no embedding or model choice (there is no BYO-model path for Answers), no retrieval API the enterprise's own AI applications can consume independent of the assistant abstraction. The entire stack is a closed box — Qlik-operated OpenSearch the customer never sees, an undocumented embedding model, and generation fixed to Amazon Bedrock (Claude), chosen by Qlik. A double-stacked cession: the retrieval opinions are Qlik's, and the reasoning core beneath them is AWS/Anthropic's, with the customer holding authority over neither. Calibration: Databricks, Palantir, and VAST are strong at 1B for general-purpose, governed retrieval the enterprise builds on — GA vector search APIs, permission-inheriting retrieval. Dell, VMware, and Nutanix are moderate with Delegated retrieval (Elastic, pgvector packaging). Qlik lands moderate from the opposite direction: real, GA, product-complete retrieval — but retrieval-as-product-feature rather than retrieval-as-infrastructure. Same altitude as the on-prem moderates, opposite DAPM texture. ### Borrowed Judgment High, and double-stacked. The retrieval opinions (indexes, chunking, citation logic) are Ceded to Qlik; Qlik in turn borrows the generative core from AWS/Anthropic with no customer surface at either link. Databricks gives you model choice on top of its captive retrieval; Palantir makes the model explicitly swappable. Qlik gives you neither the retrieval layer nor the model menu. ### Working Notes OpenSearch-as-vector-store comes from Qlik's evaluation-guide architecture docs (reasonably firm); the specific embedding model is undocumented. No GA bring-your-own-model path for Qlik Answers exists (the load-script LLM connectors are analytics sources, a different surface). Knowledge-base access control is Qlik-tenant-level (spaces, assistants), not source-ACL passthrough of the VAST/Databricks kind. ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Qlik's Heartland — CDC + Talend + Lakehouse Ingestion ### Vendor-Provided Components **CDC Replication (Qlik Replicate / QTC Ingestion)** [DAPM: Ceded] Log-based change data capture across hundreds of source/target pairs; long-GA, client-managed or as the managed replication engine in Qlik Talend Cloud. Task definitions, endpoint configs, and CDC opinions are proprietary; competitors exist but nothing lifts without rebuilding. Proprietary Qlik platform — opinions captive, no open exit. **Transformations & Pipeline Framework (Talend Heritage)** [DAPM: Ceded] No-code-to-pro-code ETL/ELT in Qlik Talend Cloud and client-managed Talend Studio. Proprietary framework; the open-source edition (Talend Open Studio) was discontinued January 2024. SQL-based transformation logic is partially portable as SQL, but the pipeline layer around it does not lift. Proprietary Qlik platform — opinions captive, no open exit. **Open Lakehouse Ingestion (Upsolver Engine)** [DAPM: Ceded] High-throughput ingestion into Iceberg with continuous optimization. Proprietary managed ingestion whose output is standard Iceberg — the same low-blast-radius Ceded as the 1A optimizer: leaving means swapping engines, not rebuilding tables. **Agentic Data Engineering (Data Product Agent, Data Quality Agent)** [DAPM: Ceded] GA June 30, 2026. Natural-language creation and evolution of pipelines and data products. The agents and what they generate are Qlik-captive. Proprietary Qlik platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 1C Dependency** Pipelines run on CPU compute; no acceleration dependency at the movement layer. ### Gap Analysis This is the layer Qlik built the company on — twice, by acquisition. The Attunity lineage gives it Qlik Replicate: log-based change data capture across hundreds of source/target pairs, arguably the category-defining CDC product, long-GA and massively deployed (client-managed or as the replication engine inside Qlik Talend Cloud). The Talend lineage gives it mature ETL/ELT: no-code-to-pro-code transformations, inline data quality, end-to-end lineage. The Upsolver lineage gives it the modern edge: high-throughput ingestion into Iceberg with continuous optimization. And as of June 30, 2026, agentic data engineering is GA — natural-language creation of pipelines and data products. The buyer gets any-source-to-any-target movement with quality and lineage threaded through, without owning a single pipeline server. Qlik pipelines are structurally pass-through — the destination is always somebody else's system (your Snowflake, your Databricks, your S3/Iceberg). So the data always lands open, which makes the capture read as light. It is not: the accumulated opinions are hundreds of CDC task definitions, endpoint configurations, transformation logic, and quality rules, all in Qlik's proprietary frameworks. The open-source escape hatch is gone — Talend Open Studio was discontinued in January 2024; there is no OSS edition to fall back to. Leaving means rebuilding the movement layer on Fivetran, Debezium, or dbt — a standard but real migration whose weight scales with task count. Calibration: Databricks 1C is strong and the most mature data-engineering layer assessed; VAST is strong (DataEngine); Palantir and Dell are moderate. Qlik belongs in the strong cohort without stretch — the one layer where its maturity genuinely matches Databricks': Databricks leads on in-platform transformation (Spark, Photon), Qlik leads on heterogeneous movement (CDC breadth across other people's systems). Different centers, same altitude. What keeps the grade honest in 2026 rather than a 2019 reputation call is the GA record: Open Lakehouse shipped, agentic data engineering shipped, and CDC remains the widest-deployed piece of the whole Qlik estate. ### Borrowed Judgment Low — the movement IP is Qlik's own (Attunity, Talend, and Upsolver acquisitions, all Qlik-owned). The enterprise inherits Qlik's judgment about replication, transformation, and ingestion, and cannot take the accumulated task definitions elsewhere. Decoupled pattern: the data lands open in the destination; the pipeline opinions stay captive in Qlik. ### Working Notes Watch-list (dated): Open Lakehouse streaming ingestion and streaming transformations were announced (late 2025) with planned GA Q1 2026; delivery not yet confirmed in product docs — batch CDC/ingestion is long-GA and carries the cell regardless. Client-managed Talend Data Fabric believed still shipping in 2026 (minor, unconfirmed). ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Managed Task & Cluster Orchestration; No GPU Plane ### Vendor-Provided Components **Qlik Cloud Managed Compute (SaaS Task/Reload Orchestration)** [DAPM: Ceded] Fully managed reload scheduling, pipeline task orchestration, and capacity management under a consumption model. Proprietary scheduling and capacity opinions, no customer surface, no lift. Proprietary Qlik platform — opinions captive, no open exit. **Open Lakehouse Cluster Orchestration (EC2 Spot in Customer VPC)** [DAPM: Ceded] Qlik provisions, scales, and cost-optimizes lakehouse clusters in the customer's own AWS account; the orchestration opinions are Qlik's even though the substrate and the bill are the customer's. Running in your VPC does not make it yours — self-deployable is not Retained, applied to topology. ### NVIDIA-Provided Components **No GPU Plane to Depend On** Layer 2A is where NVIDIA dependency concentrates for most of the map (Run:ai, GPU scheduling) — and Qlik has no GPU plane to be dependent in. The AI-compute authority question is wholly delegated to AWS/Bedrock, one row over. ### Gap Analysis There is nothing to operate, and that is the offer. Qlik Cloud is fully managed SaaS — reload scheduling, task orchestration, and capacity management happen invisibly under a consumption model. Open Lakehouse extends this into the customer's own estate: Qlik auto-provisions and autoscales lakehouse clusters on EC2 Spot instances inside the customer's VPC, choosing spare-capacity instances to keep the compute bill down. Pipeline tasks, engine reloads, cluster lifecycle: zero orchestration burden. Client-managed Qlik Sense shops carry their own ops on their own Windows servers — retained responsibility, not a Qlik capability. Two concerns. First, the standard one: this is proprietary orchestration the customer consumes without authority — no quotas surface, no fair-share policy, no infrastructure-as-code layer like Databricks' Asset Bundles. Second, the distinctive one: Open Lakehouse orchestration spends the customer's money under Qlik's judgment. The clusters run in your VPC on your EC2 bill, but the scaling and Spot strategy are Qlik's opinions — cost authority exercised on your account with, as far as the docs show, minimal knobs. And there is no GPU plane at all: no GPU scheduling, no accelerator awareness anywhere. Qlik has outsourced the entire AI-compute question to Amazon Bedrock, so the layer the 4+1 model watches most closely here simply has no Qlik surface. Calibration: Palantir 2A is moderate (Rubix — real, platform-scoped, Ceded orchestration the customer never configures), and Databricks 2A is moderate (owns workload orchestration, GPU pre-GA). Qlik sits in the same cohort by the same logic — orchestration exists, the vendor holds it, it is scoped to the vendor's own workloads — but at the thin end: Databricks has serverless compute on three clouds plus an IaC surface; Qlik has task scheduling and one cluster type on one cloud. What keeps this from being a gap is the real-dependence guardrail: the VPC lakehouse clusters run in your account, on your bill. Present-and-Ceded, per the cloud and Palantir convention. ### Borrowed Judgment Moderate. The orchestration opinions (scheduling, scaling, Spot strategy) are Qlik's and Ceded; the capacity underneath is AWS's. The enterprise inherits both without a configuration surface — and in the lakehouse case, Qlik's operational judgment is exercised directly against the customer's own EC2 spend. ### Working Notes Vendor follow-up (does not change the rating; important cost and operational question): can the customer pin instance types, cap scale, or opt out of Spot for lakehouse clusters, and what visibility exists into Qlik's scaling decisions against the customer's EC2 spend? Docs show auto-provisioning with a default single Spot instance and little else. ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Turnkey Agents & Automation; No Model Serving ### Vendor-Provided Components **Qlik Predict (No-Code ML Train/Deploy/Predict)** [DAPM: Ceded] Long-GA (formerly Qlik AutoML). Trains, deploys, and serves classical ML predictions inside Qlik Cloud. The model artifact is not where the dependence lives; the AutoML pipeline around it is — features, training, and serving are all platform-bound. Proprietary Qlik platform — opinions captive, no open exit. **Turnkey Agents (Discovery, Predict Agent, Automate Agent)** [DAPM: Ceded] GA Feb–June 2026. Qlik-built, Qlik-hosted product agents — configurable, not constructible. No framework for customer-built agents exists. Proprietary Qlik platform — opinions captive, no open exit. **Qlik Automate (Workflow Execution + Connectors)** [DAPM: Ceded] Long-GA (formerly Application Automation). Workflow automation across a large connector library, triggerable by the Automate Agent. The accumulated automation library is the real capture surface in this cell — hundreds of workflows that rebuild from scratch anywhere else. Proprietary Qlik platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No NVIDIA Runtime Dependency** There is no inference infrastructure to be NVIDIA-dependent in; generation is outsourced to Amazon Bedrock wholesale. Unlike Dell — where NVIDIA owns the 2B runtime — Qlik's runtime absence points at AWS, not NVIDIA. ### Gap Analysis Everything executes as a product feature, nothing as infrastructure. Qlik Predict (the former AutoML, long-GA) trains and deploys classical ML models no-code inside Qlik Cloud. The turnkey agents are GA and real: Discovery Agent monitoring for anomalies since February 2026, Predict Agent and Automate Agent since June 2026. Qlik Automate (the former Application Automation, long-GA) is the execution muscle — workflow automation with a large connector library, now triggerable by the Automate Agent, so insight-to-action genuinely executes end-to-end. Qlik Answers assistants are customer-configured (scoped to knowledge bases) and embeddable. For an analytics buyer, that is a complete loop: detect, predict, act — shipped, and among the non-hyperscaler rows arguably the most complete insight-to-execution story on the map. What Qlik does not ship is a runtime in the infrastructure sense. No model serving: you cannot serve an LLM or bring a model endpoint — generation is fixed Bedrock/Claude, per the 1B finding. No agent framework: you configure Qlik's agents; you cannot build and host your own on Qlik. The MCP Server points the other way — it makes Qlik a tool for agents whose runtime lives elsewhere (Claude, Copilot), an honest architectural admission that Qlik expects the agent runtime to be someone else's layer. What the enterprise accumulates here — Automate workflows, Predict models, agent configurations — is all captive, and none of it constitutes a runtime opinion that could ever move. Calibration: the strong bar at 2B is a general runtime — Databricks (any-model serving plus GA agent framework plus training) and Palantir (governed any-model agent runtime). Qlik clears none of that. The moderate cohort is Dell (runtime is NVIDIA's), Nutanix (agents in preview), VMware (foundational), VAST. Qlik fits moderate from its own angle: its agents are GA (ahead of Nutanix's preview) but they are product features, not a runtime; and the model-serving half of the layer is not offered at all, by design. Real shipped dependence keeps it out of gap; the missing runtime keeps it well out of strong. ### Borrowed Judgment High for generative execution, and double-stacked: the agents' reasoning core is Anthropic-via-AWS, chosen by Qlik, invisible to the customer. Classical ML (Qlik Predict) is Qlik IP, Ceded to Qlik. The enterprise accumulates real captive artifacts here — automation libraries especially — while holding no runtime authority at any link. ### Working Notes Watch-list (dated): Analytics Agent — planned Q3 2026, not scored. Predict model artifacts are not a portability question worth chasing: even if export existed, the feature engineering, training pipeline, and serving context around a model are all Ceded — a portable artifact with captive surroundings is the data-portability decoy one layer up. ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Bring Your Own Reasoning Plane (MCP-First) ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** No reasoning-plane product exists to carry a dependency. The zero-NVIDIA column completes here. ### Gap Analysis Qlik's honest answer to the reasoning plane is bring your own. The MCP Server (GA February 2026) is the clearest statement of posture in the portfolio: it makes Qlik a well-behaved, governed tool inside somebody else's agentic infrastructure — Claude, Copilot, whatever assistant platform the enterprise runs. Inside Qlik's own walls, the agentic experience coordinates the product agents, agents inherit the requesting user's permissions, and the Automate Agent prompts for confirmation before executing actions. For the buyer, agentic capability arrives without any reasoning-plane infrastructure to stand up. None of that is a reasoning plane. Applying the test: Intelligence-2C — which agent may act, on what, under which policy — exists only as inherited platform permissions plus hardcoded confirmation prompts. There is no productized surface: no agent registry treating agents as governed assets, no gateway product with rate limits, audit, and fallbacks, no policy engine an admin configures, no supervisor for agents the customer builds (there are no customer-built agents — the 2B finding). Infrastructure-2C — live placement of inference across model, cost, and compliance tiers — is missing twice over: with a single fixed Bedrock/Claude path there is not even a routing decision to govern. The live-placement gap is universal across the instrument, noted rather than penalized; the productized-governance gap is Qlik-specific. The enterprise that deploys Qlik agents alongside Copilot agents and homegrown agents owns the cross-system reasoning plane entirely, with Qlik participating as one MCP tool among many. Calibration: Databricks earned moderate with three productized, GA governance components — Unity Catalog agent governance with On-Behalf-Of auth, the Mosaic AI Gateway, and the Supervisor Agent — and the moderate cohort (AWS, IBM, OCI, Cisco, HPE) all have multi-component productized Intelligence-2C. Qlik has product plumbing, not products: permission inheritance and confirmation dialogs are properties of the agents scored at 2B, not a distinct policy plane. That is the Nutanix/VMware/Dell/VAST cohort — gap. Scoring implicit permission inheritance as moderate would credit at 2C what every SaaS platform does by default. The posture deserves naming without softening the score: MCP-first is arguably more architecturally honest than a thin bolt-on gateway would be. Qlik is betting the reasoning plane belongs to the enterprise — the 'bring any Layer 2C' invitation Dell makes at the bottom of the stack, made here one layer up, with the same consequence: the function stays the enterprise's to build, and the authority stays Retained by default. ### Borrowed Judgment None to borrow — and that is the finding. To the extent a reasoning plane exists in a Qlik deployment, it is external: the enterprise's assistant platform reaching in through MCP, or nothing. The enterprise owns the function by default because Qlik has not claimed it. ### Working Notes Vendor follow-up: does any GA administrative agent-governance surface exist (per-agent entitlements, agent audit console, tenant-level agent policy) beyond inherited user permissions? Product docs show none; docs are the primary source. If such a surface ships GA, the cell re-scores to thin moderate per the Databricks calibration. Watch-list (dated): Analytics Agent — planned Q3 2026. Sub-threshold signals, prose only: MCP OAuth, user-permission inheritance, human-in-the-loop confirmation prompts. ## ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Qlik's Native Layer — The Analytics Value Plane ### Vendor-Provided Components **Analytics Applications & Associative Exploration** [DAPM: Ceded] Qlik Sense apps, dashboards, set analysis, and Insight Advisor — decades-mature, first-party, GA. The engine, the expression language, and every app built in them are proprietary; the accumulated set-analysis logic is among the least portable artifacts in the entire instrument. Proprietary Qlik platform — opinions captive, no open exit. **Agentic Analytics Experience (Detect → Predict → Act)** [DAPM: Ceded] GA Feb–June 2026. The application-level loop over the 2B runtime: Discovery Agent detects, Predict forecasts, Automate Agent executes into downstream systems. Qlik-shaped, Qlik-bound. Proprietary Qlik platform — opinions captive, no open exit. **Embedded & OEM Analytics (qlik-embed)** [DAPM: Ceded] GA. Embeds Qlik apps and Answers assistants into customers' own products. The embedding library is Qlik's and the embedded content is Qlik apps; OEM customers inherit the same captivity one product-tier removed. Proprietary Qlik platform — opinions captive, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 3 Dependency** The value plane is Qlik IP with Bedrock-mediated generation. The NVIDIA column ends the row empty end-to-end — the cleanest zero-NVIDIA row in the series; Qlik's AI dependence chain runs through AWS and Anthropic instead. ### Gap Analysis This is what anyone buys Qlik for, and it is the deepest first-party value plane among the non-hyperscaler rows after Palantir. Qlik Cloud Analytics / Qlik Sense is a complete analytics application suite — apps, dashboards, the associative exploration experience where the exclusion primitive lives (what is not related, as a native engine state), Insight Advisor's natural-language insights — decades mature with an enormous deployed base. Qlik Answers puts a conversational application over unstructured knowledge. And the agentic loop closes end-to-end in GA product: Discovery Agent detects, Qlik Predict and the Predict Agent forecast, the Automate Agent executes into downstream business systems. Detect, predict, act — shipped; for the analytics domain, one of the most complete insight-to-action stories on the map. Embedded and OEM analytics (qlik-embed) extend the same value plane into customers' own products. Two concerns. First, the domain boundary: this value plane is analytics and decision workflows, not operational business applications — Palantir's Foundry apps run supply chains; Qlik answers questions about data and automates what follows. Completeness within a bounded domain is the same caveat Databricks carries, and it holds here. Second, this cell is where the row's oldest and heaviest capture sits, and it is the coupled, visible kind: enterprise Qlik estates hold years of apps, data models, and set-analysis expressions — a proprietary expression language whose accumulated logic lifts to no other platform — on top of the QVD estates scored at 1A. The row's shape inverts Databricks': Qlik's newest surface (the lakehouse) is its most open, and its oldest surface (the app estate) is its most captive. Calibration: Databricks Layer 3 is strong (first-party Genie, AI/BI, and Apps, with the analytics-centric caveat honestly stated) — Qlik matches that shape with a deeper and older first-party application estate, same caveat applying. Palantir is strong and broader (operational). VMware and Nutanix are moderate (platform-enabled, not platform-provided) — Qlik ships the applications, clearing that line easily. Not partner (Dell, Cisco, HPE, IBM): nothing here is ISV-delivered. ### Borrowed Judgment The application opinions are Qlik's, Ceded to Qlik — and the enterprise's accumulated artifacts (apps, data models, set-analysis expressions, automations) are the largest captive estate in the row, decades deep for legacy shops. The generative layer inside these applications inherits the AWS/Anthropic chain from 1B/2B: Qlik chooses, AWS serves, Anthropic reasons. ### Working Notes Everything scored here is long-GA and doc-confirmed. Qlik Answers' application surface rides on the 1B components — referenced, not double-counted. Watch-list items (Analytics Agent, Q3 2026) live at 2B/2C. ════════════════════════════════════════════════════════════════════════════════ # Salesforce (Agentforce 360 + Data 360 + Customer 360) Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.0 - Initial Assessment **Date:** July 17, 2026 **Source:** Salesforce Developer docs (Agentforce Trust Layer, Supported Models, Data 360 architecture and security, search index and retriever guides), MuleSoft Agent Fabric documentation, Salesforce Architect fundamentals (Data 360, MuleSoft Agent Fabric deep dive), Salesforce and Informatica press releases (Informatica acquisition close Nov 18 2025; 'Informatica from Salesforce' data-foundation announcement May 20 2026), Agentforce 360 GA (Spring '26, Feb 23 2026), Agentforce Observability GA (Feb 2026), AWS Big Data blog (zero-copy access to Iceberg from Data 360 via Glue REST catalog), Salesforce FY26 metrics (Q3/Q4 Data Cloud and Agentforce), published 4+1 model. Informatica folded into this row on the whose-paper rule (wholly owned since Nov 2025). Determinism distinction ('you can't prompt your way to deterministic output' / 'legal is not right') written at 2C as a universal finding and adopted as a 2C scoring-discipline rule (ratified Salesforce v1.0, sibling of 'routing is not reasoning'), applied uniformly rather than as a Salesforce-only downgrade. ## Summary Finding Salesforce is an enterprise application, agent, and data platform that floats above Layer 0 and is strong everywhere it has always lived: the data plane of storage and governance, retrieval, and movement, plus the agent runtime and the value plane. It moderates at the two layers it does not truly own. Infrastructure orchestration, where it manages its own workloads on rented cloud. And the reasoning plane, where it governs and enforces but does not reason. Authority sits in Salesforce's opinion layer wherever Salesforce ships one, which is nearly everywhere above the silicon. Salesforce is the one row on the instrument that runs both capture mechanisms at full strength, and they point in opposite directions. The value plane is the deepest coupled, visible capture in enterprise software: your business runs in Salesforce's namespace, decades deep, and no one is confused about the commitment. Zero-copy federation is the decoupled, invisible ask laid over it. Your bytes stay in Snowflake, the openness is real, and the opinions you accumulate on top are entirely Salesforce's. The coupled pole is honest. The decoupled pole underprices the commitment, because the data portability is a decoy and the dependence lives one layer up. The platform's answer to control is deterministic guardrails, not deterministic validation. Validation rules, before-save Apex, and approval processes let a human programmatically bound what an agent may commit, and that commit-boundary machinery is more than most platforms give you. It checks whether the result is legal, not whether it is right. The controls Salesforce offers for rightness are prompt-shaped: grounding, instructions, Agent Script, the reflection step. You cannot prompt your way to deterministic output. So rightness stays probabilistic, and the deterministic validator is the seam the enterprise builds itself, on a reasoning core it cannot inspect. The reasoning plane is real governance without a reasoning mechanism. MuleSoft Agent Fabric governs agents across platforms, including agents built outside Salesforce, and the sharing model enforces agent access live. All of it enforces policy the enterprise authored statically. None of it places inference or validates an outcome dynamically. The Informatica acquisition, closed November 2025, gives Salesforce arguably the richest data-provenance telemetry on the instrument: lineage, catalog, quality, and mastered entities. There is no 2C engine that consumes it to decide anything. The vendor with the best inputs to a reasoning plane still ships no plane that reasons over them. That is the Dell finding restated, and Salesforce makes it sharper. Of 31 scored components, 29 are Ceded, two are Delegated, and none are Retained. That is a more captive profile than Databricks, because Salesforce has fewer open substrates to stand on. The two open surfaces are the Iceberg federation interface and the AppExchange ecosystem. The buyer gets the most complete customer-engagement AI stack in enterprise software, agents that execute real business work under mature commit-boundary governance, and near-zero dependence on NVIDIA. In exchange they cede nearly every opinion layer Salesforce ships: the apps, the data model, the governance, the pipelines, the retrieval, the runtime, and the agents. What stays the enterprise's is the reasoning plane Salesforce declined to close. It hands you the inputs for deterministic validation and live placement, and leaves you to build both. ## ○ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Not Salesforce's Layer (By Design) ### NVIDIA-Provided Components **No NVIDIA Surface at Layer 0** Salesforce designs no silicon, networking, or acceleration fabric, and consumes no NVIDIA platform here. The AI-compute question is delegated wholly to the hyperscalers and to model providers through the LLM Gateway. The near-zero NVIDIA column starts here and holds across the row. ### Gap Analysis Layer 0 is not Salesforce's layer, and that is the pitch: the buyer never thinks about silicon. Hyperforce delivers the platform as infrastructure-as-code onto AWS, Google Cloud, Azure, and Alibaba, each region across three or more availability zones, and the buyer's only real Layer 0 decision is data residency. Agentforce 360 arrives with no infrastructure to stand up. Salesforce still operates first-party data centers for non-Hyperforce orgs, and it genuinely runs infrastructure, but that does not move the cell. The exposure test (ratified CoreWeave 1C vs. Supermicro 1C) governs: an underlay capability the vendor does not surface as a purchasable, customer-administered product is not that vendor's capability. Salesforce's compute sits invisibly behind a SaaS interface with no customer-administered surface, the same reasoning that scores a cloud's invisible internal engine as gap at the layer it silently occupies. The consequence for the 4+1 model matches Databricks, Palantir, and Qlik: a Salesforce adoption decision resolves no Layer 0 authority question. Whatever capture exists at silicon and fabric belongs to the hyperscaler underneath, a different row on this map. ### Borrowed Judgment Total at Layer 0, and irrelevant to the value proposition by design. Salesforce inherits all silicon, networking, and acceleration judgment from the host environment. The enterprise's Layer 0 authority position is set by a vendor it did not choose in the Salesforce transaction, and the docs do not hand it that choice. ### Working Notes Heroku is Salesforce-owned and sells customer-administered compute (dynos) plus Heroku AI (Managed Inference and Agents, GA), but it does not move this cell. Heroku entered a 'sustaining engineering' mode in February 2026 (stability only, no new features, no new Enterprise Account contracts), and its inference product is PaaS on rented cloud IaaS, not silicon or fabric. It stays off Layer 0 for the same reason Databricks' classic-clusters-in-your-VPC do. Nothing here is scored. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Authorization Authority, Dual Capture ### Vendor-Provided Components **Core Object Model & Sharing Model** [DAPM: Ceded] The CRM data model plus object-, field-, and record-level security, sharing rules, permission sets, and profiles. The deepest coupled capture surface in the row: decades of accumulated sharing logic that lifts nowhere. Proprietary Salesforce platform, no open exit. **Data 360 (DMOs, Harmonization, Data Spaces)** [DAPM: Ceded] Unifies customer data into a proprietary data model with harmonization and data-space governance boundaries. The harmonization opinions do not lift. Proprietary Salesforce platform, no open exit. **Data 360 Governance (Fine-Grained Policies + Dynamic Masking)** [DAPM: Ceded] The authorization authority itself: field/object/record policies and dynamic masking embedded in the semantic layer, enforced at runtime across Agentforce, analytics, and segmentation. Policies are authored in Salesforce's engine and portable nowhere. Proprietary Salesforce platform, no open exit. **Zero Copy Federation (Iceberg REST → Snowflake / Databricks / BigQuery / Redshift / S3)** [DAPM: Delegated] Federates external Iceberg tables through the AWS Glue Iceberg REST endpoint and equivalent catalogs, reading Parquet directly via the Hyper engine without moving data. The consumed interface is a genuine multi-vendor standard and the tables stay in the customer's own account, so the data-layer opinions lift. Delegated, matching Qlik Open Lakehouse and Databricks open lakehouse storage. **Informatica (MDM + Data Catalog + Data Lineage)** [DAPM: Ceded] Master data management, catalog, and end-to-end lineage, now sold as 'Informatica from Salesforce' and integrated with Data 360. General-purpose data management credited at full weight on the whose-paper rule. Master data models and lineage graphs are proprietary and do not lift. Proprietary platform, now first-party, no open exit. **Salesforce Shield (Platform Encryption, Event Monitoring, Field Audit Trail)** [DAPM: Ceded] Security implementation captive to the platform: encryption schemes, monitoring, and audit configurations do not lift to another vendor's storage. Matches Dell's built-in cyber resilience. Proprietary Salesforce platform, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 1A Dependency** The object model, Data 360 governance, zero-copy federation, and the Informatica catalog and mastering stack are Salesforce IP or open-standard interfaces. NVIDIA contributes nothing to the governance layer. ### Gap Analysis This is the layer Salesforce has been rebuilding the company around. The buyer gets the CRM object model they already run their business on, Data 360 unifying customer data across the estate, and zero-copy federation that reaches into Snowflake, Databricks, BigQuery, Redshift, and generic Iceberg catalogs without moving a byte. Q3 FY26: 32 trillion records ingested in a quarter, 15 trillion through zero-copy connectors, up 341% year over year, with federation at roughly a 28x cost reduction against batch. On top sits Informatica's master data management, catalog, and lineage. The deciding line is the one that separated Databricks and Palantir (strong) from Qlik (moderate): is the governance an authorization authority or merely curation and quality metadata? Qlik's was curation, with access enforcement delegated to AWS IAM, and that held it at moderate. Salesforce is unambiguously an authorization authority. Data 360 Governance authors fine-grained field-, object-, and record-level policies enforced at runtime, consistently across Agentforce, analytics, and segmentation, with dynamic masking embedded in the semantic definitions. That is Unity Catalog's structural property and the Ontology's, reached through the most battle-tested sharing model in enterprise software. Rule 4 (general vs. fixed-function) is the honest pressure point. On its own, Data 360 is CRM-shaped. Informatica is what dissolves the concern: general-purpose data management across products, finance, and suppliers, hybrid, multicloud, and on-prem, credited in full on the whose-paper rule. This cell is strong because Salesforce bought its way to data-plane generality in November 2025; without Informatica it would read moderate. Frontier check (rule 6): runtime authorization authority extending to agents, genuine multi-vendor open federation at scale, and a top-tier catalog, lineage, and mastering stack stands with Databricks and Palantir at the current frontier. ### Borrowed Judgment Low for the governance logic, which is Salesforce IP and Ceded to Salesforce. The zero-copy federation interface is Delegated: a genuine multi-vendor Iceberg REST standard with the tables in the customer's own account. The row's signature is that both capture mechanisms run at once here. The CRM core is coupled and visible; zero-copy is decoupled and invisible, the purest expression of 'data portability is a decoy' on the instrument, because the bytes stay in Snowflake while the access policies, harmonization, and agent grounding accumulate entirely in Salesforce. ### Working Notes Whose-paper rule folds Informatica (wholly owned since Nov 18 2025) into this row at full credit. Watch-list (not scored): Agentic Multidomain MDM ('industry's first,' Informatica World, May 2026) reads as announcement language; the Data 360 Connector and Scanner for real-time bidirectional flow with end-to-end lineage (announced May 20 2026) is GA-unconfirmed in docs. Fact question pending: how deeply Informatica is deployed as an integrated Salesforce capability today versus a co-owned but separate platform. Does not change the whose-paper credit; changes only how the cell narrates. Instrument follow-up (Keith's ruling): whether Informatica warrants a separate standalone row. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Governed Retrieval — Native Infrastructure Surface ### Vendor-Provided Components **Data 360 Vector Database + Hybrid Search (VDMO)** [DAPM: Ceded] GA. Vector and hybrid (semantic + keyword) search over chunked, embedded content in a Vector Data Model Object, permission-aware by construction. The index format, chunking, and retrieval engine are proprietary; the permission inheritance is the part that does not rebuild cheaply. Proprietary Salesforce platform, no open exit. **Custom Retrievers** [DAPM: Ceded] Customer-defined retrievers that scope how, where, and what an agent pulls, each bound to one search index. Retriever definitions are Salesforce-specific and run only here. Proprietary Salesforce platform, no open exit. **Intelligent Context** [DAPM: Ceded] GA February 2026. A workspace to refine unstructured-data extraction with prompt-based instructions before building an index. The tuning opinions are captive. Proprietary Salesforce platform, no open exit. **Atlas Advanced RAG (Query Refinement + Self-Assessed Quality)** [DAPM: Ceded] Query expansion, advanced retrieval, and self-assessment of response quality inside the reasoning engine. Proprietary retrieval reasoning; the planner and execution half of Atlas is scored at 2B, not double-counted. Proprietary Salesforce platform, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 1B Dependency** The vector database, hybrid search, custom retrievers, and Atlas retrieval reasoning are Salesforce IP. Any GPU serving the embedding models belongs to a provider's row, not Salesforce's. ### Gap Analysis RAG with no retrieval infrastructure to build. Data 360's vector database is GA, with hybrid indexes blending semantic and keyword retrieval into a Vector Data Model Object. Customers tune parsing and chunking to their content formats, select an embedding model sized to their chunk length, and define custom retrievers that scope exactly what an agent pulls. Intelligent Context (GA February 2026) adds a workspace to refine unstructured extraction with natural-language instructions before an index is committed. Atlas performs query refinement and advanced retrieval, then assesses the quality of its own response. The Qlik moderate was retrieval-as-product-feature rather than retrieval-as-infrastructure, decided on three absences: no index management, no embedding choice, no retrieval API the enterprise's own applications can consume. Salesforce has all three, which puts it in the Databricks shape: GA native vector plus hybrid search, configurable indexes, managed embedding selection, and permission-aware retrieval by construction rather than by bolt-on filter. Permission-inheriting retrieval is the structural property that earned Databricks, Palantir, and VAST strong at this layer; Salesforce reaches it through the sharing model. Frontier check (rule 6): GA native retrieval, structural permission inheritance, a configurable infrastructure surface, and retrieval-quality observability through Agentforce Observability stands with the 1B frontier and well above the Dell/VMware/Nutanix moderate cohort, whose retrieval is Delegated to Elastic or pgvector. ### Borrowed Judgment Low. Indexes, chunking strategy, the VDMO format, retriever definitions, and the embedding models are all Salesforce's and Ceded to Salesforce. The expensive-to-rebuild part is not the vectors but the permission inheritance: leaving means re-deriving decades of object-, field-, and record-level sharing semantics in a system that has never heard of them. ### Working Notes Embedding choice appears to be selection from Salesforce's own catalog (e.g., Salesforce Embedding V2 Small) rather than bring-your-own-embedding, a texture difference from Databricks that does not move the grade. The Atlas planner and execution half is scored at 2B, not double-counted here (Dell precedent: retrieval at 1B, movement at 1C). Fact question that could pull this toward moderate: can Data 360 retrieval target a third-party embedding model, and can non-Agentforce applications call retrievers through a documented API, or are retrievers only consumable from prompt templates and Agentforce? ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** MuleSoft + Informatica — Heterogeneous Movement ### Vendor-Provided Components **MuleSoft Anypoint Platform (Mule Runtime, DataWeave, Connectors)** [DAPM: Ceded] API-led integration, the Mule runtime, the DataWeave transformation language, and a large connector estate. DataWeave is proprietary and Mule flows lift nowhere; a community runtime exists but Anypoint, where the opinions live, is proprietary. Proprietary Salesforce platform, no open exit. **Informatica Cloud Data Integration + CDC + Mass Ingestion** [DAPM: Ceded] The incumbent enterprise ETL/ELT, change data capture, and mass ingestion stack, now first-party. Mappings and CDC task definitions are captive; competitors exist but nothing lifts without rebuilding. Proprietary platform, no open exit. **Data 360 Ingestion (Data Streams, Batch + Streaming Connectors)** [DAPM: Ceded] Native ingestion into the Data 360 model across batch and streaming sources. Proprietary managed ingestion into a proprietary data model. Proprietary Salesforce platform, no open exit. **MuleSoft Flex Gateway / API Manager** [DAPM: Ceded] API policy management and gateway enforcement across the integration estate. Policies and gateway configuration are Salesforce-specific. Named here for the integration function; its agent-governance evolution (the Omni Gateway) is scored at 2C, not double-counted. Proprietary Salesforce platform, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 1C Dependency** Pipelines run on CPU compute across MuleSoft and Informatica. No acceleration dependency at the movement layer. ### Gap Analysis Salesforce owns two category-leading movement platforms outright. MuleSoft Anypoint delivers API-led integration, the Mule runtime, DataWeave transformation, and a large connector estate, long-GA and deeply deployed. Informatica brings the incumbent enterprise ETL/ELT stack: Cloud Data Integration, mass ingestion, log-based change data capture, and the PowerCenter install base, running hybrid, multicloud, and on-prem. Data 360 adds its own batch and streaming ingestion, and the zero-copy path scored at 1A means the cheapest movement is often no movement at all. Qlik earned strong here on Attunity CDC plus Talend plus Upsolver, with the reasoning that different centers (in-platform transformation vs. heterogeneous movement) can sit at the same altitude. Salesforce occupies Qlik's exact center with more of it: Informatica is the incumbent Talend spent a decade competing against, and MuleSoft adds API-led integration Qlik has no answer to. Peer to Databricks and Qlik (strong), above Palantir and Dell (moderate). This is not momentum from the adjacent strongs; it is strong on either owned asset alone. The honest limit matches Palantir's 1C: the layer's purpose line includes KV cache tiering, and Salesforce has nothing there, because it does not operate at the storage-physics layer. Where Dell moves bytes between storage and the GPU cluster, Salesforce moves and governs records between systems that were never designed to talk. The general movement function is covered comprehensively; the storage-physics slice is simply not Salesforce's layer. ### Borrowed Judgment Low. The movement IP is Salesforce-owned (MuleSoft and Informatica acquisitions). Decoupled pattern: pipelines are structurally pass-through, so the data lands open in someone else's system, while the accumulated opinions (Mule flows, DataWeave, Informatica mappings and CDC task definitions, connector configs) stay captive in proprietary frameworks. For a large Informatica shop the lift to leave is a multi-year rebuild on Fivetran, Debezium, dbt, or Airflow, scaling with task count. ### Working Notes Watch-list (not scored): the Data 360 Connector and Scanner for real-time bidirectional flow (announced May 20 2026), GA-unconfirmed in docs. Informatica's IDMC framing bundles movement and governance; the split across 1C and 1A is the instrument's, following the Dell precedent that splits AIDP across 1B and 1C. Fact question: does the Mule community runtime give a customer real operational independence from Anypoint, or is the platform the only production path (read as the latter, hence Ceded rather than Delegated)? ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Managed Orchestration; No GPU Plane ### Vendor-Provided Components **Hyperforce + CloudHub 2.0 Managed Compute** [DAPM: Ceded] Infrastructure-as-code provisioning on public cloud, and fully managed Mule execution on isolated per-region EKS clusters. Invisible managed orchestration the customer never configures. Present-and-Ceded, per the cloud and Palantir convention. Proprietary Salesforce platform, no open exit. **Anypoint Runtime Fabric (Mule Control Plane on Customer Kubernetes)** [DAPM: Ceded] Deploys Mule runtimes onto the customer's own Kubernetes clusters under a Salesforce control plane. The Run:ai precedent applies: a proprietary control plane whose opinions live in Runtime Manager rather than portable Kubernetes manifests, Ceded even though it sits on Kubernetes. The customer's own cluster underneath is a different vendor's row, not Salesforce capability. Proprietary Salesforce platform, no open exit. **Governor Limits + Einstein Requests / Flex Credits** [DAPM: Ceded] Multi-tenant resource quotas and fair-share, plus AI-consumption metering. Resource-allocation authority held absolutely by the vendor and imposed on the customer, with no configuration surface. Proprietary Salesforce platform, no open exit. ### NVIDIA-Provided Components **No GPU Plane to Depend On** Layer 2A is where NVIDIA dependency concentrates for most of the map (Run:ai, GPU operators, scheduling), and Salesforce has no GPU plane to be dependent in. The AI-compute question is delegated wholly to model providers through the LLM Gateway. ### Gap Analysis There is nothing to operate, and that is the offer. Hyperforce provisions everything invisibly. CloudHub 2.0 runs Mule applications on isolated Amazon EKS clusters per region, managed entirely by Salesforce. For regulated workloads, Anypoint Runtime Fabric deploys Mule runtimes onto Kubernetes the customer creates and manages (EKS, AKS, GKE, OpenShift), isolating each application with automated failover and horizontal scaling. The buyer gets deployment orchestration across a hybrid estate without building a control plane. Calibration lands moderate, in the Qlik/Palantir/Databricks cohort by their exact logic: orchestration exists, the vendor holds it, and it is scoped to the vendor's own workloads. Not gap: the real-dependence guardrail decides it, the way it decided Qlik. Runtime Fabric runs Mule runtimes on the customer's own Kubernetes clusters, on the customer's bill, under a Salesforce control plane, which is more than Qlik's single Spot cluster type earned. Not strong: VMware and Nutanix are strong here because they orchestrate all infrastructure; Salesforce orchestrates Salesforce. Quotas and fair-share are in this layer's purpose line, and Salesforce enforces both through the most absolute exercise of 2A authority on the instrument: multi-tenant governor limits, imposed rather than configured, with the AI-era version metered as Einstein Requests and Agentforce Flex Credits. The buyer does not configure this. They consume it. ### Borrowed Judgment Moderate and Ceded, and the load-bearing consequence is traceability. The enterprise inherits an unconfigurable substrate, and because the infrastructure is a turnkey black box it cannot answer substrate-level decision provenance: which model instance, in which region, at what cost or compliance tier served a given inference. That is the Infrastructure-2C gap seen from below, and it is the honest cost of the turnkey posture. Governor limits and Flex Credits are the visible edge of that authority; the invisible part is that you cannot audit what you cannot configure. ### Working Notes Heroku is named with its gate (sustaining-engineering mode since February 2026, Enterprise contracts closed to new customers) and not scored; under rule 5 that gate keeps a real capability from carrying the cell. Fact question: does Salesforce expose any AI capacity-allocation surface beyond commercial metering (allocate Einstein Requests or Flex Credits across business units, per-agent spend caps, reserved capacity)? If yes, that is a thin quota surface worth scoring and it also touches 2C; if purely billing, the prose framing stands. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Atlas Runtime — Governed, Constructible, Any-Model ### Vendor-Provided Components **Atlas Reasoning Engine (Agent Runtime)** [DAPM: Ceded] GA. Proprietary planner, action selector, tool-execution engine, memory, and reflection loop. Agents built on it do not lift. Matches the Palantir AIP runtime. Proprietary Salesforce platform, no open exit. **Agentforce Builder + Agent Script** [DAPM: Ceded] GA (Spring '26). Conversational build-test-deploy plus a human-readable expression language for deterministic control flow and guided tool use. Agent definitions and scripts are Salesforce-specific. The determinism is real only where the script hands off to Flow or Apex; the model's job inside a step stays probabilistic. Proprietary Salesforce platform, no open exit. **Model Access (LLM Gateway / BYOLLM / LLM Open Connector)** [DAPM: Ceded] GA. Governed access to Salesforce Default (GPT-4o mix), AWS-hosted Claude, and bring-your-own models across Bedrock, Azure OpenAI, OpenAI, and Vertex, plus custom/open-source via the Open Connector. Model choice is portable; the gateway and Trust Layer integration are captive. Matches the Palantir Model Catalog (swappable model, captive governed access). Proprietary Salesforce platform, no open exit. **Flow + Apex Action Execution** [DAPM: Ceded] The deterministic execution substrate the agent invokes: Flows and Apex classes run the same way every time and enforce logic at the commit boundary. Real deterministic-code-in-the-loop, and proprietary at once: Apex runs only on Salesforce and Flows lift nowhere. Proprietary Salesforce platform, no open exit. **Einstein Trust Layer** [DAPM: Ceded] The governed inference path: masking, grounding, toxicity detection, and audit logging around every model call. A proprietary constraint layer that shapes behavior by instruction and policy, not a deterministic outcome validator. Proprietary Salesforce platform, no open exit. ### NVIDIA-Provided Components **NVIDIA Nemotron Models (Beta)** NVIDIA Nemotron (Nano/Super) appears in the supported-models list in Beta, the only NVIDIA surface in the entire row and an optional model choice rather than a structural dependency. Watch-list, not scored. The reasoning core is otherwise a frontier LLM served by OpenAI, Anthropic (via Bedrock), Google, or a bring-your-own provider through the LLM Gateway. ### Gap Analysis This is what the buyer came for in 2026, and it is a genuine first-party runtime, not a wrapper. The Atlas Reasoning Engine (GA) is Salesforce's own agent runtime: a planner that decomposes goals, an action selector, a tool-execution engine, a memory module, and a reflection step. Agents are constructed, not just configured: Agentforce Builder for conversational build-test-deploy, Agent Script for deterministic control flow, and pro-code Apex underneath. Model access is real and open: Salesforce Default (a managed GPT-4o mix), AWS-hosted Claude, and BYOLLM GA across Bedrock, Azure OpenAI, OpenAI, and Vertex, plus the LLM Open Connector for custom or open-source models. Every call routes through the Einstein Trust Layer. Agentforce Voice adds real-time execution. And the execution target is distinctive: agents invoke Flows and Apex, so the runtime executes deterministic business logic, not just API calls. The deciding precedent is Palantir, which went strong at 2B without model serving or foundation training. Palantir's strong rested on a governed, constructible, any-model agent runtime: the stateful loop, isolation, swappable model, evaluation, observability. Salesforce has the identical shape feature for feature, and adds a mature deterministic execution substrate (Flow/Apex) with commit-boundary enforcement. On the line that separated Palantir (strong) from Qlik (moderate), constructible governed runtime with swappable model versus configurable agents with a fixed model, Salesforce is unambiguously on the Palantir side. Frontier check (rule 6): peer to Palantir and Databricks. The one thinness is general model serving: Databricks sells OpenAI-compatible serving as infrastructure, while Salesforce's gateway feeds its own agents. That is why this is a clean strong and not a category-leading one. ### Borrowed Judgment Low for the runtime, governance, and execution machinery, all Salesforce IP and Ceded to Salesforce. The reasoning core is explicitly borrowed and explicitly swappable (BYOLLM), the Palantir pattern rather than the Qlik one. The honest addition, stated as a universal property rather than a Salesforce charge: the runtime governs how agents act and deterministically enforces what commits, but the correctness of the reasoning is inherited from a probabilistic model and is not deterministically validated. That gap is noted at 2C as universal, not penalized here. ### Working Notes Watch-list (Beta, not scored): NVIDIA Nemotron models. Fact question (rule 4, general vs. fixed-function): can Agentforce agents run meaningfully outside the Salesforce operating context as a general runtime (the Slack and Agentforce-in-ChatGPT surfaces suggest partly yes), or is the runtime effectively bound to acting on Salesforce objects? Read as general enough for strong given Apex, MCP, and MuleSoft tool reach; a hard boundary here is the one thing that could pull it to moderate. ## ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Governance Without a Reasoning Mechanism ### Vendor-Provided Components **MuleSoft Agent Fabric — Agent Governance (Omni / Flex Gateway)** [DAPM: Ceded] GA. Cross-platform policy enforcement at request time: auth, rate limits, token and cost caps, tool-access restrictions, and audit logging on every agent interaction, including agents built outside Salesforce. Live enforcement of statically authored policy, constraint not validation. Proprietary Salesforce platform, no open exit. **MuleSoft Agent Fabric — Agent Registry & Broker** [DAPM: Ceded] GA. A catalog of agents, tools, and MCP/A2A servers, plus a Broker that routes tasks across A2A-compliant agents. The Broker is task dispatch, flagged under routing-is-not-reasoning: it matches a task to a capability, it does not solve constraints over policy. Proprietary Salesforce platform, no open exit. **Agentforce Runtime Governance (Permission Inheritance + Commit-Boundary Enforcement)** [DAPM: Ceded] Agents inherit object-, field-, and record-level sharing at runtime, and validation rules, before-save Apex, and approval processes gate what commits. The deterministic guardrail on state, the 'legal' half. A facet of the 1A governance surface and the 2B Flow/Apex layer surfaced at the agent-action altitude, not double-counted. Proprietary Salesforce platform, no open exit. **Agentforce Observability (Command Center)** [DAPM: Ceded] GA February 2026. Agent analytics, step-by-step reasoning traceability, and health monitoring. Decision tracing, not decision validation: it reconstructs what an agent did, not whether the outcome was right. Proprietary Salesforce platform, no open exit. ### NVIDIA-Provided Components **No NVIDIA Layer 2C Dependency** The governance plane (Agent Fabric, the Trust Layer, Observability, runtime permission inheritance) is entirely Salesforce IP. NVIDIA controls no identity, governance, or routing here. ### Gap Analysis Salesforce ships one of the deepest agent-governance surfaces on the map, and MuleSoft Agent Fabric (GA) is the reason. It is a multi-component agent control plane: an Agent Registry cataloging agents, tools, and MCP/A2A servers; an Agent Broker routing tasks across A2A-compliant agents; Agent Governance (the Omni/Flex Gateway) enforcing auth, rate limits, token and cost caps, tool-access restrictions, and audit logging on every interaction; and an Agent Visualizer. Its distinctive property is cross-platform reach: it governs agents built outside Salesforce (Bedrock, Vertex, Copilot Studio), not just Agentforce agents. Add Agentforce Observability (GA February 2026), the Trust Layer, and Data 360's runtime permission inheritance, and the buyer gets a real 'govern every agent in the enterprise' surface. Applying the working notes' test and the determinism finding together: every piece of that surface is constraint, not validation, and enforcement, not reasoning. What is genuinely dynamic at this layer is policy enforcement at request time (the gateway), access evaluation at request time (permission inheritance), and task routing (the Broker). All three are live enforcement or dispatch of policy the enterprise authored statically. None places inference, and none validates an outcome. Routing is not reasoning: the Broker dispatches, it does not solve constraints. The universal finding, logged as a /reconcile candidate rather than applied as a Salesforce-only standard: you cannot prompt your way to deterministic output, so the platform's guardrails answer 'is it legal' and never 'is it right.' Calibration lands moderate, mid-cohort. Above the gap cohort (Qlik, VMware, Nutanix, Dell, VAST), which had inherited permissions and at most a thin gateway rather than a productized governance plane. Peer to Databricks, AWS, and IBM, whose moderates rested on the same shape (a runtime-enforcing gateway plus governed agent identity, enforcing static policy). Below Palantir (strong), which had a genuine reasoning mechanism in Apollo's constraint solver. Cross-platform reach is broader enforcement of the same static policy, a real feature to name but not a higher-order dynamic capability, so it does not lift the grade. The GDPR case makes the placement limit concrete. An enterprise with a mix of GDPR and non-GDPR customers wanting inference pinned to region can achieve it by static partition: configure a region-pinned BYOLLM connection and route customers to the right configuration through Data Spaces. That works, and it is a designed partition. What the platform does not offer is a per-request placement policy it evaluates by data-subject residency at request time. You architect the boundary; the platform enforces the one you built. Live placement is the universal Infrastructure-2C gap, Salesforce included. ### Borrowed Judgment Low for what is provided, all Salesforce IP and Ceded to Salesforce, with no NVIDIA dependency. But this is low borrowed judgment for partial 2C: Intelligence-2C governance is productized, GA, and unusually broad; the reasoning mechanism (an Apollo-class constraint solver) is absent; live inference placement is absent. The Informatica telemetry finding sharpens it. Salesforce now has arguably the richest data-provenance telemetry on the instrument (lineage, catalog, quality, mastered entities), which is exactly the input set a reasoning plane needs (residency for placement, confidence for autonomy gating, classification for access). Sensitivity classification reaches the runtime access and masking enforcement, and catalog and lineage are queryable to agents as context through the Agent Fabric Context Catalog. But no productized 2C engine consumes lineage-residency to place inference or quality-confidence to gate autonomy. The vendor with the best inputs to a reasoning plane ships no plane that reasons over them. The static-configuration surface plus this telemetry is more than the gap cohort ever hands you, and it is not a reasoning plane. ### Working Notes Determinism distinction ('you can't prompt your way to deterministic output' / 'legal is not right') carried here as a universal finding and adopted as a 2C scoring-discipline rule (ratified Salesforce v1.0, sibling of 'routing is not reasoning'). A cross-row /reconcile pass applies it to the other rows (it tests the Databricks/AWS/IBM moderates and Palantir's strong); like the live-placement gap it is noted as universal rather than charged to one vendor, so it is expected to confirm more grades than it moves. Sub-threshold signal, prose only: Salesforce Default's managed model mix performs opaque vendor-side model selection, not a customer-configured placement policy. Fact question that could move this to strong: a GA surface where model or region placement is a per-request policy the platform evaluates by compliance context, rather than a static connection binding, or a genuine constraint-satisfaction engine behind the Broker. Read of the docs is routing plus enforcement, which is high-moderate. Fact question that could thin it toward gap: whether Agent Fabric's cross-platform governance is fully deployed and GA versus still maturing. ## ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Salesforce's Native Layer — First-Party Value Plane ### Vendor-Provided Components **Customer 360 Applications (Sales / Service / Marketing / Commerce + Industry Clouds)** [DAPM: Ceded] The first-party application estate running core customer-facing business processes. The deepest coupled capture in enterprise software: decades of process and configuration lift nowhere. Proprietary Salesforce platform, no open exit. **Agentforce Agents (Action-Bearing Business Execution)** [DAPM: Ceded] Agents that resolve cases, qualify leads, and execute real business actions on the object model under commit-boundary governance. Agent definitions, topics, and actions are captive; the reasoning core is inherited from 2B. Proprietary Salesforce platform, no open exit. **Platform (Flow, Apex, Lightning)** [DAPM: Ceded] The application platform for custom apps and workflows. Apex and Flow run only on Salesforce and Lightning apps do not lift. Proprietary Salesforce platform, no open exit. **Slack (Agentic Work Surface)** [DAPM: Ceded] The Salesforce-owned collaboration and agent-interaction surface, now a primary channel for Agentforce. Proprietary Salesforce platform, no open exit. **AppExchange (ISV Ecosystem)** [DAPM: Delegated] The third-party application marketplace. Menu-altitude Delegated: the pre-purchase choice among substitutable ISVs is real, and each deployed app captures per its own terms. Matches Dell's Ecosystem Program and the Databricks Marketplace. ### NVIDIA-Provided Components **No NVIDIA Layer 3 Dependency** The value plane is Salesforce and ISV IP with model-provider-mediated generation. The near-zero NVIDIA column closes the row empty end to end, broken only by the Beta Nemotron model-availability thread at 2B. Salesforce's borrowed AI judgment flows to the model providers through the LLM Gateway rather than to NVIDIA. ### Gap Analysis This is what Salesforce is, and it may be the benchmark first-party value plane on the instrument. The buyer gets the deepest deployed enterprise application estate on the map: Sales, Service, Marketing, and Commerce Clouds plus the industry clouds, running core customer-facing business processes for a huge installed base. Agentforce makes that estate agentic, with action-bearing agents doing real business work (the 85%-resolution service numbers as the proof point). Slack is the conversational work surface, Flow/Apex/Lightning is the platform for custom apps, and AppExchange is the ISV ecosystem. This is the coupled, visible capture, and it is the mirror of the 1A finding. At 1A, zero-copy makes the capture decoupled and invisible; here your business runs in Salesforce's namespace, decades deep, and the lock-in was never in doubt. Org config, custom objects, Apex, Flows, page layouts, automations, and agent definitions are the largest captive estate in the row, the deepest coupled lock-in on the instrument. Calibration is strong and frontier-pegged (rule 6): peer to Palantir, Databricks, and Qlik, and broader than all three. Databricks and Qlik carry an analytics-bounded caveat; Palantir's operational apps are Forward-Deployed-Engineering-delivered and bespoke, where Salesforce's are productized across every customer-facing function. Above VMware and Nutanix (moderate, platform-enabled not platform-provided). Not partner (Dell, Cisco, HPE, IBM), because this value plane is emphatically first-party. The honest caveat matches how peers got theirs: the domain is customer engagement and the business processes around it, the largest domain any Layer 3 strong covers, but still a domain. It is not the ERP and financial core, and not arbitrary enterprise operations. ### Borrowed Judgment The application opinions are Salesforce's, Ceded to Salesforce, and the accumulated org estate is the largest captive artifact set in the row. The generative layer inside the apps inherits the model chain from 2B (portable model choice, Salesforce-governed runtime). The AppExchange ecosystem is Delegated. Coupled, visible capture throughout: the buyer knows exactly what they have committed to here, the honest inverse of the 1A decoupled ask. ### Working Notes All scored components are doc- and release-confirmed; nothing here rests on inference, and no open fact questions. ════════════════════════════════════════════════════════════════════════════════ # Supermicro DCBBS & AI Data Platforms Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.0 - Initial Assessment **Date:** July 11, 2026 **Source:** Supermicro DCBBS documentation and solutions pages (May 2026 brochure era), Seven AI Data Platform Solutions launch (March 16, 2026, with Cloudian, DDN, Everpure, IBM, Nutanix, VAST Data, WEKA), BlueField-4 STX storage server (March 18, 2026), Vera Rubin DCBBS blueprints announcements (2026), SuperCloud Software Suite documentation (SCC v3.x license brief March 2026, SCAC product page and datasheet, SDX and SCD product documentation), Broadcom Enterprise SONiC license SKUs on Supermicro paper, AI-RAN/sovereign AI announcements (March 2, 2026, Nokia/SK Telecom/Telenor), published 4+1 model. Assessment notes: this row forced two instrument-level rulings, both from the grading review. First, the whose-paper capability rule — channel delivery credits fully on the capability axis (the GPU parallel: every OEM scores strong at Layer 0 on NVIDIA silicon; the same logic credits partner data platforms productized on the vendor's paper). Second, the retraction of the OEM-channel DAPM rule — a proprietary single-implementation platform is Ceded-to-partner through any paper; two sellers of one proprietary stack is channel substitution, not authority. Delegated survives only via open/self-hostable substrates or multi-vendor standard interfaces. The 2A strong carries a standing deployment-evidence flag: it survives on documentation; named production operators are the open fact. The initial control-row hypothesis (hardware vendor, nothing above the rack) failed its documentation test — recorded as evidence the doc-first method works. ## Summary Finding Supermicro grades as the most capability-dense OEM row on this instrument — strong from Layer 0 through 2A — while owning less software IP than any peer. The mechanism is the row's defining feature, and the reason for every strong grade: Supermicro's paper productizes best-of-breed partner platforms (VAST, WEKA, DDN, Cloudian, IBM, Nutanix, Everpure) as named Supermicro products with documented single-vendor support, on top of the industry's fastest NVIDIA-generation cadence, DLC-2 liquid cooling at gigawatt scale, and a DCBBS motion that delivers whole data centers as building blocks. The capability axis credits what a customer can deploy on the vendor's paper today — the same rule that scores every OEM strong at Layer 0 on NVIDIA silicon — and by that rule Supermicro's deployable menu is the widest and strongest in the on-prem market. The authority profile is the other half of the reading, and it is not Supermicro's to hold. The capture is per-choice, decided at menu time: choose VAST and your pipeline opinions cede to VAST; choose WEKA and they cede to WEKA — through Supermicro's paper as surely as through the partner's own, because a proprietary platform's opinions have no exit regardless of the invoice. Supermicro's owned captive surfaces are few and specific: the SuperCloud suite, SuperCluster integration, and the delivered backend fabric. One contract is a procurement fact, not an authority fact — this row is the instrument's cleanest demonstration of the difference. The turnkey impression is real, and so are its seams. One paper, single-vendor support, and data centers delivered as building blocks read as an integrated stack — but a stack assembled from chosen partner platforms has boundaries the buyer owns: integration seams between engines, support demarcations behind the single contract, and no cross-platform governance or curation judgment spanning any of it. The more turnkey the procurement, the easier those seams are to underprice at purchase. Above the data layers the row thins fast, and the prose says why at each cell: 2B is NVIDIA passthrough with an OpenShift escape valve — Supermicro owns no runtime, no serving, no application content. 2C is an unclaimed gap with no designated partner. Layer 3 is an uncurated ecosystem: no program, no validated app catalog, no named accountability when a multi-vendor stack fails. The buyer who gets five strong layers on one paper also inherits the assembly, the application layer, and the reasoning plane. The 2A strong deserves its caveat stated plainly: it rests on documented function coverage — multi-vendor GPU slicing across NVIDIA, AMD, and Gaudi in SDX, multi-tenant AI cloud control in SCD — from a young software suite. On paper it survives. Named production operators are the open fact a briefing must produce. The buyer's trade: the widest, strongest validated menu in the on-prem market, with authority ceded per-choice to chosen partners and to NVIDIA rather than accumulated by the vendor — in exchange for owning everything the menu doesn't cover: the seams, the curation, the applications, and the reasoning plane. The capture mechanism is decoupled and unusually legible — visible at menu time to a buyer who knows to look, which is precisely the reading this instrument exists to provide. ## ● Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Supermicro Strength ### Vendor-Provided Components **GB300 NVL72 / HGX B300 Rack-Scale Systems** [DAPM: Retained] Shipping in volume on the commodity NVL substrate — multi-OEM, workloads move without rebuilding. First-to-market cadence is the differentiation; the substrate is the exit. **X14/H14 Building-Block Servers (NVIDIA, AMD, Intel)** [DAPM: Retained] The broadest commodity server catalog in the industry, multi-silicon. The purest Retained substrate on the instrument. **DLC-2 Liquid Cooling Stack (In-Rack + In-Row)** [DAPM: Retained] CDUs, manifolds, rear-door heat exchangers, liquid-to-air sidecars, Supermicro coolant. Physical plant accumulates no portable opinions — no lock-in surface, engineering differentiation without authority cost. **Site Infrastructure (Cooling Towers, Dry Coolers, BESS, Cabling)** [DAPM: Retained] 1MW to 50MW+ water and dry cooling towers, 1.5MW/3.1MWh battery energy storage, engineered cabling design. Physical-plant rule at site scale. **SuperCluster (Pre-Validated Multi-Rack)** [DAPM: Ceded] Plug-and-play multi-rack with integrated networking fabric and L11/L12 factory validation. Integrated-system rule: the deployment's integration opinions cannot lift to another vendor as built — the PowerRack/Gigafactory analog, precisely scoped. **SuperCluster GPU Backend Fabric (NVIDIA Spectrum-X / Quantum)** [DAPM: Ceded] The deployed fabric of the integrated cluster. From the customer's seat, fabric opinions are captive to the NVIDIA stack regardless of whose paper or bezel delivers it. **SSE Switches (Broadcom Enterprise SONiC)** [DAPM: Delegated] Whitebox switching to 800G/51.2Tbps running Broadcom's Enterprise SONiC distribution, licensed on Supermicro paper under Broadcom-branded SKUs. SONiC is a genuinely open, multi-distribution substrate — opinions built against its config model lift to community SONiC or another vendor's distribution. The one non-captive fabric lane on any OEM row. ### NVIDIA-Provided Components **GPU/Accelerator Silicon** Blackwell Ultra (B300, GB300 NVL72) shipping in volume; Vera Rubin next. Supermicro's roadmap is NVIDIA's roadmap with roughly a generation's lead on packaging cadence. **NVLink / NVSwitch** Intra-rack topology on the NVL-class systems. **Spectrum-X / Quantum InfiniBand + ConnectX/BlueField** The GPU backend fabric in SuperCluster deployments is NVIDIA's, as at every OEM. ### Gap Analysis Supermicro's Layer 0 is the fastest NVIDIA-generation cadence and the broadest building-block catalog in the market: GB300 NVL72 and HGX B300 rack-scale shipping in volume, X14/H14 server lines across NVIDIA, AMD, and Intel silicon, and DLC-2 direct liquid cooling engineered as a documented stack — in-rack CDUs to 250kW, in-row CDUs to 1.8MW, liquid-to-air sidecars for retrofit, site-level water and dry cooling towers from 1MW to 50MW+, and a 1.5MW/3.1MWh battery energy storage system. DCBBS packages all of it as data-center-scale building blocks, 5MW to 1GW, with first-party services from site survey through 4-hour-response onsite support — documented as optional attach, not structural dependency. The strong is earned on cadence, cooling engineering, and delivery scale. What Supermicro does not own is any networking software: the SSE switch line runs Broadcom's Enterprise SONiC distribution (the license SKUs on Supermicro paper are Broadcom-branded), and the GPU backend fabric is NVIDIA's. The turnkey capture is precisely scoped to SuperCluster — the pre-validated multi-rack product with integrated fabric and L11/L12 factory validation — not to DCBBS wholesale, whose blocks are commodity and whose physical plant accumulates no portable opinions. ### Borrowed Judgment Low at the system level — commodity tin, substitutable across OEMs. The fabric is where judgment is inherited: from the customer's seat, fabric opinions are captive to the Spectrum-X/Quantum stack in a SuperCluster regardless of brand posture, and the SONiC lane's opinions rest on Broadcom's distribution of an open substrate — the one non-captive fabric option on any OEM row, and it is a posture, not Supermicro IP. Supermicro contributes zero networking software judgment; even its own-bezel switches run someone else's NOS. ### Working Notes Watch-list (checked July 11, 2026, not scored): Vera Rubin NVL72 and HGX Rubin NVL8 DCBBS blueprints — customer engagements open, deployments scheduled H2 2026 aligned to NVIDIA GA; the announced Vera CPU systems and BlueField-4 context-memory AI storage system ride the same date. Scores when Supermicro documentation confirms GA. Data Center Fit-Out and Build services (bare-land to operational) are noted here and scored nowhere — deployment services, expertise transfers. Fact question resolved by documentation: DCBBS services are optional building blocks ('optional software installation,' response-time 'options'), so day-2 operation without Supermicro is supported — which is what narrows the Ceded scope to SuperCluster. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Partner-Delivered Strength ### Vendor-Provided Components **Petascale All-Flash + BlueField-4 STX Storage Servers** [DAPM: Retained] Commodity storage hardware substrate — PCIe Gen 5, E3.S, 480TB/2U; among the first BlueField-4 STX storage servers (March 18, 2026). The opinions live in whatever SDS runs on top; the tin swaps across OEMs. **Proprietary AI Data Platforms (VAST CNode-X, WEKA, DDN HyperPOD, IBM Storage Scale, Everpure, Nutanix)** [DAPM: Ceded] Named Supermicro products delivering single-implementation proprietary platforms. The lift-to-leave litmus is indifferent to the paper: these platforms' opinions — filesystems, catalogs, policies — have no exit from their owners, so the capture is Ceded to the chosen partner through Supermicro's contract as surely as through the partner's own. **S3-Interface Object Platforms (Cloudian HyperScale, MinIO AIStor)** [DAPM: Delegated] Proprietary implementations behind the multi-vendor S3 standard — object opinions (buckets, lifecycle, policies) lift to any S3 platform. The standard interface, not the channel, is what keeps these Delegated. ### NVIDIA-Provided Components **NVIDIA AI Data Platform Reference** The seven solutions are NVIDIA AIDP reference builds — the acceleration architecture (BlueField-4, STX, GPUDirect/RDMA) is NVIDIA's, the storage intelligence is the partners'. ### Gap Analysis Why strong: the capability axis credits what a customer can deploy on Supermicro paper today, and that menu is the widest and strongest in the on-prem market — seven AI Data Platform solutions launched March 16, 2026 with VAST, WEKA, DDN, Cloudian, IBM, Nutanix, and Everpure, several of them platforms rated strong on their own rows of this instrument, delivered as named Supermicro products (CNode-X, HyperPOD, HyperScale) with documented single-vendor support, on Petascale hardware (480TB in 2U) with the BlueField-4 STX storage server already in the lineup. This is the same crediting rule that scores every OEM strong at Layer 0 on NVIDIA silicon: channel delivery of differentiated capability counts in full; ownership is the other axis's question. The governance surfaces arrive with the platforms — VAST's catalog, IBM Scale's governance — so the governed-data-foundation function is deployable, though no Supermicro-owned governance layer exists, which is an authority fact the DAPM column carries, not a capability absence. The seam to name: each platform governs within its own namespace; nothing Supermicro-owned spans them. ### Borrowed Judgment Total for data management, per chosen partner. Supermicro contributes hardware, validation, and paper; the data-management judgment the enterprise inherits belongs to whichever platform it picked at menu time — and under the lift-to-leave litmus, most of those choices are one-way doors. The buyer's authority position is decided at menu time. That is this cell's finding and the row's. ### Working Notes Single-vendor support is documented ('unified architecture and single-vendor support' on the AIDP solutions page); the March 16, 2026 launch language is press-release-grade rather than product-guide-grade — flagged, scored as orderable. Fact question for a briefing: whether partner software licenses ride Supermicro SKUs uniformly (the WEKA SKU trail exists; per-partner confirmation pending). No Supermicro-owned metadata or governance catalog exists at any tier — confirm by briefing that nothing is unannounced. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Partner-Delivered Retrieval ### Vendor-Provided Components **Supermicro VAST CNode-X (Retrieval Facet)** [DAPM: Ceded] GPU compute collocated with the VAST platform's retrieval and indexing surfaces. Proprietary single-implementation platform: retrieval opinions are captive to VAST through any paper. **AIDP Retrieval — Proprietary Platforms (WEKA, DDN Infinia, IBM Scale, Everpure, Nutanix Facets)** [DAPM: Ceded] Index and context opinions accumulate in the chosen engine and have no exit from its owner. Strong capability, captive authority — the hyperscaler pattern, delivered through an OEM. **AIDP Retrieval — S3/Standard-Interface Platforms (Cloudian, MinIO)** [DAPM: Delegated] Retrieval data services behind the multi-vendor S3 standard; the standard interface keeps the opinions portable. ### NVIDIA-Provided Components **NeMo Retriever + cuVS + NIM Microservices** By Supermicro's own documented architecture, the embedding and indexing intelligence is NVIDIA software — the AIDP ingests, embeds, and indexes with NeMo Retriever, cuVS, and NIMs. The retrieval judgment is NVIDIA's; the index custody is the partner's; Supermicro's is neither. Load-bearing column. ### Gap Analysis Why strong: the documented AIDP architecture on Supermicro is a retrieval layer — continuous ingestion, embedding and indexing in place without relocating source data, and semantic query for RAG and agents ('a semantic knowledge layer AI applications can query in near real time'). The buyer picks the index and data engine from the validated partner menu and deploys it on Supermicro paper with single-vendor support; CNode-X brings GPU compute to where the data lives. Under the whose-paper crediting rule, that deployable capability — several of its engines rated strong at retrieval on their own rows — grades strong here, where peers with narrower productized menus (Dell's single Elastic lane, HPE's reference stack, Lenovo's validated use cases) graded moderate. The seam: retrieval quality observability (recall@k, latency percentiles) that a Layer 2C could consume exists nowhere on the menu — the universal peer finding — and each engine's index is its own island. ### Borrowed Judgment High, split two ways: NVIDIA holds the embedding intelligence, the chosen partner holds the index and its opinions. Supermicro contributes tin and validation. The proprietary-platform lanes are one-way doors under the litmus; the S3-standard lanes keep object-side opinions portable. ### Working Notes Facet convention: the partner platforms score their storage facet at 1A, retrieval facet here, and pipeline facet at 1C — one platform, three functions, each cell naming its facet. Fact question carried from 1A: uniform Supermicro-SKU licensing per partner. Whether NIM/NeMo licenses ride the configured solutions or the customer licenses NVAIE separately is a narrative detail, not a DAPM mover. ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Partner-Delivered Pipeline Engines ### Vendor-Provided Components **Supermicro VAST CNode-X — DataEngine (Pipeline Facet)** [DAPM: Ceded] VAST's event-driven pipeline engine, rated strong at 1C on VAST's own row, delivered as a named Supermicro product. Proprietary engine: pipeline opinions have no exit from VAST regardless of whose paper or tin. **HyperPOD — DDN Infinia (Pipeline Facet)** [DAPM: Ceded] Infinia's data-intelligence pipeline (ingest, preparation, analytics) as a named Supermicro product. Same litmus: Infinia's opinions are DDN-captive through any channel. ### NVIDIA-Provided Components **BlueField-4 STX / Context Memory Reference** The context-memory storage direction rides NVIDIA's STX architecture — watch-listed pending doc-confirmed GA on Supermicro paper, consistent with the CMX treatment on the Dell and NVIDIA rows. **Blueprints / NIM Pipeline Templates** Pipeline patterns, not pipeline infrastructure — the same templates every NVIDIA partner carries. ### Gap Analysis Why strong, and why this cell forced the instrument to say it precisely: the pipeline engines deployable from Supermicro today are the full VAST platform via CNode-X — a capability set this instrument rates strong at 1C on VAST's own row — and DDN Infinia via HyperPOD, both named Supermicro products with single-vendor support. Two calibration tests decide the grade. The CoreWeave exposure test: an underlay capability a vendor does not surface as a purchasable product is not that vendor's capability — CoreWeave runs VAST-class engines invisibly behind its service interface and scores gap; Supermicro sells the engine itself, exposed, licensed, and administered by the customer, and clears the test. The GPU parallel: channel delivery of differentiated capability credits fully on the capability axis, the same rule that scores every OEM strong at Layer 0 on NVIDIA silicon. What the grade does not do is flatter the authority position: the pipeline opinions the customer accumulates — DataEngine functions, Infinia configurations — are proprietary surfaces with no exit from VAST and DDN respectively. Strong and Ceded: complete and captive, the hyperscaler pattern on OEM paper. The seam: pipelines built in one engine do not compose with the other, and nothing Supermicro-owned bridges them. ### Borrowed Judgment High, held by the chosen engine's owner. Supermicro contributes hardware, productization, and support; VAST's and DDN's judgment governs how data moves, transforms, and triggers — and under the lift-to-leave litmus the enterprise cannot take that judgment anywhere its owner does not control. ### Working Notes Watch-list (checked July 11, 2026, not scored): the Context Memory Storage Solution (BlueField-4 STX) — 'among first to unveil' (March 18, 2026) is unveil language, and the NVIDIA context-memory software tier is pre-GA on the NVIDIA row; scores at doc-confirmed GA on Supermicro paper. Fact question, this cell's load-bearing one: whether CNode-X licensing includes the full VAST software platform (DataEngine, SyncEngine) or a storage-tier subset — the strong rests on full-platform delivery; a subset drops it. Instrument follow-up logged: Lenovo's 1C gap should be re-asked under the exposure and whose-paper rules (WEKA SDS Ready Nodes are productized on Lenovo paper). ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** First-Party GPU Cloud Orchestration ### Vendor-Provided Components **SuperCloud Composer (SCC)** [DAPM: Ceded] Unified rack-scale and liquid-cooling management — servers, networks, PDUs, CDUs, third-party systems; leak detection and facility telemetry; 20K+ hosts. Proprietary management plane: fleet and facility opinions are captive, the XClarity/OpenManage precedent with cooling depth those lack. **SuperCloud Automation Center (SCAC)** [DAPM: Delegated] Pre-built provisioning automation from firmware through Kubernetes/OpenShift, wrapping open tooling (Foreman, xCAT, Ansible/Puppet/Chef, GitOps). The Ezmeral precedent: packaging over an open automation substrate — workflow content ports to alternative packaging of the same tools. **SuperCloud Developer Console (SDX) — GPUaaS Facet** [DAPM: Ceded] GPU slicing across NVIDIA, AMD, and Intel Gaudi; multi-tenant workspace provisioning; one-click GPUaaS. Proprietary console: slicing policies and tenant opinions are captive — the Run:ai-class precedent, notable for spanning GPU vendors. **SuperCloud Director (SCD)** [DAPM: Ceded] Multi-tenant AI cloud control: bare metal, Ethernet and InfiniBand multi-tenancy, storage management, purpose-built for GPUaaS and AI-factory operations. Proprietary control plane: the operating opinions of an AI cloud accumulate here and have no exit. ### NVIDIA-Provided Components **NVIDIA Mission Control / Run:ai / AI Enterprise** The standard SuperCluster path carries NVIDIA's orchestration stack on Supermicro paper — the GPU-aware judgment inside the NVL-class deployments is NVIDIA's, as at every OEM. **Counter-finding: SDX Multi-Vendor Slicing** SDX's GPU slicing spans NVIDIA, AMD, and Intel Gaudi — the rare OEM 2A surface not wholly NVIDIA-dependent. A thinner NVIDIA column here than on the Dell or Lenovo rows, and that is itself a finding. ### Gap Analysis Why strong: the SuperCloud suite covers more of this layer's defining function — GPU scheduling, quotas, fair-share, multi-tenancy — in first-party software than any peer OEM's owned portfolio. SCC manages fleet and facility as one surface (servers, networks, PDUs, CDUs, leak detection, 20K+ hosts — liquid-cooling depth no peer's management plane has, licensed and documented at v3.x, March 2026). SCAC automates firmware-to-Kubernetes provisioning over open tooling. SDX is a GPUaaS console with GPU slicing across NVIDIA, AMD, and Intel Gaudi and one-click multi-tenant workspace provisioning. SCD is multi-tenant AI cloud control — bare metal, Ethernet and InfiniBand multi-tenancy, storage management — purpose-built for the GPU-cloud operators Supermicro already supplies. Calibration: HPE's 2A strong rests on GreenLake Intelligence's agentic cross-domain mesh while bracketing GPU-specific scheduling to NVIDIA; Supermicro has no agentic ops story but claims the GPU-cloud operations core of the layer first-party — different strengths, same band, and both above Dell's rack management and Lenovo's fleet-plus-TruScale moderates. The standing caveat is maturity: this strong rests on documented function coverage from a young suite. On paper it survives; named production operators are the open fact. ### Borrowed Judgment Moderate — lower than Dell's or Lenovo's at this layer, because Supermicro owns real GPU-sharing and multi-tenancy judgment first-party. On the standard NVL-class path, NVIDIA's Mission Control and Run:ai judgment still governs inside the cluster; the SuperCloud consoles govern around and above it, and their opinions — tenant configurations, slicing policies, facility baselines — are captive to Supermicro. ### Working Notes Standing confidence flag (the cell's grade condition): the strong survives on documentation — datasheets, the SCC v3.x license brief, documented function coverage. A briefing that cannot produce named GPU-cloud operators running SDX/SCD in production is the thing that would move this cell. Fact questions: whether SDX's multi-vendor slicing is Supermicro engineering or licensed third-party technology; who operates SCD day-2 in practice (customer or Supermicro DCaaS motion). ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** NVIDIA Passthrough + Open Substrate ### Vendor-Provided Components **SDX Execution Workspaces (2B Facet)** [DAPM: Ceded] Per-tenant training/fine-tuning/inference/benchmarking environments provisioned by the proprietary console. Workspace and pipeline-provisioning opinions live in SDX and do not lift. **Nutanix AI Platform (AIDP Solution, Runtime Facet)** [DAPM: Ceded] Partner inference/agentic runtime delivered on Supermicro paper. Proprietary single-vendor platform: opinions are captive to Nutanix through any channel — the litmus is indifferent to the invoice. **OpenShift / Kubernetes-as-a-Service via SCAC** [DAPM: Delegated] Documented provisioning workflows delivering the instrument's benchmark open substrate — K8s manifests and OpenShift opinions port anywhere. ### NVIDIA-Provided Components **NVIDIA AI Enterprise + NIMs** The model-serving runtime of the standard path, validated across the SuperCluster and platform configs on Supermicro paper. The runtime lane is a passthrough of NVIDIA's judgment. **Dynamo** Distributed inference with KV-aware routing on the NVL-class racks — single-variable cache-locality optimization, not placement policy. **OpenShell / NemoClaw** Alpha — not GA, watch-listed, consistent with their treatment on every peer row. ### Gap Analysis Why moderate, and why the label says passthrough: Supermicro owns no runtime — no model serving, no agent execution, no application content, and no AI-application services lane of the Lenovo AI Discover/Fast Start kind. What ships is real and multi-path: the NVIDIA lane (AI Enterprise + NIMs validated on Supermicro paper — a passthrough of NVIDIA's serving, optimization, and guardrail judgment), the Nutanix AI Platform delivered as one of the seven AIDP solutions, and a genuinely open lane — OpenShift/Kubernetes-as-a-Service provisioned by documented SCAC workflows. SDX adds per-tenant execution workspaces (training, fine-tuning, inference, benchmarking) above the runtimes without being one. Calibrated against Dell (blueprints + services), HPE (PCAI + frameworks), and Lenovo (four paths + first-party library), this is the thinnest first-party moderate of the four — the deployable menu clears the bar; nothing Supermicro-owned distinguishes it. ### Borrowed Judgment High on the standard path: NVIDIA's runtime judgment inherited whole, through the dependency column rather than a scored component, exactly as on the peer rows. The OpenShift lane is the open-substrate escape valve — opinions built there lift to any OpenShift deployment. SDX contributes provisioning judgment only, and its opinions are Supermicro-captive. ### Working Notes Watch-list: OpenShell/NemoClaw remain alpha (dated on the NVIDIA row). Fact questions: NVAIE Supermicro-SKU confirmation (presumed from the platform pattern; part-number verification pending); whether any first-party or SKU'd agentic content exists anywhere in the portfolio — research says none, and a briefing confirming the absence firms both this cell and Layer 3. Facet note: SDX's GPUaaS/slicing facet is scored at 2A; the execution-workspace facet here. ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Enterprise Responsibility ### NVIDIA-Provided Components **AI-Q Reference Architecture** Multi-agent workflow scaffolding. Does not make placement decisions. **Dynamo KV-Aware Routing** Performance-aware routing, single variable. Not multi-variable policy optimization. **OpenShell Governance** Runtime sandboxing — 2B constraint enforcement, not 2C placement reasoning. Alpha — not GA. ### Gap Analysis The 4+1 model defines Layer 2C as a required function — policy-driven decisions about where compute runs relative to data, which model serves which request, and how cost, compliance, and latency are arbitrated at request time. Supermicro does not provide it, and 'routing is not reasoning' disposes of every candidate faster than on any peer row: SCD is multi-tenant operations control (a 2A function), SDX is provisioning, Dynamo routes on cache locality, AI-Q is scaffolding. There is no policy engine, no Intelligence-2C governance surface, and — like Lenovo, unlike HPE — no designated 2C partner anywhere in the ecosystem. The shape of the absence differs from Dell's, whose gap is briefing-confirmed deliberate strategy: Supermicro reads as a scope boundary — the company sells to GPU-cloud operators who build their own control planes above SCD — but absent a stated position it gets the omission framing. The enterprise, or the operator running Supermicro gear, retains the reasoning plane, and the seams finding lands hardest here: five strong layers of assembled capability with no judgment layer spanning them. ### Borrowed Judgment None to borrow — the enterprise retains full responsibility for this function, mostly without naming it as a function. On Supermicro infrastructure that responsibility is compounded by the menu model: each chosen platform brings its own policies, and nothing arbitrates across them. ### Working Notes Fact question for a briefing: whether SCD's roadmap reaches toward policy-driven placement — its tenancy scope is adjacent, and it is the natural home if Supermicro ever claims this layer. Today's documentation shows operations control only. The live-placement gap is universal across the OEM rows — an instrument-wide finding, not a Supermicro-specific defect. Strategy-versus-omission is the same open question carried on the Lenovo row. ## ◇ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Uncurated Partner Ecosystem ### Vendor-Provided Components **Per-Solution Application Deliveries (Nutanix Agentic AI Solution, Application Facet)** [DAPM: Delegated] Application-tier capability arriving through the AIDP solution set, substitutable per use case at the ecosystem altitude. The chosen platform's own capture applies per choice — Nutanix's runtime opinions are scored Ceded at its 2B component. **Vertical Solution Collaborations (AI-RAN, Sovereign AI — Nokia, SK Telecom, Telenor)** [DAPM: Delegated] Per-vertical delivery partnerships with production deployments. Substitutable partners; the accumulated opinions sit with the operator and the telco partner, not Supermicro. ### NVIDIA-Provided Components **NIM Agent Blueprints + NeMo** The application substrate — identical blueprints ship on Dell, HPE, and Lenovo, differentiating nothing. ### Gap Analysis Applications reach Supermicro infrastructure through three doors, none of them a program: NVIDIA's blueprint and NIM patterns, per-solution partner deliveries (the Nutanix 'Agentic AI Solution for AI Factories' is the closest thing to an application product on Supermicro paper), and vertical collaborations — AI-RAN and sovereign-AI builds with Nokia, SK Telecom, and Telenor, with real deployments behind them (Norway's sovereign AI cloud, SK Telecom's mega-cluster). The biggest door is implicit: Supermicro's core customers are GPU clouds and AI factories that bring their own application layer entirely. Why partner rather than gap: the layer is addressed, and validated per-solution deliveries exist on Supermicro paper. Why uncurated is the finding: there is no ISV program, no validated app catalog, no interoperability testing, and no named accountability when a multi-vendor application stack fails — Dell, HPE, and Cisco all run structured programs at their partner grade, and Lenovo broke to moderate on first-party agents. Supermicro is the fourth OEM shape, and the weakest L3 posture among them. This is where the turnkey-with-seams trade bills the buyer: peers at least decide which ISVs get validated; Supermicro's buyer inherits the curation judgment too. ### Borrowed Judgment Fully distributed, which is architecturally correct at Layer 3 — with the Supermicro-specific sharpening that no curation judgment exists either. Whatever application platform the enterprise or operator chooses captures them on that platform's own terms; Supermicro is structurally indifferent to the choice. ### Working Notes The absence of a formal ISV/application program is confirmed across the 2026 announcement corpus; a briefing revealing one would move the narrative, not likely the grade. The Nutanix platform's runtime facet is scored at 2B; the application-delivery facet here. ════════════════════════════════════════════════════════════════════════════════ # VAST AI Operating System Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.5 - 2B Maturity-Discount Reconciliation **Date:** June 21, 2026 **Source:** VAST Forward 2026, VAST AI OS white paper, analyst coverage, published 4+1 model. v1.5 (instrument reconciliation): 2B moderate→strong — AgentEngine is GA (H2 2025) and a complete, owned agent runtime; the prior moderate rested on a maturity/less-proven discount, which is not a capability axis (GA + complete owned runtime = strong, the same bar Databricks 2B cleared). Maturity may become a scored attribute as the agentic market matures. ## Summary Finding VAST is the most architecturally distinct vendor in this assessment series because it deliberately collapses traditionally separate infrastructure layers into a single platform. The VAST AI Operating System unifies storage (DataStore), metadata and vector database (DataBase + Catalog), global namespace (DataSpace), serverless compute (DataEngine), retrieval (InsightEngine), agent runtime (AgentEngine), governance (PolicyEngine), and model lifecycle (TuningEngine) under one authority boundary. Where Dell assembles 5–7 partner technologies to span Layers 1A through 2B, VAST provides a vertically integrated alternative with no inter-layer seams. VAST is making the most aggressive middle-out bid for Layer 2C of any infrastructure vendor: Polaris (placement abstraction) and DataSpace (namespace) ship today, with PolicyEngine (inline policy enforcement) and TuningEngine designed to close the loop. But the policy-reasoning engine is announced for End of 2026, not shipping - today VAST ships placement abstraction without the reasoning plane, which is routing, not reasoning. So for an architect building now, VAST Layer 2C is a gap with a credible, dated roadmap: the clearest building-toward-it story in the series, not yet a deployable control plane. The DAPM trade-off is stark: VAST eliminates seams by collapsing authority into one vendor. The enterprise gains architectural coherence and eliminates integration risk. But the entire data plane, retrieval plane, agent runtime, AND emerging governance plane become a single Ceded dependency. If you run VAST, you run VAST for everything. There is no substitutability at any individual layer. The $30B valuation, $4B+ in bookings, $500M+ CARR, and 1,000+ enterprise deployments validate market traction. The CoreWeave anchor ($1.17B commercial agreement) validates hyperscale credibility. The CNode-X partnership with NVIDIA (GPU-accelerated servers through Cisco and Supermicro OEMs) validates compute integration. But the product-market fit question for the 4+1 model is whether enterprises will accept total data-plane vendor dependency in exchange for architectural simplicity. VAST is building the thing the 4+1 model says is missing. Whether the enterprise will buy it from a storage vendor is the open question. ## ◑ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Software-Defined ### Vendor-Provided Components **DASE Architecture (Foundation)** [DAPM: Ceded] Disaggregated Shared-Everything: separates stateless compute (CNodes) from persistent storage (DNodes/DBoxes) over an NVMe-over-Fabrics network. Every CNode mounts every SSD in the cluster at boot time — no node ‘owns’ any particular data. Eliminates east-west chatter between nodes, shard coordination, and per-node metadata bottlenecks. ACID transactional semantics via global Element Store. Scales linearly from TB to EB. This is VAST’s core IP and the foundation for every layer above. Proprietary VAST architecture — opinions captive, no open exit. **CNodes (Stateless Compute Servers)** [DAPM: Ceded] Process all user requests: file/object serving (NFS, S3), database queries, erasure coding, data reduction, vector search. Fully stateless — containerized VAST software. Add CNodes to scale performance independently from capacity. Can run on dedicated servers or as containers within EBoxes. ConnectX-7/8 network adapters for 200GbE+ RDMA connectivity. Proprietary VAST architecture — opinions captive, no open exit. **DNodes / DBoxes (Persistent Storage)** [DAPM: Ceded] NVMe-oF storage shelves connecting SCM and hyperscale flash SSDs to the NVMe fabric. DBoxes are fully redundant (no single point of failure) — redundant DNodes, NICs, fans, power. DNode containers route NVMe commands between SSDs and CNodes. Add DBoxes to scale capacity independently from performance. Proprietary VAST architecture — opinions captive, no open exit. **EBox (Everything Box)** [DAPM: Ceded] Converged form factor: runs CNode + DNode containers on the same industry-standard x86 server. Introduced in VAST 5.2. Each EBox runs three containers (1 CNode, 2 redundant DNodes). Enables deployment on hyperscaler server configs, cloud VM instances, and OEM servers (Cisco, Supermicro). The CNode in an EBox does NOT own the SSDs in that EBox — all SSDs are equally accessible by all CNodes across the cluster. Proprietary VAST architecture — opinions captive, no open exit. **CNode-X (GPU-Accelerated, 2026)** [DAPM: Delegated] New node type adding GPU acceleration directly into VAST clusters. First time GPUs are embedded in the data platform rather than consuming it externally. Supermicro config: CloudDC AS-1116CS-TN (storage) + SYS-212GB-FNR 2U (compute) with 2x NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, AMD EPYC 9005 CPUs. Follows NVIDIA AI Data Platform reference architecture. OEM partners: Cisco, Supermicro (shipping); HPE, Lenovo (in progress). Dell notably absent. **NVMe-over-Fabrics Network** [DAPM: Delegated] High-bandwidth RDMA networking (200GbE+ via NVIDIA ConnectX-7/8) connecting all CNodes to all DNodes/SSDs. BlueField-3 DPUs + Spectrum-X switches for zero-copy data paths. This is the fabric that makes ‘shared everything’ possible — every compute node sees the entire storage namespace over NVMe-oF. ### NVIDIA-Provided Components **NVIDIA GPU Silicon (CNode-X)** RTX PRO 6000 Blackwell Server Edition GPUs in CNode-X configurations. GPU acceleration for vector search, data manipulation, inference, and analytics within the VAST platform. **NVIDIA ConnectX-7/8 NICs** Network adapters in every CNode providing RDMA connectivity to the NVMe fabric. This is the physical layer that enables DASE’s shared-everything model. **NVIDIA BlueField-3 DPUs** Data processing units in CNode-X for storage-side compute offload. Enable zero-copy data paths between GPU memory and NVMe storage. **NVIDIA Spectrum-X Switches** Ethernet switches for the NVMe fabric and GPU cluster networking. Same Spectrum silicon that Dell brands as PowerSwitch. **NVIDIA Libraries (cuVS, cuDF, DOCA)** Integrated directly into VAST software services on CNode-X. GPU acceleration for vector search (cuVS), data manipulation (cuDF), and networking (DOCA). ### Gap Analysis VAST’s Layer 0 is architecturally the inverse of Dell’s. Dell manufactures and sells the physical servers (PowerEdge, PowerRack) but depends on NVIDIA for the software runtime. VAST designs the software architecture (DASE) but depends on OEM partners (Cisco, Supermicro, HPE, Lenovo) for the physical servers. The DASE architecture is VAST’s genuine Layer 0 differentiator. No other vendor has a shared-everything model where every compute node can directly access every SSD over NVMe-oF with ACID guarantees. Dell’s PowerScale and ObjectScale are traditional storage architectures (even if highly performant). VAST’s DASE fundamentally changes how compute and storage relate — they share the same data structures, the same namespace, and the same transactional model. The CNode-X evolution is architecturally significant: it dissolves the boundary between ‘storage infrastructure’ and ‘compute infrastructure.’ In Dell’s architecture, PowerEdge servers run NVIDIA’s inference runtime and PowerScale provides separate storage — data moves between them. In VAST’s CNode-X architecture, GPUs embedded in the data platform accelerate data services AND serve inference — no data movement because compute and storage are the same system. The EBox model is worth noting for the procurement story: ‘Gemini model’ pricing means certified hardware is supplied at cost from the manufacturer, with VAST software as a capacity-based subscription. VAST guarantees software compatibility with new and older hardware for up to 10 years. This is a fundamentally different commercial model than Dell’s (buy the server, license NVIDIA AI Enterprise separately, integrate yourself). The NVIDIA dependency at Layer 0 is real but different from Dell’s. Dell depends on NVIDIA for GPU silicon AND the entire software stack above it (Run:ai, NemoClaw, OpenShell, AI Enterprise). VAST depends on NVIDIA for GPU silicon and networking silicon but retains authority over the software architecture. If NVIDIA changes its GPU roadmap, both Dell and VAST are affected. But if NVIDIA changes its software roadmap (NemoClaw, AI Enterprise licensing), only Dell is affected — VAST’s software is its own. ### Borrowed Judgment Moderate but architecturally different from Dell’s. VAST borrows hardware judgment from OEM partners (Cisco/Supermicro for servers) and silicon judgment from NVIDIA (GPUs, NICs, DPUs, switches). But VAST retains software architecture judgment (DASE, containerized CNodes, NVMe fabric design) entirely. Dell borrows software judgment from NVIDIA (Run:ai, NemoClaw, OpenShell, AI Enterprise) while retaining hardware judgment (PowerEdge design, thermal engineering, rack integration). The DAPM distinction: Dell’s borrowed judgment at Layer 0 is silicon-level (structural, everyone shares it). VAST’s borrowed judgment at Layer 0 is hardware-manufacturing-level (Cisco/Supermicro build the boxes). Neither vendor is fully independent at Layer 0, but their dependencies are in different dimensions. ### Working Notes Dell is notably absent from VAST’s OEM partner list (Cisco, Supermicro shipping; HPE, Lenovo in progress). This is a competitive signal — Dell positions itself as VAST’s primary competitor in AI data platforms (PowerScale/ObjectScale/MetadataIQ vs. DataStore/DataBase/Catalog). Dell building CNode-X configurations would be akin to VMware selling on Hyper-V. The containerized update model is a Layer 0 operational differentiator: VAST updates CNodes by spinning up a new container version alongside the old one and switching in seconds. Traditional server updates require node reboots and downtime. This reduces the operational overhead of infrastructure management — a function that Dell handles through OpenManage Enterprise and firmware lifecycle processes. The Gemini procurement model (hardware at cost, software as subscription) means VAST’s revenue is software-driven. Dell’s revenue is hardware-driven with NVIDIA software licenses as pass-through. This affects how each vendor invests in software capabilities — VAST’s business model incentivizes software differentiation; Dell’s incentivizes hardware volume. ## ● Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** VAST Strength ### Vendor-Provided Components **VAST DataStore** [DAPM: Ceded] High-performance file, object, and block storage on all-flash NVMe. Universal Storage heritage — no separate engines for file vs. object vs. block. NFS v3/v4.1, SMB 2.1/3.1, S3 REST API. An Element is not an NFS file or an S3 object — it’s a superset of all of them. Write a file as an object, read it as SMB, read it over NFS — same underlying data, no copies, no translation layers. Protocol-independent storage with protocol-dependent access. **VAST DataBase** [DAPM: Ceded] Purpose-built database for AI: structured data (EDW tables), contextual metadata, vectors for similarity search, real-time streams (Kafka-compatible), catalogs, and logs. Vectors live alongside structured records in a single query path. Supports open formats like Parquet. Handles tables alongside unstructured data in the same transactional system. **VAST Catalog** [DAPM: Ceded] Indexes metadata attributes of all data on cluster — files, objects, directories. Queryable via Web UI, VMS API, CLI, or connected third-party query engines. Automated classification based on multi-protocol access patterns. The discovery surface that downstream engines (InsightEngine, PolicyEngine) consume. **VAST Element Store (Foundation)** [DAPM: Ceded] Low-level key-value store on B-tree data structure underlying all services. Organizes physical storage into a global namespace holding Elements (files/objects, tables, block volumes/LUNs). Every Element is automatically enriched with metadata, security data, and data reduction data. Provides Element-level access control, encryption, snapshots, clones, and replication. ACID transactional semantics with decentralized locks at file/object/table level. No eventual-consistency pitfalls — if an update is committed, subsequent queries see it. **Security & Compliance (Built-in)** [DAPM: Ceded] Inline encryption (at rest and in transit). Immutable snapshots (critical for ransomware recovery and audit trails). MFA. Granular access control at Element level. U.S. Government STIG-aligned security configuration. HIPAA, SOC 2, GDPR-ready compliance posture. CrowdStrike integration monitors data access, admin activity, and workload behavior continuously — security embedded at the data layer, not overlaid. Security governance follows data end-to-end: access controls on source files propagate automatically through embeddings to inference outputs. ### NVIDIA-Provided Components **NVIDIA cuVS (via CNode-X)** GPU-accelerated vector indexing and search. Integrated directly into VAST DataBase on CNode-X rather than as an external accelerator. Enhances performance but is not required for core storage or database operations — VAST DataStore/DataBase run without GPUs on standard CNodes/EBoxes. ### Gap Analysis This is VAST’s strongest layer and its primary 4+1 differentiator. Where Dell requires PowerScale (file) + ObjectScale (object) + Exascale (combined) + MetadataIQ (metadata) + Lightning FS (parallel) + Trust3 AI (governance partner) as separate components with integration seams and multiple authority boundaries, VAST provides a single platform where all data types, metadata, vectors, and security controls share the same Element Store with ACID guarantees. The multiprotocol Element Store is the architectural fact that changes the 4+1 mapping. In Dell’s architecture, data stored as files (PowerScale/NFS) must be accessed differently than data stored as objects (ObjectScale/S3). MetadataIQ indexes across both but the storage engines are separate systems with separate governance surfaces. In VAST, any Element can be accessed as a file, an object, or a table row — same data, same permissions, same metadata — because the Element Store is protocol-independent. For AI pipelines where data flows between ingestion (S3), preprocessing (NFS), training, and inference, this removes an entire class of integration complexity. The governance catalog question from the 4+1 model — ‘is the metadata rich enough to drive Layer 2C placement decisions?’ — has a clearer answer with VAST than with any other vendor in this assessment series. Because the Catalog, DataBase, PolicyEngine, and security controls are all part of the same platform operating on the same Element Store, the metadata that PolicyEngine queries for governance decisions is the same metadata that DataStore manages. No API boundary, no schema translation, no integration seam between ‘where the data lives’ and ‘where governance decisions are made.’ The security-follows-data model is significant: access controls on source files propagate automatically through embeddings to inference outputs. Dell’s Trust3 AI provides similar governance intent but as a partner overlay on separate storage engines — the propagation must cross system boundaries. VAST’s propagation is structural because storage, metadata, and security share the same data structures. The trade-off remains stark: if you choose VAST for Layer 1A, you’ve also chosen VAST for Layer 1B (InsightEngine) and Layer 1C (DataEngine). The layers are not independently substitutable. Dell’s modular approach lets you swap Elastic for Weaviate at Layer 1B without touching Layer 1A. VAST does not offer that optionality. This is the fundamental DAPM trade-off: architectural coherence vs. component substitutability. ### Borrowed Judgment Low — the lowest in either vendor assessment. VAST owns the storage engine, the database engine, the metadata catalog, the Element Store, and the security/compliance stack. NVIDIA dependency is limited to GPU acceleration (cuVS on CNode-X) which enhances but isn’t required for core storage operations. VAST’s storage platform runs without GPUs on standard CNodes/EBoxes; CNode-X adds GPU acceleration on top. Compare to Dell’s Layer 1A: Dell owns PowerScale, ObjectScale, Exascale, Lightning FS, and MetadataIQ (Retained) but depends on NVIDIA for acceleration (cuVS, SuperNICs, NeMo Retriever connector) and Trust3 AI for governance (Delegated). Dell’s authority at Layer 1A is Retained with two Delegated dependencies. VAST’s authority at Layer 1A is entirely self-contained — but Ceded to VAST as a vendor dependency. The DAPM distinction: Dell’s enterprise Retains Layer 1A authority. VAST’s enterprise Cedes Layer 1A authority to VAST. Both are ‘low borrowed judgment’ in different ways — Dell borrows less from partners because it built the storage; VAST borrows less from NVIDIA because its storage is GPU-independent. But the enterprise’s relationship to the vendor is different: Dell = you own it; VAST = VAST owns it, you subscribe. ### Working Notes The Element Store enrichment model is the architectural reason VAST’s governance catalog is inherently richer than Dell’s MetadataIQ: every Element is automatically enriched with metadata, security data, and data reduction data at write time — not indexed after the fact. This is structural metadata richness vs. bolt-on tagging. The practical consequence: VAST’s metadata is always current (enriched inline with writes). Dell’s MetadataIQ indexes asynchronously, which means metadata currency depends on indexing lag. The compliance story is maturing: U.S. Government STIG-aligned security configuration guide (v1.5), HIPAA compliance documentation, and the Kiteworks/Cybersecurity Insiders 2026 forecast showing 63% of organizations cannot enforce purpose limitations on AI agents. VAST’s PolicyEngine (Layer 2C) is designed to address this gap — but it depends on Layer 1A’s Element-level security model as the enforcement surface. CrowdStrike integration at the data layer (vs. Dell’s perimeter-level integration) means threat detection and automated response happen where the data is, not at the network edge. For AI workloads where data is continuously accessed by agents, data-layer monitoring is architecturally more appropriate than perimeter monitoring. ## ● Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** VAST Strength ### Vendor-Provided Components **VAST InsightEngine** [DAPM: Ceded] Framework for building, managing, and automating real-time AI pipelines. Built on DataEngine using event-driven triggers and serverless functions. Processes data the moment it lands: automated chunking, embedding (via NVIDIA NIM), and storage in DataBase’s native vector store. Eliminates traditional batch processing delays — near-instant availability for AI retrieval and inference. Described as the first solution to securely ingest, process, and retrieve all enterprise data (files, objects, tables, streams) in real time. Model agnostic — compatible with any model using the OpenAI API spec. **VAST Vector Search (Native)** [DAPM: Ceded] Built directly into VAST DataBase on DASE architecture. No sharding, no memory-bound indexes. Scales linearly by adding stateless compute nodes — throughput grows without data reorganization. Vectors, metadata, and raw content live side-by-side in the Element Store. Single query path for vector search + SQL filters + metadata predicates + joins. A video RAG workflow can retrieve semantically relevant clips while filtering by time range, camera ID, location, or policy — in a single execution path. Scales from millions to trillions of vectors with consistent performance. **Permission-Aware Retrieval** [DAPM: Ceded] InsightEngine ties the permissions of vector database rows to the permissions of source data. RAG requests only return embeddings and chunks for data the requesting user is authorized to access. This is not a bolt-on ACL check — it’s structural because vectors inherit the security model of their source Elements. Access controls propagate from source files through embeddings to inference outputs. **SyncEngine (Enterprise Data Ingestion)** [DAPM: Ceded] Ingests data from enterprise systems (Google Drive, Jira, Confluence, S3 buckets, file systems) while preserving identity and access semantics. The answer to ‘what about my non-VAST data?’ — brings external enterprise data into the VAST namespace with governance intact. Maintains durable searchable index and triggers enrichment pipelines automatically on ingestion. **VAST Catalog (Discovery Surface)** [DAPM: Ceded] Metadata indexed across all data types. Queryable via Web UI, VMS API, CLI, or third-party engines. Provides the discovery and filtering surface that InsightEngine and retrieval pipelines consume. ### NVIDIA-Provided Components **NVIDIA NIM (Embedding Models)** InsightEngine triggers NVIDIA NIM embedding agents as data is written. Embeddings stored in DataBase within milliseconds. NIM provides the embedding intelligence; VAST provides the pipeline automation and storage. Model agnostic — NIM is the default but any OpenAI-compatible model works. **NVIDIA cuVS (GPU-Accelerated Search)** GPU-accelerated vector indexing and search integrated into DataBase via CNode-X. Enhances search speed but the retrieval logic, permission model, and pipeline orchestration are all VAST’s. ### Gap Analysis Layer 1B is where the architectural difference between VAST and Dell is sharpest — and where the 4+1 model’s layer boundaries are most challenged. Dell’s Layer 1B is a three-party dependency: Dell (PowerScale/ObjectScale storage + MetadataIQ metadata), Elastic (Elasticsearch 9.4 search intelligence), and NVIDIA (cuVS acceleration). Three authority boundaries, three integration seams, three vendors to coordinate for a single retrieval pipeline. VAST’s Layer 1B is a single-authority system: InsightEngine (pipeline), Vector Search (retrieval), DataBase (vector + structured storage), and Catalog (discovery) — all VAST IP running on the same Element Store. NVIDIA provides embedding models (NIM) and search acceleration (cuVS) but the retrieval logic, permission propagation, and pipeline orchestration are entirely VAST’s. The ‘data decay’ concept from VAST’s engineering is relevant to the 4+1 model: the gap between what the data IS and what the system BELIEVES the data to be widens over time in batch-indexed systems. InsightEngine addresses this by triggering embedding generation the moment new data is written — embeddings are always current with source data. Dell’s incremental indexing (ingesting only updated files) addresses the same problem but across system boundaries — MetadataIQ indexes PowerScale/ObjectScale, then the Data Search Engine re-indexes based on metadata changes. VAST’s approach is structurally tighter because the trigger, the embedding, and the vector storage share the same platform. The permission-aware retrieval is the most significant security finding at Layer 1B. Vector database rows inherit permissions from source data Elements. RAG queries only return authorized content. This is structural security (same Element Store, same permission model) rather than an overlay check. Dell’s retrieval through Elastic does not natively propagate PowerScale/ObjectScale permissions into search results — that’s a separate integration concern. SyncEngine’s enterprise data ingestion (Google Drive, Jira, Confluence) with preserved identity and access semantics is VAST’s answer to the heterogeneous enterprise. But it’s also an honest constraint: enterprise data that doesn’t enter the VAST namespace isn’t retrievable through InsightEngine. Dell’s Elastic-based search can potentially index data from more sources without requiring ingestion into Dell storage. The 4+1 Layer 1B question: does VAST expose retrieval quality observability (recall@k, latency percentiles, cache hit rates) that a Layer 2C could use for placement decisions? VAST’s integrated architecture makes this more feasible than Dell’s (InsightEngine and PolicyEngine share the same platform), but published materials don’t detail retrieval quality metrics as a first-class observable. The pipeline automation and real-time embedding are well documented; the retrieval quality feedback loop is not. ### Borrowed Judgment Low. VAST owns the retrieval logic (InsightEngine + Vector Search), the storage substrate (DataStore + Element Store), the metadata catalog (Catalog + DataBase), and the permission model (Element-level security propagated to vector rows). NVIDIA provides embedding models (NIM) and search acceleration (cuVS) but neither is required for the core retrieval function to work — InsightEngine is model agnostic and Vector Search runs on standard CNodes without GPU acceleration. Compare to Dell’s Layer 1B: Dell’s retrieval quality depends on Elastic (search intelligence), NVIDIA (acceleration), and Dell (storage + metadata). If Elastic changes Elasticsearch licensing or features, Dell’s retrieval story changes. VAST has no equivalent third-party dependency at Layer 1B. The DAPM classification is Ceded to VAST — the enterprise doesn’t own the retrieval engine. But the authority is unified in one vendor rather than split across three. For the enterprise architect, this means one vendor relationship to manage for retrieval instead of three — but total dependency on that one vendor. ### Working Notes The single-query-path architecture (vector search + SQL + metadata predicates + joins in one execution path) remains the most significant Layer 1B capability in this assessment series. No other vendor provides this without stitching together separate systems. The model-agnostic design is worth emphasizing: InsightEngine works with NVIDIA NIM by default but supports any model compatible with the OpenAI API spec. This means VAST’s Layer 1B retrieval pipeline is not locked to NVIDIA’s embedding models — unlike the Layer 2B runtime where NemoClaw/OpenShell creates a stronger NVIDIA dependency in Dell’s stack. The real-time embedding trigger (process data the moment it lands) vs. Dell’s incremental indexing (index only updated files) represents two approaches to the data currency problem. VAST’s is architecturally tighter but also more compute-intensive — every write triggers an embedding pipeline. Dell’s is more conservative on compute but introduces metadata lag. For agentic AI workloads where agents make decisions based on current data, VAST’s approach is architecturally more appropriate. For cost-constrained environments where batch indexing is acceptable, Dell’s approach is more efficient. ## ● Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** VAST Strength ### Vendor-Provided Components **VAST DataEngine** [DAPM: Ceded] Serverless compute layer executing containerized functions, triggers, and analytical engines directly on CNodes. Three core execution modes: (1) Event Triggers via Kafka-compatible Event Broker, (2) Serverless Functions as lightweight Python in stateless containers, (3) Containerized Engines for native VAST and partner workloads. Functions, events, and pipelines defined as Kubernetes custom resources. Completely event-driven — unlike centralized orchestrators (Airflow, Slurm) that poll resources, DataEngine reacts to events at any scale without bottlenecks. Co-located with data: reduces latency, simplifies management, improves security. **VAST SyncEngine** [DAPM: Ceded] Scalable data discovery and migration service within DataEngine. Indexes and synchronizes data from external sources (S3 buckets, file systems, SaaS platforms like Google Drive, Jira, Confluence) into the DataSpace. Maintains durable searchable index. Triggers enrichment pipelines automatically as new data is onboarded. Preserves identity and access semantics from source systems. **VAST Event Broker** [DAPM: Ceded] Native Kafka-compatible streaming ingestion built into the platform (not an external dependency). Stores message topics as VAST Tables for global ordering and real-time SQL querying of live streams. Events are immediately accessible as tables in the DataBase. Detects S3 object creation, tagging, and deletion events and triggers downstream workflows. Decoupled design: functions are never hardwired to data sources — the broker ensures events flow to the right consumer. **DataEngine CLI + SDK** [DAPM: Ceded] Command-line interface (vastde) for managing functions, pipelines, triggers, compute clusters, container registries. Python SDK with DataEngine context object for function development. Multiple output formats (human-readable, JSON, YAML), dry-run mode, integrated monitoring with built-in access to logs and traces. Low-code interface also available for non-developers. **Composability Architecture** [DAPM: Ceded] The layers compose into end-to-end workflows: SyncEngine discovers and onboards data → DataEngine processes via event-driven functions and Event Broker → InsightEngine contextualizes with chunking and embedding models → DataBase stores vector embeddings for querying → AgentEngine executes multi-step tasks leveraging tools and data across the ecosystem. Each stage is a DataEngine pipeline stage, not a separate system. ### NVIDIA-Provided Components **NVIDIA cuDF (via CNode-X)** GPU-accelerated data manipulation for pipeline processing. Integrated into DataEngine on CNode-X configurations. Accelerates the transformation and enrichment stages of data pipelines. **NVIDIA CUDA Libraries** CNode-X integrates NVIDIA libraries directly into DataEngine services. 44% faster queries, 80% lower costs per VAST claims. GPU acceleration is additive — DataEngine runs on standard CNodes without GPUs; CNode-X adds performance. **NVIDIA NIM (via InsightEngine)** Embedding model execution triggered by DataEngine events. NIM provides the embedding intelligence for the InsightEngine pipeline stage. Model agnostic — any OpenAI API-compatible model works. ### Gap Analysis (GA-gate) TuningEngine in the lifecycle span below is End of 2026 (see watch-list); the layer strength is the shipping DataEngine, SyncEngine, and Event Broker. VAST’s DataEngine is the most architecturally integrated Layer 1C in this assessment series. The gap analysis requires comparing three dimensions: orchestration model, data movement, and lifecycle completeness. Orchestration model: DataEngine is completely event-driven using Kubernetes custom resources. Dell’s Dataloop is a no-code/low-code engine acquired from a startup. Apache Airflow and Slurm (common alternatives) use centralized orchestrators that poll resources. DataEngine’s event-driven model avoids the scaling bottlenecks of centralized orchestration — it reacts to events rather than scheduling them. This is a meaningful architectural advantage for continuous AI pipelines where data flows never stop. Data movement: DataEngine executes co-located with data on CNodes. Dell’s Dataloop orchestrates across API boundaries between separate storage (PowerScale), search (Elastic), and compute (NVIDIA) systems. The practical difference: VAST’s pipelines don’t move data between systems because storage and compute share the same platform. Dell’s pipelines necessarily move data between PowerScale/ObjectScale and the GPU cluster. Dell’s KV Cache offload (NVIDIA CMX, 19x TTFT improvement) is a sophisticated answer to this data movement problem — but it’s solving a problem that VAST’s architecture doesn’t have. Lifecycle completeness: DataEngine + SyncEngine + Event Broker + InsightEngine + AgentEngine + TuningEngine spans from data ingestion through embedding through agent execution through model improvement. Dell spans from data orchestration (Dataloop) through search (Elastic) through analytics (Starburst) — but the agent runtime (NemoClaw) and model lifecycle are separate NVIDIA-owned layers. VAST’s lifecycle is more complete but less proven. Dell’s lifecycle has more mature individual components but more seams. The composability architecture is worth noting for the 4+1 model: VAST’s layers (SyncEngine → DataEngine → InsightEngine → DataBase → AgentEngine) compose as pipeline stages within a single platform. The 4+1 model defines these as separate layers (1A, 1B, 1C, 2B). VAST’s architecture challenges the layer separation by making the boundaries internal to one platform rather than external between systems. The layers still exist functionally, but the authority boundaries don’t — VAST owns all of them. ### Borrowed Judgment Low. DataEngine, SyncEngine, Event Broker, composability architecture, and the DataEngine CLI/SDK are all VAST IP. TuningEngine (end of 2026) will be VAST IP. NVIDIA provides GPU acceleration (cuDF, CUDA, NIM) but the orchestration logic, event processing, pipeline management, and lifecycle coordination are entirely VAST’s. Compare to Dell’s Layer 1C: Dell owns the Dataloop orchestration engine (Retained — its strongest software move) but depends on NVIDIA for acceleration (cuDF, CMX for KV cache), Starburst for analytics, and NVIDIA Blueprints/NIMs for pipeline templates. Dell’s Layer 1C has four authority boundaries: Dell (Dataloop), NVIDIA (acceleration + templates), Starburst (analytics), and the customer (pipeline configuration). VAST’s Layer 1C has one: VAST. The Kubernetes foundation is a shared dependency: both Dell and VAST depend on K8s. But VAST embeds K8s custom resources into its platform (functions, events, pipelines as CRDs), while Dell depends on external K8s distributions (Red Hat OpenShift AI, Canonical Ubuntu). VAST’s K8s dependency is internal; Dell’s is external. ### Working Notes The DataEngine CLI (vastde) and Python SDK represent a developer experience that Dell’s Dataloop doesn’t yet match in published documentation. The CLI manages functions, pipelines, triggers, compute clusters, and container registries with dry-run mode and integrated observability. Dell’s Dataloop is positioned as no-code/low-code — targeting data engineers. VAST’s DataEngine targets both low-code users and developers/MLOps engineers with full Python SDK access. Different audience emphasis. The built-in observability (logs, traces within the same UI, without external Grafana or tracing infrastructure) is an operational differentiator. Dell’s observability for data pipelines depends on partner tools and external monitoring stacks. The TuningEngine + PolicyEngine combination represents VAST’s ‘thinking machine’ vision: systems that observe, reason, act, evaluate, and improve automatically. This is the most ambitious lifecycle claim from any vendor in this assessment series. Dell’s equivalent would require assembling Dataloop (orchestration) + NVIDIA NeMo (model training) + NemoClaw (agent execution) + manual feedback loops — four systems from three vendors with no automated improvement loop. Watch-list (pending GA, not scored): VAST TuningEngine — model-tuning automation, End of 2026. 1C strength stands on DataEngine/SyncEngine/Event Broker (GA). ## ◑ Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** Partial ### Vendor-Provided Components **Polaris Control Plane** [DAPM: Ceded] Global control plane purpose-built for AI data infrastructure spanning public cloud, neocloud, and on-prem. Kubernetes-based architecture with lightweight agent on every VAST node. Intent-driven: administrators define desired state, Polaris coordinates cloud-native services to achieve and maintain it. Automates provisioning, cloud marketplace integration (subscription/entitlement), centralized upgrade orchestration, expansion, and node replacement. Multi-cluster: converts distributed infrastructure into a single operational platform. Available as VAST-managed, partner-managed, or customer-managed. **Polaris Enterprise Controls** [DAPM: Ceded] Enterprise identity integration, role-based access control, and audit logging. Cloud-style operational consistency across hybrid and multicloud environments. Supports sovereign deployments. Multi-tenant by design. **DataEngine Built-in Scheduler** [DAPM: Ceded] Built-in scheduler and cost-optimizer within DataEngine. Deploys serverless functions on CPU, GPU, and DPU architectures. Manages function lifecycle, container orchestration, and resource allocation. Event-driven scheduling rather than centralized job queuing. **Polaris + DataSpace Coordination** [DAPM: Ceded] Polaris abstracts INFRASTRUCTURE location (where clusters are deployed and maintained). DataSpace abstracts DATA location (how data is presented across locations via global namespace). Together they provide the ‘where should this run relative to this data’ coordination that is a prerequisite for Layer 2C. ### NVIDIA-Provided Components **NVIDIA GPU Operator (on CNode-X)** GPU lifecycle management for containerized VAST services running on CNode-X. Manages driver installation, device plugins, monitoring for the GPUs embedded in the data platform. **GPU Scheduling Scope Distinction** CNode-X GPUs accelerate DATA PLATFORM services (vectorization, SQL via Sirius/cuDF, vector search via cuVS, inference within InsightEngine). They are NOT general-purpose GPU compute for external training/inference workloads. VAST does not provide Run:ai-equivalent fair-share GPU scheduling because the use case is different — GPU resources serve the platform’s own services, not multi-tenant external workloads competing for GPU time. ### Gap Analysis Layer 2A is where the architectural models diverge most sharply between Dell and VAST, and where the 4+1 model’s layer definitions need careful application. Dell’s Layer 2A problem: GPU-aware workload scheduling for multi-tenant inference and training. Multiple teams compete for scarce GPU resources. Run:ai provides fair-share scheduling, quotas, and RBAC. Dell has no proprietary capability here. VAST’s Layer 2A reality: GPU resources in CNode-X serve the data platform’s own services (vector search, SQL acceleration, embedding generation, data pipeline processing). These are not multi-tenant external workloads competing for GPU time — they are platform services running on platform-embedded GPUs. The scheduling question is different: VAST’s DataEngine scheduler allocates serverless function execution across CNodes, not GPU time-slices across competing users. Polaris provides genuine Layer 2A capability that Dell lacks entirely: intent-driven, multi-cluster, multi-cloud infrastructure provisioning and lifecycle management. Dell’s OpenManage manages a single rack. Polaris manages a fleet of VAST deployments across geographies, clouds, and on-prem sites as one system. The three management modes (VAST-managed, partner-managed, customer-managed) provide operational flexibility that has no Dell equivalent. However, if the enterprise deploys separate GPU clusters for training/inference (not embedded in VAST), those clusters still need GPU-aware scheduling. VAST doesn’t provide this — the customer would need NVIDIA Run:ai or equivalent for the GPU compute layer that sits alongside (not within) the VAST data platform. In a typical deployment, VAST handles the data plane (storage, retrieval, orchestration), and a separate GPU cluster (Dell PowerEdge, HPE ProLiant, etc. with Run:ai) handles training/inference. Polaris manages the VAST fleet; something else manages the GPU cluster. The Polaris + DataSpace coordination is the most interesting Layer 2A finding: Polaris abstracts infrastructure location while DataSpace abstracts data location. This separation of concerns — managing where infrastructure runs vs. managing where data lives — is exactly the architectural prerequisite for Layer 2C placement reasoning. No other vendor in this assessment separates these abstractions as cleanly. ### Borrowed Judgment Low for infrastructure orchestration (Polaris is VAST IP). Moderate for GPU-specific scheduling (depends on NVIDIA GPU Operator for CNode-X, and the customer still needs Run:ai or equivalent for separate GPU clusters). Compare to Dell: Dell’s Layer 2A has HIGH borrowed judgment because all GPU-aware orchestration is NVIDIA-controlled (GPU Operator, Run:ai, AI Enterprise, MIG/MPS). VAST’s Layer 2A has MODERATE borrowed judgment because Polaris handles infrastructure orchestration (VAST IP) while GPU scheduling is a narrower dependency limited to CNode-X management. The key distinction: Dell NEEDS Run:ai because GPU scheduling is the core Layer 2A function in Dell’s architecture. VAST’s Polaris handles the broader infrastructure orchestration function, and GPU scheduling is a specific sub-function for CNode-X hardware management, not the defining Layer 2A capability. Polaris is included in VAST AI OS at no additional charge. Dell’s customers pay separately for NVIDIA Run:ai and AI Enterprise licenses. This pricing distinction reflects the authority distinction: VAST bundles infrastructure orchestration because it owns it; Dell passes through NVIDIA licensing because it doesn’t. ### Working Notes The three management modes (VAST-managed, partner-managed, customer-managed) have DAPM implications: • VAST-managed: fully Ceded — VAST operates the infrastructure • Partner-managed: Delegated to the partner (CSP, MSP) • Customer-managed: the enterprise operates Polaris but the software is still VAST’s Even in customer-managed mode, the control plane software is VAST IP. The enterprise operates it but doesn’t own it. This is analogous to running VMware vCenter on your own hardware — you operate it, but the software authority belongs to the vendor. Polaris is available now as part of VAST cloud deployments, with expanded multi-cluster orchestration planned in future releases. The ‘expanded multi-cluster orchestration’ roadmap item is worth tracking — this is where Polaris evolves from infrastructure management (Layer 2A) into placement reasoning (Layer 2C). 89% of APAC enterprises deploy workloads across multiple public clouds; 72% operate hybrid cloud models (cited by VAST). This validates Polaris’s multi-cloud orchestration thesis. ## ● Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** VAST Strength — AgentEngine (GA) ### Vendor-Provided Components **VAST AgentEngine** [DAPM: Ceded] AI agent deployment and orchestration system running natively within the DataEngine. Described as ‘the application management layer of the VAST AI OS, designed specifically for the Agentic AI Era’ and ‘the final piece of the puzzle, rounding out the core services needed to run agentic applications on AI hardware.’ Low-code runtime that simplifies programming and coordinates multi-agent workflows, model invocation, and tool usage. Brings agents to life directly within DataEngine — reasoning occurs where the data lives. **Agent Runtime Capabilities** [DAPM: Ceded] Robust runtime for long-running containerized services. Lifecycle management for agent deployment, scaling, and versioning. Persistent state via execution-scoped or persistent scratch space spanning files, objects, or tables. Secure service discovery for agents to find and communicate with other agents and tools. Support for agents with multiple personas and security credentials. Fault-tolerant backbone with Kafka-backed queuing for durability and ordering. **MCP Toolbox** [DAPM: Ceded] Model Context Protocol toolbox enabling agents to orchestrate multiple tools and services together for higher-order workflows. Composability lets teams build complex agentic applications without stitching fragile systems. Aligns with the open MCP standard (donated to Linux Foundation Dec 2025). **Permission-Aware Execution** [DAPM: Ceded] Each agent call is defined with the user’s identities and permissions. Tools and data are accessed responsibly under the same permission model that governs the underlying Element Store. Multi-tenant operation with strong governance controls. This is structural — agent permissions inherit from the data platform’s security model, not from a bolt-on ACL layer. **Observability & Compliance** [DAPM: Ceded] Logs and traces every agent action, tool call, and data access. End-to-end observability creates transparency for compliance, stakeholder trust, and performance refinement. Built into the same platform — no external Grafana, tracing systems, or observability infrastructure required. **Data Event Integration** [DAPM: Ceded] Every add, move, change, or delete in storage emits an event that feeds into agent workflows via DataEngine. Ties data operations directly to agentic AI actions in real time. Agents can react to data changes as they happen — no polling, no batch triggers. ### NVIDIA-Provided Components **NVIDIA NIM Containers (on CNode-X)** VAST can serve NVIDIA NIM inference containers on CNode-X infrastructure, providing access to NVIDIA’s model ecosystem. NIM and AgentEngine are not mutually exclusive — NIM provides optimized model serving; AgentEngine provides the orchestration and lifecycle management around it. **NVIDIA CUDA (Inference Acceleration)** Model inference on CNode-X is GPU-accelerated via CUDA. VAST provides the orchestration and governance; NVIDIA provides the compute acceleration. This separation is cleaner than Dell’s stack where NVIDIA owns both orchestration (NemoClaw) and acceleration. ### Gap Analysis VAST AgentEngine is the most significant Layer 2B finding in the comparative analysis and the starkest architectural contrast with Dell’s approach. Dell’s Layer 2B: Dell does not appear to own the core agent runtime. NemoClaw provides the agent execution stack (NVIDIA). OpenShell provides the sandboxed runtime (NVIDIA). NeMo Guardrails provide safety constraints (NVIDIA). Cohere North provides agent workflow orchestration (ISV partner). DataRobot provides agent lifecycle management (ISV partner). Dell provides hardware, packaging, and professional services. Five authorities for one layer. VAST’s Layer 2B: AgentEngine provides agent deployment, orchestration, runtime, lifecycle management, persistent state, MCP toolbox, permission-aware execution, observability, and data event integration — all within a single platform. One authority for the entire layer. The architectural differentiators: 1. Agents execute where the data lives. In VAST, there is no data movement between storage and inference runtime because they are the same platform. In Dell’s architecture, data flows from PowerScale/ObjectScale through NVIDIA’s inference runtime, crossing authority and system boundaries. 2. Permission-aware by structure. Agent permissions inherit from the Element Store’s security model. Dell’s agent permissions depend on NeMo Guardrails (runtime constraint) and the ISV partner’s governance logic (Cohere North, DataRobot). 3. Real-time data event triggers. Every storage operation emits events that feed into agent workflows. Dell’s agents consume data through retrieval (Elastic search) and pipeline (Dataloop) stages, not through direct data event feeds. 4. Built-in observability. Logs and traces every agent action without external monitoring infrastructure. Dell’s observability depends on partner tools. 5. MCP Toolbox. Native MCP support for agent-to-tool orchestration. Aligns with open standards. Dell’s agentic platform uses ISV-specific tool orchestration (Cohere North, DataRobot SDKs). Maturity is a watch-item, not a capability discount today: AgentEngine is generally available and a complete, owned agent runtime — which is what the score reflects — and proven-ness at enterprise scale may become a differentiator as the agentic market matures. AgentEngine is newer and less proven than NVIDIA’s NemoClaw/OpenShell stack. NVIDIA’s open-source foundation (OpenClaw), broad model support (Nemotron, NIM), and ecosystem validation (Dell, HPE, Lenovo, CSPs) represent a larger installed base. Jensen’s ‘operating system for personal AI’ framing signals NVIDIA’s long-term commitment to the runtime layer. VAST’s AgentEngine must prove equivalent reliability, security, and model compatibility at enterprise scale. The commoditization signal from Augment Code’s 2026 multi-agent orchestration analysis: ‘The runtime layer (tool registries, state management, retry logic) is being commoditized by open standards and hyperscaler investment. Teams building custom implementations will find platform solutions commoditizing that work within 12–18 months.’ This applies to both VAST AgentEngine and NVIDIA NemoClaw — the question is whether the runtime will differentiate or whether the governance layer above it (Layer 2C) becomes the battleground. ### Borrowed Judgment Low for agent orchestration, lifecycle, and governance (AgentEngine, MCP Toolbox, permission model, observability are all VAST IP). Moderate for model inference (depends on NVIDIA CUDA acceleration on CNode-X). NIM containers can run on the platform but are not required — AgentEngine is model-agnostic. Compare to Dell: Dell’s Layer 2B borrowed judgment is TOTAL for the runtime (NemoClaw/OpenShell/NIM/Dynamo/AI Enterprise are all NVIDIA) and DELEGATED for orchestration (Cohere North/DataRobot/ClearML). Dell’s one Retained asset is professional services. VAST’s Layer 2B authority structure: VAST owns the runtime, the orchestration, the lifecycle management, the permission model, and the observability. NVIDIA provides inference acceleration. The enterprise Cedes to one vendor (VAST) rather than to three+ vendors (NVIDIA + Cohere + DataRobot + ClearML in Dell’s case). The DAPM summary: both Dell and VAST require the enterprise to Cede Layer 2B authority. The difference is whether you Cede to a fragmented set of authorities (Dell’s model) or to a unified authority (VAST’s model). The 4+1 model doesn’t prescribe which is better — it requires the enterprise architect to make the choice explicitly. ### Working Notes The positioning contrast is telling: NVIDIA says NemoClaw is ‘the operating system for personal AI.’ VAST says AgentEngine is ‘the application management layer of the AI Operating System.’ NVIDIA positions the agent runtime as an OS — the platform everything else runs on. VAST positions it as a layer within a larger OS that includes storage, retrieval, governance, and infrastructure orchestration. The difference: NVIDIA’s vision is runtime-centric (the agent runtime IS the platform). VAST’s vision is data-centric (the data platform includes the agent runtime as one of its services). For the Enterprise AI Control Plane working document: AgentEngine + PolicyEngine (Layer 2C) + the Element Store permission model (Layer 1A) form a coherent governance chain within VAST’s stack. In Dell’s stack, the equivalent governance chain would be: NeMo Guardrails (NVIDIA) + ???(no Layer 2C) + Trust3 AI (partner) + MetadataIQ (Dell) — four authorities, one missing layer. ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Gap — placement abstraction, not a reasoning plane ### NVIDIA-Provided Components **NVIDIA NeMo Data Designer (via TuningEngine)** TuningEngine integrates with NeMo Data Designer for training and fine-tuning Nemotron models. This is NVIDIA’s model training framework accessed through VAST’s tuning pipeline — NVIDIA provides the training technology, VAST provides the orchestration and governance around it. **No NVIDIA Layer 2C Governance Dependency** PolicyEngine, Polaris, and DataSpace are VAST IP. NVIDIA does not provide or control the governance, placement, or policy reasoning layer. The NeMo Data Designer integration is a training tool dependency, not a governance dependency. This is structurally different from Dell, where the closest Layer 2C functions (Dynamo routing, NeMo Guardrails) are NVIDIA-owned. ### Gap Analysis Deployable today, this is a gap. Polaris (placement abstraction) and DataSpace (namespace) ship, but placement abstraction without a policy-reasoning engine is routing, not reasoning - and PolicyEngine, TuningEngine, and the closed loop are announced for End of 2026, not shipping (see watch-list). What follows describes the architecture VAST is building toward, dated end-2026; it is not an option an architect can deploy today. This is the most significant finding in the comparative assessment. Layer 2C is where the 4+1 model’s thesis is tested most directly: does any vendor provide policy-driven placement decisions across models, data, agents, and infrastructure? Applying the ‘Routing Is Not Reasoning’ test to VAST’s stack: PolicyEngine goes BEYOND constraint enforcement (what NeMo Guardrails does at Layer 2B). Guardrails say ‘the agent cannot do X.’ PolicyEngine says ‘the agent can do X only if conditions Y and Z are met, as determined by AI-derived context, and the action will be logged with tamper-proof traceability.’ This is active policy reasoning, not static constraint enforcement. The pre-execution enforcement model means governance decisions happen BEFORE actions execute, not after — a fundamentally different posture than post-hoc audit. Polaris goes BEYOND infrastructure provisioning (what Kubernetes does at Layer 2A). Polaris doesn’t just deploy clusters — it abstracts infrastructure location so that workloads can be placed based on GPU availability and compliance requirements without changing application behavior. Combined with DataSpace abstracting data location, Polaris + DataSpace provide the multi-variable placement reasoning that the 4+1 model defines as Layer 2C. TuningEngine goes BEYOND model serving (what NemoClaw does at Layer 2B). It manages the continuous improvement of models based on production outcomes, governed by PolicyEngine. This closes the loop from execution through evaluation through improvement through re-governance — no other vendor has this cycle in a single platform. The multi-variable test from the 4+1 model: can the system make placement decisions based on cost + compliance + latency + data residency + model capability simultaneously? • Cost: DataEngine’s built-in cost-optimizer can deploy on CPU/GPU/DPU based on cost • Compliance: PolicyEngine enforces permissions and audit requirements before execution • Latency: DataSpace provides remote data caches for geo-distributed access • Data residency: DataSpace’s per-site consistency management respects data locality • Model capability: TuningEngine manages model versions and fitness for purpose VAST is the only vendor in this series that addresses all five variables, even if the integration between them is not yet fully productized (PolicyEngine and TuningEngine ship end of 2026). Caveats remain critical: • PolicyEngine and TuningEngine are announced, not shipped. GA end of 2026. • Polaris is available but multi-cluster placement orchestration is in ‘expanded capabilities planned in future releases.’ • The ‘AI-derived context’ in PolicyEngine is described but the AI decision-making model is not detailed — how the AI determines permissions from context is a black box. • The closed loop (observe-reason-act-evaluate-improve) is architecturally described but not production-validated at enterprise scale. Compare to Google Cloud: Inference Gateway + DWS + Knowledge Catalog is productized and shipping. Google’s Layer 2C is narrower (inference placement and scheduling) but real. VAST’s is broader (governance + placement + learning) but pre-GA. Compare to Dell: No productized Dell-owned Layer 2C is evident. Dell has the Dell + Intel ‘control plane’ signal but no product. VAST is building what Dell hasn’t started. Compare to Kamiwaza: Policy-driven Inference Mesh + Distributed Data Engine + ReBAC is the most comparable Layer 2C approach — explicit multi-variable policy optimization for inference placement. Kamiwaza and VAST are approaching from different directions (Kamiwaza from inference, VAST from data) toward the same Layer 2C function. ### Borrowed Judgment Low — the lowest Layer 2C borrowed judgment score in the assessment series because VAST is the only vendor building a proprietary Layer 2C. PolicyEngine, Polaris, DataSpace, and TuningEngine are all VAST IP. The NVIDIA dependency is limited to NeMo Data Designer for model training within TuningEngine — a training tool, not a governance dependency. The DAPM classification is Ceded to VAST. The enterprise doesn’t own the control plane — VAST does. But the 4+1 model’s DAPM framework reveals an important distinction: • Dell at Layer 2C: ABSENT — no authority exists to Cede or Retain. The enterprise operates without governance. • VAST at Layer 2C: CEDED — authority exists and is Ceded to VAST. The enterprise has governance but doesn’t own it. • Google at Layer 2C: CEDED — authority exists and is Ceded to Google. Productized and shipping. Absent is worse than Ceded. Having governance you don’t own is better than having no governance at all. The enterprise architect’s decision is not ‘should we have Layer 2C?’ (the 4+1 model says yes). It’s ‘who should hold the Layer 2C authority?’ VAST’s answer: VAST holds it, as part of a vertically integrated platform where governance, execution, data, and infrastructure are one system. Dell’s non-answer: nobody holds it. Build it yourself or operate without it. Google’s answer: Google holds it, as part of a cloud platform where the enterprise Cedes everything. ### Working Notes Analyst and media framing validates the Layer 2C classification: • Blocks and Files: ‘VAST broadens AI platform push with control plane’ — explicitly using control plane language • Moor Insights: Polaris positions VAST as ‘the persistent operational layer for AI’ • theCUBE Research (Strechay): ‘VAST looks at it as going up the stack’ • TipRanks: PolicyEngine + TuningEngine transform the AI OS into a ‘consolidated stack merging data, compute, policy enforcement, and model training into one infrastructure fabric’ • SiliconANGLE: Polaris is ‘complementary to DataSpace — DataSpace focuses on how data is presented across locations; Polaris focuses on how clusters are deployed and maintained’ This validates the Enterprise AI Control Plane working document’s Pattern 4: Storage Vendors Reaching Up. VAST is the most aggressive example. The three-vector convergence: • Dell: bottom-up (infrastructure OEM, hasn’t started Layer 2C, Dell+Intel signal only) • Google: top-down (cloud provider, Inference Gateway + DWS + Knowledge Catalog, productized) • VAST: middle-out (data platform, PolicyEngine + Polaris + DataSpace + TuningEngine, emerging) The Microsoft Agent Governance Toolkit (April 2026) and the AgentGuardian research paper validate that agent governance is becoming a recognized engineering discipline, not just a VAST-specific product category. The 4+1 model’s Layer 2C aligns with this emerging discipline. Watch-list (pending GA, not scored): PolicyEngine, TuningEngine, and the Closed Loop are all End of 2026 - VAST's policy-reasoning 2C is architected and dated but not shipping. Deployable today: Polaris (placement abstraction) plus DataSpace (namespace) - routing and plumbing, not a policy-reasoning plane (routing is not reasoning), so the layer is a gap until PolicyEngine ships. ## ◑ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Focused Ecosystem ### Vendor-Provided Components **VAST Cosmos Community (Unified Partner Program)** [DAPM: Delegated] Global partner ecosystem formalized at VAST Forward 2026 with distinct partner tracks: Software Partners (ISVs) build integrations and validated solutions on the AI OS. Hardware/Platform Partners (Cisco, Supermicro, HPE, Lenovo) deliver infrastructure foundations. Cloud Partners (CSPs, neoclouds, hyperscalers). Channel Partners (resellers, SIs, advisory). Developer Community with technical resources, learning pathways, hands-on labs, and contribution opportunities for integrations and blueprints. **TwelveLabs Partnership (Video AI)** [DAPM: Delegated] First customer-managed deployment path for TwelveLabs’ video foundation models (Marengo for embeddings/search, Pegasus for deep video understanding). Extends video intelligence beyond public cloud into on-prem and sovereign environments. Target: media companies, financial services (surveillance-based fraud detection), government agencies (data sovereignty). TwelveLabs gains on-prem deployment; VAST gains a compelling vertical anchor use case. **CrowdStrike Integration (AI Lifecycle Security)** [DAPM: Delegated] Strategic partnership connecting VAST AI OS telemetry to CrowdStrike Falcon platform. Coordinated detection across data ingestion, model training, and runtime inference environments. Deeper than perimeter security — integrated into PolicyEngine with event-driven automation for threat detection and response. Unified policy management, agent governance, encryption, and regulatory reporting. **CoreWeave (Hyperscale Anchor)** [DAPM: Delegated] $1.17B commercial agreement. CoreWeave’s EVP of Product and Engineering presented at VAST Forward 2026 on operating at thousands-of-GPU scale. CoreWeave validates that the VAST platform handles the data coordination requirements of hyperscale AI: predictable data movement, reuse/staging without latency spikes, and consistent flow that enables schedulers to make reliable decisions. The constraint at CoreWeave scale: ‘adding more GPUs doesn’t recover performance if data flow breaks — it amplifies the inefficiency.’ **NVIDIA NIM + Open Models (on VAST)** [DAPM: Delegated] VAST can serve NVIDIA NIM containers and open models on its platform via AgentEngine and CNode-X. Provides access to NVIDIA’s model ecosystem without requiring NemoClaw/OpenShell. AgentEngine is the primary runtime surface; NIM is an optional model delivery mechanism. ### NVIDIA-Provided Components **NVIDIA Model Ecosystem Access** CNode-X and AgentEngine provide the substrate for running NVIDIA NIM containers, Nemotron models, and NVIDIA Blueprints. VAST’s model-agnostic architecture means NVIDIA models are one option among many, not the mandatory runtime path. ### Gap Analysis VAST’s Layer 3 ecosystem is structurally different from Dell’s — not just smaller. Dell’s ecosystem is broad and horizontal: OpenAI (dev productivity), Palantir (operational AI), Google (sovereign compute), ServiceNow (workflow automation), SpaceXAI (enterprise assistant), Hugging Face (model hub), Mistral, Reflection, Poolside, UneeQ, Fogsphere. 5,000+ deployment customers. Dell’s ecosystem breadth compensates for infrastructure gaps — ISV partners provide Layer 2B functions (Cohere North, DataRobot, ClearML) that Dell’s platform lacks. VAST’s ecosystem is focused and vertical: CoreWeave (hyperscale validation), TwelveLabs (video AI), CrowdStrike (AI lifecycle security), plus the Cosmos Community structure for SIs, CSPs, and ISVs. VAST’s ecosystem depth reflects platform self-sufficiency — ISV partners provide application use cases, not infrastructure functions, because the platform handles Layers 1A through 2C internally. The Dell comparison at Layer 3 reveals the DAPM trade-off between the two architectures: • Dell’s Layer 3 partners provide both application logic AND infrastructure-level functions (agent orchestration, governance, GPU scheduling). The partner ecosystem is load-bearing — remove Cohere North and Dell loses agent workflow orchestration. • VAST’s Layer 3 partners provide application logic ONLY. The partner ecosystem is additive — remove TwelveLabs and VAST loses a vertical use case but not a platform capability. This is the cleanest DAPM distinction in the comparative assessment: Dell’s ecosystem is structurally necessary. VAST’s ecosystem is strategically valuable. The ecosystem risk is real: enterprise buyers evaluate ecosystem breadth as a proxy for platform maturity. Dell’s partner roster with OpenAI, Google, Palantir, and ServiceNow signals broad market acceptance and Fortune 500 validation. VAST’s smaller roster signals focused applicability — strong in hyperscale and data-intensive verticals (media, financial services, government), less proven in general enterprise workflows. The CoreWeave validation deserves specific attention: at thousands-of-GPU scale, CoreWeave’s constraint is not compute but data coordination. ‘Adding more GPUs doesn’t recover performance if data flow breaks — it amplifies the inefficiency.’ This is a direct validation of the 4+1 model’s thesis that the data plane (Layers 1A/1B/1C) is the binding constraint for AI at scale, not the compute plane (Layer 0). ### Borrowed Judgment Distributed across partners at Layer 3, which is architecturally correct. VAST’s structural advantage: because the platform provides its own Layers 1A through 2C, Layer 3 partners bring only application logic and vertical use cases. They don’t need to bring infrastructure capabilities, and they don’t need to fill platform gaps. Compare to Dell: Dell’s Layer 3 borrowed judgment is distributed across partners who provide BOTH application logic AND infrastructure-level functions. The distinction is whether partners are additive (VAST) or load-bearing (Dell). The CrowdStrike integration is worth classifying separately: it operates at both Layer 3 (application-level security) and Layer 1A (data-layer monitoring). The integration connects AI OS telemetry to CrowdStrike Falcon for coordinated detection across the full AI lifecycle. This crosses layer boundaries — appropriate because security is a cross-cutting concern, not a single-layer function. ### Working Notes The $30B valuation, $4B+ bookings, and $500M+ CARR signal investor confidence. The CoreWeave $1.17B agreement anchors the customer base at hyperscale. But the enterprise deployment base is smaller than Dell’s 5,000+ and skews toward neoclouds and data-intensive verticals rather than traditional enterprise IT. The developer community track within Cosmos (hands-on labs, learning pathways, blueprint contributions) parallels Dell’s Enterprise Hub on Hugging Face but with a different emphasis: Dell’s Hub is about model access; VAST’s Cosmos is about building on the platform. The TwelveLabs partnership demonstrates the ‘customer-managed deployment’ model that is becoming the pattern for AI model companies moving from cloud-only to hybrid: the model runs on customer infrastructure (VAST AI OS) rather than the model company’s cloud. Dell’s equivalent is OpenAI Codex connecting to the Dell AI Data Platform and SpaceXAI Grok on-premises. The pattern is the same; the deployment substrate is different. ════════════════════════════════════════════════════════════════════════════════ # VMware Private AI Foundation with NVIDIA Mapped to the 4+1 Layer AI Infrastructure Model **Version:** v1.4 - Interface-Portability Reconciliation **Date:** June 20, 2026 **Source:** VMware Explore 2025, VCF 9.0/9.1 announcements, Broadcom press releases, VCF Private AI blog series, published 4+1 model. v1.4 (instrument reconciliation): 1B Vector Database (pgvector) Ceded→Delegated (OSS, opinions lift via pg_dump); 2A vSphere Supervisor + VKS Ceded→Delegated (conformant K8s, consistent with cloud managed-K8s treatment). ## Summary Finding VMware Private AI Foundation with NVIDIA occupies a structurally unique position in this assessment series: it is neither an infrastructure OEM (Dell, HPE), a hyperscaler (AWS, Google Cloud), nor a data platform vendor (VAST). It is a virtualization and private cloud platform — the abstraction layer that sits between physical infrastructure and workloads. Broadcom’s strategic thesis is that VCF is ‘the permanent abstraction layer between AI software and physical chips,’ and the Private AI Foundation extends that thesis into AI workloads specifically. The 4+1 model reveals both the power and the limits of this position. VCF’s strength is Layer 2A — infrastructure orchestration is VMware’s heritage and its deepest IP. VCF Automation, vSphere Supervisor, VKS, vSAN, NSX/vDefend, and VCF Operations collectively provide the most mature unified orchestration surface for mixed workloads (VMs, containers, AI) of any on-prem vendor assessed. No other vendor in this series manages GPU-accelerated AI workloads, Kubernetes clusters, and traditional VMs from a single control plane with equivalent operational maturity. At Layer 0, VMware’s multi-accelerator management (AMD, NVIDIA, Intel) requires careful contextualization. GPU vendor choice is not unique to VMware — HPE’s GX5000 supports NVIDIA and AMD blades in the same rack, and hyperscalers fully abstract accelerators at the service layer (a developer calling Vertex AI or Bedrock never sees which silicon powers the response). VMware’s actual differentiator is the level of architectural control: operators manage GPU placement, isolation, and scheduling through familiar vSphere primitives (vGPU profiles, vmclasses, DRS, resource pools). The control plane is borrowed judgment — it is VMware’s opinionated virtualization model applied to acceleration — but that opinion provides stronger knobs that appeal to operators already comfortable with virtualization management. Where hyperscalers abstract the accelerator away from the architect, VMware puts the architect in the driver’s seat through a familiar console. But the closer the stack gets to AI-specific functions — model serving, retrieval, agent execution, governance — the more authority shifts to NVIDIA (Layer 2B runtime via NVIDIA AI Enterprise), to open-source components (pgvector, Elasticsearch), or to capabilities that are emerging but not yet at the depth of purpose-built alternatives. Private AI Services (Model Runtime, Agent Builder, Data Indexing/Retrieval, Vector Database, Model Store) are genuine platform capabilities delivered as part of the VCF subscription, but they are foundational AI services, not the deep data lifecycle or agent orchestration that Dell (Dataloop), HPE (Ezmeral/Kamiwaza), or VAST (DataEngine/AgentEngine) provide. Layer 1C (data pipelines) is absent. Layer 2C (reasoning plane) has building blocks — MCP Server Governance, GPU/Model Metrics, Intelligent Assist — but none passes the ‘Routing Is Not Reasoning’ test. The installed base is the strategic moat: nine of the top ten Fortune 500 companies have committed to VCF, with 100M+ cores licensed worldwide. For the enormous VMware installed base, Private AI Foundation is the lowest-friction path to on-prem AI — no new infrastructure vendor, no new management plane, no new operational model, and no incremental cost beyond GPU hardware. The 4+1 question is whether lowest-friction adoption translates to sufficient architectural depth when agentic AI workloads demand governance, policy-driven placement, and cross-agent orchestration that VCF does not yet provide. VMware Private AI Foundation is the enterprise’s most natural on-ramp to private AI. Hock Tan’s ‘permanent abstraction layer’ framing is a Layer 2A statement, not a Layer 2C statement — the abstraction layer manages resources; the reasoning plane governs them. VMware has the former; it does not have the latter. Whether it becomes the enterprise’s durable AI platform depends on whether Broadcom invests in the Layer 1A governance depth, Layer 1C pipeline capability, and Layer 2C reasoning plane that the 4+1 model identifies as structurally necessary — or whether the Broadcom acquisition thesis (cash generation from the installed base) constrains that investment. VMware’s unique structural advantage is that VCF sees everything from the hypervisor up across all OEM hardware — the data to build a multi-vendor reasoning plane exists. The engineering commitment does not yet. ## ◑ Layer 0 · Compute: Compute & Network Fabric *Raw compute, networking, and acceleration fabric* **Status:** Hardware-Agnostic Abstraction ### Vendor-Provided Components **VMware vSphere 8/9 (Hypervisor)** [DAPM: Ceded] Industry-standard virtualization layer. GPU passthrough and vGPU support via NVIDIA AI Enterprise integration. vSphere Supervisor manages both VMs and Kubernetes workloads from a single control plane. vMotion for live migration of AI workloads. DRS for automated load balancing. The hypervisor is Broadcom’s foundational IP — the abstraction layer between physical infrastructure and all workloads above. **Multi-Vendor Hardware Support** [DAPM: Retained] VCF runs on Dell PowerEdge, HPE ProLiant, Lenovo ThinkSystem, Cisco UCS, Supermicro, NEC, Fujitsu, and others. Hardware-agnostic by design — VCF does not manufacture or specify compute. This is the fundamental architectural difference from Dell AI Factory or HPE Private Cloud AI: VMware abstracts hardware, OEMs provide it. The enterprise retains hardware vendor choice. **NVIDIA GPU Integration (vGPU + Passthrough)** [DAPM: Delegated] VCF 9.1 supports NVIDIA Blackwell architecture, NVSwitch on HGX platform, GPUDirect RDMA over InfiniBand for distributed LLM inference across multiple HGX servers. Enhanced DirectPath I/O for ConnectX-7 NICs and BlueField-3 DPUs. vGPU profiles mapped to vSphere Namespaces as vmclasses for multi-tenant GPU isolation — memory strictly partitioned per tenant with zero side-channel risk across GPU framebuffer. **NSX / vDefend Networking** [DAPM: Ceded] Software-defined networking with micro-segmentation, Zero Trust enforcement via Distributed Firewall (Antrea CNI for Kubernetes), and in-memory malware defense. Avi Load Balancer provides virtualized load balancing for AI inference endpoints and agentic applications — eliminates hardware appliance requirements. Post-quantum cryptography support. **Multi-Accelerator Management (AMD + NVIDIA + Intel)** [DAPM: Ceded] VCF 9.1 manages AMD, NVIDIA, and Intel accelerators through the same virtualization control plane — vGPU profiles, vmclasses, DRS policies, resource pools. The architect retains granular control over GPU placement, isolation, and scheduling using familiar vSphere primitives. This is not unique as a multi-vendor GPU capability (HPE GX5000 supports NVIDIA Rubin + AMD MI430X in the same rack; hyperscalers fully abstract accelerators at the service layer). VMware’s differentiator is the level of architectural control: operators already comfortable with virtualization management get GPU scheduling knobs they know how to turn. The control plane is opinionated — it applies VMware’s virtualization model to acceleration — but that opinion is the value for VMware-native shops. ### NVIDIA-Provided Components **NVIDIA AI Enterprise (NVAIE)** Enterprise AI software suite providing vGPU drivers, GPU Operator for Kubernetes, and validated AI frameworks. NVAIE licenses purchased separately from VCF. Deeply integrated but independently licensed — the same NVIDIA dependency Dell and HPE share. **NVIDIA GPU Silicon + Networking** Blackwell, H100/H200, ConnectX-7/8, BlueField-3 DPUs, NVSwitch, InfiniBand. VMware validates and integrates but does not manufacture or specify GPU silicon. **NVIDIA NIM Microservices** Pre-built inference microservices deployable on Private AI Foundation. Nemotron models and community models available through Model Store. ### Gap Analysis VMware’s Layer 0 position is fundamentally different from every other vendor in this assessment: VMware provides the abstraction layer, not the physical infrastructure. Dell, HPE, and VAST own or specify hardware. Google and AWS own data centers. VMware sits above all of them. This creates a unique DAPM profile: the enterprise Retains hardware vendor choice (can switch from Dell to HPE to Lenovo without changing the management plane) but Delegates GPU runtime to NVIDIA and inherits whatever GPU integration VMware has validated. The abstraction is VMware’s value proposition and its architectural constraint — VMware can only support GPU features that the hypervisor can virtualize or pass through. Multi-accelerator support requires nuanced comparison across the assessed vendors. GPU vendor choice is NOT unique to VMware: • Hyperscalers (Google, AWS) fully abstract accelerators at the service layer — developers call Vertex AI or Bedrock and never see whether TPUs, Trainium, or NVIDIA GPUs power the response. Accelerator choice is Ceded to the cloud provider. Simplest developer experience, least architectural control. • HPE GX5000 supports NVIDIA Rubin and AMD MI430X GPU blades in the same rack architecture. Multi-vendor at the hardware level. • Dell AI Factory is NVIDIA-only. ‘Dell AI Platform with AMD’ is a separate branding, separate software stack — not a unified runtime. • VAST CNode-X is NVIDIA-only. VMware’s actual differentiator is the level of architectural control over acceleration: VCF manages AMD, NVIDIA, and Intel accelerators through familiar virtualization primitives (vGPU profiles mapped to vmclasses, DRS for GPU workload balancing, vMotion for live migration, resource pools for multi-tenant isolation). The enterprise architect retains granular control over GPU placement, scheduling, and isolation using tools they already operate. The control plane is borrowed judgment — it is VMware’s opinionated virtualization model applied to GPU resources — but that opinion provides stronger knobs that appeal specifically to operators already comfortable with vSphere management. Where hyperscalers abstract the accelerator away from the architect, VMware puts the architect in the driver’s seat through a familiar console. The vDefend security story is architecturally significant for AI: micro-segmentation at the packet level between every AI component (model server, vector database, embedding service, API gateway) with Terraform-codified firewall rules. This is infrastructure-layer Zero Trust for AI workloads — a capability that Dell and HPE don’t provide at equivalent depth from the platform layer. VAST’s CrowdStrike integration operates at a different level (application/data-layer security vs. network-layer micro-segmentation). ### Borrowed Judgment Multi-directional: VMware borrows GPU silicon judgment from NVIDIA (same as everyone) and hardware engineering judgment from OEM partners (Dell, HPE, Lenovo build the servers). But VMware retains the abstraction layer — the hypervisor, the networking, the security model, the orchestration. This is the inverse of Dell’s position: Dell retains hardware judgment and borrows software judgment from NVIDIA. VMware retains software judgment and borrows hardware judgment from OEMs. Critically, the VMware control plane is itself borrowed judgment for the enterprise: the architect gains granular GPU management through vSphere primitives, but those primitives encode VMware’s opinions about how acceleration should be virtualized, scheduled, and isolated. The enterprise borrows VMware’s virtualization worldview in exchange for operational familiarity. This is a different trade-off than the hyperscalers (where the enterprise cedes acceleration decisions entirely) or bare-metal (where the enterprise retains full control but builds everything). VMware occupies the middle: more control than cloud, less effort than bare-metal, but through an opinionated lens. The NVIDIA AI Enterprise dependency is real but no deeper than Dell’s or HPE’s: all three require NVAIE for GPU virtualization and AI framework support. VMware’s co-engineering relationship with NVIDIA on VCF integration is comparable to HPE’s Private Cloud AI co-engineering. ### Working Notes The Broadcom acquisition context is impossible to ignore at Layer 0: Gartner projects VMware’s virtualization market share will fall from 70% (2024) to 40% (2029) due to pricing changes. Nutanix CEO has publicly targeted 165,000 of VMware’s approximately 300,000 customers. Broadcom has converted 90%+ of the top 10,000 VMware customers to VCF subscriptions with 200-500% price increases reported. This creates a unique installed-base dynamic: Private AI Foundation’s market opportunity is less about winning new customers than about retaining existing ones by making VCF indispensable for AI workloads. If the enterprise is already paying for VCF, Private AI Services come at no additional cost — a fundamentally different go-to-market than Dell (buy new PowerEdge + NVIDIA), HPE (buy new Private Cloud AI), or VAST (deploy new AI OS). The air-gapped deployment support is significant for regulated industries and government — same capability Dell and HPE emphasize, delivered through the existing VCF automation framework rather than a purpose-built AI appliance. ## ◑ Layer 1A · Storage: Data Storage & Governance *Durable, governed data foundation — the Governance Catalog that Layer 2C queries* **Status:** Platform Storage, Not AI-Native ### Vendor-Provided Components **vSAN (HCI Storage)** [DAPM: Ceded] Hyper-converged storage integrated into VCF. Block and file storage natively. Native Object Storage (S3-compatible) in tech preview with VCF 9.1.x — brings S3 interface natively into the platform without third-party licensing. vSAN deduplication and compression for cost reduction. Unified storage policies and multi-tenant self-service access. **vSAN for Recovery + Ransomware Recovery** [DAPM: Ceded] Sovereign, in-place ransomware recovery using native snapshot capabilities. Deep snapshot chains and integrated replication workflows. On-prem recovery without external dependencies. **Model Store (Private AI Services)** [DAPM: Ceded] Curated LLM repository with integrated RBAC access control. MLOps teams and data scientists can securely manage and provide LLMs with governance and security for enterprise data and IP. NVIDIA models, Nemotron, and community models available. Proprietary VMware/Broadcom platform — opinions captive, no open exit. **External Storage Integration** [DAPM: Delegated] VCF supports Dell PowerScale, Dell ObjectScale, NetApp ONTAP, Pure Storage, HPE Alletra, and other enterprise storage via vSphere APIs. The storage layer is not limited to vSAN — enterprises can bring existing storage investments. This is a heterogeneous storage approach vs. Dell’s vertically integrated storage (PowerScale/ObjectScale/Exascale) or VAST’s collapsed storage (Element Store). ### NVIDIA-Provided Components **No Direct NVIDIA Layer 1A Dependency** NVIDIA does not provide storage or governance components in the VMware Private AI stack. Storage is VMware-owned (vSAN) or enterprise-chosen (external arrays). ### Gap Analysis VMware’s Layer 1A is fundamentally different from every other assessed vendor because VMware is not a storage company. vSAN provides competent hyper-converged storage with the VCF 9.1 addition of native S3 object storage, but this is general-purpose platform storage — not AI-optimized data infrastructure. Compare to Dell: MetadataIQ indexes billions of files with AI-specific metadata enrichment. Exascale provides 10+ PB/rack unified file+object+fast-file. Trust3 AI provides storage-layer governance for sensitive data discovery. Compare to HPE: Data Fabric v8.1 provides policy-based data placement with Apache Polaris catalog for Iceberg tables. Alletra B10000 provides real-time agentic storage support with semantic understanding. Compare to VAST: Element Store collapses file, object, table, and vector into a single governed data structure with inline metadata enrichment. VMware’s storage story is ‘bring your existing storage’ — which is pragmatic for the installed base but means Layer 1A governance (metadata richness, data lineage, policy-based placement) depends entirely on whichever external storage vendor the enterprise has deployed. VMware itself provides no AI-specific governance catalog, no metadata enrichment, no data lineage tracking. The Model Store capability is a Layer 1A function worth noting: RBAC-governed model repository is a governance primitive that Dell’s AI Factory lacks as a platform-native capability. But Model Store governs models, not data — it does not address the broader question of which data feeds which model under what compliance constraints. ### Borrowed Judgment Low for vSAN (VMware-owned). High for AI-specific storage governance — entirely dependent on whichever external storage vendor the enterprise deploys. If the enterprise runs Dell storage, it inherits Dell’s governance capabilities (MetadataIQ). If it runs NetApp, it inherits NetApp’s. VMware provides no abstraction or unification of storage governance across heterogeneous backends. This is the inverse of VMware’s Layer 0 strength: at Layer 0, VMware abstracts heterogeneous hardware into a unified management plane. At Layer 1A, VMware does NOT abstract heterogeneous storage governance into a unified governance plane. The storage abstraction stops at provisioning and capacity management — it does not extend to metadata, lineage, or policy. ### Working Notes The native S3 Object Storage in VCF 9.1.x (tech preview) is a strategic move: S3 compatibility is the lingua franca of AI data pipelines. Every vendor in this assessment provides S3 access (Dell ObjectScale, HPE Alletra X10000, VAST DataStore, AWS S3, Google Cloud Storage). VMware adding native S3 to vSAN reduces the dependency on external object storage for AI workloads. The Tanzu Marketplace integration provides a curated path to certified middleware and data services — this is VMware’s approach to ecosystem curation at the data layer, comparable in intent (not depth) to HPE’s Unleash AI program or VAST’s Cosmos Community. SQL Server DBaaS as a first-class VCF citizen is a pragmatic enterprise play — most enterprises have SQL Server deployments, and making it a platform service reduces the friction of data access for AI workloads. ## ◑ Layer 1B · Retrieval: Context Management & Retrieval *Low-latency retrieval for RAG — vector/hybrid search, context windows* **Status:** Foundational RAG Services ### Vendor-Provided Components **Data Indexing & Retrieval (Private AI Services)** [DAPM: Ceded] Index and maintain multiple data sources, making them readily available for consumption by AI applications. Integrated with Model Runtime for RAG workflows. Keeps indexed data current as sources change. Proprietary VMware/Broadcom platform — opinions captive, no open exit. **Vector Database (Private AI Services)** [DAPM: Delegated] pgvector on PostgreSQL delivered via Data Services Manager with VMware enterprise-level support. Enables domain-specific, up-to-date context for AI models. PostgreSQL 16.8 with pgvector 0.8.0 extension in Private AI Services 2.1. pgvector is OSS and the consumed interface is standard PostgreSQL — embeddings, schema, and index definitions lift via pg_dump to any Postgres platform; DSM operates the database, the open substrate keeps the retrieval opinions portable. **RAG Pipeline Integration** [DAPM: Delegated] NVIDIA NIM RAG Blueprint v2.5.0 validated on VCF — production-grade, multi-model RAG pipeline. Pre-built catalog items in VCF Automation for deploying complete RAG workflows. Elasticsearch supported as external vector database for advanced retrieval scenarios. ### NVIDIA-Provided Components **NVIDIA NIM + NeMo Retriever** Inference microservices and retrieval-augmented generation components. RAG Blueprints provide pre-built retrieval patterns. Same capabilities available on Dell and HPE platforms. **NVIDIA AI Enterprise RAG Stack** Validated software stack for RAG workflows on VMware Private AI Foundation. GPU-accelerated embedding generation and retrieval. ### Gap Analysis VMware’s Layer 1B provides functional RAG capabilities through Private AI Services — Data Indexing/Retrieval and Vector Database are genuine platform services, not just partner integrations. But the retrieval stack is foundational, not differentiated. pgvector on PostgreSQL is a competent vector database for moderate-scale use cases but lacks the performance characteristics of purpose-built alternatives. Compare to Dell’s Data Search Engine (Elasticsearch 9.4 with GPU-accelerated hybrid search, MetadataIQ integration). Compare to VAST’s InsightEngine (native to the data platform, no data movement for retrieval). Compare to HPE’s Alletra X10000 with KV cache storage support for inference state persistence. The Data Indexing & Retrieval service addresses the core RAG requirement — keeping context current as data sources change — but without the metadata richness or governance integration that Dell (MetadataIQ + Elastic), HPE (Data Fabric + Kamiwaza), or VAST (Catalog + InsightEngine) provide. No retrieval quality observability (recall@k, latency percentiles) is evident — the same gap identified in the Dell assessment. A Layer 2C placement engine would need retrieval quality metrics to make informed routing decisions. ### Borrowed Judgment Moderate. VMware owns the Data Indexing/Retrieval and Vector Database services. RAG pipeline patterns depend on NVIDIA NIM/Blueprints (same dependency as Dell and HPE). Elasticsearch as external vector database option introduces the same Elastic dependency Dell has — search intelligence is Elastic’s, not VMware’s. The pgvector choice is notable: PostgreSQL is the most widely deployed enterprise database. By building on pgvector, VMware reduces adoption friction (most enterprises already have PostgreSQL expertise) at the cost of retrieval performance ceiling. Dell chose Elasticsearch (higher performance, more complex). VAST built its own (highest integration, most proprietary). VMware chose the most pragmatic option. ### Working Notes The OpenWebUI integration with Private AI Services RAG demonstrates VMware’s approach to Layer 1B: provide the retrieval infrastructure, let the enterprise choose the user-facing application layer. This is consistent with VMware’s platform philosophy — VMware provides infrastructure services, not applications. The RAG Blueprint validation on VCF (multi-model, production-grade, 8x NVIDIA H100 80GB GPUs) provides a concrete reference architecture that enterprises can deploy from VCF Automation catalog items. This is operationally simpler than assembling equivalent RAG infrastructure on bare-metal Dell or HPE hardware. ## ○ Layer 1C · Pipelines: Data Movement & Pipelines *Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering* **Status:** Gap ### NVIDIA-Provided Components **NVIDIA Blueprints** Pre-built AI application patterns deployable through VCF Automation. Pipeline templates, not pipeline infrastructure — same as Dell and HPE. **NVIDIA CMX (Future)** KV cache management for context memory offload. When integrated with VCF, could provide the same KV cache tiering Dell has validated (19x TTFT improvement). ### Gap Analysis Layer 1C is VMware’s most significant gap relative to other assessed vendors. VMware provides no equivalent to: • Dell’s Data Orchestration Engine (Dataloop): No-code/low-code AI data lifecycle management, Dell’s most meaningful software acquisition. • HPE’s Ezmeral Unified Analytics: Enterprise-hardened ML pipeline stack (Airflow, Kubeflow, Ray, Feast, MLflow, Spark). • HPE’s Data Fabric: Policy-based data placement with compliance tagging and data lineage. • VAST’s DataEngine: Serverless data transformation with CLI/SDK, built-in observability, triggers, and automated pipelines. VCF Automation provides deployment pipelines (standing up AI infrastructure) but not data pipelines (moving, transforming, and governing data through ML workflows). The enterprise running VMware Private AI Foundation must bring its own data pipeline orchestration — Airflow, Kubeflow, or a commercial alternative — and deploy it on VKS. This is architecturally consistent with VMware’s platform philosophy: VCF provides infrastructure services, not application-layer data engineering tools. But it leaves a functional gap that competitors have filled. An enterprise choosing VMware for AI inherits a Layer 1C assembly problem that Dell (Dataloop), HPE (Ezmeral), or VAST (DataEngine) partially or fully solve. ### Borrowed Judgment High. The enterprise must borrow data pipeline judgment from whatever tools it deploys on VKS — Apache Airflow community, Kubeflow community, or a commercial vendor (Dataloop, Databricks, etc.). VMware provides no opinion on data pipeline architecture, no integration between pipeline metadata and infrastructure governance, and no data lineage capability. Compare to Dell: Dell acquired Dataloop specifically to address Layer 1C. Compare to HPE: HPE assembled Ezmeral through four acquisitions (BlueData, MapR, Ampool, Arrikto). Compare to VAST: VAST built DataEngine as a native platform capability. VMware has made no equivalent investment in data pipeline IP. ### Working Notes The gap is real but may be strategic: VMware has historically succeeded by providing infrastructure primitives that partner ecosystems build on, rather than by building application-layer tooling. The question is whether AI data pipelines are infrastructure (VMware should own them) or applications (VMware should enable them). The Tanzu Marketplace could address this gap through curated data pipeline services — certified Airflow, MLflow, or Kubeflow deployments validated for VCF. This would be a Delegated approach (partner provides the capability, VMware validates the deployment) rather than a Retained approach (VMware builds the capability). Architecturally similar to HPE’s Unleash AI ecosystem model. The KV cache story is notably absent: Dell has validated NVIDIA CMX with 19x TTFT improvement on PowerScale. HPE has native KV cache storage support in Alletra X10000. VAST collocates cache and compute in CNode-X. VMware has not yet announced equivalent KV cache tiering capabilities. (Enhanced NVMe memory tiering in VCF 9.1 addresses memory-bound performance, not data-pipeline orchestration.) ## ● Layer 2A · Orchestration: Infrastructure Orchestration *GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization* **Status:** VMware Heritage Strength ### Vendor-Provided Components **VCF Automation (formerly vRealize/Aria Automation)** [DAPM: Ceded] Self-service catalog with pre-built AI workload templates. Quickstart deployment for Private AI Foundation. Infrastructure-as-code with Terraform integration. Multi-tenant resource provisioning with RBAC. Live Application Stack Blueprints for versioned, redeployable application topologies. Day 2 operations for AI Blueprints. Proprietary VMware/Broadcom platform — opinions captive, no open exit. **vSphere Supervisor + VKS** [DAPM: Delegated] Unified management of VMs, containers, and AI workloads from a single control plane. VKS (vSphere Kubernetes Service) 3.6 supports up to 500 Kubernetes clusters per Supervisor. Simplified Container-as-a-Service for application teams. VM Fast-Deploy for accelerated provisioning. vSphere Elastic Provisioning for zero-touch fleet expansion. GitOps-based infrastructure management. VKS is conformant Kubernetes — manifests and workloads lift to another conformant cluster without rebuilding, so the K8s opinions are portable (scored as the cloud managed-K8s services are); the proprietary fleet and lifecycle management is captured in the Ceded VCF Automation / Operations / SDDC Manager components. **VCF Operations (formerly vRealize/Aria Operations)** [DAPM: Ceded] Private AI Model and GPU Metrics — utilization, memory pressure, and model-level visibility on the same console as the rest of the estate. Real-Time Operational Observability turns telemetry into action. Customizable dashboards for AI model and agent performance. Capacity management and compliance monitoring for AI workloads. Proprietary VMware/Broadcom platform — opinions captive, no open exit. **SDDC Manager** [DAPM: Ceded] Full-stack lifecycle management for VCF. Automated deployment, patching, and upgrades across vSphere, vSAN, NSX, and VKS. Single-pane fleet management. Expanded fleet size and upgrade scale in 9.1. This is the operational backbone — the equivalent of HPE’s GreenLake or Dell’s APEX management, but with 20+ years of enterprise maturity. Proprietary VMware/Broadcom platform — opinions captive, no open exit. **Advanced Cyber Compliance (ACC)** [DAPM: Ceded] Continuous compliance enforcement with automated drift detection and remediation. Hardened infrastructure images. Integrated with vDefend for security posture management. Disaster recovery via vSAN for Recovery. Proprietary VMware/Broadcom platform — opinions captive, no open exit. **MCP Server Governance (VCF 9.1)** [DAPM: Ceded] IT operations can centrally manage and control access to MCP tools and associated servers across their environment. Ensures user groups can only access approved MCP tools. Security guardrails for MCP servers via vDefend and Avi Load Balancer. This is a Layer 2A governance function with 2C implications — controlling which agents can access which tools. Proprietary VMware/Broadcom platform — opinions captive, no open exit. ### NVIDIA-Provided Components **NVIDIA GPU Operator** Kubernetes operator for GPU lifecycle management. Manages GPU drivers, container runtime, device plugins. Standard across all NVIDIA-integrated platforms. **NVIDIA vGPU Manager** GPU virtualization profiles and multi-tenant GPU allocation. Memory partitioning and time-sliced compute scheduling. Managed through vSphere Supervisor. ### Gap Analysis Layer 2A is VMware’s strongest layer — arguably the strongest Layer 2A of any vendor in this assessment series. The reason is operational maturity: VCF has been managing enterprise infrastructure for two decades. No other vendor assessed has equivalent depth in lifecycle management, multi-tenant orchestration, compliance automation, and unified VM/container/AI workload management. Specific 2A differentiators vs. other assessed vendors: • Unified workload management: Dell manages AI workloads separately from traditional workloads (OpenManage for servers, Run:ai for GPUs, separate tools for each). HPE manages AI through GreenLake Intelligence + OpsRamp + Private Cloud AI (three systems). VMware manages AI, containers, and VMs from ONE control plane (vSphere Supervisor). VAST manages only VAST workloads. • Operational maturity: VCF’s Day 2 operations (patching, upgrades, compliance, capacity planning) for AI workloads inherit the same proven processes used for the enterprise’s existing VM fleet. New operational model required? Zero. Dell and HPE AI stacks require new operational processes. VAST requires an entirely new operational discipline. • MCP Server Governance is a notable 2A/2C bridge: centrally controlling which user groups can access which MCP tools is an infrastructure-level governance function that no other on-prem vendor provides as a platform native capability. Google’s Agent Gateway provides equivalent capability in cloud. The GPU scheduling gap remains: NVIDIA GPU Operator and vGPU Manager handle GPU allocation, but policy-driven GPU scheduling (which workload gets which GPU based on cost, compliance, and performance constraints) is not a VCF-native function. This is the same gap Dell has with Run:ai — the scheduling intelligence is NVIDIA’s, not the platform vendor’s. ### Borrowed Judgment Low — the lowest of any layer in the VMware assessment. VCF Automation, vSphere, VKS, vSAN, NSX, VCF Operations, and SDDC Manager are all Broadcom/VMware IP. GPU scheduling is the primary borrowed judgment (NVIDIA GPU Operator + vGPU Manager), but this is the same dependency every on-prem vendor shares. Compare to Dell Layer 2A: Dell splits 2A between OpenManage (Dell-owned) and Run:ai (NVIDIA-owned, acquired). VMware retains more 2A authority than Dell. Compare to HPE Layer 2A: HPE’s GreenLake Intelligence is HPE-owned 2A with MCP-based agent communication. VMware’s VCF Operations is VMware-owned 2A with emerging MCP support. Both retain 2A authority; different architectural approaches (HPE: agentic mesh; VMware: traditional orchestration evolving toward agentic). ### Working Notes The Intelligent Assist for VCF (tech preview) signals VMware’s evolution toward agentic infrastructure management. An AI-driven support assistant that diagnoses and resolves issues by consulting Broadcom’s knowledge base is functionally similar to HPE’s GreenLake Intelligence domain agents or Dell’s CloudIQ — but at an earlier stage of development. The 100M+ licensed cores installed base gives VMware an operational data advantage no other on-prem vendor can match: patterns learned from managing the world’s largest virtualization fleet can inform AI workload optimization in ways that newer platforms cannot. Whether Broadcom invests in leveraging this data advantage for AI-specific intelligence is an open question. ## ◑ Layer 2B · Runtime: Application Runtime & Execution *Model serving, agent execution, inference APIs, distributed inference* **Status:** Platform-Native + NVIDIA-Dependent ### Vendor-Provided Components **Model Runtime (Private AI Services)** [DAPM: Ceded] Run inference and embedding models as a service across the organization. API Gateway allows users and AI applications to interact with models directly via API. Multi-accelerator support — same model deployment on AMD and NVIDIA GPUs without refactoring. Model endpoints configurable through VCF Automation UI. Multi-tenant Models-as-a-Service enables secure model sharing across business units to lower costs and reduce power consumption. Proprietary VMware/Broadcom platform — opinions captive, no open exit. **Agent Builder (Private AI Services)** [DAPM: Ceded] Build AI agents in a user-friendly playground, leveraging models and knowledge bases created using other Private AI services. Integrated with Model Runtime and Data Indexing/Retrieval for end-to-end agent development. This is a platform-native agent construction surface — not as deep as VAST’s AgentEngine or Google’s Agent Studio/ADK, but integrated into the VCF operational model. Proprietary VMware/Broadcom platform — opinions captive, no open exit. **Deep Learning VMs** [DAPM: Delegated] Pre-configured virtual machines with validated AI/ML software stacks: PyTorch, TensorFlow, Miniconda. Software stack validated in advance on NVIDIA GPUs — data scientists start developing immediately without compatibility validation. Provisioned through VCF Automation self-service catalog. **VKS AI Clusters** [DAPM: Ceded] GPU-capable Kubernetes worker nodes for cloud-native AI/ML workloads. Triton Inference Server deployable from catalog. Distributed LLM inference with GPUDirect RDMA over InfiniBand for models that exceed single-server capacity (DeepSeek-R1, Llama 3.1-405B). This is infrastructure-layer runtime support, not an opinionated agent execution framework. **Tanzu Platform (Application Runtime)** [DAPM: Ceded] PaaS-layer application runtime. ‘You provide code, we put it into production.’ Governed agentic coding with Tanzu. Developers can self-publish AI agents and MCP servers, sharing AI applications and tools across the enterprise. MCP server publishing makes Tanzu a distribution surface for enterprise agent tooling. ### NVIDIA-Provided Components **NVIDIA AI Enterprise Runtime** NIM inference microservices, model optimization, GPU-accelerated frameworks. The core AI runtime dependency — VMware’s Model Runtime wraps NVIDIA inference capabilities in a platform-managed service. **NVIDIA NIM Agent Blueprints** Pre-built agentic workflows (RAG, PDF extraction, digital twins). Same blueprints available on Dell, HPE, Cisco, Lenovo. Non-differentiating for VMware at 2B. **NVIDIA Triton Inference Server** Multi-framework model serving. Deployable as VCF Automation catalog item on GPU-capable VKS clusters. ### Gap Analysis VMware’s Layer 2B is the most architecturally interesting in this assessment because it combines platform-native AI services (Model Runtime, Agent Builder) with NVIDIA runtime dependency — a hybrid Retained/Delegated model. The Model Runtime is a genuine platform capability: model serving as a managed VCF service with API Gateway, multi-tenant isolation, and multi-accelerator support. This is structurally different from Dell’s 2B (entirely NVIDIA-dependent — NemoClaw/OpenShell) and closer to HPE’s 2B (HPE provides the deployment platform, NVIDIA provides the execution runtime, with HPE-owned bracketing governance above and below). The Agent Builder is notable as a platform-native agent construction surface. Compare to alternatives: • Dell: No platform-native agent builder. Relies on NVIDIA NIM/NemoClaw post-deployment. • HPE: CrewAI pre-installed (partner framework). Deloitte Zora AI (partner application). • VAST: AgentEngine — deeply integrated, proprietary agent runtime. • Google: Agent Studio (no-code), ADK (code-first), Agent Designer. • AWS: Bedrock Agents (no-code), Strands SDK (code-first). VMware’s Agent Builder is simpler than the hyperscaler offerings but it’s integrated into the VCF operational model — agents built here inherit VCF’s security (vDefend microsegmentation), governance (MCP server controls), and observability (GPU/model metrics). That operational integration is VMware’s differentiator. The Tanzu Platform MCP server publishing capability is a Layer 2B/2C bridge worth tracking: enabling developers to self-publish MCP servers creates an enterprise-internal agent tool marketplace governed by IT. This is a distributed model for agent capability deployment that differs from Google’s centralized Agent Registry or HPE’s curated Unleash AI ecosystem. ### Borrowed Judgment Moderate. VMware owns Model Runtime, Agent Builder, and the Tanzu application runtime. But inference execution depends on NVIDIA AI Enterprise (same structural dependency as Dell and HPE). The multi-accelerator support (AMD + NVIDIA) provides a runtime alternative that Dell AI Factory customers don’t have (Dell’s AMD track is a separate stack), though HPE’s GX5000 also supports multi-vendor GPU blades and hyperscalers abstract accelerators entirely. VMware’s value is that the architect controls which accelerator serves which workload through familiar virtualization primitives — the control plane is opinionated (VMware’s virtualization model applied to GPUs) but provides operational knobs that vSphere-native teams already understand. The NVIDIA dependency at 2B is real but partially mitigated by VMware’s abstraction: Model Runtime provides a VMware-managed API surface. If NVIDIA changes its NIM/NemoClaw architecture, VMware absorbs the integration change — the enterprise’s API doesn’t change. This is the same ‘bracketing’ architecture HPE uses (GreenLake governance above and below the NVIDIA runtime), expressed differently (VMware API abstraction wrapping the NVIDIA runtime). ### Working Notes The multi-accelerator Model Runtime is significant but requires context: running the same AI model on AMD and NVIDIA GPUs without refactoring is a runtime-level abstraction that VMware provides through familiar virtualization management tools. HPE’s GX5000 supports multi-vendor GPU blades in the same rack, and hyperscalers abstract accelerators entirely at the service layer (Vertex AI, Bedrock). VMware’s differentiator is not multi-accelerator support per se but the level of architectural control — the operator manages GPU placement and scheduling through vSphere primitives they already know, with stronger knobs than cloud providers offer. Dell’s AMD track (Dell AI Platform with AMD) remains a separate branding with a separate software stack, not a unified runtime. The Tanzu-mediated MCP server publishing is an emerging capability that could become significant for agentic AI: if every enterprise developer can publish MCP servers through Tanzu, and IT governs access through VCF 9.1’s MCP server governance, VMware creates a platform for enterprise agent tooling that is neither centralized (Google) nor delegated to partners (HPE Unleash AI) but distributed-and-governed. Whether this pattern scales depends on enterprise developer adoption of Tanzu. ## ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane *Policy-driven placement and resource coordination — the Autonomy Layer* **Status:** Emerging Signals Only ### NVIDIA-Provided Components **No NVIDIA Layer 2C on VMware** NVIDIA provides no agent governance, policy-driven placement, or reasoning plane components in the VMware stack. Same gap as Dell — NVIDIA’s AI-Q is workflow scaffolding, OpenShell is constraint enforcement, Dynamo is performance routing. None is Layer 2C. ### Gap Analysis Applying the ‘Routing Is Not Reasoning’ test: MCP Server Governance = access control. Intelligent Assist = IT operations automation. GPU/Model Metrics = observability telemetry. None provides policy-driven decisions about where compute runs relative to data, which model serves which request, and how cost/compliance/latency are arbitrated in real time. VMware’s Layer 2C position is comparable to Dell’s: not yet evident as a productized capability. The signals (MCP governance, metrics observability, Intelligent Assist) suggest the building blocks exist but have not been composed into a reasoning plane. Compare to other vendors’ Layer 2C status: • Dell: Absent. Dell+Intel ‘actively addressing’ but no product announced. • HPE: Delegated to Kamiwaza (multi-layer orchestration partner via Unleash AI). GreenLake Intelligence provides IT ops 2C. • VAST: PolicyEngine + Polaris — the most aggressive middle-out Layer 2C build. • Google: Agent Identity + Gateway + Registry + Orchestration + Observability — the most complete productized 2C. • AWS: AgentCore with implicit 2C through managed service placement decisions. VMware’s unique 2C opportunity: VCF manages the entire infrastructure estate. If Broadcom builds a Layer 2C that queries vSAN governance metadata, GPU utilization telemetry, vDefend security posture, and Model Runtime performance metrics to make autonomous placement decisions, it would have the broadest infrastructure visibility of any on-prem 2C — because VCF sees everything from the hypervisor up. The data to build 2C exists in VCF Operations. The governance primitives exist in MCP Server Governance and ACC. The placement engine does not. ### Borrowed Judgment Inverted: there IS no judgment to borrow because no Layer 2C exists. Same structural position as Dell. The enterprise must build custom 2C logic, bring a partner (Kamiwaza, potentially), or operate without it. Most will choose option 3. The MCP Server Governance capability is an interesting partial answer: it provides governance over agent-tool interactions without providing placement intelligence. This is access-control-as-governance — necessary but not sufficient for a reasoning plane. ### Working Notes VMware has a structural advantage in building Layer 2C that no other on-prem vendor possesses: VCF is the control plane for the enterprise’s entire virtualized estate. Dell manages Dell hardware. HPE manages HPE hardware. VAST manages VAST storage. VMware manages EVERYTHING virtualized — across Dell, HPE, Lenovo, Cisco, and any other OEM’s hardware. A VMware Layer 2C would be the first multi-vendor infrastructure reasoning plane — making placement decisions across heterogeneous hardware from a single governance surface. No other vendor can build this because no other vendor has the cross-vendor infrastructure visibility. Whether Broadcom invests in this opportunity is an open question. The Broadcom acquisition thesis prioritizes cash generation from the installed base, not R&D investment in new platform capabilities. Layer 2C is a significant engineering investment. The $30B annual infrastructure software segment gives Broadcom the resources; the question is whether the strategic priority exists. Hock Tan’s framing of VCF as ‘the permanent abstraction layer between AI software and physical chips’ is a Layer 2A statement, not a Layer 2C statement. The abstraction layer manages resources. The reasoning plane governs them. VMware has the former; it does not have the latter. ## ◑ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane *AI-powered business capabilities — business logic, workflow automation* **Status:** Platform-Enabled, Not Platform-Provided ### Vendor-Provided Components **Private AI Services (Integrated)** [DAPM: Ceded] Model Runtime + Agent Builder + Data Indexing/Retrieval + Vector Database + Model Store + GPU Monitoring — all included in VCF subscription at no additional cost. This is not a Layer 3 application stack — it is an integrated set of AI platform services that enables Layer 3 development. The distinction matters: VMware provides the tools to BUILD AI applications, not the applications themselves. Proprietary VMware/Broadcom platform — opinions captive, no open exit. **NVIDIA Blueprints + NIM on VCF** [DAPM: Delegated] Pre-built AI application patterns deployable through VCF Automation catalog. Multimodal PDF Extraction, Digital Twins, RAG pipelines. Same blueprints available on Dell, HPE, Cisco, Lenovo — non-differentiating for VMware. **Tanzu Platform (Agent Distribution)** [DAPM: Ceded] Developers self-publish AI agents and MCP servers to the enterprise via Tanzu. IT maintains governance and oversight. Tanzu Marketplace provides curated path to certified middleware, data services, and AI tooling. This is an enterprise app store model for AI capabilities. **ISV + OEM Ecosystem** [DAPM: Delegated] VMware Private AI Foundation validated on Dell, HPE, Lenovo, Cisco, Supermicro, NEC, Fujitsu hardware. ISV ecosystem spans the entire VMware partner network — thousands of validated applications across every industry. AI-specific ISV validation is emerging but not yet at the curation depth of HPE’s Unleash AI (26+ selected ISV partners) or Dell’s AI Ecosystem Program (OpenAI, Palantir, Google, ServiceNow). **OpenWebUI Integration** [DAPM: Delegated] Open-source AI user interface integrated with VCF Private AI Services RAG. Provides a ChatGPT-like interface for enterprise users to interact with privately-hosted models. Demonstrates the ‘platform enables applications’ model. ### NVIDIA-Provided Components **NVIDIA Model Ecosystem** Nemotron models, community models, NVIDIA NIM containers available through Model Store. NVIDIA provides the model layer; VMware provides the serving and governance layer. ### Gap Analysis VMware’s Layer 3 is structurally different from every other assessed vendor because VMware is explicitly a platform, not an application provider. VMware provides the tools to build and deploy AI applications (Private AI Services) but does not build the applications themselves. This is the correct architectural position for an infrastructure platform vendor — and it’s the same position Dell occupies (Dell doesn’t build AI applications; it partners with OpenAI, Palantir, ServiceNow). The difference is ecosystem depth: • Dell’s AI ecosystem: OpenAI, Palantir, Google, ServiceNow, SpaceXAI, Hugging Face, 5,000+ deployment customers. Explicitly curated for AI. • HPE’s Unleash AI: 26+ selected ISV partners with validated interoperability. Kamiwaza orchestration. CrewAI pre-installed. Purpose-built for AI. • VAST’s Cosmos Community: CoreWeave, TwelveLabs, CrowdStrike with distinct partner tracks. Focused and vertical. • VMware’s AI ecosystem: Inherits the broader VMware partner ecosystem (thousands of ISVs) but without AI-specific curation depth. Private AI Foundation validation is available on major OEM hardware, but AI-specific ISV partnerships are not yet at the maturity of Dell or HPE programs. The Tanzu-mediated MCP server publishing could evolve into VMware’s distinctive Layer 3 model: instead of curating an external ISV ecosystem (HPE’s approach) or partnering with AI application vendors (Dell’s approach), VMware enables the enterprise’s own developers to build and distribute AI agents internally. This is an internally-generated Layer 3 rather than an externally-sourced one. The VCF installed base is the Layer 3 enabler: 100M+ cores means Private AI Services reach more enterprise infrastructure than any competitor’s AI platform. The AI applications built on VMware will be built by the enterprise’s own developers, using VMware’s tools, on VMware’s platform. Whether that bottom-up, developer-driven approach generates Layer 3 applications as quickly as Dell’s top-down partnerships (OpenAI, Palantir) or HPE’s curated ecosystem (Unleash AI) is the open question. ### Borrowed Judgment Distributed across the enterprise’s own development teams and chosen partners. VMware provides the platform; the enterprise provides the application logic. This is the most explicit Retained model for Layer 3 in this assessment — the enterprise builds its own AI applications rather than consuming a vendor’s or partner’s. The trade-off: maximum control (Retained), maximum effort (the enterprise must build everything above the platform services layer). Dell and HPE offer Delegated shortcuts (partner applications). VMware offers Retained responsibility. ### Working Notes The Private AI Foundation at no additional cost for VCF subscribers is a strategic masterstroke for customer retention: every VCF customer already has access to Model Runtime, Agent Builder, Vector Database, Data Indexing/Retrieval, and Model Store. The marginal cost of trying Private AI is zero (beyond GPU hardware). This is the lowest-barrier entry to on-prem AI of any vendor assessed. The 9/10 Fortune 500 commitment to VCF means Private AI Foundation has the largest potential enterprise deployment footprint of any on-prem AI platform. Whether that potential converts to actual AI workload deployment depends on whether enterprises find Private AI Services sufficient for production AI or whether they choose purpose-built alternatives (Dell AI Factory, HPE Private Cloud AI, VAST AI OS) for deeper capabilities. The competitive dynamic is unusual: VMware doesn’t compete with Dell or HPE at Layer 0 (VMware runs ON their hardware). VMware competes with them at Layers 1-3 (management, orchestration, AI services). An enterprise could run Dell hardware + VMware VCF + VMware Private AI Services — getting Dell’s Layer 0 with VMware’s Layers 2A/2B. Or Dell hardware + Dell AI Factory — getting Dell’s Layer 0 with Dell/NVIDIA’s Layers 2A/2B. The choice is between VMware’s operational maturity and Dell/HPE’s AI-specific depth. --- *Layer2C · AI Infrastructure Decision Intelligence · The CTO Advisor LLC · thectoadvisor.com*