Executive Summary: Cloudera (Anywhere Cloud + Data Services + Cloudera AI)

Cloudera is a data platform that runs wherever the enterprise puts it, and the map reads it strongest at the two layers that were always its own. Layer 1A is strong on an Apache Iceberg lakehouse with the Shared Data Experience (SDX) governing it and Octopai lineage across the estate; Layer 1C is strong on DataFlow (Apache NiFi 2.0 with 450-plus connectors), Streaming (Kafka, Flink), and Data Engineering (Spark, Airflow) on every deployment path; Layer 2B is strong on Cloudera AI Inference serving any model on the customer's own GPUs since February 2026 and an Agent Studio that reached general availability on premises in June 2026. Two layers are moderate (2A on platform-scoped Kubernetes with GPU quota pools; Layer 3 on AI features in Data Visualization and a Copilot), and three are gaps: no substrate (the customer supplies it on every path), no generally available retrieval product (RAG Studio and Semantic Search are Technical Preview), and no reasoning plane beyond Agent Studio's own guardrails and audit.

The capture is the least coupled on the map, because so much of the platform is Apache. Iceberg and its REST catalog, Ranger and Atlas policies, NiFi flows, Kafka topics, Flink jobs, Spark code, and the models the enterprise brings all leave with the enterprise, and on premises the enterprise runs the open engines itself. What's Cloudera's is the packaging and the opinions on top: the distribution and operators (licensed), Flow Designer and ReadyFlows, SQL Stream Builder and Schema Registry, the Lakehouse Optimizer, the Data Catalog, the Octopai lineage service, the AI Inference service around NVIDIA's runtimes, the AI Registry and Workbench, and Agent Studio's workflows. Twenty-five components read Retained 6, Delegated 6, Ceded 13; the Retained count would be five if Keith rules that tool code written inside Agent Studio is platform-bound.

The buyer's trade: an open lakehouse and movement layer they can take with them, plus any-model serving and an agent builder on their own hardware, in exchange for a retrieval stack still in preview, a control plane (Anywhere Cloud, launched August 19, 2026) whose availability isn't yet documented, and a governance story that governs data rather than agents. What would move a cell: RAG Studio or Semantic Search reaching GA (1B), Anywhere Cloud documentation showing customer-administered placement across environments (2A), or an agent registry or gateway in SDX (2C). The VAST Data AI factory partnership is a reference architecture two sales teams sell, not a substrate on Cloudera's paper.

Layer-by-layer status: Layer 0 (Enterprise Responsibility on Every Path; Cloudera Ships Software, Not Substrate), Layer 1A (Open Lakehouse on Iceberg + SDX Governance + Lineage, Everywhere the Data Sits), Layer 1B (Retrieval Is in Technical Preview: RAG Studio, Semantic Search, and Tool Calling Unscored), Layer 1C (DataFlow, Streaming, and Data Engineering: Any Source to Any Destination on Every Path), Layer 2A (Platform-Scoped Kubernetes with GPU Quota Pools; Anywhere Cloud Control Plane Launched, Undocumented), Layer 2B (AI Inference (NIM, Hugging Face on vLLM, Triton; Any Model; On-Prem and Cloud) + Agent Studio (GA On Premises)), Layer 2C (Data Governance Reaches the Models; Agent Governance Lives Inside Agent Studio), Layer 3 (+1) (AI Visuals and Annotation in Data Visualization, a Copilot in the Workbench; SQL Assistant in Preview).

Assessment framework: 4+1 Layer AI Infrastructure Model. Scoring model: Decision Authority Placement Model (DAPM) — Retained, Delegated, or Ceded. Published by The CTO Advisor LLC (DBA The Advisor Bench). Author: Keith Townsend. Date assessed: September 5, 2026. Version: v1.0 - 4+1 v2: Authority Split.

Cloudera (Anywhere Cloud + Data Services + Cloudera AI)

Mapped to the 4+1 Layer AI Infrastructure Model

v1.0 - 4+1 v2: Authority SplitAssessed September 5, 2026Source: Cloudera documentation (Cloudera AI: AI Studios overview, Agent Studio overview, key features, tool components, MCP integration guide, guardrails, shared responsibility model, RAG Studio guide and configuration, AI Inference concepts and model-server inventory, AI Registry, quota management and GPU setup on premises; Cloudera Data Services on premises 1.5.5 release notes, 1.5.5 SP3 release summary, ECS installation and hardware requirements; Data Engineering GPU prerequisites and quota management; Management Console quota management; Data Warehouse Hue SQL AI Assistant overview; Data Visualization 8.0.0 and 8.0.9 release notes; Data Access documentation; Flow Management and Streams Messaging operator docs; Lakehouse Optimizer APIs) and product pages (AI Inference Service, AI Studios, AI Workbench, AI Assistants, DataFlow, Streaming, SDX, Unified Data Fabric, Data Lineage, Anywhere Cloud, pricing); Cloudera press releases (Anywhere Cloud, August 19, 2026; VAST Data partnership, July 14, 2026; EVOLVE26, March 23, 2026; fiscal 2026 results, February 10, 2026; AI Inference and Data Warehouse with Trino on premises, February 9, 2026; Trino, SDX, and Data Lineage integration, November 20, 2025; Lakehouse Optimizer and Iceberg REST Catalog, September 25, 2025; Data Services private AI, August 6, 2025; Taikun acquisition, August 4, 2025; AI Inference with NVIDIA NIM, October 8, 2024; AI Assistants, June 24, 2024; CrewAI, December 10, 2024); Cloudera blog (Agent Studio and NVIDIA; AI Studios; AI Workbench MCP Server, December 4, 2025); the cloudera/CAI_STUDIO_AGENT repository; Snowflake's Openflow documentation for the NiFi calibration. Peer-reviewed cell by cell through the labs claims ledger (cloudera-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (cloudera-row-agy); totals and escalated items in labs/reviews/cloudera-judgment.md.
ACTIVE ASSESSMENT
Strength
Moderate
Gap
Partner
Layer 0 · ComputeCompute & Network FabricEnterprise Responsibility on Every Path; Cloudera Ships Software, Not Substratedecides: absent · Absent

Raw compute, networking, and acceleration fabric

Vendor-Provided

NVIDIA-Provided

NVIDIA Under Cloudera AI Inference on the Customer's Hardware

Cloudera AI Inference on premises runs NVIDIA NIM microservices and NVIDIA Triton on the customer's NVIDIA GPUs (Blackwell named in the February 2026 release); on ECS the NVIDIA device plugin advertises GPUs once the customer installs drivers and the container toolkit, and on OpenShift the customer installs the Node Feature Discovery and GPU Operators. NVIDIA AI Enterprise is embedded in the service. Cloudera sells none of the hardware.

Gap Analysis

Cloudera sells no silicon, fabric, or arrays, and every path runs on infrastructure the enterprise administers. On premises, Cloudera Data Services install on the customer's servers through Cloudera's Embedded Container Service (ECS) or the customer's Red Hat OpenShift, with GPU nodes the customer provisions; on public cloud, Cloudera provisions its services into the customer's AWS, Azure, or Google Cloud account and bills per Cloudera Compute Unit on top of the cloud bill the customer pays directly; Anywhere Cloud (announced August 19, 2026) adds sovereign, edge, and air-gapped targets to the same model. The VAST Data partnership (July 14, 2026) is a joint AI factory reference architecture, available through both companies' sales teams, in which VAST sells its AI Operating System, Cloudera sells its containerized data services, and the NVIDIA hardware is the customer's. The buyer never buys a Layer 0 from Cloudera on any path. Applying the exposure test: the substrate is the customer's on every path, and Cloudera's control plane sees it only as nodes and quotas. Nothing is inherited invisibly. Calibration: Elastic, Kamiwaza, and MongoDB read gap with authority Absent for vendors whose customers self-manage; Nutanix and VMware read moderate on a hardware-agnostic abstraction the enterprise buys; Databricks and Snowflake read gap with authority Ceded because their customers consume a fleet they can't see. Cloudera's customers always see it. Gap, Absent.

Borrowed Judgment

None to borrow. The enterprise supplies and operates the hardware or the cloud account on every path, and Cloudera's reference architectures (VAST, NVIDIA) are recommendations each partner sells its own part of, not a substrate on Cloudera's paper. Absent: nothing offered at the layer's function, nothing inherited.

Working Notes

Reviewed for awareness: whether the joint Cloudera and VAST AI factory, sold through both sales teams, is Layer 0 capability on Cloudera's paper (ruled no: each company sells its own component and neither resells the GPUs). Watch-list, notes only: Anywhere Cloud's sovereign and air-gapped deployment targets. Public evidence that moves the cell: a Cloudera-owned or Cloudera-sold compute product.

Layer 1A · StorageData Storage & GovernanceOpen Lakehouse on Iceberg + SDX Governance + Lineage, Everywhere the Data Sitsdecides: vendor · Delegated

Durable, governed data foundation — the Governance Catalog that Layer 2C queries

Vendor-Provided

Apache Iceberg Lakehouse + Cloudera Iceberg REST CatalogDelegated

Open-format tables on HDFS, Ozone, or cloud object storage exposed through the standard Iceberg REST catalog interface to any engine. The format and the catalog interface are multi-vendor standards: Delegated.

HDFS + Apache Ozone Storage, Self-RunRetained

The distributed file and object stores the enterprise runs on its own hardware under Cloudera's distribution. Apache 2.0 projects the enterprise operates: Retained on the open-source seam.

SDX Self-Run on Premises: Apache Ranger (RBAC, ABAC, Masking) + Apache Atlas (Lineage, Classification) + KnoxRetained

Open-source governance the enterprise operates on its own infrastructure; policies and lineage run on any Ranger and Atlas. Retained on the open-source seam.

SDX Operated by Cloudera in the Customer's Cloud AccountDelegated

The same Ranger, Atlas, and Knox services run by Cloudera's control plane on cloud environments. Provider-operated open source the enterprise can substitute: Delegated.

Cloudera Lakehouse Optimizer (Iceberg Maintenance Policies and Automation)Ceded

Cloudera's policy and automation service over Iceberg tables with its own APIs, beyond the standard. Ceded.

Cloudera Data Catalog (Discovery, Enriched Metadata)Ceded

Cloudera's catalog UI over SDX metadata. Cloudera's surface: Ceded.

Cloudera Data Lineage (Octopai SaaS; 60+ External Integrations) + Unified Data Fabric Knowledge GraphCeded

Automated cross-system lineage harvesting as a SaaS service and the Trino-plus-SDX knowledge graph (November 2025). Cloudera's service and index: Ceded.

NVIDIA-Provided

No NVIDIA Layer 1A Dependency

Nothing at this layer runs on accelerators.

Gap Analysis

This is the layer Cloudera was built for. The store is open: Apache Iceberg tables on the Hadoop Distributed File System (HDFS), Apache Ozone, or the cloud object stores, with Cloudera's Iceberg REST Catalog exposing them to any engine that speaks the standard, and the Lakehouse Optimizer (September 2025) handling compaction and maintenance under Cloudera-defined policies. Governance is the Shared Data Experience (SDX): Apache Ranger for role- and attribute-based access, masking, and row filtering; Apache Atlas for lineage and classification; Knox for perimeter; the Cloudera Data Catalog for discovery and enriched metadata; Kerberos, auto-TLS, and key management underneath, one policy set enforced across on-premises, cloud, and edge. Cloudera Data Lineage (formerly Octopai, a SaaS service also sold on the AWS and Azure marketplaces) adds automated harvesting of lineage across more than sixty external systems, and the November 20, 2025 update wove Trino federation, SDX, and Data Lineage into a unified data fabric that maps the estate into a searchable knowledge graph. The buyer gets a system of record for analytics and AI with a catalog Layer 2C can query, and open formats under all of it. The architect's concern is which parts lift, and the answer is most of the foundation and little of the tooling. Iceberg tables and the REST catalog interface are standards; Ranger and Atlas are Apache projects whose policies and lineage graphs run on any Ranger and Atlas; the Data Catalog, the Lakehouse Optimizer's policies, and the Octopai lineage service are Cloudera's, and what accumulates in them is discovery metadata, optimization policy, and cross-system lineage. Calibration: Snowflake and Databricks read strong on governed lakehouses with catalogs (Horizon, Unity); NetApp and VAST strong on data foundations with metadata catalogs; Elastic moderate without a catalog. Cloudera has the catalog, the lineage, the classification, and the open format. Strong.

Borrowed Judgment

Low, split by path. The Iceberg REST catalog is a multi-vendor standard: Delegated. Ranger, Atlas, and Knox the enterprise runs itself on premises are open-source governance it operates without Cloudera: Retained; the same services Cloudera operates in the customer's cloud account read Delegated. HDFS and Ozone the enterprise runs are Retained on the open-source seam. The Lakehouse Optimizer, the Cloudera Data Catalog, and Cloudera Data Lineage (Octopai) are Cloudera's: Ceded. The runtime tradeoff is Cloudera's platform executing the policies and classifications the enterprise wrote, with every knob exposed: vendor decides, visible, overridable, Delegated, the Databricks reading.

Working Notes

Public evidence that moves the cell: nothing upward from strong. Cloudera Data Lineage is SaaS-only; it reads on-premises estates but doesn't deploy there.

Layer 1B · RetrievalContext Management & RetrievalRetrieval Is in Technical Preview: RAG Studio, Semantic Search, and Tool Calling Unscoreddecides: absent · Absent

Low-latency retrieval for RAG — vector/hybrid search, context windows

Vendor-Provided

NVIDIA-Provided

NVIDIA NIM Embedding and Reranking Models Through Cloudera AI Inference

The embedding and reranker models Cloudera AI Inference serves run as NVIDIA NIM microservices on the customer's GPUs, scored at 2B as serving; nothing generally available at this layer consumes them on Cloudera's paper.

Gap Analysis

Cloudera's retrieval layer exists, and none of it has left preview. RAG Studio, one of the four AI Studios in the AI Workbench, assembles retrieval-augmented chat over enterprise documents with Qdrant embedded by default (OpenSearch and ChromaDB as alternatives), embeddings and generation from Cloudera AI Inference, Amazon Bedrock, Azure OpenAI, or OpenAI, NiFi ingestion, automatic scoring of answers for faithfulness and relevance, and tool calling to Model Context Protocol (MCP) servers; Cloudera Semantic Search is OpenSearch with vector indexing run as a Cloudera service. The AI Studios overview page carries a Technical Preview badge and says the feature isn't recommended for production; the Data Access documentation marks Semantic Search Technical Preview; the RAG Studio guide marks tool calling Technical Preview. The June 2026 release that made Agent Studio generally available on premises says nothing about RAG Studio. What the enterprise can do today on Cloudera's paper is serve embedding and reranking models through AI Inference (scored at 2B) and deploy open-source or partner vector stores of its own (the Anywhere Cloud marketplace lists Pinecone, Milvus, and Qdrant, without product documentation). The buyer gets a retrieval toolkit to try, and builds production retrieval themselves. Applying the GA-gate: preview surfaces go to notes and never into scored components, and the layer's function has no generally available Cloudera product behind it. Trino federation supplies live structured context to agents, which is a data-access capability scored at 1A. Calibration: CoreWeave reads gap with authority Absent as enterprise-provided retrieval; Elastic, Databricks, and MongoDB strong on native retrieval engines; Cohere and Cloudflare moderate on shipped, generally available primitives. Cloudera has primitives in preview. Gap, Absent, with the preview surfaces watch-listed and dated.

Borrowed Judgment

None to borrow at the layer's function today: the retrieval the enterprise runs in production on Cloudera is its own assembly of open-source vector stores and served models. Absent: nothing generally available offered, nothing inherited.

Working Notes

Grade moved during review: drafted moderate on RAG Studio and Semantic Search, moved to gap when both surfaced as Technical Preview in the current docs (GA-gate). Watch-list, notes only, and these are what move the cell: RAG Studio (Technical Preview in the cloud AI Studios overview; not named in the 1.5.5 SP3 GA release), Cloudera Semantic Search (Technical Preview), RAG Studio tool calling and hosted or external MCP tools (Technical Preview; reranking not available with Azure OpenAI), the Anywhere Cloud marketplace's vector databases (no product documentation). When RAG Studio reaches GA, its components split by ownership: embedded or customer-run Qdrant and ChromaDB Retained, provider-operated stores Delegated, proprietary embedding spaces (OpenAI, Azure, Bedrock, NVIDIA-owned models) Ceded under the carve-out, open-weight embeddings self-served Delegated. Public evidence that moves the cell: a GA badge on RAG Studio or Semantic Search.

Layer 1C · PipelinesData Movement & PipelinesDataFlow, Streaming, and Data Engineering: Any Source to Any Destination on Every Pathdecides: vendor · Delegated

Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering

Vendor-Provided

Apache NiFi, Kafka, Flink, and Spark Artifacts and Self-Run Engines (Flows, Topics, Jobs, DAGs)Retained

The flows, topics, jobs, and code the enterprise authors, and the Apache engines it runs on its own hardware or Kubernetes. Retained on the open-source seam.

Cloudera Distribution, Operators, and Management (Flow Management Operator, Streams Messaging Operator, Cloudera Manager)Ceded

The licensed packaging and lifecycle management around the Apache engines; the operators require an active Cloudera license and Cloudera-credentialed images. Cloudera's: Ceded.

Cloudera DataFlow and Streaming as Managed Cloud ServicesDelegated

NiFi, Kafka, and Flink operated by Cloudera's control plane in the customer's cloud account. Managed services over open engines the enterprise can substitute: Delegated.

Flow Designer + ReadyFlows + DataFlow Functions (Serverless NiFi on Lambda, Azure Functions, Cloud Functions)Ceded

Cloudera's authoring surface, templates, and serverless packaging around NiFi. Cloudera's: Ceded.

SQL Stream Builder + Schema Registry + Streams Replication Manager + SurveyorCeded

Cloudera's Flink SQL authoring, proprietary schema management, cross-cluster replication, and Kafka observability. Cloudera's: Ceded.

Cloudera Data Engineering (Managed Spark + Airflow Workflow Orchestrator)Delegated

Managed Spark with Airflow-based orchestration running the enterprise's jobs and DAGs. Open engines under Cloudera's service: Delegated.

NVIDIA-Provided

No Current NVIDIA Layer 1C Dependency

GPU acceleration for Data Engineering Spark jobs was a Technical Preview through Data Services 1.5.5 SP1 and was removed in SP2 and later, citing security concerns and Spark end of support; unscored.

Gap Analysis

Cloudera's movement layer is Apache's, packaged three ways. Cloudera DataFlow runs Apache NiFi 2.0 on premises through Flow Management, in the cloud as a managed service, and as a Kubernetes operator, with more than 450 connectors across Kafka, Iceberg, Delta Lake, BigQuery, MongoDB, Salesforce, and the rest, Flow Designer for authoring, ReadyFlows for author-once-deploy-anywhere templates, and DataFlow Functions to run flows serverless on AWS Lambda, Azure Functions, and Google Cloud Functions. Cloudera Streaming is Apache Kafka (Streams Messaging), Apache Flink with SQL Stream Builder (Streaming Analytics), Schema Registry, Streams Replication Manager for cross-cluster replication, and Surveyor for Kafka observability. Cloudera Data Engineering runs Spark with Airflow orchestration (the Workflow Orchestrator), and change data capture (CDC) arrives through NiFi's CDC processors and Debezium in Kafka Connect. Lineage is Atlas inside the platform and Octopai across it. The buyer gets an any-to-any movement layer, deployable on the same paths as the rest of the platform. The architect's concern is the same as Snowflake's Openflow, which is also Apache NiFi: the flows are NiFi flows and export as NiFi flows; the authoring surface, the templates, the serverless packaging, and the licensed operators and distribution are Cloudera's. No KV-cache tiering, and cost-aware movement is replication policy rather than a placement engine. Calibration: Snowflake reads strong on Openflow plus streaming plus declarative pipelines; Databricks strong on Lakeflow; Qlik strong on any-to-any CDC; VAST strong on DataEngine; NetApp moderate on a fixed ingest pipeline. Cloudera is Snowflake's shape with more deployment paths and the same upstream. Strong.

Borrowed Judgment

Low, and on the open-source seam more than any peer. NiFi flows, Kafka topics and Connect configurations, Flink jobs, and Spark code run on the Apache projects outside Cloudera: the artifacts and the self-run engines read Retained; Cloudera's distribution, operators, and management (the Flow Management Operator requires an active Cloudera license and Cloudera images; Schema Registry is Cloudera proprietary) read Ceded; the managed cloud DataFlow and Streaming services read Delegated. The runtime tradeoff is NiFi's processors (record readers, schema inference, connector behavior) and the Spark and Flink engines executing flows, jobs, and DAGs the enterprise wrote, with every knob exposed: vendor decides, visible, overridable, Delegated, the Snowflake Openflow reading for the same Apache NiFi architecture.

Working Notes

Watch-list, notes only: GPU acceleration for Data Engineering (Technical Preview through 1.5.5 SP1, removed in SP2 and later). Inference flagged: whether DataFlow Functions carries a GA badge on all three serverless targets (documented for each; no date found). Reviewed for awareness: the authority reading follows Snowflake's Openflow (vendor / Delegated, the same NiFi architecture) rather than the Elastic and MongoDB code / Retained reading; the split is logged for the next reconcile. Public evidence that moves the cell: nothing upward from strong.

Layer 2A · OrchestrationInfrastructure OrchestrationPlatform-Scoped Kubernetes with GPU Quota Pools; Anywhere Cloud Control Plane Launched, Undocumenteddecides: vendor · Ceded

GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization

Vendor-Provided

ECS Kubernetes Substrate (Cloudera Embedded Container Service; Kubernetes as the Consumed Interface)Delegated

The Kubernetes Cloudera ships for on-premises Data Services, consumed through the standard Kubernetes API with the NVIDIA device plugin advertising customer GPUs. A distribution behind a multi-vendor interface: Delegated, the VKS and NKP reading.

Cloudera Manager + ECS Lifecycle + Quota Management (Hierarchical CPU, GPU, and Memory Pools for Cloudera AI)Ceded

Cloudera's management plane and quota layer over the substrate, with GPU allocation per pool and manual eviction when quotas bind. Cloudera's management opinions: Ceded.

Red Hat OpenShift, Customer-Run, as the Alternative SubstrateRetained

Customer-operated OpenShift with the NFD and GPU Operators hosting the same Data Services. A self-run Kubernetes substrate the enterprise operates: Retained.

Cloudera Control Plane for Cloud Environments (Provisioning into the Customer's AWS, Azure, or Google Cloud Account)Ceded

Environment creation and data service scaling inside the customer's cloud account, billed per Cloudera Compute Unit. Cloudera's control plane: Ceded.

NVIDIA-Provided

NVIDIA Device Plugin Advertises Customer GPUs to the Scheduler

On ECS the NVIDIA device plugin advertises GPUs once drivers are installed; on OpenShift the customer installs the Node Feature Discovery and GPU Operators. Quota pools then allocate those GPUs to Cloudera AI workbenches.

Gap Analysis

Cloudera orchestrates its own services on Kubernetes it ships or the customer runs, and it schedules GPUs inside that boundary. On premises, Cloudera Data Services run on the Embedded Container Service (ECS), Cloudera's own Kubernetes distribution, or on the customer's OpenShift; Quota Management defines hierarchical resource pools with CPU, GPU, and memory limits by business unit, user, application, or data service, with GPU allocation set per pool by slider and a utilization dashboard, and when an administrator lowers a quota below current usage the pending workloads stay blocked until the administrator evicts running ones. Data Engineering's share of GPU quota is Technical Preview and unscored. In the cloud, the Cloudera control plane provisions environments into the customer's AWS, Azure, or Google Cloud account and scales the data services within them. Anywhere Cloud, with Taikun (acquired August 4, 2025) as the compute layer, is the unified control plane meant to deploy all of it from public cloud to air-gapped, announced August 19, 2026 with ADMIRAL, ExxonMobil, IQVIA, IXEN.ai, and Mastercard as design partners and no availability statement. The buyer gets Kubernetes they don't have to build and GPU quotas across the AI workbenches they run on it. The architect's concern is scope. The scheduler and the quotas govern Cloudera's workloads only; a shared enterprise GPU fleet running other platforms isn't Cloudera's to place, and Anywhere Cloud's control plane is launched without documentation. Hierarchical quota pools are configuration on Kubernetes' scheduling judgment, and eviction when they bind is capacity management, not an override. Calibration: Nutanix and VMware read strong on fleet-wide orchestration heritage; Snowflake moderate on managed compute with GA GPU pools, platform-scoped; Elastic moderate on ECE and ECK; NetApp moderate on a data control plane without GPU scheduling. Cloudera is Snowflake's shape on the customer's own metal. Moderate.

Borrowed Judgment

Bounded. Quota pools, node roles, and environment sizing are configuration the enterprise sets on Kubernetes' scheduling judgment as Cloudera packages it; evicting a running workload to free quota is capacity management inside those bounds, not a reversal of a placement decision. Vendor decides, visible, not overridable, Ceded, the Elastic, Snowflake, and Nutanix reading. Kubernetes is the consumed interface, so ECS's substrate reads Delegated the way VMware's VKS and Nutanix's NKP do, with Cloudera's management plane Ceded; customer-run OpenShift reads Retained as a substrate the enterprise operates.

Working Notes

Watch-list, notes only, and this is what moves the cell: Anywhere Cloud with Taikun as a unified control plane (announced August 19, 2026; design partners named; no availability statement or product documentation yet); Data Engineering GPU quota sharing (Technical Preview). Public evidence that moves the cell: Anywhere Cloud product documentation showing customer-administered placement across environments.

Layer 2B · RuntimeApplication Runtime & ExecutionAI Inference (NIM, Hugging Face on vLLM, Triton; Any Model; On-Prem and Cloud) + Agent Studio (GA On Premises)decides: model · Delegated

Model serving, agent execution, inference APIs, distributed inference

Vendor-Provided

Cloudera AI Inference Service (NIM, Hugging Face on KServe with vLLM, Triton with ONNX; On-Prem and Cloud; Scale-to-Zero; Auth and Encryption)Ceded

The serving service with an OpenAI-compatible API for LLMs and Open Inference Protocol APIs for classic models, on the customer's GPUs or cloud. Cloudera's service around NVIDIA's and open runtimes: Ceded.

Model Access via OpenAI-Compatible and Open Inference Protocol Interfaces; NIM, vLLM, and Triton RuntimesDelegated

Standard-shaped inference interfaces over NVIDIA's NIM and Triton (under NVIDIA AI Enterprise through Cloudera's paper) and the open Hugging Face and vLLM path. Delegated on the inference-interface and inference-runtime rulings.

Customer-Owned Models and Fine-Tunes Loaded into the RegistryRetained

The weights the enterprise trained or fine-tuned and loads into the AI Registry to serve. The enterprise's own artifacts: Retained. Third-party open weights served on the platform (Nemotron, Hugging Face community models) read Delegated with the runtime chip.

Cloudera AI Registry + AI Workbench (Projects, Jobs, Notebooks, Copilot)Ceded

The model registry and the workbench the studios live in. Cloudera's: Ceded.

Agent Studio (GA On Premises, June 2026: CrewAI Workflows, Local stdio MCP Servers, Native Function Calling, Workflow Guardrails, Auditing, Service-Account Deployment) + Fine Tuning and Synthetic Data StudiosCeded

The agent builder and its sibling studios, shipped as a prebuilt runtime image. Cloudera's studios, and the agents and workflows built in them: Ceded.

Customer-Authored Tool Code in Agent Studio (Python Tool Classes, Isolated in a Namespace on the Customer's Hardware)Retained

The tools the enterprise writes as plain Python in CrewAI tool classes, run in a bubblewrap namespace inside the Workbench on the customer's own infrastructure. Written Retained; escalated against the narrowed customer-tools ruling (open runtime on customer hardware, inside a vendor product).

NVIDIA-Provided

NVIDIA NIM and Triton Runtimes Inside Cloudera AI Inference; Agent Studio Designed With NVIDIA

Two of the inference service's three runtime paths are NVIDIA's (NIM for NVIDIA-packaged models, Triton with the ONNX backend for deep-learning models), under NVIDIA AI Enterprise licensing, with Blackwell named for on-premises; the third, Hugging Face Model Server on KServe with vLLM, isn't. Agent Studio was built with NVIDIA.

Gap Analysis

Cloudera runs models and agents where the data is, and since February 9, 2026 that includes the customer's data center. Serving: Cloudera AI Inference deploys and scales models through three runtimes, NVIDIA NIM for NVIDIA-packaged models (including Nemotron), Hugging Face Model Server on KServe with vLLM for transformer text-generation and embedding tasks, and NVIDIA Triton with the Open Neural Network Exchange (ONNX) backend for deep-learning and classic models, behind an OpenAI-compatible API for large language models (LLMs) and Open Inference Protocol APIs for traditional models, with scale-to-zero autoscaling, authentication, authorization, and encryption, and the Cloudera AI Registry storing and versioning the models; it runs on premises and in the cloud from the same control plane. Agents: Agent Studio (designed with NVIDIA, built on CrewAI, which joined Cloudera's ecosystem in December 2024) reached general availability on premises in Data Services 1.5.5 SP3 (June 2026) as a low-code to high-code builder for multi-agent workflows with custom tools the enterprise writes in Python, local MCP servers run as stdio processes inside the Workbench, native function calling, workflow-level guardrails that intercept prompts, tool calls, and outputs and can block them, two-level auditing, deployment as workflows under a service account, and built-in observability, shipped as a prebuilt runtime image for air-gapped installs; the cloud AI Studios overview still carries a Technical Preview badge, so the row scores the on-premises path. Fine Tuning Studio and Synthetic Data Studio sit beside it; the AI Workbench MCP Server (December 2025, Apache 2.0, unsupported) lets external agents drive Workbench projects and jobs. The buyer gets any-model serving on their own GPUs and an agent factory beside the lakehouse. The architect's concern splits along the runtime seam and the studio seam. Two of three serving runtimes are NVIDIA's through Cloudera's paper and the third is open, so the model layer lifts with the models the enterprise brings; the AI Registry, the Workbench, and the studios are Cloudera's, and the agents and workflows built in Agent Studio exist only there. The tool code the enterprise writes inside it is plain Python in a CrewAI tool class, isolated in a bubblewrap namespace on the customer's own hardware, and whether that reads Retained or Ceded is escalated. Calibration: Nutanix reads moderate on platform-native serving with agents pre-GA; VAST strong on AgentEngine; Kamiwaza moderate on a runtime with workrooms; Databricks strong on serving plus GA agents; VMware moderate on a model runtime plus an agent builder. Cloudera has any-model serving on both paths plus a GA agent builder with documented gates. Strong.

Borrowed Judgment

Split three ways. Model access through the OpenAI-compatible and Open Inference Protocol interfaces is Delegated on the inference-interface ruling, and the runtimes are NVIDIA's (NIM, Triton) through Cloudera's paper or open (Hugging Face on KServe with vLLM), Delegated on the Nutanix inference-runtime reading; the models the enterprise trained or fine-tuned are Retained, and third-party open weights it serves (Nemotron, Hugging Face) are Delegated; the AI Registry, the Workbench, Agent Studio, and its sibling studios are Cloudera's, Ceded, and so are the agents and workflows built in them. Customer-authored tool code is written Retained and escalated: it's the enterprise's Python on the enterprise's hardware inside Cloudera's runtime. The model decides which tool to call; the workflow-level guardrails intercept the call before the tool runs and can block it, and the enterprise's own tool code executes in its own environment: model decides, visible, overridable, Delegated, the Kamiwaza, VAST, and Nutanix reading for self-hosted platforms.

Working Notes

Escalated: customer-authored tool code in Agent Studio (Retained written; Ceded argued as platform-bound; the same open-runtime question as Cloudflare's Workers code). Watch-list, notes only: RAG Studio and RAG Studio tool calling (Technical Preview); remote MCP servers (Agent Studio supports local stdio servers only); the cloud AI Studios overview's Technical Preview badge; a per-action human approval step (guardrails block or mask; no approval prompt is documented). Public evidence that moves the cell: nothing upward from strong.

Layer 2C · ReasoningAgentic Infrastructure — The Reasoning PlaneData Governance Reaches the Models; Agent Governance Lives Inside Agent Studiodecides: absent · Absent

Policy-driven placement and resource coordination — the Autonomy Layer

Vendor-Provided

NVIDIA-Provided

No NVIDIA Layer 2C Dependency

No reasoning-plane product exists to carry a dependency.

Gap Analysis

Cloudera's governance is SDX, and SDX governs data. Ranger policies follow the tables into Trino, Spark, and the AI Workbench; the AI Registry versions models with access control; AI Inference authenticates and authorizes callers; Data Visualization logs and traces AI queries. Inside Agent Studio, and only there, the agent-governance legs appear: workflow-level guardrails that intercept prompt inputs, tool calls, and tool and MCP outputs and block or mask them; two-level auditing of who created, changed, or deleted workflows, agents, tools, and MCP servers, and of executions; deployment of workflows under a dedicated service account rather than a user; built-in observability of the studio's own workflows. Applying the test: Intelligence-2C asks which agent may act, under whose identity, within what policy, across the estate. Cloudera answers it for the agents built in Agent Studio, as properties of that runtime, the way North's autonomy policies and Cowork's permissions answer it for theirs. There's no registry treating agents as governed assets across systems, no request-time gateway with policy over tool calls from other runtimes, no first-class agent identity governed centrally (the service account is an execution identity for deployed workflows), no cross-agent orchestration outside one Agent Studio workflow, and no observability of agents that aren't Cloudera's. Infrastructure-2C, placement across model, cost, and compliance tiers, is absent as everywhere. Calibration: NetApp reads gap on the same shape (data guardrails aren't agent governance); Cohere gap with North's autonomy policies, RBAC, and agent directory inside one product; Anthropic gap on runtime permissions; Nutanix moderate on an agent gateway over any runtime's traffic; Elastic moderate on registry, orchestration, and observability legs in Kibana; Databricks moderate on Unity Catalog's on-behalf-of authorization plus a gateway. Cloudera has Databricks' data governance and Cohere's in-product agent controls, and neither peer's plane. Gap, Absent.

Borrowed Judgment

None to borrow at the layer's function: data policy is scored at 1A and Agent Studio's guardrails, audit, and service-account execution at 2B, where they strengthen the override reading. To the extent a reasoning plane exists in a Cloudera estate, the enterprise built it. Absent: nothing offered, nothing inherited.

Working Notes

Reviewed for awareness: whether Agent Studio's GA guardrails, auditing, and service-account execution lift the cell to thin moderate (ruled subthreshold on the Cohere and Anthropic precedent: properties of one runtime, not a plane over agents). Public evidence that moves the cell: an agent registry, a policy gateway over tool calls from other runtimes, or a centrally governed agent principal in SDX or the AI Workbench.

Layer 3 (+1) · ApplicationsAI Application Layer — The Value PlaneAI Visuals and Annotation in Data Visualization, a Copilot in the Workbench; SQL Assistant in Previewdecides: vendor · Ceded

AI-powered business capabilities — business logic, workflow automation

Vendor-Provided

Cloudera Data Visualization AI Visual (GA January 31, 2025) + AI Annotation (GA December 16, 2025) with AI Query Logging, Redaction, and TraceabilityCeded

Conversational analytics and generated insights inside the BI layer, on premises and cloud. Cloudera's application: Ceded.

Cloudera Copilot in the AI Workbench (GA in ML Runtimes 2024.10.1)Ceded

Coding assistance in Workbench sessions. Cloudera's assistant: Ceded.

NVIDIA-Provided

No NVIDIA Layer 3 Dependency

The assistants call whatever models the platform serves.

Gap Analysis

Cloudera's own applications are the visualization layer and the assistants, and the GA ones are narrower than the 2024 announcement. Cloudera Data Visualization's AI Visual (conversational analytics over governed data) reached general availability in 8.0.0 on January 31, 2025, and AI Annotation (generated insights and summaries for visuals and dashboards) in 8.0.9 on December 16, 2025, with consistent logging, redaction, and traceability for AI queries; Data Visualization runs on premises since May 2025. Cloudera Copilot reached general availability in the AI Workbench runtimes (2024.10.1). The Hue SQL AI Assistant in Cloudera Data Warehouse, announced with the others in June 2024, remains Technical Preview and unscored. The RAG Studio chatbots and Agent Studio workflows the enterprise builds are applications too, but the enterprise's, on Cloudera's runtime; the Anywhere Cloud marketplace's partner applications (PuppyGraph and others) have no product documentation yet. The buyer gets conversational analytics and a coding assistant over governed data, and builds the rest. Applying the test: the layer asks for AI-powered business capabilities, and what ships first-party and generally available is analytics and developer assistance, not a business process. Real and narrow. Calibration: Databricks reads strong on AI/BI Genie plus apps plus a marketplace; Qlik strong on an analytics value plane; Nutanix and VMware moderate on platform-enabled applications; NetApp gap as a foundation beneath others' apps. Cloudera is between Nutanix and Databricks: first-party AI features in a BI layer and a copilot, platform-enabled beyond. Moderate.

Borrowed Judgment

Low and not compounding. The AI Visual's answers are over the enterprise's governed data and the Copilot's suggestions are advice; Data Visualization dashboards and the AI features' opinions are Cloudera's and rebuild elsewhere; the applications built in the studios are the enterprise's on Cloudera's runtime. Vendor decides, visible, not overridable, Ceded, the first-party tooling reading.

Working Notes

Watch-list, notes only: the Hue SQL AI Assistant (Technical Preview); Anywhere Cloud marketplace partner applications (no product documentation; when documented, open-source self-run engines read Retained, proprietary partner applications Ceded to their owners, and Delegated only behind a multi-vendor consumed interface). Public evidence that moves the cell: a first-party application that runs a business process.

Summary Finding

Cloudera is a data platform that runs wherever the enterprise puts it, and the map reads it strongest at the two layers that were always its own. Layer 1A is strong on an Apache Iceberg lakehouse with the Shared Data Experience (SDX) governing it and Octopai lineage across the estate; Layer 1C is strong on DataFlow (Apache NiFi 2.0 with 450-plus connectors), Streaming (Kafka, Flink), and Data Engineering (Spark, Airflow) on every deployment path; Layer 2B is strong on Cloudera AI Inference serving any model on the customer's own GPUs since February 2026 and an Agent Studio that reached general availability on premises in June 2026. Two layers are moderate (2A on platform-scoped Kubernetes with GPU quota pools; Layer 3 on AI features in Data Visualization and a Copilot), and three are gaps: no substrate (the customer supplies it on every path), no generally available retrieval product (RAG Studio and Semantic Search are Technical Preview), and no reasoning plane beyond Agent Studio's own guardrails and audit.

The capture is the least coupled on the map, because so much of the platform is Apache. Iceberg and its REST catalog, Ranger and Atlas policies, NiFi flows, Kafka topics, Flink jobs, Spark code, and the models the enterprise brings all leave with the enterprise, and on premises the enterprise runs the open engines itself. What's Cloudera's is the packaging and the opinions on top: the distribution and operators (licensed), Flow Designer and ReadyFlows, SQL Stream Builder and Schema Registry, the Lakehouse Optimizer, the Data Catalog, the Octopai lineage service, the AI Inference service around NVIDIA's runtimes, the AI Registry and Workbench, and Agent Studio's workflows. Twenty-five components read Retained 6, Delegated 6, Ceded 13; the Retained count would be five if Keith rules that tool code written inside Agent Studio is platform-bound.

The buyer's trade: an open lakehouse and movement layer they can take with them, plus any-model serving and an agent builder on their own hardware, in exchange for a retrieval stack still in preview, a control plane (Anywhere Cloud, launched August 19, 2026) whose availability isn't yet documented, and a governance story that governs data rather than agents. What would move a cell: RAG Studio or Semantic Search reaching GA (1B), Anywhere Cloud documentation showing customer-administered placement across environments (2A), or an agent registry or gateway in SDX (2C). The VAST Data AI factory partnership is a reference architecture two sales teams sell, not a substrate on Cloudera's paper.

4+1 Layer AI Infrastructure Model · Vendor Assessment Series · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com