# Dataiku (DSS 15 + LLM Mesh + Agent Hub + Govern + Dataiku Cloud) — 4+1 Layer AI Infrastructure Assessment

> Mapped to the 4+1 Layer AI Infrastructure Model  
> Version: v1.0 - 4+1 v2: Authority Split · Date: September 10, 2026  
> Source: Dataiku DSS 15 reference documentation (release notes 15.0.0 August 14, 2026 and 15.0.1 September 10, 2026; Generative AI and LLM Mesh: LLM connections, Running Hugging Face models, Guardrails, Cost Control, Rate Limiting, Fine-tuning, LLM Mesh API, Evaluating LLMs; Adding Knowledge to LLMs: vector stores, search settings, document-level security, GraphRAG, Automated RAG Optimization, RAG guardrails; AI Agents: introduction, Simple and Structured Visual Agents, Code Agents, Agent Skills, managed tools, human approval, MCP server, A2A, tracing, Agent Review, Agent Evaluation, Agent Hub, Agent Chat; Semantic Models; Catalog; Data Lineage; Metastore catalog; Streaming data; Elastic AI computation, managed Kubernetes clusters, EKS, DGX Systems; Installing and setting up, Cloud Stacks for AWS, custom install; API Node and API Deployer, deploying on Kubernetes; Production deployments and bundles; MLOps and Unified Monitoring; AI Governance: items, process features, deployment policies; Process Mining, RFx Accelerator, Manufacturing Operations; Stories; Dataiku Applications; Dataiku AI and Cobuild; Security; Agent Hub release notes 1.6.0 August 21, 2026; Trace Explorer release notes); Dataiku Knowledge Base (Compute and Resource Quotas on Dataiku Cloud: compute engines, elastic AI compute capacity, containerized execution configurations for standard or GPU compute; Dataiku Solutions catalog); Dataiku product pages (plans and features, LLM Guard Services, Agent Management, Expert-to-Agent); Dataiku press releases ($350M ARR, October 17, 2025; Cobuild general availability, June 2026; Expert-to-Agent, May 2026); the dataiku-api-client-python repository (Apache 2.0). Peer-reviewed cell by cell through the labs claims ledger (dataiku-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (dataiku-row-agy); totals and escalated items in reviews/dataiku-judgment.md.  
> Published by: The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com  
> Author: Keith Townsend

[Full interactive assessment](https://layer2c.com/assessment/dataiku) · [Methodology](https://layer2c.com/methodology) · [What Is Layer 2C?](https://layer2c.com/what-is-layer-2c)

## Executive Summary

Dataiku is a software-only AI platform that sits on top of whatever the enterprise already owns, and the map reads it as strong at the two layers it was built for and at the one it rebuilt itself around: data pipelines (Layer 1C), the runtime that serves models and runs agents (Layer 2B), and the applications business users touch (Layer 3). Governance, retrieval, and orchestration are real but partial: a catalog, lineage, quality rules, and a feature store over data stored elsewhere (1A moderate); Knowledge Banks that orchestrate retrieval across Chroma, Elasticsearch, Pinecone, pgvector, Snowflake Cortex Search, Databricks AI Search, and others without a retrieval engine of Dataiku's own (1B moderate); Kubernetes clusters Dataiku creates and drives on the enterprise's cloud accounts (2A moderate); and an agent-governance plane with four of the five legs the instrument asks for, missing only a first-class agent identity (2C moderate). Layer 0 is a gap by design: Dataiku sells no compute, network, or storage, on any of its three deployment shapes.

The capture is decoupled and mostly invisible, which is the dangerous kind. Every openness claim is true: data stays on the enterprise's connections in Parquet, Iceberg, and the warehouse's own tables; models come from the enterprise's own OpenAI, Anthropic, Bedrock, Vertex, Mistral, Snowflake Cortex, or Databricks accounts through the large language model (LLM) Mesh; open Hugging Face weights run on the enterprise's own graphics processing unit (GPU) under vLLM; the Python client is Apache 2.0. What accumulates in Dataiku is the opinion layer above all of that: the Flow and its visual recipes, dataset definitions, lineage, quality rules, Knowledge Bank configurations, agent definitions and Structured Visual Agent block logic, managed-tool descriptors, guardrail pipelines, quotas, Govern workflows and sign-offs, and Agent Hub configurations. None of it lifts to another platform. Of the row's components, the great majority read Ceded; the Delegated chips are the standard interfaces Dataiku sits behind (managed Kubernetes it provisions, provider model application programming interface (API), open weights it serves); the four Retained chips are things the enterprise possesses without Dataiku (the transformation code inside its recipes, a Kubernetes cluster it built and attached, open vector stores it runs, the fine-tuned and imported model weights it owns), and every Dataiku-shaped object, including customer-authored tools bound to Dataiku's classes, stays behind.

The buyer's trade: one governed workbench where analysts, data scientists, and business users build pipelines, models, and agents against the enterprise's existing estate, with decision-time gates the instrument credits (human approval on tool calls, deterministic blocks in Structured Visual Agents, configured guardrails before and after model calls, sign-off that blocks a bundle from deploying), in exchange for a platform whose accumulated logic has no exit. The decision-authority readings say the same thing from the other side: Delegated at every scored layer, because the enterprise can see and override the runtime call, and never Retained, because the engine making the call is always Dataiku's. What would move cells: a first-class agent principal (2C), Agent Management shipping in October 2026 as announced (2C), a native retrieval engine or hybrid search independent of the backing store (1B), and streaming leaving its Experimental badge (1C).

## Layer Status

| Layer | Status | Classification |
|---|---|---|
| Layer 0 · Compute | ○ Not Dataiku's Layer (By Design); Dataiku Cloud on Dataiku's Accounts, Cloud Stacks and Custom Installs on Yours | Compute & Network Fabric |
| Layer 1A · Storage | ◑ Catalog, Column-Level Lineage, Quality Rules, and a Feature Store Over Data Stored on the Enterprise's Own Connections | Data Storage & Governance |
| Layer 1B · Retrieval | ◑ Knowledge Banks Over Ten Vector Stores + Semantic Models + Document-Level Security; No Retrieval Engine of Its Own | Context Management & Retrieval |
| Layer 1C · Pipelines | ● The Flow: Visual and Code Recipes on Any Engine, Any Source to Any Destination, Scenarios, Lineage; Streaming Experimental | Data Movement & Pipelines |
| Layer 2A · Orchestration | ◑ Elastic AI: Kubernetes Clusters Dataiku Creates and Drives, Execution Configs With GPU Limits, Fleet Manager; No Fair-Share Scheduler | Infrastructure Orchestration |
| Layer 2B · Runtime | ● LLM Mesh Gateway + Local Open-Weight Serving on vLLM + Visual, Structured, and Code Agents With Human Approval + API Node Model Serving | Application Runtime & Execution |
| Layer 2C · Reasoning | ◑ Gateway, Govern Registry With Sign-Off and a Deployment Gate on Bundles, Multi-Agent Orchestration, Observability; No First-Class Agent Identity; Agent Management Due October 2026 | Agentic Infrastructure — The Reasoning Plane |
| Layer 3 (+1) · Applications | ● Agent Hub for Business Users, Three Business Applications, Stories and Dashboards, Cobuild; Forty-Plus Unsupported Solutions Unscored | AI Application Layer — The Value Plane |

## DAPM Portability Profile (components)

| Classification | Count | Meaning |
|---|---|---|
| Retained | 4 | I possess the capability and can operate it independently of this provider |
| Delegated | 4 | Someone else provides the capability, but I can substitute that provider without reconstructing my accumulated opinions |
| Ceded | 31 | Changing providers requires reconstructing those opinions |

**Decision authority (per layer, gaps included)**

| Reading | Layers | Meaning |
|---|---|---|
| Retained | 0 | The enterprise, or code it writes or controls, decides |
| Delegated | 7 | Vendor or model decides; the enterprise can see and override |
| Ceded | 0 | Vendor or model decides; no override, often invisible |
| Absent | 1 | Nothing offered, nothing inherited |

## Strongest Layers

- **Layer 1C** (Data Movement & Pipelines) — The Flow: Visual and Code Recipes on Any Engine, Any Source to Any Destination, Scenarios, Lineage; Streaming Experimental
- **Layer 2B** (Application Runtime & Execution) — LLM Mesh Gateway + Local Open-Weight Serving on vLLM + Visual, Structured, and Code Agents With Human Approval + API Node Model Serving
- **Layer 3 (+1)** (AI Application Layer — The Value Plane) — Agent Hub for Business Users, Three Business Applications, Stories and Dashboards, Cobuild; Forty-Plus Unsupported Solutions Unscored

## Gap Areas

- **Layer 0** (Compute & Network Fabric) — Not Dataiku's Layer (By Design); Dataiku Cloud on Dataiku's Accounts, Cloud Stacks and Custom Installs on Yours

## Layer-by-Layer Detail

### ○ Layer 0 · Compute: Compute & Network Fabric

*Raw compute, networking, and acceleration fabric*  
**Status:** Not Dataiku's Layer (By Design); Dataiku Cloud on Dataiku's Accounts, Cloud Stacks and Custom Installs on Yours

**Decision authority:** Absent (decides: absent; visible: n/a; overridable: n/a; boundary: vendor)

**Gap Analysis:** Dataiku ships software, on three shapes, and none of them is a substrate the enterprise buys from Dataiku. Dataiku Cloud is the hosted plan: Dataiku operates the instance and the Elastic AI compute in its own cloud accounts, with a subscription quota measured in virtual CPU (vCPU), gigabytes of RAM, and parallel activities, Multi-Purpose Credits to burst past it, and GPU containerized execution configurations available only with credits and only in the instance types the hosting region offers. Cloud Stacks deploys Dataiku through Fleet Manager into the enterprise's own AWS, Azure, or Google Cloud tenant, where Dataiku has no access to the data and the virtual machines and Kubernetes clusters are the enterprise's bill. Custom installs run on the enterprise's own Linux servers, on premises or in any cloud, with NVIDIA GPUs the enterprise supplies for local model inference (compute capability 7.5 or higher; a CUDA 13 compatible driver on Kubernetes nodes, and the full CUDA 13 toolkit when models run on the Data Science Studio (DSS) server itself). The buyer gets to keep whatever infrastructure it already has.

The architect's concern is the usual one for a software-only row: the compute decisions on the Dataiku Cloud path are Dataiku's and invisible, and on the other two paths they are the enterprise's and never Dataiku's. Neither is a Dataiku capability at this layer.

Calibration: Elastic, MongoDB, Redis, Databricks, and Snowflake read gap on the same shape (a software or software as a service (SaaS) vendor whose hosted plan runs on a hyperscaler the vendor picks); SAP reads moderate because SAP operates its own data centers and sells the IaaS. Dataiku is the former. Gap, with authority Absent under rule 7 (the Elastic, MongoDB, and Redis reading): the layer is inherited from the hyperscaler or from the enterprise's own estate, not absorbed by Dataiku.

**Borrowed Judgment:** None to Dataiku as a substrate. On Dataiku Cloud the enterprise consumes managed execution capacity (a quota of vCPUs, RAM, and parallel activities, credits, container sizes, GPU configurations by instance type) and inherits the hyperscaler's substrate through Dataiku's account, picking nothing below the container; on Cloud Stacks and custom installs the substrate is the enterprise's own choice and the judgment belongs to whoever it bought the servers, GPUs, and network from. No customer-administered Layer 0 is offered on any path: Absent, the Elastic and Redis reading, with the Cloud path's absorbed substrate named rather than scored.

### ◑ Layer 1A · Storage: Data Storage & Governance

*Durable, governed data foundation — the Governance Catalog that Layer 2C queries*  
**Status:** Catalog, Column-Level Lineage, Quality Rules, and a Feature Store Over Data Stored on the Enterprise's Own Connections

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**Catalog (Catalog Collections, Assets Including Models, Agents, Tools, and Semantic Models; AI Search; Database Explorer)** [DAPM: Ceded]  
The organization-wide index of Dataiku assets and indexed external tables, with curated collections, roles, and metadata completeness. A Dataiku-native catalog structure with no export path to another catalog: Ceded.

**Column-Level Data Lineage Across Projects + Manual Lineage + Flow Document Generator** [DAPM: Ceded]  
Direct column lineage inferred from recipes across datasets and projects, traversed upstream and downstream, with manual remapping where inference stops, and exportable to PDF or images. The graph is computed from Dataiku Flows and lives only there: Ceded.

**Metrics, Checks, and Data Quality Rules** [DAPM: Ceded]  
Dataset-level metrics, checks, and quality rules evaluated in the Flow and in scenarios, with status surfaced in the Catalog. Rules are authored in Dataiku's rule types against Dataiku dataset definitions: Ceded.

**Feature Store (Curated Features from the Flow)** [DAPM: Ceded]  
Features curated and shared across projects for machine learning, with lineage back to the producing recipes. A Dataiku object over Dataiku datasets: Ceded.

**Dataset Definitions, Schemas, Meanings, and Connection Settings (Project Metadata Over the Enterprise's Own S3, Azure Blob, GCS, SQL Warehouses, Iceberg and Hive Catalogs, Glue and Hive Metastores)** [DAPM: Ceded]  
How Dataiku knows the enterprise's data: every dataset's storage, format, schema, storage types, meanings, and partitioning, and every connection's settings, stored as project configuration and exported only as a Dataiku project archive. The bytes stay where the enterprise put them and never move, which is the reassuring fact; the definitions rebuild on exit: Ceded. The storage itself belongs to the storage and warehouse vendors' rows.

**Gap Analysis:** Dataiku isn't a data store; it governs data that lives on the enterprise's connections. Datasets default to those connections (uploaded files and local-connection datasets sit in the DSS data directory, which on Dataiku Cloud is Dataiku-hosted, and Dataiku Cloud is contractually a data processor for that content): cloud storage (S3, Azure Blob, Google Cloud Storage) in Parquet and CSV with fast-path reads and writes, SQL warehouses (Snowflake, BigQuery, Databricks, Redshift, Trino, Synapse, ClickHouse from DSS 15), Apache Iceberg through REST and Hive catalogs, Hive and Glue metastores, and a long list of SaaS sources. Over that estate Dataiku puts the Catalog (DSS 15 widened it from a data catalog to a catalog of assets: datasets, indexed external tables, saved and fine-tuned models, retrieval-augmented LLMs, agents, agent tools, semantic models, with curated Catalog Collections, AI Search in natural language, a Database Explorer for SQL and metastore connections, and a Most Used Datasets view), column-level Data Lineage across datasets and projects with upstream and downstream traversal and manual lineage for what the Flow can't infer, metrics, checks, and Data Quality rules, and a Feature Store curated from the Flow. Security is project-centered: per-project group permissions, per-user connection credentials, discoverable and private projects, and User Isolation Framework on the host, with object-level exceptions (Catalog Collections expose a dataset in READ or metadata-only DISCOVER mode without project access, shared objects give another project read-only use, and project-level API keys carry per-dataset read, write, schema, and metadata permissions). The buyer gets one place to find, trust, and trace what the organization has built, without moving any of it.

The architect's concern is where the governance lives and where it stops. The dataset definitions, catalog entries, lineage graph, quality rules, and feature definitions are Dataiku project metadata in the DSS configuration directory; they describe data on the enterprise's own storage but they don't leave with it, and the export path is a Dataiku project archive. And the enforcement is coarse: Dataiku governs who can open a project or use a connection, not which rows or columns a user sees inside a warehouse table; masking, row filters, and tags stay with the warehouse's own governance. Document-level security exists, but for Knowledge Banks (1B), not for datasets.

Calibration: Palantir and Cloudera read strong on a governance authority that enforces at interaction time (the Ontology's markings; Ranger's row and column policies) over storage the vendor also provides or runs; Databricks and Snowflake strong on catalogs that own the tables' grants. Qlik reads moderate on an open lakehouse the enterprise stores plus a captive trust layer of data products and quality. Dataiku is Qlik's shape without the lakehouse: a captive catalog and lineage layer over data it neither stores nor enforces on. Moderate.

**Borrowed Judgment:** Moderate and captive above the data. The bytes are the enterprise's, in the enterprise's own storage and warehouses, in formats every engine reads; that is the reassuring part and it's true, and it's the storage vendors' portability, not Dataiku's. Everything Dataiku holds at this layer (dataset definitions, catalog structure, lineage graph, quality rules, feature definitions) is Dataiku's model of the estate and rebuilds on exit: Ceded, every chip. The runtime call at this layer is small (lineage inference from the Flow, schema and meaning detection on datasets, quality checks the enterprise wrote), made by Dataiku's engine and overridable per dataset (schema, storage types, meanings, manual lineage): vendor decides, visible, overridable, Delegated, the Databricks and Cloudera reading.

### ◑ Layer 1B · Retrieval: Context Management & Retrieval

*Low-latency retrieval for RAG — vector/hybrid search, context windows*  
**Status:** Knowledge Banks Over Ten Vector Stores + Semantic Models + Document-Level Security; No Retrieval Engine of Its Own

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**Knowledge Banks (Embed Documents and Embed Dataset Recipes; Managed Chroma, FAISS, Milvus Local, Qdrant; Filtering, Query Strategy, Hybrid Search Where Supported, Native and Model-Based Reranking; Multimodal; Document-Level Security)** [DAPM: Ceded]  
The retrieval object and its configuration: chunking, metadata, update mode, search settings, security tokens, and the embedded index files in a Dataiku managed folder. Rebuilt from source on exit, but the opinions are Dataiku's: Ceded.

**Open Vector Stores the Enterprise Runs (Elasticsearch, OpenSearch, Milvus Remote, pgvector on PostgreSQL), Reached Through Their Own Interfaces** [DAPM: Retained]  
The index lives in an open-source engine the enterprise operates, and an unmanaged Knowledge Bank can point at an index built without Dataiku. Retained where the enterprise runs the engine itself; Delegated where it subscribes to a managed variant (Amazon OpenSearch Service, Elastic Cloud, a hosted PostgreSQL). The Cloudera self-run reading.

**Proprietary Retrieval Services Through Dataiku's Connections (Azure AI Search With Semantic Ranker, Pinecone, Vertex Vector Search, Snowflake Cortex Search, Databricks AI Search)** [DAPM: Ceded]  
Single-vendor retrieval services whose index formats, ranking, and search services are the owner's; Snowflake's own row reads Cortex Search Ceded. Ceded to the owner through Dataiku's paper under the channel-substitution rule, the Elastic third-party-endpoints reading.

**Hosted Embedding and Reranking Models Through the LLM Mesh (OpenAI, Azure OpenAI, Bedrock Including Cohere Embed and Rerank, Vertex, Mistral Embed, Microsoft Foundry, NVIDIA Inference Microservices (NIM), Databricks BGE)** [DAPM: Ceded]  
Embeddings from the enterprise's own provider accounts. The vectors are useless without the same model at query time and no second vendor serves the space: Ceded to the model's owner through Dataiku's paper under the carve-out.

**Local Hugging Face Embedding and Reranking Models on the Enterprise's GPUs or CPUs** [DAPM: Delegated]  
Open embedding, reranking, and image-embedding weights the enterprise pulls from Hugging Face (or preloads in the model cache when air-gapped) and serves in its own containers. The weights and the vectors recompute anywhere the same model runs; the serving surface is Dataiku's: Delegated on the Arctic Embed precedent.

**Semantic Models + Semantic Model Query Tool (Entities, Attributes, Relationships, Metrics Over SQL Datasets; LLM-Assisted Generation; Versioning)** [DAPM: Ceded]  
Business context that turns natural-language questions into SQL for agents, the Snowflake semantic-view shape. Defined in Dataiku's semantic model format with no second implementer: Ceded.1.3, July 16, 2026, compatible with DSS 14.4 and above), and whether 'Lab' in a Dataiku plugin name is a maturity badge that keeps the surface out of a scored cell awaits ruling; the reference documentation carries no preview badge.

**Gap Analysis:** Retrieval in Dataiku is a Knowledge Bank: the Embed Documents and Embed Dataset recipes chunk and vectorize a corpus with any embedding model reachable through the LLM Mesh and write the index to a vector store. Managed Knowledge Banks use a store Dataiku creates, Chroma by default with Milvus (local), Qdrant, and FAISS as no-setup alternatives, or a dedicated store through a connection: Azure AI Search, Elasticsearch, OpenSearch (managed or serverless), Milvus (remote), Pinecone, pgvector, Vertex Vector Search. Unmanaged Knowledge Banks read an index the enterprise already maintains, adding Snowflake Cortex Search (DSS 15 can create the Cortex Search Service from a Snowflake dataset and keep it synchronized) and Databricks AI Search. On top: static, dynamic, and agent-inferred filtering; a Smart query strategy that decides whether to retrieve and reformulates the query; similarity, threshold, maximal marginal relevance (MMR) diversity, and hybrid search (hybrid only where the store supports it: Azure AI Search, Databricks AI Search, Elasticsearch, Milvus, Snowflake Cortex Search); native rankers (Azure Semantic Ranker, Elasticsearch and Milvus reciprocal rank fusion) and model-based rerankers (Hugging Face open rerankers such as BGE-Reranker, Cohere Rerank through Bedrock and Microsoft Foundry); document-level security by intersecting security tokens on documents with the caller's groups, passed by Agent Hub as the end user's identity; multimodal Knowledge Banks that retrieve embedded images; GraphRAG (a plugin recipe that builds a knowledge graph and a search tool over it); Automated retrieval-augmented generation (RAG) Optimization (a plugin recipe that tunes chunking and search parameters with Optuna against an evaluation set); RAG guardrails on relevancy and faithfulness (Advanced LLM Mesh add-on). Semantic Models add structured context: entities, attributes, relationships, and metrics over SQL datasets that the Semantic Model Query tool uses to write precise SQL for agents. The buyer gets retrieval over its own documents and tables, with model and store choice, behind the same permissions its users already carry.

The architect's concern is that Dataiku is the orchestrator, not the engine. There is no Dataiku retrieval engine; the embedded stores are open-source libraries Dataiku wraps, and hybrid search, native ranking, and scale are whatever the backing store provides. The Knowledge Bank configuration (chunking, metadata, filters, search settings, security tokens, the semantic model) is Dataiku's object, and the vectors belong to whichever model made them.

Calibration: Elastic, MongoDB, and Snowflake read strong on native hybrid engines with generally available models of their own; Cloudflare reads moderate on a managed index plus open models; Cohere moderate on frontier models plus a fixed-function bundle; Qlik moderate on turnkey RAG in a closed box. Dataiku is more general than Qlik (any store, any model, any chunking) and less than Elastic (no engine, hybrid search delegated to the store). Rule 6 pegs strong to the native engines. Moderate.

**Borrowed Judgment:** Split at the store, captive at the configuration. Indexes in open engines the enterprise runs (Elasticsearch, OpenSearch, Milvus, pgvector) are the enterprise's: Retained, or Delegated on a managed variant. Indexes in proprietary services (Azure AI Search, Pinecone, Vertex Vector Search, Snowflake Cortex Search, Databricks AI Search) belong to those owners through Dataiku's paper: Ceded. The embedded Chroma, FAISS, Milvus (local), and Qdrant indexes live in Dataiku managed folders and rebuild from source: the Knowledge Bank that owns them is Dataiku's, Ceded. Hosted embeddings through the Mesh belong to their model's space under the carve-out, Ceded to the owner; open Hugging Face embedding and reranking weights the enterprise serves itself are Delegated on the Arctic Embed precedent. The runtime call (retrieve or not, which filter, which rerank) is Dataiku's engine executing settings the enterprise chose per tool and per request: vendor decides, visible, overridable, Delegated, the Elastic reading.

### ● Layer 1C · Pipelines: Data Movement & Pipelines

*Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering*  
**Status:** The Flow: Visual and Code Recipes on Any Engine, Any Source to Any Destination, Scenarios, Lineage; Streaming Experimental

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**The Flow: Visual Recipes, Engine Selection (DSS, SQL Pushdown, Spark, Containers), Partitioning, Scenarios, Metrics and Checks, Data Quality Gates** [DAPM: Ceded]  
The pipeline graph and every visual transformation in it, with the engine chosen per recipe and builds scheduled and gated by scenarios. Dataiku's format with no second implementer: Ceded.

**Code Recipes and Notebooks (Python With Pandas and Polars, R, SQL, Spark, PySpark, Scala) + Code Studios** [DAPM: Retained]  
The enterprise's own transformation code, run on the DSS host, in containers, on Spark, or pushed down as SQL to the enterprise's warehouse. The logic lifts (Pandas, Polars, Spark, the warehouse's SQL) and the Dataiku binding is a thin read-and-write wrapper: Retained (ruled September 15, 2026; the 2B customer tools that subclass Dataiku classes stay Ceded).

**Connections and Connectors (Cloud Storage, SQL Warehouses, Iceberg REST and Hive Catalogs, ClickHouse, SharePoint and SaaS Sources, Plugin Connectors) + Metastore Integration (Hive, Glue, DSS Virtual Metastore)** [DAPM: Ceded]  
Dataiku's connector implementations and the connection settings the enterprise accumulates in them, with fast paths on the major warehouses and object stores. The sources are the enterprise's and each source's own captivity belongs to its vendor's row; the connectors and their settings are Dataiku's: Ceded, the Snowflake Openflow Connectors and Databricks Lakeflow Connect reading.

**Column-Level Data Lineage + Flow Document Generator (Facet at 1C)** [DAPM: Ceded]  
Lineage is in this layer's own definition: the graph traced across every recipe and project, exportable, with manual lineage for external steps. Computed from and stored in Dataiku: Ceded.

**Project Deployer + Bundles + Automation Node (Promotion With Connection Remapping, Deployment Hooks and Policies)** [DAPM: Ceded]  
Versioned promotion of a Flow from design to production with the connections remapped per stage. A Dataiku packaging format: Ceded.

**Gap Analysis:** This is the layer Dataiku was founded on and it's complete. The Flow is a directed graph of datasets and recipes: visual recipes (Prepare with its processor library, Join, Group, Window, Stack, Split, Pivot, Sync, Distinct, Sort, Top N, Sample, Generate Features) and code recipes in Python, R, SQL, Spark, PySpark, and Scala, with notebooks and Code Studios beside them. Each recipe runs on an engine the platform proposes and the user can override: the DSS engine, in-database SQL pushdown to Snowflake, BigQuery, Databricks, Redshift, Synapse, Trino, PostgreSQL, and the rest, Spark (including Spark on Kubernetes), or a container. Sources and destinations are any connection (cloud storage, SQL, Iceberg, SaaS, the metastore), partitioning is first class, DSS 15 added Polars alongside Pandas, fast-path reads on Snowflake, S3, Azure Blob, and GCS (Polars on the three object stores), and fast-path writes on those four plus, when enabled on the connection, BigQuery, Databricks, Redshift, Trino, and Synapse. Scenarios schedule and trigger builds with metrics, checks, and Data Quality gates; column-level lineage is computed across the whole estate; Application-as-recipe packages a Flow as a reusable step; the Project Deployer promotes bundles from Design to Automation nodes with connection remapping. The buyer gets one pipeline surface that spans analysts and engineers and pushes the work down to the engines it already pays for.

The architect's concern is that the Flow is the product. The visual recipes, engine choices, scenario logic, and quality gates are Dataiku's format; the code recipes carry the enterprise's logic wrapped in Dataiku's input and output calls; the lineage is computed from the Flow and exists only there. Streaming (Kafka, Amazon Simple Queue Service (SQS), HTTP server-sent events, continuous recipes) carries an Experimental warning and is off by default, so the strong reads on batch.

Calibration: Cloudera reads strong on NiFi, Kafka, Flink, Spark, and Airflow spanning any source to any destination; Qlik strong on change data capture (CDC) plus Talend transformations; Databricks strong on Lakeflow and Spark; Snowflake strong on Openflow and declarative pipelines. Dataiku matches the general-purpose test of rule 4 (arbitrary code, insertable stages, any destination) with a decade of production deployment behind it, and adds column-level lineage that most peers score elsewhere. No scope gate applies to the batch surface. Strong on the batch surface, with streaming the named exception. Ruled September 15, 2026: strong stands on rule 4's general-purpose test; streaming has never been a requirement for strong at 1C (AWS, Azure, GCP, Qlik, and VAST read strong without it, VAST on event-driven DataEngine pipelines), and the reviewers' rule 6 argument assumed a frontier the instrument doesn't hold. Streaming stays named as Experimental.

**Borrowed Judgment:** Moderate and captive at the Flow. The logic inside code recipes is the enterprise's and ports (SQL runs on the enterprise's own warehouse; Python runs on Pandas, Polars, or Spark), so that chip reads Retained (ruled September 15, 2026: a thin input-and-output wrapper over an open runtime isn't a binding): the wrapper is Dataiku's, the substance is not. Everything else is Ceded: visual recipes, engine settings, scenarios, quality gates, connectors, lineage, bundles. The runtime call is Dataiku's engine choosing an execution engine and inferring schemas, with the user able to override the engine per recipe and the cluster per project: vendor decides, visible, overridable, Delegated, the Databricks and Snowflake reading under the override rule.

### ◑ Layer 2A · Orchestration: Infrastructure Orchestration

*GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization*  
**Status:** Elastic AI: Kubernetes Clusters Dataiku Creates and Drives, Execution Configs With GPU Limits, Fleet Manager; No Fair-Share Scheduler

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**Managed Kubernetes Clusters Dataiku Creates and Drives in the Enterprise's Cloud Account (EKS, AKS, GKE Through Dataiku Plugins)** [DAPM: Delegated]  
Standard Kubernetes provisioned by Dataiku's plugin into the enterprise's tenant, sized and autoscaled by settings the administrator gives it. Managed Kubernetes behind the Kubernetes API: Delegated, the Cloudera ECS and VMware VKS reading.

**Customer-Deployed Kubernetes Clusters, Attached (Unmanaged Clusters the Enterprise Built and Runs)** [DAPM: Retained]  
Any Kubernetes 1.10 or later the enterprise created itself and attaches through kubectl contexts; Dataiku documents that creating such clusters 'is not handled by DSS'. Self-run Kubernetes the enterprise operates independently of Dataiku: Retained, the Cloudera customer-run OpenShift reading. OpenShift 4 and DGX clusters attach the same way but carry an experimental, Tier 2 support badge and are unscored.

**Elastic AI Control: Containerized Execution Configurations (CPU, Memory, GPU Custom Limits, Namespaces, Group Access), Cluster Plugins and Lifecycle (Create, Attach, Start, Stop, Dynamic Clusters), Image Builds, Spark-on-Kubernetes** [DAPM: Ceded]  
How Dataiku places its work: pod specs, namespace policies, cluster selection per project and scenario, base images. Dataiku configuration with no second implementer: Ceded.

**Fleet Manager (Cloud Stacks on AWS, Azure, GCP: Instance Templates, Provisioning, Upgrades, Snapshots, Load Balancers)** [DAPM: Ceded]  
The lifecycle of the Dataiku nodes themselves in the enterprise's cloud tenant. Dataiku's provisioning model: Ceded.

**Dataiku Cloud Elastic AI Compute (Quota in vCPUs, RAM, Parallel Activities; Multi-Purpose Credits; GPU Configurations by Instance Type)** [DAPM: Ceded]  
Fully managed by Dataiku on Dataiku's accounts; the enterprise picks container sizes inside a pool it can't see below. Bounds are configuration: Ceded.

**API Deployer Infrastructures (API Nodes, Kubernetes Deployments With Replicas and GPU Reservations, External machine learning (ML) Platforms) + Unified Monitoring** [DAPM: Ceded]  
Where API services run, how many replicas, and their health, with Prometheus integration. Dataiku's deployment model: Ceded.

**Gap Analysis:** Dataiku pushes work off its own host onto Kubernetes, and it will build the cluster for you. Elastic AI computation covers Python and R recipes and notebooks, visual recipes on the containerized DSS engine, Spark on Kubernetes, model training and scoring, webapps, local Hugging Face model instances, and API services. Managed clusters (Elastic Kubernetes Service (EKS), Azure Kubernetes Service (AKS), Google Kubernetes Engine (GKE) through Dataiku plugins) are created, attached, started, stopped, and autoscaled from Administration, with node pools the administrator sizes and a global default cluster that each project can override; scenarios can pin a named cluster or create a dynamic one that shuts down when the run ends; unmanaged clusters the enterprise built attach the same way (OpenShift 4 and DGX Kubernetes clusters too, but both pages carry an experimental, Tier 2 support warning and are unscored). Containerized execution configurations set CPU and memory limits, GPU custom limits, namespaces (dynamic namespace policies per project or user), and which groups may use them; image build configurations publish the base images and DSS 15 rebuilds code environment images automatically after an upgrade. Fleet Manager (Cloud Stacks) provisions and upgrades the Dataiku nodes themselves on AWS, Azure, and GCP with instance templates, snapshots, and load balancers. On Dataiku Cloud the whole thing is hidden: a quota of vCPUs, RAM, and parallel activities, container sizes to pick from, and GPU configurations by instance type with credits. The API Deployer runs API services as Kubernetes deployments with replicas and GPU reservations; local Hugging Face models autoscale instances from zero against a target requests-per-instance. The buyer gets compute that appears when a recipe needs it and disappears after.

The architect's concern is that this is workload placement, not resource governance. There is no fair-share or priority scheduler, no GPU partitioning, no cross-tenant quota beyond Kubernetes limits and Dataiku Cloud's subscription pool; a GPU is a custom limit on a pod. The cluster is the enterprise's (or Dataiku's, on Cloud), but the execution configurations, cluster plugins, namespace policies, and Fleet Manager templates are Dataiku's.

Calibration: Cloudera reads moderate on a platform-scoped Kubernetes with hierarchical quota pools; Databricks and Snowflake moderate on managed compute with GPU pools and no customer-visible scheduler; Palantir moderate on Rubix. Dataiku has Cloudera's shape with less quota machinery and more cluster lifecycle automation. Moderate.

**Borrowed Judgment:** Split at the cluster. A Kubernetes cluster the enterprise built and attaches is its own: Retained. A cluster Dataiku's plugin creates in the enterprise's account is standard Kubernetes reached through the Kubernetes API: Delegated. Everything Dataiku layers on either (execution configurations, cluster plugins and lifecycle, namespace policies, Spark-on-Kubernetes configs, Fleet Manager templates, the Dataiku Cloud quota pool, the Deployer's infrastructures) is Dataiku's: Ceded. The runtime call is Dataiku's engine creating pods with the limits it was given and the autoscaler sizing nodes inside bounds, but the enterprise can pin a recipe to an execution configuration, a project to a cluster, and a scenario to a named or dynamic cluster; a documented per-object pin is an override under the ratified rule: vendor decides, visible, overridable, Delegated.

### ● Layer 2B · Runtime: Application Runtime & Execution

*Model serving, agent execution, inference APIs, distributed inference*  
**Status:** LLM Mesh Gateway + Local Open-Weight Serving on vLLM + Visual, Structured, and Code Agents With Human Approval + API Node Model Serving

**Decision authority:** Delegated (decides: model; visible: true; overridable: true; boundary: model)

**LLM Mesh Gateway (Connections to Sixteen Provider Types; Guardrails Pipeline; Cost Quotas; RPM and TPM Rate Limits; Audit and Tracing; Prompt Studio and Prompt Recipe; Persisted Conversations; Web Search Tool)** [DAPM: Ceded]  
The control surface every model call passes through. Connections, guardrail pipelines, quotas, rate rules, and prompt recipes are Dataiku configuration with no second implementer: Ceded.

**Chat, Completion, and Responses Model Access Through the Enterprise's Own Provider Accounts, In and Out Through OpenAI-Compatible Interfaces** [DAPM: Delegated]  
OpenAI, Anthropic, Bedrock, Vertex, Mistral, Foundry, Azure OpenAI, Cohere, Snowflake Cortex, and Databricks text and multimodal models on the enterprise's own keys, with Dataiku exposing an OpenAI-compatible Chat Completions and Responses API in front. Standard interfaces on both sides: Delegated, the S3-of-inference reading. Embedding models through the same connections are the carve-out and read Ceded at Layer 1B.

**Local Hugging Face Model Serving on the Enterprise's GPUs or CPUs (vLLM; Presets and Custom Models; Autoscaling Instances; Model Cache)** [DAPM: Delegated]  
Open weights the enterprise downloads or caches and serves in its own containers, with an inference engine (vLLM) that runs anywhere. The artifact and the engine are open; the serving surface and instance controls are Dataiku's: Delegated on the open-weights precedent.

**Customer-Owned Local Fine-Tunes and Imported Models (Safetensors From the Fine-Tuning Recipe or Python Fine-Tuning; MLflow and External Models Loaded for Serving)** [DAPM: Retained]  
Weights the enterprise produced or imported and possesses as files it can serve outside Dataiku (the Python fine-tuning path must produce safetensors; MLflow models import and export). The artifact is the enterprise's: Retained, the Cloudera customer-owned models reading.

**NVIDIA NIM Connection (Hosted Catalog or Self-Hosted Through the NIM Operator; NVIDIA AI Enterprise Subscription)** [DAPM: Ceded]  
Proprietary NIM containers on an NVIDIA license through Dataiku's connection. Ceded to NVIDIA under the channel-substitution rule, the SUSE and Elastic reading; the OpenAI-compatible endpoint in front is the Delegated chip above.

**Agents: Simple Visual, Structured Visual (Blocks, CEL, Human Confirmation), Code Agents, Agent Skills, Managed Tools, Human Approval, Tracing, Agent Review and Evaluate Agent (Advanced Add-On), MCP Server and A2A Exposure, Slack and Teams** [DAPM: Ceded]  
The agent harness and every definition in it. Dataiku objects, run and scaled by Dataiku, exported nowhere: Ceded.

**Customer-Authored Inline Code Tools, Plugin Tools, and Code Agent Classes (Python Against Dataiku's Tool Descriptor and BaseLLM Interfaces; LangChain and LangGraph Inside)** [DAPM: Ceded]  
The enterprise writes the code, but it subclasses Dataiku's classes and runs on Dataiku's runtime; the logic inside a LangGraph graph would port, the tool and agent objects would not. Bound to a single-vendor SDK: Ceded, the Palantir and Qlik reading, pending the open-runtime ruling.

**Fine-Tuning Recipe on Hosted Providers (OpenAI, Azure OpenAI, AWS Bedrock; Advanced LLM Mesh Add-On; Deployment Lifecycle Managed by DSS)** [DAPM: Ceded]  
Fine-tuned models that live in the provider's account and are deployed and billed there. Ceded to the provider through Dataiku's paper; local Hugging Face fine-tunes are the Retained exception, scored in the customer-owned models chip.

**API Node and API Deployer (Visual, Python, R, and MLflow Model Endpoints; Python and R Functions; SQL and Lookup Endpoints; Versioning; JSON Web Token (JWT) and OAuth2; Kubernetes With GPU)** [DAPM: Ceded]  
Classical real-time serving with versioned endpoints and OpenAPI documentation. The service packaging is Dataiku's, though models export to Python, Java, MLflow, and PMML: Ceded.

**Gap Analysis:** Dataiku's runtime has three parts, and each is production. The LLM Mesh is the gateway: administrator-defined connections to Anthropic, AWS Bedrock (with cross-region inference profiles and Bedrock Guardrails), SageMaker, Microsoft Foundry, Azure OpenAI (including through API Management (APIM)), Azure ML, Cohere, Databricks Mosaic AI, Google Vertex, Mistral, NVIDIA NIM, OpenAI, Snowflake Cortex, Stability AI, and custom plugin providers; DSS 15 added Claude Sonnet 5 and Fable 5, Nova 2 embeddings, and GPT-5.6 presets on Bedrock. Guardrail pipelines configured per connection, per agent, or at the point of use (personally identifiable information (PII), prompt injection, toxicity, topic boundaries, forbidden terms, custom code; bias detection is an unsupported add-on plugin) check the query and, when enabled, the response; every call is metered against cost quotas that can block, throttled by requests- or tokens-per-minute rules, and traced. Applications reach it through the LLM Mesh API, an OpenAI-compatible API that serves Chat Completions and Responses for any model in the Mesh, and Persisted Conversations that keep history server side with human-in-the-loop tool validation. Local Hugging Face models are the self-hosted half: open weights (the presets include Qwen3, GPT-OSS 20B, Llama 3.3 70B, Gemma 4, Mistral Small 4, Kimi Linear, DeepSeek V4 Flash) served on the enterprise's GPUs under vLLM, with instances that scale from zero to a maximum against a target load; fine-tuning (Advanced LLM Mesh add-on) covers OpenAI, Azure OpenAI, Bedrock, and local models, and a Python recipe can produce a safetensors model of its own. Agents are the third part: Simple Visual Agents pick tools autonomously; Structured Visual Agents compose blocks (agentic loop with exit conditions, LLM request, manual and mandatory tool call, delegate to another agent, routing, for-each, parallel, reflection, Python code, long-term memory, context compression, generate artifact) with Common Expression Language (CEL) expressions and Jinja templates and, since DSS 15, human confirmation of tool calls inside the blocks; Code Agents implement a class Dataiku hosts, scales, and audits, with LangChain and LangGraph inside if wanted; Agent Skills package instructions in the SKILL.md format; twenty-five-plus managed tools (Knowledge Bank search, semantic model query, dataset lookup and append, remote and local Model Context Protocol (MCP), OpenAPI with endpoint discovery, model predict, API endpoint, image generation, send message, generate artifact, SQL question answering, Google Search and Calendar, Jira, Salesforce, ServiceNow, Snowflake Cortex Analyst and Search, Databricks Genie and Vector Search, LLM-side web search) plus inline code and plugin tools, each with optional human approval that pauses the agent before the call and can let the user edit the arguments. Agents are exposed as MCP tools through Dataiku's own MCP server per project, as Agent2Agent (A2A) endpoints (specification 0.3 and 1.0), in Slack and Teams, and as virtual LLMs anywhere the Mesh is used; Agent Review (LLM-as-judge traits plus human review) and the Evaluate Agent recipe (Advanced add-on) test them. Classical serving is the API node: visual, Python, R, and MLflow models, arbitrary functions, SQL queries, and dataset lookups as versioned REST endpoints on API nodes or Kubernetes, with models exportable to Python, Java, MLflow, and Predictive Model Markup Language (PMML). The buyer gets model access, self-hosted open weights, and governed agents on one runtime it already uses for pipelines.

The architect's concern is the SDK. Agent definitions, block logic, tool descriptors, guardrail pipelines, quotas, and prompt recipes are Dataiku objects; inline tools and Code Agents subclass Dataiku's classes; human approval isn't available for tools called through the MCP server or inside for-each, parallel, and reflection sub-sequences. The models are the enterprise's; the harness is not.

Calibration: Databricks, Snowflake, and Palantir read strong on serving plus generally available agent runtimes with evaluation and tracing; Cloudera strong on inference plus Agent Studio. Elastic was held at moderate on first-release maturity with no evaluation surface and tracing in preview; Dataiku has generally available tracing, evaluation, and review, several releases of the agent surface, and self-hosted serving the Elastic row lacked. The Advanced LLM Mesh add-on gates fine-tuning, Agent Review, Evaluate Agent, RAG guardrails, and custom quotas; the gateway, local serving, every agent type, the managed tools, human approval, tracing, MCP and A2A exposure, and the API node are in the standard product. Strong, written. Ruled September 15, 2026: a vendor's own purchasable add-on on the same paper is a price, not a rule 5 gate (the whose-paper rule credits what the enterprise buys with the vendor's support); the add-on is named, and the grade stands.

**Borrowed Judgment:** Split at the model, captive at the harness. Model access through the enterprise's own provider accounts behind the Mesh's OpenAI-compatible surface is Delegated (the inference-interface ruling); open Hugging Face weights the enterprise serves itself are Delegated, and the fine-tunes and imported models the enterprise owns outright are Retained; NIM is Ceded to NVIDIA through Dataiku's paper; hosted fine-tunes belong to their providers, Ceded. The Mesh configuration, agent definitions, tools, and API services are Dataiku's, Ceded, and so is customer-authored tool and agent code, because it binds to Dataiku's tool descriptor and BaseLLM classes rather than to a multi-vendor interface (the Palantir and Qlik SDK reading). The decision to act is the model's, and the enterprise can put deterministic code between that decision and its effect: human approval on a tool persists in the tool's settings and pauses the agent before every call; Structured Visual Agent blocks route, gate, and run Python before the next step; configured guardrails execute before the call and, where enabled, after it. Model decides, visible, overridable, Delegated, the Anthropic hooks and Mistral per-function approval reading.

### ◑ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane

*Policy-driven placement and resource coordination — the Autonomy Layer*  
**Status:** Gateway, Govern Registry With Sign-Off and a Deployment Gate on Bundles, Multi-Agent Orchestration, Observability; No First-Class Agent Identity; Agent Management Due October 2026

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**LLM Mesh Request-Time Gateway (Guardrail Pipelines per Connection, Agent, and Use; Blocking Cost Quotas; Rate Limits; Audit; API-Key Permissions and Caller-Ticket Impersonation; Document-Level Security Tokens)** [DAPM: Ceded]  
The gateway leg. Every call is inspected, metered, throttled, and logged in Dataiku's pipeline. Dataiku configuration: Ceded.

**Dataiku Govern GenAI Registry (Agents, Augmented LLMs, Fine-Tuned LLMs and Versions; Standard and Custom Workflows; Feedback and Final Approval Sign-Off; Timeline; Advanced License) + Deployment Governance Policies on Bundles and Model Versions** [DAPM: Ceded]  
The registry leg, with a persisted human gate that the Deployer enforces deterministically per infrastructure for project bundles and model versions. Govern templates and sign-off definitions are Dataiku's: Ceded.

**Cross-Agent Orchestration (Agent Hub Orchestrating LLM in Tools or Manual Mode; Structured Visual Agent Delegate, Routing, Parallel, and Reflection Blocks; Agents as Tools; A2A and MCP Endpoints)** [DAPM: Ceded]  
The orchestration leg, decided by a model over agent descriptions or by blocks the enterprise arranged. Dataiku's definitions: Ceded.

**Observability (LLM Mesh Traces With Guardrail and Cost Spans; Trace Explorer; LangSmith Push; Interaction Logging; Evaluate Agent and Evaluation Stores With Alerting; Agent Review)** [DAPM: Ceded]  
The observability leg. Traces are a nested JSON the enterprise can push to LangSmith or store in its own datasets, but the evaluation stores and reviews are Dataiku's: Ceded. Unified Monitoring (project and endpoint health, Prometheus) is operations telemetry beside this leg, not part of it.

**Gap Analysis:** Read against the five legs the instrument asks of an Intelligence-2C plane, Dataiku has four in production. Gateway: every agent and model call passes the LLM Mesh, where guardrail pipelines configured per connection, per agent, and at the point of use run on the request and, when enabled, the response, cost quotas block when exhausted, requests- and tokens-per-minute rules throttle per model and provider, and every call is audited and traced; agents run with the permissions of the API key that called them, and Agent Hub can pass the DSS caller ticket so tools execute as the end user, with document-level security tokens carried in the completion context. Registry: Dataiku Govern (Advanced license) syncs agents, augmented LLMs, and fine-tuned LLMs with their versions into a GenAI registry beside models and bundles, wraps each in a workflow with feedback and final approval sign-offs, records a timeline of every change, and lets the Deployer refuse to deploy an unapproved project bundle or model version (a policy per infrastructure: prevent, warn, or always deploy); an agent ships to production inside a project bundle, so the gate covers the agentic project, but the agent version's own registry approval isn't what the Deployer checks. Cross-agent orchestration: Agent Hub's orchestrating LLM routes a conversation across enterprise agents as tools or in parallel; Structured Visual Agents delegate to other agents, route, fan out, and reflect; agents are tools to other agents; A2A and MCP endpoints expose them outward. Observability: nested traces from every Mesh call with guardrail and cost spans, the Trace Explorer webapp, LangSmith push, interaction logging datasets, the Evaluate Agent recipe writing trajectory and answer metrics into evaluation stores with alerting (Advanced add-on), Agent Review's trait-based tests, with Unified Monitoring beside them as project and endpoint health telemetry rather than agent observability. Missing: first-class agent identity. An agent has no principal of its own; it borrows the API key's permissions or impersonates the calling user. The buyer gets governance that reaches from build (sign-off) to run (guardrails, quotas) to review (evaluation) inside one platform.

The architect's concern is that most of this plane decides by instruction or by model, not by deterministic code. The Agent Hub router is an LLM choosing an agent from descriptions; the PII, toxicity, prompt-injection, and topic detectors are models (bias detection is an unsupported plugin behind the add-on); forbidden-terms matching, quotas, rate limits, and the deployment policy are the deterministic parts. That is guardrails on state and a commit gate, which the instrument credits, and not outcome validation, which it notes as universally absent. There is no placement reasoning: which model serves a request is whatever the agent or recipe names.

Calibration: GCP holds strong on all five legs; Azure strong with a deterministic policy engine as the orchestration leg; Snowflake, Databricks, SAP, and Kamiwaza read moderate with a leg named as missing (Snowflake: gateway and orchestration; SAP: identity and observability). Dataiku is the fullest moderate on the map, four legs, and the rule is explicit that a partial plane is moderate. Moderate, identity the named leg.

**Borrowed Judgment:** High and captive. Guardrail pipelines, quota scopes, rate rules, Govern templates and sign-off definitions, deployment policies, Agent Hub routing configuration, and evaluation recipes are Dataiku objects: Ceded. The runtime call is split: the deterministic parts (forbidden terms, quotas, limits, the deploy policy, a final approval) execute rules the enterprise wrote, and the router and the detectors are models. Quotas, rate rules, and allow-lists are bounds and don't decide direction; the persisted gates do: a Govern final approval or rejection that the Deployer enforces before a bundle or model version ships, human approval that pauses an agent before a tool's effect, and guardrail actions the enterprise sets per connection and per agent. Vendor decides, visible, overridable, Delegated, the Snowflake reading; the missing gateway leg there is the missing identity leg here, and both are capability findings, not authority ones.

### ● Layer 3 (+1) · Applications: AI Application Layer — The Value Plane

*AI-powered business capabilities — business logic, workflow automation*  
**Status:** Agent Hub for Business Users, Three Business Applications, Stories and Dashboards, Cobuild; Forty-Plus Unsupported Solutions Unscored

**Decision authority:** Delegated (decides: vendor; visible: true; overridable: true; boundary: vendor)

**Agent Hub + Agent Chat (Enterprise Agent Library, Multi-Agent Chat, My Agents From Documents, SharePoint, and Datasets, Chart Generation, Document Guardrails, Admin Controls; Slack and Teams; Answers as the Predecessor)** [DAPM: Ceded]  
The consumer portal for governed agents and the business user's own agent builder, delivered as a Dataiku plugin webapp. Dataiku's product: Ceded.

**Business Applications: Process Mining, RFx Accelerator, Manufacturing Operations (Enterprise Edition or Enterprise AI Package; Designer License Types)** [DAPM: Ceded]  
Self-contained applications for named business processes, installed and operated inside DSS. Dataiku's applications on Dataiku's runtime: Ceded.

**Stories, Dashboards, Workspaces, and Dataiku Applications (Packaged Project Apps With Tiles and Forms)** [DAPM: Ceded]  
The analytics and app-packaging surfaces business users consume. Dataiku formats: Ceded.

**Cobuild (Generally Available June 18, 2026: Objective to Governed Project; Recipes, Evaluations, Dashboards; Per-User Instructions; Admin Prompt Controls) + Dataiku AI Assistants (SQL Assistant, AI Search, Stories AI, AI Explain)** [DAPM: Ceded]  
The building agent and the embedded assistants that act inside Dataiku. The models are the enterprise's choice through Dataiku AI Services or the LLM Mesh (Anthropic, OpenAI, Google, Snowflake Cortex, Databricks AI Gateway, Bedrock, Foundry, open-source and on-premises models); the prompts and the project objects they act on are Dataiku's: Ceded.

**Gap Analysis:** The value plane is where Dataiku's agent story lands for people who don't build. Agent Hub (a Dataiku plugin, version 1.6.0 on August 21, 2026, requiring DSS 14.6 or later) is the portal: a curated library of enterprise agents, a single chat that orchestrates several of them, My Agents that business users build themselves from uploaded documents, SharePoint libraries, and datasets, chart generation with versioning, document guardrails, and an admin panel that decides which agents, models, and tools each Hub exposes; Agent Chat puts a single agent behind a shareable interface; Dataiku Answers, the earlier packaged chat and RAG webapp, is documented as mostly replaced by the Hub. Three Business Applications ship as self-contained products: Process Mining (DSS 14.5 or later) discovers process flows from event logs, analyzes variants, finds bottlenecks, and drills to cases; the RFx Accelerator (DSS 14.6 or later) extracts the structure of RFPs and RFIs, drafts answers from a curated knowledge base, and exports the questionnaire; Manufacturing Operations delivers a Parameters Analyzer and a Manufacturing Event Tracker. Forty-plus Dataiku Solutions install as accelerators across retail (demand forecast, market basket, markdown optimization, site selection), banking and insurance (anti-money-laundering (AML) alerts triage, credit scoring, credit risk stress testing, claims modeling), pharma (pharmacovigilance, clinical trial intelligence, cohort discovery), manufacturing (batch performance, maintenance planning, production quality control), finance (financial forecasting, reconciliation), and governance itself (EU AI Act Readiness, ISO 42001 Readiness, LLM Provider Due Diligence). Stories are refreshing data presentations with generative features, dashboards and workspaces carry the analytics, and Dataiku Applications package any project as an app with tiles and forms. Cobuild (generally available June 18, 2026) is the building agent: it turns a business objective into a project, and in 15.0.1 creates evaluation recipes and dashboards, with per-user instructions and an administrator switch to block project-level custom prompts. Expert-to-Agent (May 2026) is the packaging name for this stack, not a separate product. The buyer gets a place for the business to use, and build, agents on governed data, plus packaged applications for a handful of named processes.

The architect's concern is thin domain depth and what carries the grade. The Business Applications are three, gated to the Enterprise edition or an Enterprise AI Package and to Designer license types; the Solutions are templated projects with sample data that Dataiku says it doesn't support and offers at the customer's own risk, so they are named here for reach and not scored (rule 1 credits what ships with the vendor's support); the Hub is a plugin the administrator installs; everything is a Dataiku project and stays one.

Calibration: Databricks reads strong on Genie, dashboards, Apps, and Databricks One; Snowflake strong on CoWork, CoCo, Observe, and Streamlit; Qlik strong on the analytics value plane; SAP strong on business applications with Joule; Cloudera moderate on visuals and a workbench copilot. Dataiku has Databricks' shape (a consumer portal for agents plus analytics surfaces) with packaged business applications SAP-style at small scale. The MongoDB and Redis question (is a developer assistant a value plane) doesn't decide this cell; Cobuild is a component, not the grade. Strong, written, on Agent Hub with My Agents, Answers, the analytics surfaces, Dataiku Applications, Cobuild, and the three Business Applications; the edition gate is named. Ruled September 15, 2026: strong stands; an edition tier of the same product is a price, not a rule 5 gate, and the Solutions were already unscored; moderate had been argued on the edition gate.

**Borrowed Judgment:** High and captive. Agent Hub configurations, My Agents, Business Applications, Stories, dashboards, and Dataiku Applications are Dataiku projects and webapps with no other host: Ceded. The runtime call in the Hub and its agents is Dataiku's application logic wrapped around a model's choices (which agent answers, which tool it calls), and the enterprise can gate the effect on the Hub and Agent Chat paths: human approval on a managed tool persists in the tool's settings and pauses the agent before the call (not for tools reached through the MCP server, and not inside for-each, parallel, or reflection sub-sequences); the Hub's document guardrails and agent allow-lists are administrator rules. Vendor decides, visible, overridable, Delegated, the SUSE Liz reading for an application whose function is acting; MongoDB's read-only assistant is the contrast, and the Databricks, Snowflake, Qlik, and SAP Layer 3 cells predate the override rule.

---
*Layer2C · AI Infrastructure Decision Intelligence · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com*
