Dataiku is a software-only AI platform that sits on top of whatever the enterprise already owns, and the map reads it as strong at the two layers it was built for and at the one it rebuilt itself around: data pipelines (Layer 1C), the runtime that serves models and runs agents (Layer 2B), and the applications business users touch (Layer 3). Governance, retrieval, and orchestration are real but partial: a catalog, lineage, quality rules, and a feature store over data stored elsewhere (1A moderate); Knowledge Banks that orchestrate retrieval across Chroma, Elasticsearch, Pinecone, pgvector, Snowflake Cortex Search, Databricks AI Search, and others without a retrieval engine of Dataiku's own (1B moderate); Kubernetes clusters Dataiku creates and drives on the enterprise's cloud accounts (2A moderate); and an agent-governance plane with four of the five legs the instrument asks for, missing only a first-class agent identity (2C moderate). Layer 0 is a gap by design: Dataiku sells no compute, network, or storage, on any of its three deployment shapes.
The capture is decoupled and mostly invisible, which is the dangerous kind. Every openness claim is true: data stays on the enterprise's connections in Parquet, Iceberg, and the warehouse's own tables; models come from the enterprise's own OpenAI, Anthropic, Bedrock, Vertex, Mistral, Snowflake Cortex, or Databricks accounts through the large language model (LLM) Mesh; open Hugging Face weights run on the enterprise's own graphics processing unit (GPU) under vLLM; the Python client is Apache 2.0. What accumulates in Dataiku is the opinion layer above all of that: the Flow and its visual recipes, dataset definitions, lineage, quality rules, Knowledge Bank configurations, agent definitions and Structured Visual Agent block logic, managed-tool descriptors, guardrail pipelines, quotas, Govern workflows and sign-offs, and Agent Hub configurations. None of it lifts to another platform. Of the row's components, the great majority read Ceded; the Delegated chips are the standard interfaces Dataiku sits behind (managed Kubernetes it provisions, provider model application programming interface (API), open weights it serves); the four Retained chips are things the enterprise possesses without Dataiku (the transformation code inside its recipes, a Kubernetes cluster it built and attached, open vector stores it runs, the fine-tuned and imported model weights it owns), and every Dataiku-shaped object, including customer-authored tools bound to Dataiku's classes, stays behind.
The buyer's trade: one governed workbench where analysts, data scientists, and business users build pipelines, models, and agents against the enterprise's existing estate, with decision-time gates the instrument credits (human approval on tool calls, deterministic blocks in Structured Visual Agents, configured guardrails before and after model calls, sign-off that blocks a bundle from deploying), in exchange for a platform whose accumulated logic has no exit. The decision-authority readings say the same thing from the other side: Delegated at every scored layer, because the enterprise can see and override the runtime call, and never Retained, because the engine making the call is always Dataiku's. What would move cells: a first-class agent principal (2C), Agent Management shipping in October 2026 as announced (2C), a native retrieval engine or hybrid search independent of the backing store (1B), and streaming leaving its Experimental badge (1C).
Layer-by-layer status: Layer 0 (Not Dataiku's Layer (By Design); Dataiku Cloud on Dataiku's Accounts, Cloud Stacks and Custom Installs on Yours), Layer 1A (Catalog, Column-Level Lineage, Quality Rules, and a Feature Store Over Data Stored on the Enterprise's Own Connections), Layer 1B (Knowledge Banks Over Ten Vector Stores + Semantic Models + Document-Level Security; No Retrieval Engine of Its Own), Layer 1C (The Flow: Visual and Code Recipes on Any Engine, Any Source to Any Destination, Scenarios, Lineage; Streaming Experimental), Layer 2A (Elastic AI: Kubernetes Clusters Dataiku Creates and Drives, Execution Configs With GPU Limits, Fleet Manager; No Fair-Share Scheduler), Layer 2B (LLM Mesh Gateway + Local Open-Weight Serving on vLLM + Visual, Structured, and Code Agents With Human Approval + API Node Model Serving), Layer 2C (Gateway, Govern Registry With Sign-Off and a Deployment Gate on Bundles, Multi-Agent Orchestration, Observability; No First-Class Agent Identity; Agent Management Due October 2026), Layer 3 (+1) (Agent Hub for Business Users, Three Business Applications, Stories and Dashboards, Cobuild; Forty-Plus Unsupported Solutions Unscored).
Assessment framework: 4+1 Layer AI Infrastructure Model. Scoring model: Decision Authority Placement Model (DAPM) — Retained, Delegated, or Ceded. Published by The CTO Advisor LLC (DBA The Advisor Bench). Author: Keith Townsend. Date assessed: September 10, 2026. Version: v1.0 - 4+1 v2: Authority Split.
Dataiku is a software-only AI platform that sits on top of whatever the enterprise already owns, and the map reads it as strong at the two layers it was built for and at the one it rebuilt itself around: data pipelines (Layer 1C), the runtime that serves models and runs agents (Layer 2B), and the applications business users touch (Layer 3). Governance, retrieval, and orchestration are real but partial: a catalog, lineage, quality rules, and a feature store over data stored elsewhere (1A moderate); Knowledge Banks that orchestrate retrieval across Chroma, Elasticsearch, Pinecone, pgvector, Snowflake Cortex Search, Databricks AI Search, and others without a retrieval engine of Dataiku's own (1B moderate); Kubernetes clusters Dataiku creates and drives on the enterprise's cloud accounts (2A moderate); and an agent-governance plane with four of the five legs the instrument asks for, missing only a first-class agent identity (2C moderate). Layer 0 is a gap by design: Dataiku sells no compute, network, or storage, on any of its three deployment shapes.
The capture is decoupled and mostly invisible, which is the dangerous kind. Every openness claim is true: data stays on the enterprise's connections in Parquet, Iceberg, and the warehouse's own tables; models come from the enterprise's own OpenAI, Anthropic, Bedrock, Vertex, Mistral, Snowflake Cortex, or Databricks accounts through the large language model (LLM) Mesh; open Hugging Face weights run on the enterprise's own graphics processing unit (GPU) under vLLM; the Python client is Apache 2.0. What accumulates in Dataiku is the opinion layer above all of that: the Flow and its visual recipes, dataset definitions, lineage, quality rules, Knowledge Bank configurations, agent definitions and Structured Visual Agent block logic, managed-tool descriptors, guardrail pipelines, quotas, Govern workflows and sign-offs, and Agent Hub configurations. None of it lifts to another platform. Of the row's components, the great majority read Ceded; the Delegated chips are the standard interfaces Dataiku sits behind (managed Kubernetes it provisions, provider model application programming interface (API), open weights it serves); the four Retained chips are things the enterprise possesses without Dataiku (the transformation code inside its recipes, a Kubernetes cluster it built and attached, open vector stores it runs, the fine-tuned and imported model weights it owns), and every Dataiku-shaped object, including customer-authored tools bound to Dataiku's classes, stays behind.
The buyer's trade: one governed workbench where analysts, data scientists, and business users build pipelines, models, and agents against the enterprise's existing estate, with decision-time gates the instrument credits (human approval on tool calls, deterministic blocks in Structured Visual Agents, configured guardrails before and after model calls, sign-off that blocks a bundle from deploying), in exchange for a platform whose accumulated logic has no exit. The decision-authority readings say the same thing from the other side: Delegated at every scored layer, because the enterprise can see and override the runtime call, and never Retained, because the engine making the call is always Dataiku's. What would move cells: a first-class agent principal (2C), Agent Management shipping in October 2026 as announced (2C), a native retrieval engine or hybrid search independent of the backing store (1B), and streaming leaving its Experimental badge (1C).
Raw compute, networking, and acceleration fabric
Dataiku is certified DGX-Ready Software and documents DGX Kubernetes clusters as an unmanaged cluster type, but the page carries a warning that DGX support is experimental and covered by Tier 2 support, and Multi-Instance GPU hasn't been tested. A dependency fact for the customer who already owns DGX, not a Dataiku substrate.
Dataiku ships software, on three shapes, and none of them is a substrate the enterprise buys from Dataiku. Dataiku Cloud is the hosted plan: Dataiku operates the instance and the Elastic AI compute in its own cloud accounts, with a subscription quota measured in virtual CPU (vCPU), gigabytes of RAM, and parallel activities, Multi-Purpose Credits to burst past it, and GPU containerized execution configurations available only with credits and only in the instance types the hosting region offers. Cloud Stacks deploys Dataiku through Fleet Manager into the enterprise's own AWS, Azure, or Google Cloud tenant, where Dataiku has no access to the data and the virtual machines and Kubernetes clusters are the enterprise's bill. Custom installs run on the enterprise's own Linux servers, on premises or in any cloud, with NVIDIA GPUs the enterprise supplies for local model inference (compute capability 7.5 or higher; a CUDA 13 compatible driver on Kubernetes nodes, and the full CUDA 13 toolkit when models run on the Data Science Studio (DSS) server itself). The buyer gets to keep whatever infrastructure it already has. The architect's concern is the usual one for a software-only row: the compute decisions on the Dataiku Cloud path are Dataiku's and invisible, and on the other two paths they are the enterprise's and never Dataiku's. Neither is a Dataiku capability at this layer. Calibration: Elastic, MongoDB, Redis, Databricks, and Snowflake read gap on the same shape (a software or software as a service (SaaS) vendor whose hosted plan runs on a hyperscaler the vendor picks); SAP reads moderate because SAP operates its own data centers and sells the IaaS. Dataiku is the former. Gap, with authority Absent under rule 7 (the Elastic, MongoDB, and Redis reading): the layer is inherited from the hyperscaler or from the enterprise's own estate, not absorbed by Dataiku.
None to Dataiku as a substrate. On Dataiku Cloud the enterprise consumes managed execution capacity (a quota of vCPUs, RAM, and parallel activities, credits, container sizes, GPU configurations by instance type) and inherits the hyperscaler's substrate through Dataiku's account, picking nothing below the container; on Cloud Stacks and custom installs the substrate is the enterprise's own choice and the judgment belongs to whoever it bought the servers, GPUs, and network from. No customer-administered Layer 0 is offered on any path: Absent, the Elastic and Redis reading, with the Cloud path's absorbed substrate named rather than scored.
Dataiku Cloud GPU compute is documented in the Knowledge Base as a containerized execution configuration type that requires Multi-Purpose Credits and a GPU instance type from the region's list. Streaming, DGX, and the rest of the Layer 0 adjacent facts are carried where they score. Public evidence that moves the cell: Dataiku selling or operating compute the enterprise administers as its own, which nothing in the current documentation describes.
Durable, governed data foundation — the Governance Catalog that Layer 2C queries
The organization-wide index of Dataiku assets and indexed external tables, with curated collections, roles, and metadata completeness. A Dataiku-native catalog structure with no export path to another catalog: Ceded.
Direct column lineage inferred from recipes across datasets and projects, traversed upstream and downstream, with manual remapping where inference stops, and exportable to PDF or images. The graph is computed from Dataiku Flows and lives only there: Ceded.
Dataset-level metrics, checks, and quality rules evaluated in the Flow and in scenarios, with status surfaced in the Catalog. Rules are authored in Dataiku's rule types against Dataiku dataset definitions: Ceded.
Features curated and shared across projects for machine learning, with lineage back to the producing recipes. A Dataiku object over Dataiku datasets: Ceded.
How Dataiku knows the enterprise's data: every dataset's storage, format, schema, storage types, meanings, and partitioning, and every connection's settings, stored as project configuration and exported only as a Dataiku project archive. The bytes stay where the enterprise put them and never move, which is the reassuring fact; the definitions rebuild on exit: Ceded. The storage itself belongs to the storage and warehouse vendors' rows.
Catalog, lineage, quality, and the feature store run on the Dataiku backend and the enterprise's databases. Nothing here touches a GPU.
Dataiku isn't a data store; it governs data that lives on the enterprise's connections. Datasets default to those connections (uploaded files and local-connection datasets sit in the DSS data directory, which on Dataiku Cloud is Dataiku-hosted, and Dataiku Cloud is contractually a data processor for that content): cloud storage (S3, Azure Blob, Google Cloud Storage) in Parquet and CSV with fast-path reads and writes, SQL warehouses (Snowflake, BigQuery, Databricks, Redshift, Trino, Synapse, ClickHouse from DSS 15), Apache Iceberg through REST and Hive catalogs, Hive and Glue metastores, and a long list of SaaS sources. Over that estate Dataiku puts the Catalog (DSS 15 widened it from a data catalog to a catalog of assets: datasets, indexed external tables, saved and fine-tuned models, retrieval-augmented LLMs, agents, agent tools, semantic models, with curated Catalog Collections, AI Search in natural language, a Database Explorer for SQL and metastore connections, and a Most Used Datasets view), column-level Data Lineage across datasets and projects with upstream and downstream traversal and manual lineage for what the Flow can't infer, metrics, checks, and Data Quality rules, and a Feature Store curated from the Flow. Security is project-centered: per-project group permissions, per-user connection credentials, discoverable and private projects, and User Isolation Framework on the host, with object-level exceptions (Catalog Collections expose a dataset in READ or metadata-only DISCOVER mode without project access, shared objects give another project read-only use, and project-level API keys carry per-dataset read, write, schema, and metadata permissions). The buyer gets one place to find, trust, and trace what the organization has built, without moving any of it. The architect's concern is where the governance lives and where it stops. The dataset definitions, catalog entries, lineage graph, quality rules, and feature definitions are Dataiku project metadata in the DSS configuration directory; they describe data on the enterprise's own storage but they don't leave with it, and the export path is a Dataiku project archive. And the enforcement is coarse: Dataiku governs who can open a project or use a connection, not which rows or columns a user sees inside a warehouse table; masking, row filters, and tags stay with the warehouse's own governance. Document-level security exists, but for Knowledge Banks (1B), not for datasets. Calibration: Palantir and Cloudera read strong on a governance authority that enforces at interaction time (the Ontology's markings; Ranger's row and column policies) over storage the vendor also provides or runs; Databricks and Snowflake strong on catalogs that own the tables' grants. Qlik reads moderate on an open lakehouse the enterprise stores plus a captive trust layer of data products and quality. Dataiku is Qlik's shape without the lakehouse: a captive catalog and lineage layer over data it neither stores nor enforces on. Moderate.
Moderate and captive above the data. The bytes are the enterprise's, in the enterprise's own storage and warehouses, in formats every engine reads; that is the reassuring part and it's true, and it's the storage vendors' portability, not Dataiku's. Everything Dataiku holds at this layer (dataset definitions, catalog structure, lineage graph, quality rules, feature definitions) is Dataiku's model of the estate and rebuilds on exit: Ceded, every chip. The runtime call at this layer is small (lineage inference from the Flow, schema and meaning detection on datasets, quality checks the enterprise wrote), made by Dataiku's engine and overridable per dataset (schema, storage types, meanings, manual lineage): vendor decides, visible, overridable, Delegated, the Databricks and Cloudera reading.
Warehouse-side governance (Snowflake Horizon, Unity Catalog, Ranger) is consumed through connections and belongs to those rows, not this one; so does the portability of the underlying object stores and warehouses. Per-user credentials on connections are the mechanism by which the warehouse's own policies apply to Dataiku users. Uploaded files default to the DSS data directory's uploads folder but can be placed on any file-based connection with managed datasets allowed; on Dataiku Cloud that directory is Dataiku-hosted. Public evidence that moves the cell: Dataiku enforcing row- or column-level policies on datasets at query time across engines, or Dataiku selling a governed store of its own.
Low-latency retrieval for RAG — vector/hybrid search, context windows
The retrieval object and its configuration: chunking, metadata, update mode, search settings, security tokens, and the embedded index files in a Dataiku managed folder. Rebuilt from source on exit, but the opinions are Dataiku's: Ceded.
The index lives in an open-source engine the enterprise operates, and an unmanaged Knowledge Bank can point at an index built without Dataiku. Retained where the enterprise runs the engine itself; Delegated where it subscribes to a managed variant (Amazon OpenSearch Service, Elastic Cloud, a hosted PostgreSQL). The Cloudera self-run reading.
Single-vendor retrieval services whose index formats, ranking, and search services are the owner's; Snowflake's own row reads Cortex Search Ceded. Ceded to the owner through Dataiku's paper under the channel-substitution rule, the Elastic third-party-endpoints reading.
Embeddings from the enterprise's own provider accounts. The vectors are useless without the same model at query time and no second vendor serves the space: Ceded to the model's owner through Dataiku's paper under the carve-out.
Open embedding, reranking, and image-embedding weights the enterprise pulls from Hugging Face (or preloads in the model cache when air-gapped) and serves in its own containers. The weights and the vectors recompute anywhere the same model runs; the serving surface is Dataiku's: Delegated on the Arctic Embed precedent.
Business context that turns natural-language questions into SQL for agents, the Snowflake semantic-view shape. Defined in Dataiku's semantic model format with no second implementer: Ceded.1.3, July 16, 2026, compatible with DSS 14.4 and above), and whether 'Lab' in a Dataiku plugin name is a maturity badge that keeps the surface out of a scored cell awaits ruling; the reference documentation carries no preview badge.
Embedding, reranking, and image-embedding models from Hugging Face run in Dataiku's containerized execution on Kubernetes nodes with NVIDIA GPUs (compute capability 7.5 or higher, CUDA 13 driver) or on CPU. The GPU path is the enterprise's hardware; hosted embeddings through the LLM Mesh need none.
Retrieval in Dataiku is a Knowledge Bank: the Embed Documents and Embed Dataset recipes chunk and vectorize a corpus with any embedding model reachable through the LLM Mesh and write the index to a vector store. Managed Knowledge Banks use a store Dataiku creates, Chroma by default with Milvus (local), Qdrant, and FAISS as no-setup alternatives, or a dedicated store through a connection: Azure AI Search, Elasticsearch, OpenSearch (managed or serverless), Milvus (remote), Pinecone, pgvector, Vertex Vector Search. Unmanaged Knowledge Banks read an index the enterprise already maintains, adding Snowflake Cortex Search (DSS 15 can create the Cortex Search Service from a Snowflake dataset and keep it synchronized) and Databricks AI Search. On top: static, dynamic, and agent-inferred filtering; a Smart query strategy that decides whether to retrieve and reformulates the query; similarity, threshold, maximal marginal relevance (MMR) diversity, and hybrid search (hybrid only where the store supports it: Azure AI Search, Databricks AI Search, Elasticsearch, Milvus, Snowflake Cortex Search); native rankers (Azure Semantic Ranker, Elasticsearch and Milvus reciprocal rank fusion) and model-based rerankers (Hugging Face open rerankers such as BGE-Reranker, Cohere Rerank through Bedrock and Microsoft Foundry); document-level security by intersecting security tokens on documents with the caller's groups, passed by Agent Hub as the end user's identity; multimodal Knowledge Banks that retrieve embedded images; GraphRAG (a plugin recipe that builds a knowledge graph and a search tool over it); Automated retrieval-augmented generation (RAG) Optimization (a plugin recipe that tunes chunking and search parameters with Optuna against an evaluation set); RAG guardrails on relevancy and faithfulness (Advanced LLM Mesh add-on). Semantic Models add structured context: entities, attributes, relationships, and metrics over SQL datasets that the Semantic Model Query tool uses to write precise SQL for agents. The buyer gets retrieval over its own documents and tables, with model and store choice, behind the same permissions its users already carry. The architect's concern is that Dataiku is the orchestrator, not the engine. There is no Dataiku retrieval engine; the embedded stores are open-source libraries Dataiku wraps, and hybrid search, native ranking, and scale are whatever the backing store provides. The Knowledge Bank configuration (chunking, metadata, filters, search settings, security tokens, the semantic model) is Dataiku's object, and the vectors belong to whichever model made them. Calibration: Elastic, MongoDB, and Snowflake read strong on native hybrid engines with generally available models of their own; Cloudflare reads moderate on a managed index plus open models; Cohere moderate on frontier models plus a fixed-function bundle; Qlik moderate on turnkey RAG in a closed box. Dataiku is more general than Qlik (any store, any model, any chunking) and less than Elastic (no engine, hybrid search delegated to the store). Rule 6 pegs strong to the native engines. Moderate.
Split at the store, captive at the configuration. Indexes in open engines the enterprise runs (Elasticsearch, OpenSearch, Milvus, pgvector) are the enterprise's: Retained, or Delegated on a managed variant. Indexes in proprietary services (Azure AI Search, Pinecone, Vertex Vector Search, Snowflake Cortex Search, Databricks AI Search) belong to those owners through Dataiku's paper: Ceded. The embedded Chroma, FAISS, Milvus (local), and Qdrant indexes live in Dataiku managed folders and rebuild from source: the Knowledge Bank that owns them is Dataiku's, Ceded. Hosted embeddings through the Mesh belong to their model's space under the carve-out, Ceded to the owner; open Hugging Face embedding and reranking weights the enterprise serves itself are Delegated on the Arctic Embed precedent. The runtime call (retrieve or not, which filter, which rerank) is Dataiku's engine executing settings the enterprise chose per tool and per request: vendor decides, visible, overridable, Delegated, the Elastic reading.
Semantic Models: ruled September 15, 2026, component kept ('Lab' in the plugin name is a product name, not a GA-gate badge); the reference documentation has a first-class Semantic Models section with no availability badge, the editor ships as the Semantic Models Lab plugin (1.0.0 May 12, 2026; 1.1.3 July 16, 2026), and the grade is moderate with or without the component. GraphRAG, Automated RAG Optimization, and Trace Explorer are plugins installed from the plugin store; no preview or beta badge on any of them. Inference flagged: whether Dataiku Cloud offers every listed vector store connection. Public evidence that moves the cell: a Dataiku-owned retrieval engine with hybrid search and ranking independent of the backing store, or Dataiku-trained embedding models.
Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering
The pipeline graph and every visual transformation in it, with the engine chosen per recipe and builds scheduled and gated by scenarios. Dataiku's format with no second implementer: Ceded.
The enterprise's own transformation code, run on the DSS host, in containers, on Spark, or pushed down as SQL to the enterprise's warehouse. The logic lifts (Pandas, Polars, Spark, the warehouse's SQL) and the Dataiku binding is a thin read-and-write wrapper: Retained (ruled September 15, 2026; the 2B customer tools that subclass Dataiku classes stay Ceded).
Dataiku's connector implementations and the connection settings the enterprise accumulates in them, with fast paths on the major warehouses and object stores. The sources are the enterprise's and each source's own captivity belongs to its vendor's row; the connectors and their settings are Dataiku's: Ceded, the Snowflake Openflow Connectors and Databricks Lakeflow Connect reading.
Lineage is in this layer's own definition: the graph traced across every recipe and project, exportable, with manual lineage for external steps. Computed from and stored in Dataiku: Ceded.
Versioned promotion of a Flow from design to production with the connections remapped per stage. A Dataiku packaging format: Ceded.
Recipes run on the DSS engine, in SQL pushdown, on Spark, or in CPU containers. GPU containers are an option for code recipes that want them, not a requirement.
This is the layer Dataiku was founded on and it's complete. The Flow is a directed graph of datasets and recipes: visual recipes (Prepare with its processor library, Join, Group, Window, Stack, Split, Pivot, Sync, Distinct, Sort, Top N, Sample, Generate Features) and code recipes in Python, R, SQL, Spark, PySpark, and Scala, with notebooks and Code Studios beside them. Each recipe runs on an engine the platform proposes and the user can override: the DSS engine, in-database SQL pushdown to Snowflake, BigQuery, Databricks, Redshift, Synapse, Trino, PostgreSQL, and the rest, Spark (including Spark on Kubernetes), or a container. Sources and destinations are any connection (cloud storage, SQL, Iceberg, SaaS, the metastore), partitioning is first class, DSS 15 added Polars alongside Pandas, fast-path reads on Snowflake, S3, Azure Blob, and GCS (Polars on the three object stores), and fast-path writes on those four plus, when enabled on the connection, BigQuery, Databricks, Redshift, Trino, and Synapse. Scenarios schedule and trigger builds with metrics, checks, and Data Quality gates; column-level lineage is computed across the whole estate; Application-as-recipe packages a Flow as a reusable step; the Project Deployer promotes bundles from Design to Automation nodes with connection remapping. The buyer gets one pipeline surface that spans analysts and engineers and pushes the work down to the engines it already pays for. The architect's concern is that the Flow is the product. The visual recipes, engine choices, scenario logic, and quality gates are Dataiku's format; the code recipes carry the enterprise's logic wrapped in Dataiku's input and output calls; the lineage is computed from the Flow and exists only there. Streaming (Kafka, Amazon Simple Queue Service (SQS), HTTP server-sent events, continuous recipes) carries an Experimental warning and is off by default, so the strong reads on batch. Calibration: Cloudera reads strong on NiFi, Kafka, Flink, Spark, and Airflow spanning any source to any destination; Qlik strong on change data capture (CDC) plus Talend transformations; Databricks strong on Lakeflow and Spark; Snowflake strong on Openflow and declarative pipelines. Dataiku matches the general-purpose test of rule 4 (arbitrary code, insertable stages, any destination) with a decade of production deployment behind it, and adds column-level lineage that most peers score elsewhere. No scope gate applies to the batch surface. Strong on the batch surface, with streaming the named exception. Ruled September 15, 2026: strong stands on rule 4's general-purpose test; streaming has never been a requirement for strong at 1C (AWS, Azure, GCP, Qlik, and VAST read strong without it, VAST on event-driven DataEngine pipelines), and the reviewers' rule 6 argument assumed a frontier the instrument doesn't hold. Streaming stays named as Experimental.
Moderate and captive at the Flow. The logic inside code recipes is the enterprise's and ports (SQL runs on the enterprise's own warehouse; Python runs on Pandas, Polars, or Spark), so that chip reads Retained (ruled September 15, 2026: a thin input-and-output wrapper over an open runtime isn't a binding): the wrapper is Dataiku's, the substance is not. Everything else is Ceded: visual recipes, engine settings, scenarios, quality gates, connectors, lineage, bundles. The runtime call is Dataiku's engine choosing an execution engine and inferring schemas, with the user able to override the engine per recipe and the cluster per project: vendor decides, visible, overridable, Delegated, the Databricks and Snowflake reading under the override rule.
Ruled September 15, 2026: strong stands on general-purpose batch pipelines with column-level lineage; moderate had been argued because streaming is Experimental, and streaming isn't a criterion for strong at 1C on the instrument. Streaming data: the DSS 15 documentation says 'Support for Streaming is Experimental and not fully-supported' and the feature must be enabled by an administrator; unscored. Ruled September 15, 2026 (component): code recipes and notebooks read Retained on the thin-wrapper line; Delegated had been written and Ceded argued on the single-vendor-SDK sentence. Dataiku's own plugins (connectors, recipes) are published on GitHub under Apache 2.0, and the Python API client is Apache 2.0; the platform they plug into is not. Public evidence that moves the cell: a streaming badge change adds a component and settles the grade question.
GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization
Standard Kubernetes provisioned by Dataiku's plugin into the enterprise's tenant, sized and autoscaled by settings the administrator gives it. Managed Kubernetes behind the Kubernetes API: Delegated, the Cloudera ECS and VMware VKS reading.
Any Kubernetes 1.10 or later the enterprise created itself and attaches through kubectl contexts; Dataiku documents that creating such clusters 'is not handled by DSS'. Self-run Kubernetes the enterprise operates independently of Dataiku: Retained, the Cloudera customer-run OpenShift reading. OpenShift 4 and DGX clusters attach the same way but carry an experimental, Tier 2 support badge and are unscored.
How Dataiku places its work: pod specs, namespace policies, cluster selection per project and scenario, base images. Dataiku configuration with no second implementer: Ceded.
The lifecycle of the Dataiku nodes themselves in the enterprise's cloud tenant. Dataiku's provisioning model: Ceded.
Fully managed by Dataiku on Dataiku's accounts; the enterprise picks container sizes inside a pool it can't see below. Bounds are configuration: Ceded.
Where API services run, how many replicas, and their health, with Prometheus integration. Dataiku's deployment model: Ceded.
Managed EKS, AKS, and GKE clusters created by Dataiku's plugins install the NVIDIA driver DaemonSet on GPU node pools (DSS 15 runs it only on GPU nodes); a containerized execution configuration requests GPUs with a nvidia.com/gpu custom limit (1, 2, 4, or 8 recommended for vLLM tensor parallelism). Scheduling of those GPUs is Kubernetes'; Dataiku sets the request.
Dataiku pushes work off its own host onto Kubernetes, and it will build the cluster for you. Elastic AI computation covers Python and R recipes and notebooks, visual recipes on the containerized DSS engine, Spark on Kubernetes, model training and scoring, webapps, local Hugging Face model instances, and API services. Managed clusters (Elastic Kubernetes Service (EKS), Azure Kubernetes Service (AKS), Google Kubernetes Engine (GKE) through Dataiku plugins) are created, attached, started, stopped, and autoscaled from Administration, with node pools the administrator sizes and a global default cluster that each project can override; scenarios can pin a named cluster or create a dynamic one that shuts down when the run ends; unmanaged clusters the enterprise built attach the same way (OpenShift 4 and DGX Kubernetes clusters too, but both pages carry an experimental, Tier 2 support warning and are unscored). Containerized execution configurations set CPU and memory limits, GPU custom limits, namespaces (dynamic namespace policies per project or user), and which groups may use them; image build configurations publish the base images and DSS 15 rebuilds code environment images automatically after an upgrade. Fleet Manager (Cloud Stacks) provisions and upgrades the Dataiku nodes themselves on AWS, Azure, and GCP with instance templates, snapshots, and load balancers. On Dataiku Cloud the whole thing is hidden: a quota of vCPUs, RAM, and parallel activities, container sizes to pick from, and GPU configurations by instance type with credits. The API Deployer runs API services as Kubernetes deployments with replicas and GPU reservations; local Hugging Face models autoscale instances from zero against a target requests-per-instance. The buyer gets compute that appears when a recipe needs it and disappears after. The architect's concern is that this is workload placement, not resource governance. There is no fair-share or priority scheduler, no GPU partitioning, no cross-tenant quota beyond Kubernetes limits and Dataiku Cloud's subscription pool; a GPU is a custom limit on a pod. The cluster is the enterprise's (or Dataiku's, on Cloud), but the execution configurations, cluster plugins, namespace policies, and Fleet Manager templates are Dataiku's. Calibration: Cloudera reads moderate on a platform-scoped Kubernetes with hierarchical quota pools; Databricks and Snowflake moderate on managed compute with GPU pools and no customer-visible scheduler; Palantir moderate on Rubix. Dataiku has Cloudera's shape with less quota machinery and more cluster lifecycle automation. Moderate.
Split at the cluster. A Kubernetes cluster the enterprise built and attaches is its own: Retained. A cluster Dataiku's plugin creates in the enterprise's account is standard Kubernetes reached through the Kubernetes API: Delegated. Everything Dataiku layers on either (execution configurations, cluster plugins and lifecycle, namespace policies, Spark-on-Kubernetes configs, Fleet Manager templates, the Dataiku Cloud quota pool, the Deployer's infrastructures) is Dataiku's: Ceded. The runtime call is Dataiku's engine creating pods with the limits it was given and the autoscaler sizing nodes inside bounds, but the enterprise can pin a recipe to an execution configuration, a project to a cluster, and a scenario to a named or dynamic cluster; a documented per-object pin is an override under the ratified rule: vendor decides, visible, overridable, Delegated.
OpenShift ('Openshift support is experimental and covered by Tier 2 support'; only unmanaged operation) and DGX Systems ('NVIDIA DGX support is experimental'; Multi-Instance GPU (MIG) untested) are unscored on their badges. Cross-row note for the reconcile: Cloudera 2A reads Ceded (its quota-pool bounds are configuration, and eviction isn't an override); Dataiku's per-recipe, per-project, and per-scenario cluster pins are the SUSE Virtualization shape the override rule moved to Delegated. The 2A column hasn't been re-read under the rule (open item from the September 5 batch). Public evidence that moves the cell: a Dataiku-native GPU scheduler with fair-share or priority across projects, or a generally available (GA) badge on OpenShift or DGX.
Model serving, agent execution, inference APIs, distributed inference
The control surface every model call passes through. Connections, guardrail pipelines, quotas, rate rules, and prompt recipes are Dataiku configuration with no second implementer: Ceded.
OpenAI, Anthropic, Bedrock, Vertex, Mistral, Foundry, Azure OpenAI, Cohere, Snowflake Cortex, and Databricks text and multimodal models on the enterprise's own keys, with Dataiku exposing an OpenAI-compatible Chat Completions and Responses API in front. Standard interfaces on both sides: Delegated, the S3-of-inference reading. Embedding models through the same connections are the carve-out and read Ceded at Layer 1B.
Open weights the enterprise downloads or caches and serves in its own containers, with an inference engine (vLLM) that runs anywhere. The artifact and the engine are open; the serving surface and instance controls are Dataiku's: Delegated on the open-weights precedent.
Weights the enterprise produced or imported and possesses as files it can serve outside Dataiku (the Python fine-tuning path must produce safetensors; MLflow models import and export). The artifact is the enterprise's: Retained, the Cloudera customer-owned models reading.
Proprietary NIM containers on an NVIDIA license through Dataiku's connection. Ceded to NVIDIA under the channel-substitution rule, the SUSE and Elastic reading; the OpenAI-compatible endpoint in front is the Delegated chip above.
The agent harness and every definition in it. Dataiku objects, run and scaled by Dataiku, exported nowhere: Ceded.
The enterprise writes the code, but it subclasses Dataiku's classes and runs on Dataiku's runtime; the logic inside a LangGraph graph would port, the tool and agent objects would not. Bound to a single-vendor SDK: Ceded, the Palantir and Qlik reading, pending the open-runtime ruling.
Fine-tuned models that live in the provider's account and are deployed and billed there. Ceded to the provider through Dataiku's paper; local Hugging Face fine-tunes are the Retained exception, scored in the customer-owned models chip.
Classical real-time serving with versioned endpoints and OpenAPI documentation. The service packaging is Dataiku's, though models export to Python, Java, MLflow, and PMML: Ceded.
Text generation, embedding, reranking, and image models run in Dataiku containers on NVIDIA GPUs (compute capability 7.5 or higher, CUDA 13 driver) with vLLM as the default engine and transformers as a deprecated fallback; presets size models to 24 GB and 48 GB GPUs, and DSS 15 added CPU inference for some models (AVX-512 recommended). CUDA packages are pulled from NVIDIA's repository at environment build.
NIM microservices from the NVIDIA API catalog or self-hosted through the NIM Kubernetes operator, requiring an NVIDIA AI Enterprise subscription or Developer Program membership, for chat, multimodal, and embedding models with streaming and tool calling.
Dataiku's runtime has three parts, and each is production. The LLM Mesh is the gateway: administrator-defined connections to Anthropic, AWS Bedrock (with cross-region inference profiles and Bedrock Guardrails), SageMaker, Microsoft Foundry, Azure OpenAI (including through API Management (APIM)), Azure ML, Cohere, Databricks Mosaic AI, Google Vertex, Mistral, NVIDIA NIM, OpenAI, Snowflake Cortex, Stability AI, and custom plugin providers; DSS 15 added Claude Sonnet 5 and Fable 5, Nova 2 embeddings, and GPT-5.6 presets on Bedrock. Guardrail pipelines configured per connection, per agent, or at the point of use (personally identifiable information (PII), prompt injection, toxicity, topic boundaries, forbidden terms, custom code; bias detection is an unsupported add-on plugin) check the query and, when enabled, the response; every call is metered against cost quotas that can block, throttled by requests- or tokens-per-minute rules, and traced. Applications reach it through the LLM Mesh API, an OpenAI-compatible API that serves Chat Completions and Responses for any model in the Mesh, and Persisted Conversations that keep history server side with human-in-the-loop tool validation. Local Hugging Face models are the self-hosted half: open weights (the presets include Qwen3, GPT-OSS 20B, Llama 3.3 70B, Gemma 4, Mistral Small 4, Kimi Linear, DeepSeek V4 Flash) served on the enterprise's GPUs under vLLM, with instances that scale from zero to a maximum against a target load; fine-tuning (Advanced LLM Mesh add-on) covers OpenAI, Azure OpenAI, Bedrock, and local models, and a Python recipe can produce a safetensors model of its own. Agents are the third part: Simple Visual Agents pick tools autonomously; Structured Visual Agents compose blocks (agentic loop with exit conditions, LLM request, manual and mandatory tool call, delegate to another agent, routing, for-each, parallel, reflection, Python code, long-term memory, context compression, generate artifact) with Common Expression Language (CEL) expressions and Jinja templates and, since DSS 15, human confirmation of tool calls inside the blocks; Code Agents implement a class Dataiku hosts, scales, and audits, with LangChain and LangGraph inside if wanted; Agent Skills package instructions in the SKILL.md format; twenty-five-plus managed tools (Knowledge Bank search, semantic model query, dataset lookup and append, remote and local Model Context Protocol (MCP), OpenAPI with endpoint discovery, model predict, API endpoint, image generation, send message, generate artifact, SQL question answering, Google Search and Calendar, Jira, Salesforce, ServiceNow, Snowflake Cortex Analyst and Search, Databricks Genie and Vector Search, LLM-side web search) plus inline code and plugin tools, each with optional human approval that pauses the agent before the call and can let the user edit the arguments. Agents are exposed as MCP tools through Dataiku's own MCP server per project, as Agent2Agent (A2A) endpoints (specification 0.3 and 1.0), in Slack and Teams, and as virtual LLMs anywhere the Mesh is used; Agent Review (LLM-as-judge traits plus human review) and the Evaluate Agent recipe (Advanced add-on) test them. Classical serving is the API node: visual, Python, R, and MLflow models, arbitrary functions, SQL queries, and dataset lookups as versioned REST endpoints on API nodes or Kubernetes, with models exportable to Python, Java, MLflow, and Predictive Model Markup Language (PMML). The buyer gets model access, self-hosted open weights, and governed agents on one runtime it already uses for pipelines. The architect's concern is the SDK. Agent definitions, block logic, tool descriptors, guardrail pipelines, quotas, and prompt recipes are Dataiku objects; inline tools and Code Agents subclass Dataiku's classes; human approval isn't available for tools called through the MCP server or inside for-each, parallel, and reflection sub-sequences. The models are the enterprise's; the harness is not. Calibration: Databricks, Snowflake, and Palantir read strong on serving plus generally available agent runtimes with evaluation and tracing; Cloudera strong on inference plus Agent Studio. Elastic was held at moderate on first-release maturity with no evaluation surface and tracing in preview; Dataiku has generally available tracing, evaluation, and review, several releases of the agent surface, and self-hosted serving the Elastic row lacked. The Advanced LLM Mesh add-on gates fine-tuning, Agent Review, Evaluate Agent, RAG guardrails, and custom quotas; the gateway, local serving, every agent type, the managed tools, human approval, tracing, MCP and A2A exposure, and the API node are in the standard product. Strong, written. Ruled September 15, 2026: a vendor's own purchasable add-on on the same paper is a price, not a rule 5 gate (the whose-paper rule credits what the enterprise buys with the vendor's support); the add-on is named, and the grade stands.
Split at the model, captive at the harness. Model access through the enterprise's own provider accounts behind the Mesh's OpenAI-compatible surface is Delegated (the inference-interface ruling); open Hugging Face weights the enterprise serves itself are Delegated, and the fine-tunes and imported models the enterprise owns outright are Retained; NIM is Ceded to NVIDIA through Dataiku's paper; hosted fine-tunes belong to their providers, Ceded. The Mesh configuration, agent definitions, tools, and API services are Dataiku's, Ceded, and so is customer-authored tool and agent code, because it binds to Dataiku's tool descriptor and BaseLLM classes rather than to a multi-vendor interface (the Palantir and Qlik SDK reading). The decision to act is the model's, and the enterprise can put deterministic code between that decision and its effect: human approval on a tool persists in the tool's settings and pauses the agent before every call; Structured Visual Agent blocks route, gate, and run Python before the next step; configured guardrails execute before the call and, where enabled, after it. Model decides, visible, overridable, Delegated, the Anthropic hooks and Mistral per-function approval reading.
Ruled September 15, 2026: strong stands; moderate had been argued on rule 5 because Agent Review, Evaluate Agent, fine-tuning, RAG guardrails, and custom quotas require the Advanced LLM Mesh add-on, and the ruling reads an add-on as a price, not a gate. Customer-authored tool code: written Ceded on the single-vendor SDK binding; the LangGraph logic inside a Code Agent is the same open-runtime question as Cloudera's CrewAI tool classes and Cloudflare's Workers code (open ruling from the September 5 batch) and follows that ruling. Dataiku Code Assistant for Code Studios is deprecated in DSS 15 in favor of OpenAI Codex, Claude Code, OpenCode, Gemini command-line interface (CLI), and GitHub Copilot, which Dataiku now documents as coding agents inside Code Studios. Human approval isn't supported on tools exposed through the MCP server or on tools called by agents exposed through it, nor inside For Each, Parallel, and Reflection sub-sequences; the A2A server timeout is 30 minutes. Bias detection is a 'Not supported' plugin behind the add-on. Inference flagged: the exact list of models in the local GPU and CPU preset catalogs changes by release. Public evidence that moves the cell: nothing upward from strong.
Policy-driven placement and resource coordination — the Autonomy Layer
The gateway leg. Every call is inspected, metered, throttled, and logged in Dataiku's pipeline. Dataiku configuration: Ceded.
The registry leg, with a persisted human gate that the Deployer enforces deterministically per infrastructure for project bundles and model versions. Govern templates and sign-off definitions are Dataiku's: Ceded.
The orchestration leg, decided by a model over agent descriptions or by blocks the enterprise arranged. Dataiku's definitions: Ceded.
The observability leg. Traces are a nested JSON the enterprise can push to LangSmith or store in its own datasets, but the evaluation stores and reviews are Dataiku's: Ceded. Unified Monitoring (project and endpoint health, Prometheus) is operations telemetry beside this leg, not part of it.
Guardrails, quotas, Govern workflows, and tracing run on the Dataiku nodes. The only GPU-adjacent item is the guardrail detectors' models, which can be hosted or local.
Read against the five legs the instrument asks of an Intelligence-2C plane, Dataiku has four in production. Gateway: every agent and model call passes the LLM Mesh, where guardrail pipelines configured per connection, per agent, and at the point of use run on the request and, when enabled, the response, cost quotas block when exhausted, requests- and tokens-per-minute rules throttle per model and provider, and every call is audited and traced; agents run with the permissions of the API key that called them, and Agent Hub can pass the DSS caller ticket so tools execute as the end user, with document-level security tokens carried in the completion context. Registry: Dataiku Govern (Advanced license) syncs agents, augmented LLMs, and fine-tuned LLMs with their versions into a GenAI registry beside models and bundles, wraps each in a workflow with feedback and final approval sign-offs, records a timeline of every change, and lets the Deployer refuse to deploy an unapproved project bundle or model version (a policy per infrastructure: prevent, warn, or always deploy); an agent ships to production inside a project bundle, so the gate covers the agentic project, but the agent version's own registry approval isn't what the Deployer checks. Cross-agent orchestration: Agent Hub's orchestrating LLM routes a conversation across enterprise agents as tools or in parallel; Structured Visual Agents delegate to other agents, route, fan out, and reflect; agents are tools to other agents; A2A and MCP endpoints expose them outward. Observability: nested traces from every Mesh call with guardrail and cost spans, the Trace Explorer webapp, LangSmith push, interaction logging datasets, the Evaluate Agent recipe writing trajectory and answer metrics into evaluation stores with alerting (Advanced add-on), Agent Review's trait-based tests, with Unified Monitoring beside them as project and endpoint health telemetry rather than agent observability. Missing: first-class agent identity. An agent has no principal of its own; it borrows the API key's permissions or impersonates the calling user. The buyer gets governance that reaches from build (sign-off) to run (guardrails, quotas) to review (evaluation) inside one platform. The architect's concern is that most of this plane decides by instruction or by model, not by deterministic code. The Agent Hub router is an LLM choosing an agent from descriptions; the PII, toxicity, prompt-injection, and topic detectors are models (bias detection is an unsupported plugin behind the add-on); forbidden-terms matching, quotas, rate limits, and the deployment policy are the deterministic parts. That is guardrails on state and a commit gate, which the instrument credits, and not outcome validation, which it notes as universally absent. There is no placement reasoning: which model serves a request is whatever the agent or recipe names. Calibration: GCP holds strong on all five legs; Azure strong with a deterministic policy engine as the orchestration leg; Snowflake, Databricks, SAP, and Kamiwaza read moderate with a leg named as missing (Snowflake: gateway and orchestration; SAP: identity and observability). Dataiku is the fullest moderate on the map, four legs, and the rule is explicit that a partial plane is moderate. Moderate, identity the named leg.
High and captive. Guardrail pipelines, quota scopes, rate rules, Govern templates and sign-off definitions, deployment policies, Agent Hub routing configuration, and evaluation recipes are Dataiku objects: Ceded. The runtime call is split: the deterministic parts (forbidden terms, quotas, limits, the deploy policy, a final approval) execute rules the enterprise wrote, and the router and the detectors are models. Quotas, rate rules, and allow-lists are bounds and don't decide direction; the persisted gates do: a Govern final approval or rejection that the Deployer enforces before a bundle or model version ships, human approval that pauses an agent before a tool's effect, and guardrail actions the enterprise sets per connection and per agent. Vendor decides, visible, overridable, Delegated, the Snowflake reading; the missing gateway leg there is the missing identity leg here, and both are capability findings, not authority ones.
Watch-list (dated): Dataiku Agent Management, a cross-platform agent control tower connecting to agents on n8n, Bedrock, Salesforce, and Dataiku through native APIs and log streams, with key performance indicator (KPI) tracking, drift detection, and governance workflows; the product page says it 'will be available in October 2026'. Not scored; would add to the registry and observability legs (reach across platforms isn't a leg of its own). Govern's GenAI registry needs the Advanced license (a license gate, named; the registry is generally available). Bias detection is a 'Not supported' plugin behind the Advanced LLM Mesh add-on and isn't credited. A reviewer argued the authority reading should be decides model because the Agent Hub router is an LLM; rejected on the Snowflake and Databricks 2C readings (model-driven orchestration legs, decides vendor) and logged for the reconcile. Persisted Conversations use project permissions, and the documentation says end_user_id 'is not an access-control boundary'. Public evidence that moves the cell: a first-class agent principal with its own scoped credentials, or Agent Management generally available in the documentation.
AI-powered business capabilities — business logic, workflow automation
The consumer portal for governed agents and the business user's own agent builder, delivered as a Dataiku plugin webapp. Dataiku's product: Ceded.
Self-contained applications for named business processes, installed and operated inside DSS. Dataiku's applications on Dataiku's runtime: Ceded.
The analytics and app-packaging surfaces business users consume. Dataiku formats: Ceded.
The building agent and the embedded assistants that act inside Dataiku. The models are the enterprise's choice through Dataiku AI Services or the LLM Mesh (Anthropic, OpenAI, Google, Snowflake Cortex, Databricks AI Gateway, Bedrock, Foundry, open-source and on-premises models); the prompts and the project objects they act on are Dataiku's: Ceded.
Agent Hub, the Business Applications, Solutions, Stories, and Cobuild run on the Dataiku nodes and call models through the Mesh. Dataiku's AI Blueprint for manufacturers with NVIDIA (June 25, 2026) is a go-to-market fact, not a runtime dependency.
The value plane is where Dataiku's agent story lands for people who don't build. Agent Hub (a Dataiku plugin, version 1.6.0 on August 21, 2026, requiring DSS 14.6 or later) is the portal: a curated library of enterprise agents, a single chat that orchestrates several of them, My Agents that business users build themselves from uploaded documents, SharePoint libraries, and datasets, chart generation with versioning, document guardrails, and an admin panel that decides which agents, models, and tools each Hub exposes; Agent Chat puts a single agent behind a shareable interface; Dataiku Answers, the earlier packaged chat and RAG webapp, is documented as mostly replaced by the Hub. Three Business Applications ship as self-contained products: Process Mining (DSS 14.5 or later) discovers process flows from event logs, analyzes variants, finds bottlenecks, and drills to cases; the RFx Accelerator (DSS 14.6 or later) extracts the structure of RFPs and RFIs, drafts answers from a curated knowledge base, and exports the questionnaire; Manufacturing Operations delivers a Parameters Analyzer and a Manufacturing Event Tracker. Forty-plus Dataiku Solutions install as accelerators across retail (demand forecast, market basket, markdown optimization, site selection), banking and insurance (anti-money-laundering (AML) alerts triage, credit scoring, credit risk stress testing, claims modeling), pharma (pharmacovigilance, clinical trial intelligence, cohort discovery), manufacturing (batch performance, maintenance planning, production quality control), finance (financial forecasting, reconciliation), and governance itself (EU AI Act Readiness, ISO 42001 Readiness, LLM Provider Due Diligence). Stories are refreshing data presentations with generative features, dashboards and workspaces carry the analytics, and Dataiku Applications package any project as an app with tiles and forms. Cobuild (generally available June 18, 2026) is the building agent: it turns a business objective into a project, and in 15.0.1 creates evaluation recipes and dashboards, with per-user instructions and an administrator switch to block project-level custom prompts. Expert-to-Agent (May 2026) is the packaging name for this stack, not a separate product. The buyer gets a place for the business to use, and build, agents on governed data, plus packaged applications for a handful of named processes. The architect's concern is thin domain depth and what carries the grade. The Business Applications are three, gated to the Enterprise edition or an Enterprise AI Package and to Designer license types; the Solutions are templated projects with sample data that Dataiku says it doesn't support and offers at the customer's own risk, so they are named here for reach and not scored (rule 1 credits what ships with the vendor's support); the Hub is a plugin the administrator installs; everything is a Dataiku project and stays one. Calibration: Databricks reads strong on Genie, dashboards, Apps, and Databricks One; Snowflake strong on CoWork, CoCo, Observe, and Streamlit; Qlik strong on the analytics value plane; SAP strong on business applications with Joule; Cloudera moderate on visuals and a workbench copilot. Dataiku has Databricks' shape (a consumer portal for agents plus analytics surfaces) with packaged business applications SAP-style at small scale. The MongoDB and Redis question (is a developer assistant a value plane) doesn't decide this cell; Cobuild is a component, not the grade. Strong, written, on Agent Hub with My Agents, Answers, the analytics surfaces, Dataiku Applications, Cobuild, and the three Business Applications; the edition gate is named. Ruled September 15, 2026: strong stands; an edition tier of the same product is a price, not a rule 5 gate, and the Solutions were already unscored; moderate had been argued on the edition gate.
High and captive. Agent Hub configurations, My Agents, Business Applications, Stories, dashboards, and Dataiku Applications are Dataiku projects and webapps with no other host: Ceded. The runtime call in the Hub and its agents is Dataiku's application logic wrapped around a model's choices (which agent answers, which tool it calls), and the enterprise can gate the effect on the Hub and Agent Chat paths: human approval on a managed tool persists in the tool's settings and pauses the agent before the call (not for tools reached through the MCP server, and not inside for-each, parallel, or reflection sub-sequences); the Hub's document guardrails and agent allow-lists are administrator rules. Vendor decides, visible, overridable, Delegated, the SUSE Liz reading for an application whose function is acting; MongoDB's read-only assistant is the contrast, and the Databricks, Snowflake, Qlik, and SAP Layer 3 cells predate the override rule.
Ruled September 15, 2026: strong stands (an edition tier of the same product is a price, not a gate); moderate had been argued on the edition gate and the unsupported Solutions. Business Applications access line, verbatim from the documentation: 'available starting with DSS 14.4 as part of the Dataiku Enterprise edition or an Enterprise AI Package', with Data Designer, Advanced Analytics Designer, or Full Designer license types; a license gate under rule 5, named in the cell. Dataiku Solutions (forty-plus installable accelerators across retail, banking, insurance, pharma, manufacturing, finance, and governance readiness for the EU AI Act and ISO 42001) are unscored: the catalog says 'Dataiku does not support Dataiku Solutions and makes no representations or warranties', and rule 1 credits only what ships with the vendor's support. Answers is 'mostly replaced by Agent Hub' per its own page; Agent Connect is deprecated and replaced by Agent Hub. The Dataiku Code Assistant is deprecated in DSS 15. A reviewer argued the authority reading should be Ceded on the Databricks, Snowflake, Qlik, and SAP Layer 3 cells; rejected on the override rule and the SUSE Layer 3 reading, and logged for the reconcile of the pre-rule Layer 3 column. Public evidence that moves the cell: nothing upward from strong.
Dataiku DSS 15 reference documentation (release notes 15.0.0 August 14, 2026 and 15.0.1 September 10, 2026; Generative AI and LLM Mesh: LLM connections, Running Hugging Face models, Guardrails, Cost Control, Rate Limiting, Fine-tuning, LLM Mesh API, Evaluating LLMs; Adding Knowledge to LLMs: vector stores, search settings, document-level security, GraphRAG, Automated RAG Optimization, RAG guardrails; AI Agents: introduction, Simple and Structured Visual Agents, Code Agents, Agent Skills, managed tools, human approval, MCP server, A2A, tracing, Agent Review, Agent Evaluation, Agent Hub, Agent Chat; Semantic Models; Catalog; Data Lineage; Metastore catalog; Streaming data; Elastic AI computation, managed Kubernetes clusters, EKS, DGX Systems; Installing and setting up, Cloud Stacks for AWS, custom install; API Node and API Deployer, deploying on Kubernetes; Production deployments and bundles; MLOps and Unified Monitoring; AI Governance: items, process features, deployment policies; Process Mining, RFx Accelerator, Manufacturing Operations; Stories; Dataiku Applications; Dataiku AI and Cobuild; Security; Agent Hub release notes 1.6.0 August 21, 2026; Trace Explorer release notes); Dataiku Knowledge Base (Compute and Resource Quotas on Dataiku Cloud: compute engines, elastic AI compute capacity, containerized execution configurations for standard or GPU compute; Dataiku Solutions catalog); Dataiku product pages (plans and features, LLM Guard Services, Agent Management, Expert-to-Agent); Dataiku press releases ($350M ARR, October 17, 2025; Cobuild general availability, June 2026; Expert-to-Agent, May 2026); the dataiku-api-client-python repository (Apache 2.0). Peer-reviewed cell by cell through the labs claims ledger (dataiku-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (dataiku-row-agy); totals and escalated items in reviews/dataiku-judgment.md.