Executive Summary: Google Cloud AI Infrastructure

Google Cloud is the only vendor in this assessment series that owns a frontier foundation model — and that single fact restructures the entire 4+1 analysis. Google built Gemini, trains Gemini on its own TPUs, optimizes its silicon for Gemini’s training requirements, and weaves Gemini’s intelligence into every layer of its cloud platform. This creates a model-integrated stack: an architecture where the frontier model is not a component plugged into infrastructure but the intelligence that pervades the infrastructure.

No other vendor assessed possesses this vertical integration. Google owns every layer of the 4+1 model with proprietary IP: custom silicon (TPUs), custom networking (Virgo), proprietary storage (Colossus/Spanner/BigQuery), its own runtime and frameworks (JAX/Pathways), its own frontier models (Gemini), and a unified orchestration surface (Gemini Enterprise Agent Platform).

The DAPM implication is not merely that every layer is Ceded — it is that every layer is ceded to a unified intelligence. With AWS, authority is distributed across multiple vendors’ judgment (AWS infrastructure, Anthropic model reasoning, ISV application logic). That distribution creates complexity but also structural checks. With Google Cloud + Gemini, the enterprise concentrates authority in one vendor’s judgment across every layer — from silicon to application.

This is the deepest expression of vertical integration in enterprise technology since the mainframe era. The enterprise gains end-to-end optimization that no multi-vendor assembly can match. But the 4+1 framework makes visible what the integration obscures: the enterprise has no fallback position at any layer. Google Distributed Cloud (GDC) addresses data sovereignty without addressing judgment sovereignty — GDC still runs Google’s software stack and Google’s models.

The structural question: does concentrating all layers of authority and all layers of model judgment in a single vendor deliver enough value to justify the governance position — and has the enterprise made that concentration explicit rather than inheriting it by default?

Layer-by-layer status: Layer 0 (TPU + GPU Full Stack), Layer 1A (Governed Catalog; Gemini Context Graph in Preview), Layer 1B (Model Prep & Managed Retrieval), Layer 1C (BigQuery-Native Pipelines), Layer 2A (GKE + Sovereign Extensions), Layer 2B (Model-Integrated Stack), Layer 2C (Complete Intelligence-2C Plane, No Placement), Layer 3 (+1) (Open Model Layer, Captive Platform).

Assessment framework: 4+1 Layer AI Infrastructure Model. Scoring model: Decision Authority Placement Model (DAPM) — Retained, Delegated, or Ceded. Published by The CTO Advisor LLC (DBA The Advisor Bench). Author: Keith Townsend. Date assessed: October 5, 2026. Version: v1.19 - 4+1 v2: Layer 1A Re-read on the GA Gate.

Disclosure: The CTO Advisor has a commercial relationship with this vendor. Assessment findings are independent. See disclosure policy.

Google Cloud AI Infrastructure

Mapped to the 4+1 Layer AI Infrastructure Model

v1.19 - 4+1 v2: Layer 1A Re-read on the GA GateAssessed October 5, 2026Sources & revision history
ACTIVE ASSESSMENT

Summary Finding

Google Cloud is the only vendor in this assessment series that owns a frontier foundation model — and that single fact restructures the entire 4+1 analysis. Google built Gemini, trains Gemini on its own TPUs, optimizes its silicon for Gemini’s training requirements, and weaves Gemini’s intelligence into every layer of its cloud platform. This creates a model-integrated stack: an architecture where the frontier model is not a component plugged into infrastructure but the intelligence that pervades the infrastructure.

No other vendor assessed possesses this vertical integration. Google owns every layer of the 4+1 model with proprietary IP: custom silicon (TPUs), custom networking (Virgo), proprietary storage (Colossus/Spanner/BigQuery), its own runtime and frameworks (JAX/Pathways), its own frontier models (Gemini), and a unified orchestration surface (Gemini Enterprise Agent Platform).

The DAPM implication is not merely that every layer is Ceded — it is that every layer is ceded to a unified intelligence. With AWS, authority is distributed across multiple vendors’ judgment (AWS infrastructure, Anthropic model reasoning, ISV application logic). That distribution creates complexity but also structural checks. With Google Cloud + Gemini, the enterprise concentrates authority in one vendor’s judgment across every layer — from silicon to application.

This is the deepest expression of vertical integration in enterprise technology since the mainframe era. The enterprise gains end-to-end optimization that no multi-vendor assembly can match. But the 4+1 framework makes visible what the integration obscures: the enterprise has no fallback position at any layer. Google Distributed Cloud (GDC) addresses data sovereignty without addressing judgment sovereignty — GDC still runs Google’s software stack and Google’s models.

The structural question: does concentrating all layers of authority and all layers of model judgment in a single vendor deliver enough value to justify the governance position — and has the enterprise made that concentration explicit rather than inheriting it by default?

Strength
Moderate
Gap
Partner
Layer 0 · ComputeCompute & Network Fabric3 LABSTPU + GPU Full Stackdecides: vendor · Ceded▼

Raw compute, networking, and acceleration fabric

Vendor-Provided

TPU Single-Host Slices Through an Open Runtime (TPU7x Ironwood, v6e Trillium, v5p VMs Under JAX, PyTorch/XLA, vLLM)Retained

Generally available TPUs rent as single-host VMs: v6e from 1 chip (ct6e-standard-1t, 4t, 8t), Ironwood and v5p at 4 chips per VM. Workloads written in JAX or in PyTorch through XLA also run on NVIDIA GPUs, so the opinions lift without a rebuild: Retained under the Layer 0 runtime rule.

TPU Single-Host Slices Through TPU-Specific Code (Pallas Kernels Written for TPU, Sharding Tuned to TPU Memory)Ceded

Kernels and sharding written for the TPU's memory hierarchy don't run elsewhere, and leaving means rewriting them: Ceded under the Layer 0 runtime rule. Most TPU workloads hold both paths, which is why the slices are recorded as a split.

TPU Pod-Scale Capacity (Multi-Host Slices on the Inter-Chip Interconnect)Ceded

Multi-host slices up to full pods, wired by Google's inter-chip interconnect. Ceded on the integrated-system rule, the NVL72 reading: pod-scale capacity on a captive chip-to-chip interconnect can't be lifted as deployed. Google designs TPUs to train Gemini, so the enterprise doesn't direct the optimization priorities; that's narrative, not what places the chip. Read the same way as Trainium3 UltraServers on AWS's row.

NVIDIA GPU Instances on Google Cloud (A3, A3 Mega, A3 Ultra; G4 Fractional GPUs)Retained

A3/A3 Mega/A3 Ultra and fractional G4 VMs (industry-first RTX PRO 6000 Blackwell vGPU). Standard NVIDIA instances: a CUDA or PyTorch workload moves to another cloud's NVIDIA instances or an OEM server without a rebuild. Retained under the Layer 0 runtime rule, read from the customer's seat; the CUDA dependency is NVIDIA's and sits in the NVIDIA column. Third-party models (Claude, Llama, Mistral) run on NVIDIA, not TPU: the playing field is structurally uneven.

NVIDIA NVL-Class Rack-Scale Capacity on Google Cloud (A5X Bare Metal on Vera Rubin NVL72)Ceded

A5X bare metal on Vera Rubin NVL72 (among the first cloud providers to deploy). Scale: 80,000 Rubin GPUs single-site, 960,000 across multisite. A4 Ultra NVL72 was in preview at last check and isn't scored. Ceded on the integrated-system rule, the Dell XE9812 and Lenovo NVL-class reading: a rack-scale system built around a captive NVLink fabric can't be lifted as deployed.

Virgo Network FabricCeded

Purpose-built AI-optimized DC fabric. 134,000 TPU 8t chips connected at 47 Pb/s non-blocking bi-section bandwidth per DC. 4x bandwidth per accelerator, 40% lower unloaded latency vs prior gen. Also available for A5X. Designed for Gemini’s training topology.

Google Distributed Cloud (GDC)Delegated

On-prem deployment of Google Cloud services. Connected and air-gapped configs. 4 racks to hundreds. NVIDIA Blackwell GPUs + Gemini Flash models on-prem. Managed GDC Provider initiative (Clarence, Gulf Energy, T-Systems, WWT). NATO deployment. Customer provides facility; Google provides and operates HW+SW. Addresses data sovereignty but not judgment sovereignty.

Managed Lustre (Parallel File System for Training Data)Ceded

10 TB/s bandwidth (10x YoY, claimed 20x faster than other hyperscalers), 80 PB capacity. RDMA-enabled. Storage substrate feeding TPU and GPU clusters, GA June 30, 2025; moved from 1C under the M2 storage ruling (October 5, 2026).

NVIDIA-Provided

NVIDIA GPU Silicon

Vera Rubin NVL72, Blackwell B200/B300, H100/H200. 1M+ NVIDIA GPUs. NVIDIA instances serve third-party models that can’t run on TPU.

NVIDIA NIXL + Networking

NIXL for disaggregated inference. ConnectX/BlueField for GPU networking. Google manages the NVIDIA integration layer.

◆ Gap Analysis

No Layer 0 capability gap — Google’s portfolio is the broadest of any single cloud provider. The gap is governance: the enterprise has no authority over any Layer 0 component beyond selecting instance types. The silicon-model feedback loop is structurally unique: Google designs TPUs to train Gemini, not primarily to sell cloud compute. TPU roadmap decisions reflect Gemini’s training topology, not enterprise customer workload requirements. The enterprise inherits optimization it didn’t direct. AWS’s Trainium is designed for customer workloads. NVIDIA designs for the broadest market. Google designs for Gemini and makes TPUs available to customers. The multi-accelerator matching problem (TPU vs NVIDIA vs Axion CPU) creates a workload-to-silicon decision that recurs per-workload in cloud vs once at procurement on-prem. No productized policy engine automates that matching. Fluid Compute (Layer 2A) begins to address it but doesn’t consult governance metadata. GDC follows the same inverted operating model as AWS AI Factories: Google operates infrastructure the customer houses. Unlike Dell PowerRack or HPE ProLiant (enterprise-owned hardware), GDC is Google-operated even when customer-hosted.

◆ Borrowed Judgment

The silicon-model feedback loop: Google’s TPU roadmap is driven by Gemini’s training requirements. If Google decides TPU 9 should optimize for MoE architectures because that’s where Gemini is heading, every enterprise TPU workload inherits that architectural bet. Borrowed judgment at the silicon layer — a concept with no parallel in the Dell or HPE assessments. Virgo as borrowed network judgment: the enterprise inherits Google’s network optimization decisions without visibility or control. Cannot audit bandwidth sharing across tenants or prioritization of Google’s own Gemini training traffic.

◆ Working Notes

The dual-architecture hedge (TPU + NVIDIA) gives Google pricing leverage and architectural independence. The enterprise benefits indirectly but does not control whether NVIDIA GPU instances remain first-class citizens as Google optimizes for its own silicon. Lab-measured (July 2026, reconcile v1.6): the gemma4-tpu-inference lab measured the self-serve path onto this silicon for a bring-your-own mid-size mixture-of-experts model. v5e was reachable, with the serving quota bump auto-approved. Trillium v6e capacity was dry in all three attempted zones on the measured date with quota clean, so the constraint was capacity, not policy. On v5e this model's two global key-value heads cap pure tensor parallelism at 2, a sharding constraint common to mixture-of-experts models on any multi-accelerator substrate, and at that width bf16 does not fit the per-chip memory budget, so the deployable path is quantizing off-box or adopting Google's own serving stack. Scoped to the on-demand self-serve lane, a single project, and the open-source vLLM-TPU stack; no TPU serving latency was measured in this lab. The strong grade stands. The silicon and fabric are real; the friction finding lands on the authority axis, not the capability axis. Named, not scored (October 5, 2026 GA-gate): TPU 8t (training, 9,600-chip superpods) and TPU 8i (inference, 288GB HBM, 384MB on-chip SRAM), announced April 22, 2026 at Google Cloud Next. The product page lists them as Coming Soon and Compute Engine's TPU documentation covers TPU7x, v6e, and v5p only. At GA they read like the GA generations: single-host slices split by runtime, pods Ceded on the integrated-system rule.

Validated in Labs

Hands-on builds on real hardware that test where authority actually holds at this layer.

Lab 012Buy the harness, not the tier

Open weights on capacity I already carry clear 17 of 22 repairs for nothing, and the gate names the five they miss. Renting more hardware for the rest is slower and dearer than the API. The minimum that finishes the job is a mini model in an agentic harness, at $3.14. Two Opus generations finish the same 22 for four times that.

Lab 008Renting the chip was the easy part

Renting the TPU is fast; serving a bring-your-own model on it means adopting Google’s stack or quantizing off-box. The friction was the finding.

Lab 007The CPU exit is a batch lane, not a serving lane

Across 22 measured configurations, zero met the interactive latency bar; the Xeon lane delivers throughput or interactive latency, not both.

Layer 1A · StorageData Storage & Governance1 LABGoverned Catalog; Gemini Context Graph in Previewdecides: vendor · Delegated▼

Durable, governed data foundation — the Governance Catalog that Layer 2C queries

Vendor-Provided

Cloud Storage + Rapid Tier (Colossus)Ceded

Standard object storage at planetary scale. Rapid tier uses Colossus — Google’s internal distributed storage platform (previously powering Search, Gmail, YouTube, Gemini training). Sub-millisecond read/write. The enterprise gets the same storage engine that holds Google’s training data.

Smart Storage Object ContextsCeded

Structured, mutable, IAM-governed context on every Cloud Storage object (GA): tags, classifications, and extracted entities the enterprise writes or annotation pipelines attach, discoverable by downstream systems. Automated annotation at write time with Gemini is preview and not scored.

Knowledge Catalog (Formerly Dataplex Universal Catalog)Ceded

Governance catalog across BigQuery and Cloud Storage: automatic discovery (GA), column-level lineage (GA), data profiling and quality scans, data products (GA), Iceberg REST cataloging (GA), dbt metadata import (GA), and policy tags with row and column security. Gemini context-graph features (unstructured insights, lookupContext, relationships discovery) are preview and not scored.

BigQuery StorageCeded

Serverless columnar storage for structured/semi-structured. Managed Iceberg tables. Separates storage and compute. BigQuery spans storage, analytics, ML, and governance in a single service.

Lakehouse Runtime Catalog and Governance (Formerly BigLake)Ceded

Unified fabric across Cloud Storage and BigQuery: single schema, multiple engines, row and column governance through the BigQuery Storage API across access paths including open-source engines, multi-cloud reach through BigQuery Omni. The governance and runtime catalog are Google's own surface. BigQuery writes and automatic table management for managed Iceberg tables are preview and not scored. Moved from 1C October 5, 2026: a catalog grades at 1A.

Apache Iceberg REST Catalog Interface (Lakehouse)Delegated

Iceberg tables served through the Apache Iceberg REST catalog, a multi-vendor standard interface Spark, Trino, and other engines speak; Lakehouse namespace and table commands GA April 21, 2026, automated cataloging in Knowledge Catalog GA April 16, 2026. A Google-managed implementation behind a standard interface: the tables and their catalog contract lift to another Iceberg REST catalog.

NVIDIA-Provided

Assessment pending

◆ Gap Analysis

Knowledge Catalog (Dataplex Universal Catalog, renamed April 10, 2026) answers what Layer 2C needs to know about the data, generally available: automatic discovery of Cloud Storage data (GA June 14, 2026), column-level lineage (GA, including Dataproc May 15, 2026), data profiling and data quality scans, data products (GA May 25, 2026), automated cataloging of the Lakehouse Iceberg REST catalog (GA April 16, 2026), and dbt metadata import (GA September 24, 2026), with policy tags and row and column security enforced across BigQuery access paths. Classification and sensitivity reach the catalog for structured assets: Knowledge Catalog aspects carry them, and Sensitive Data Protection publishes profile insights as aspects for BigQuery tables, Cloud SQL tables, and Vertex AI datasets built from BigQuery. They stop at Cloud Storage. The Sensitive Data Protection integration is unavailable for Cloud Storage data because Knowledge Catalog doesn't ingest Cloud Storage buckets that way, so object classification there is what the enterprise writes into object contexts or reaches through BigQuery object and external tables. That is owned storage plus a catalog a reasoning plane can query for lineage, quality, and the classification of structured data, which is rule 9's strong, calibrated to Azure (Purview), Snowflake (Horizon), and Databricks (Unity Catalog). The Gemini half is mostly preview. Unstructured data insights (Gemini entity and relationship extraction, preview April 16, 2026), unstructured data profiling (preview June 11, 2026), data relationships discovery (preview April 20, 2026), the lookupContext method that hands agents an LLM-ready context bundle (preview June 4, 2026), metadata connectors that import from Salesforce, SAP, Workday, ServiceNow, and Palantir, and Smart Storage's automated annotations are named here and not scored. Gemini-assisted documentation and insights for data products ship GA with data products, which doesn't move the grade. Shipping today, the catalog governs and describes the data; Gemini building the context graph agents ground on is the roadmap. When it ships, that design would have Gemini decide what the enterprise's data means before any application-level inference. Google is one of the clearest hyperscaler implementations of the catalog-to-agent-grounding pattern; its differentiation is ambition and tight BigQuery integration, not uniqueness, since Purview, Unity Catalog, and Horizon carry governed metadata, classification, and lineage stories of their own. Governance reach: the shipping pattern imports metadata into Knowledge Catalog (connectors, many in preview), with partner tools able to view Knowledge Catalog metadata. It is not yet a neutral federated control plane: an enterprise running PowerScale, S3, and Cloud Storage can't use it across all three without making Google Cloud the metadata authority.

◆ Borrowed Judgment

Two halves. Policy execution is the enterprise's: policy tags, row and column security, data quality rules, glossaries, and custom aspects are written by the enterprise and enforced as written. Discovery, profiling, and classification are Google's judgment, and the enterprise can override the result per object: aspects can be created, updated, and deleted, Cloud Storage object contexts can be set, replaced, or cleared, and Sensitive Data Protection's automatic sensitivity tags overwrite manual values only if the enterprise opts in. The grade-carrying half is the catalog's description of the data, so the layer reads vendor-decides, visible, overridable: Delegated, the Purview, Horizon, and Unity Catalog classification reading. The borrowed judgment grows when the Gemini half ships. A context graph built by Gemini would become every grounded agent's representation of what the data means, and a single ingest-time annotation from Smart Storage would persist as a governance fact. That is the cell to re-read at GA. Comparison to AWS: Lake Formation enforces customer-defined policies and reads Delegated on that basis. Google reads the same today; the difference is the roadmap, not the shipping product.

◆ Working Notes

Named, not scored (preview; the dates are preview starts, not GA dates, so none is a watch-list item): unstructured data insights (April 16, 2026), data relationships discovery (April 20), the Data Lineage MCP server (May 27), lookupContext (June 4), unstructured data profiling (June 11), Oracle and MySQL connectors (June 22), lineage ingestion control (June 23), SQL Server and PostgreSQL connectors (July 9), governance workflows (July 24), dbt Core and MetricFlow import (August 26), data domains (September 7), BigQuery graph metadata (September 14), Databricks, AWS, and Snowflake Iceberg REST catalog support (April 16), third-party metadata connectors, BigQuery writes to Lakehouse Iceberg tables, and Smart Storage automated annotations. The Layer 1A / 2C boundary question: a model-built context graph that decides what context agents receive is architecturally 1A and functionally 2C. When lookupContext and unstructured insights reach GA, read whether the catalog ranks and routes context (2C) or describes data (1A).

Layer 1B · RetrievalContext Management & Retrieval1 LABModel Prep & Managed Retrievaldecides: vendor · Delegated▼

Low-latency retrieval for RAG — vector/hybrid search, context windows

Vendor-Provided

BigQuery MLCeded

In-database ML training and inference. Supports linear/logistic regression, K-means, time series, XGBoost, DNNs, imported TF/PyTorch models. Collapses the boundary between data preparation (1B) and AI runtime (2B) by running ML directly in the warehouse.

Dataflow + Managed Service for Apache Spark (Formerly Dataproc)Delegated

Managed Apache Beam (batch and stream) and managed Spark/Hadoop for large-scale processing. Open-source frameworks with Google-managed execution: a managed service behind a standard interface, so Beam pipelines and Spark jobs lift to another runner or platform without rebuilding. Calibrates to AWS Glue ETL, Databricks Apache Spark, and IBM’s ML Pipeline Stack, all Delegated on the same substrate.

Data Agent Kit (Open-Source)Delegated

MCP-based agents packaged as tools and skills. Supports Claude Code, Gemini CLI, Codex, VS Code. Enables intent-driven development: practitioners define goals, agents handle implementation. Creates governance recursion: the agent builds the pipeline that prepares the data that feeds the agent.

Feature Store on Gemini Enterprise Agent Platform (Formerly Vertex AI Feature Store)Ceded

Managed feature serving for online/offline ML models and agents. Consistent feature serving across training and inference. 1B→2B bridge.

Agent Search (Formerly Vertex AI Search; Discovery Engine API)Ceded

Turnkey managed RAG and enterprise search via the Discovery Engine API (formerly Vertex AI Agent Builder, now part of Gemini Enterprise Agent Platform). Create a data store over Cloud Storage, BigQuery, or websites and Google handles parsing, chunking, embedding, indexing, ranking, and grounded retrieval end to end. The direct analog to Bedrock Knowledge Bases, but on a proprietary managed datastore with no swappable open backend, so more captive than Bedrock KB: the retrieval opinions have no open exit.

BigQuery Vector SearchCeded

VECTOR_SEARCH over an IVF vector index, GA, with hybrid keyword-plus-vector retrieval in-warehouse. (The ScaNN index type is still in preview.) A self-orchestrate primitive: the enterprise builds the retrieval pipeline, but the store is BigQuery, a single-vendor API with no open exit.

Firestore Vector SearchCeded

K-nearest-neighbor vector search on transactional Firestore data, GA, with composite pre-filters and LangChain and LlamaIndex integrations. A self-orchestrate primitive on a single-vendor Firestore API: convenient for operational RAG, captive at the store layer.

NVIDIA-Provided

Assessment pending

◆ Gap Analysis

No meaningful capability gap. Most mature Layer 1B in the assessment series. BigQuery ML eliminates data-to-model handoff. LookML Agent automates semantic model construction. Data Agent Kit enables agent-driven pipeline development. The gap is governance over model-generated data artifacts. When LookML Agent generates a semantic model, when Data Agent Kit writes a pipeline, when BigQuery ML trains a model — who reviews the output for correctness? These are AI-generated artifact governance problems. Google provides no productized capability for governing model-generated data artifacts at scale. Data Agent Kit’s explicit support for Claude Code and non-Google tooling is strategically significant — the one point in the stack where third-party model access is genuinely equal. Retrieval is the surface this layer is named for, and Google covers it comprehensively but captively. Agent Search (formerly Vertex AI Search; the Discovery Engine API) is the turnkey managed RAG product, the direct analog to Bedrock Knowledge Bases: ingest, chunk, embed, index, and grounded retrieval handled end to end. BigQuery VECTOR_SEARCH and Firestore Vector Search are the self-orchestrate primitives. The governance fork mirrors AWS: consume Agent Search and Google owns the retrieval decisions (Ceded), or self-orchestrate on BigQuery or Firestore vector and Retain the retrieval-strategy opinions in code. The twist against AWS: every native Google vector store is single-vendor (BigQuery, Firestore, and Vector Search on Gemini Enterprise Agent Platform are all Ceded), so self-orchestration Retains the pipeline logic but still lands on a captive store unless the enterprise deliberately chooses AlloyDB or Cloud SQL pgvector, the one portable escape. Even the turnkey surface is more captive than AWS: Bedrock Knowledge Bases is Delegated because it orchestrates over swappable open backends, while Agent Search runs a proprietary managed datastore with no open exit.

◆ Borrowed Judgment

The semantic model as borrowed judgment: when LookML Agent generates definitions, every analytics query and agent interaction using those definitions inherits Gemini’s interpretation of business logic. Powerful (automates weeks of manual semantic modeling) and risky (embeds model judgment in the analytical foundation). The pipeline-building agent as borrowed judgment: Data Agent Kit agents write Dataflow jobs and BigQuery transformations. The enterprise inherits the agent’s data engineering judgment — join strategies, filter logic, null handling. Previously human expertise, now model-generated. The retrieval surfaces deepen the capture: Agent Search, BigQuery vector, and Firestore vector are all single-vendor APIs, so the enterprise inherits Google opinions on chunking, ranking, and grounding with no portable exit short of pgvector. Decision authority (October 5, 2026): Agent Search (formerly Vertex AI Search) parses, chunks, and ranks, but each request can set boostSpec, filter, and a relevance filter that overrides the global threshold. Vendor decides, visible, overridable, Delegated. Self-orchestrated BigQuery or Firestore vector would read code and Retain the retrieval logic.

◆ Working Notes

The comparison to AWS SageMaker Unified Studio: AWS provides a single governed environment across services (service-wide integration). Google achieves integration through BigQuery spanning storage, analytics, ML, and governance (service-deep integration). Google = tighter integration at cost of BigQuery lock-in. AWS = service diversity at cost of integration complexity. Watch-list (preview, not scored): LookML Agent (Gemini-powered semantic-model generation) - moved off the scored components per the GA-gate until it reaches GA; it remains a notable forward signal for model-generated semantic models.

Layer 1C · PipelinesData Movement & PipelinesBigQuery-Native Pipelinesdecides: code · Retained▼

Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering

Vendor-Provided

BigQuery-Native Pipelines (Data Transfer Service, Dataform SQL Workflows, BigQuery Pipelines and Data Preparation, Continuous Queries)Ceded

Scheduled ingest from S3, Azure Blob, Redshift, Teradata, Snowflake, Salesforce, and Oracle through the Data Transfer Service; SQL transform workflows with dependencies, schedules, and assertions in Dataform; BigQuery pipelines with tables, views, data sources, and data quality tests as tasks (GA July 30, 2026); Gemini-assisted data preparation over Cloud Storage and Drive files (GA March 23, 2026); continuous queries streaming results to Pub/Sub, Bigtable, and Spanner (stateful operations are Preview and not scored). The workflows are written against BigQuery and don't lift to another warehouse without a rebuild: Ceded.

AI Data Prep in BigQuery (AI.GENERATE_EMBEDDING; Autonomous Embedding Generation)Ceded

Embeddings generated in SQL against Google or open embedding models (GA March 6, 2026), and autonomous embedding generation that keeps a declared embedding column current as rows change (GA June 17, 2026). Ceded on the embeddings carve-out: the vectors are useless without the same model at query time.

Dataflow (Apache Beam Batch and Streaming; Dataflow ML With RunInference and MLTransform)Delegated

Managed runner for the enterprise's Beam pipelines, with inference, embedding, and enrichment steps inside the same pipeline. Beam code runs on other runners (Flink, Spark), so the provider swaps without a rebuild: Delegated, the AWS Glue Spark substrate reading.

Datastream (Change Data Capture to BigQuery, Cloud Storage, and Iceberg)Ceded

Change data capture from operational databases with per-object and partial backfill. Datastream detects new source columns and adds them to the destination on its own. Sources marked Preview (Salesforce Marketing Cloud, Dataverse, ServiceNow) aren't scored.

Managed Service for Apache Airflow (Formerly Cloud Composer 3)Delegated

Managed Airflow running the enterprise's DAGs (GA December 16, 2024; Airflow 3 GA April 15, 2026). The DAGs lift to any Airflow: Delegated, the MWAA reading.

Pub/Sub (BigQuery and Cloud Storage Subscriptions; AI Inference Single Message Transform)Ceded

Streaming ingest with subscriptions that write straight to BigQuery and Cloud Storage, and per-message model enrichment through the AI Inference transform (GA April 6, 2026). Import topics are named, not scored, pending per-source badges.

Knowledge Catalog Lineage and Data Quality on Pipelines (Formerly Dataplex Universal Catalog)Ceded

Lineage recorded automatically from BigQuery, Dataflow, Airflow, Spark, and pipeline jobs (GA; Looker lineage is Preview), and scheduled data quality scans with alerts on pipeline output. The day-two surface; the catalog itself is scored at 1A.

NVIDIA-Provided

Assessment pending

◆ Gap Analysis

BigQuery carries the strong on its own. It ingests on a schedule, transforms with dependencies and assertions, runs real-time SQL, and prepares data for AI inside the engine, with embeddings kept current as rows change. That's the Snowflake and Databricks shape, where the warehouse ships the pipeline. Dataflow, Datastream, managed Airflow, and Pub/Sub add depth around it, and lineage and data quality on the pipelines pass the day-two test. The grade doesn't need them. The cell was rebuilt October 5, 2026. It used to rest on five chips that couldn't carry it: two storage tiers, a catalog, a Preview federation feature, and an enrichment feature with no product documentation. The products that actually move data on Google Cloud appeared only in prose. Capability didn't change. The evidence did. Decision authority: BigQuery executes what the enterprise configured. Data preparation's Gemini suggestions happen at design time, and at runtime the accepted steps run as written; autonomous embedding maintains a column the enterprise declared. Code decides, visible, overridable, Retained, the judgment test and its defaults clause (October 4 and 5, 2026). Datastream is the exception: it reads the source and adds columns on its own. It adds depth but doesn't carry the grade.

◆ Borrowed Judgment

Low at runtime, higher at design time. The pipelines run what the enterprise wrote, but the enterprise writes them in BigQuery's dialect, and Gemini's preparation suggestions shape what gets written. The judgment Google keeps is in Datastream's schema handling and in the embedding models the AI prep calls. The same embed-and-index fork as 1B applies to the RAG ingest pipeline: consume Agent Search (formerly Vertex AI Search) and the ingest, chunk, and embed movement is managed for you (Ceded), or self-orchestrate it in Dataflow or BigQuery and Retain the pipeline logic, though the chosen store stays captive.

◆ Working Notes

Named, not scored (October 5, 2026): Cross-Cloud Lakehouse is Preview under Pre-GA terms, with catalog federation and intelligent caching also Preview; Smart Storage auto-annotate at ingest appears in a Google blog with no product documentation. RAG Engine ingestion is GA only in europe-west3 and europe-west4 (allowlist in the US regions, Preview elsewhere) and isn't cited as a carrier. Re-homed under the M2 storage ruling: Managed Lustre to Layer 0, BigLake (renamed Lakehouse, April 2026) to 1A; Rapid Tier and Smart Storage were already scored at 1A. The asymmetry between data plane federation and control plane federation (absent at 2C) is a structural finding. Google invests in making data accessible across clouds but not in making agent governance portable across clouds. Data accessibility without orchestration portability draws workloads toward GCP as the governance center.

Layer 2A · OrchestrationInfrastructure OrchestrationGKE + Sovereign Extensionsdecides: vendor · Ceded▼

GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization

Vendor-Provided

GKE + GKE Agent SandboxDelegated

Managed K8s with AI-era extensions. Agent Sandbox: gVisor-based secure isolation, 300 sandboxes/second/cluster with sub-second time to first instruction. Infrastructure built for the agentic era, not retrofitted. GKE is a managed service behind the standard Kubernetes API — manifests lift to another conformant cluster, so Delegated (operation delegated to Google; the standard interface keeps the opinions portable).

Fluid ComputeCeded

GCE + GKE dynamically shifting workloads in real-time. CPUs for branchy agent logic, secure sandboxes, RL, SLM inference, RAG. GPU/TPU for training and large-model inference. Proto-Layer 2C: routes based on workload characteristics, not business context.

GDC (Sovereignty Analysis)Delegated

On-prem Google Cloud services: GKE, Agent Platform, managed storage, Gemini Flash, Blackwell GPUs. Air-gapped for sensitive workloads. Addresses data sovereignty (where computation happens) but NOT judgment sovereignty (whose model drives computation). GDC runs Gemini on-prem — same recursive dependency, inside the enterprise perimeter.

Capacity ManagementCeded

CUDs (1yr/3yr), on-demand, preemptible/spot, Dynamic Workload Scheduler. Same pattern as AWS — capacity acquisition (2A), not workload placement (2C).

NVIDIA-Provided

NVIDIA GPU Operator (on GKE)

Available for NVIDIA instances on GKE. Google manages the GPU integration layer.

◆ Gap Analysis

GKE is the most mature managed K8s for AI workloads. GKE Agent Sandbox has no equivalent in Dell, HPE, or AWS portfolios — 300 sandboxes/second is built for agentic workload density. Fluid Compute sits at the 2A/2C boundary. Its dynamic workload shifting is more than capacity acquisition — runtime decisions about which compute type serves which workload. But less than full 2C — routes on workload characteristics, not business context (data residency, compliance tags, cost targets). The Fluid Compute → Knowledge Catalog connection does not exist: workload placement does not consult governance metadata. Same Infrastructure Layer 2C gap every vendor has. GDC: the full sovereignty analysis reveals that data sovereignty ≠ judgment sovereignty. Knowledge Catalog on GDC uses Gemini. Smart Storage on GDC uses Gemini. Agent Platform on GDC uses Google’s runtime. The enterprise gains physical sovereignty but retains the same judgment concentration. The ‘self-driving cloud’ narrative implies Gemini-powered infrastructure operations — autonomous root-cause analysis on infrastructure telemetry. If the Reasoning Plane is itself Gemini-powered, the operational intelligence and application intelligence are the same intelligence. Capability and deployment are separate readings. The grade states what an architect can deploy on this vendor's paper; it does not assert that a given customer consumes it. Customers routinely run self-managed Kubernetes, Slurm, or a third-party distribution on Compute Engine rather than adopting GKE, in which case Google is bought as substrate and the 2A judgment belongs to whoever runs the control plane. GKE's maturity is a statement about the option on offer, not about occupancy of the layer.

◆ Borrowed Judgment

GKE consumed via the standard Kubernetes interface: the enterprise's manifests and operators lift to any Kubernetes (Delegated — managed K8s service; the enterprise could switch without rebuilding). Google's optional Gemini-driven scheduling optimization is the only borrowed-judgment layer, and only if the enterprise opts in. GDC as borrowed judgment in sovereign packaging: physical control over facility, Google’s judgment in software, model, governance, and operations. Data sovereignty with judgment concentration. Fluid Compute as proto-2C borrowed judgment: when Fluid Compute routes agent work to CPU vs GPU, that routing is Google’s judgment about optimal compute matching. Enterprise doesn’t configure the routing policy.

◆ Working Notes

Data sovereignty vs judgment sovereignty: the 4+1 framework should distinguish between where data resides (which GDC addresses) and whose model’s reasoning shapes the AI system (which GDC does not address). An enterprise running GDC air-gapped has data sovereignty while fully ceding judgment sovereignty.

Layer 2B · RuntimeApplication Runtime & Execution2 LABSModel-Integrated Stackdecides: model · Delegated▼

Model serving, agent execution, inference APIs, distributed inference

Vendor-Provided

Gemini Enterprise Agent PlatformCeded

Unified platform for building, scaling, governing, optimizing agents. Subsumes all Vertex AI services. Agent Studio (low-code), ADK (code-first, Python/Go/Java/TypeScript, model-agnostic, open-source), Model Garden (200+ models incl Gemini, Claude, Llama, Gemma), Agent Runtime, Agent-to-Agent Orchestration, Agent Identity (GA), Agent Gateway, Agent Observability, Agent Registry, Memory Bank, Antigravity (desktop app + CLI).

Model Serving via the OpenAI-Compatible Vertex EndpointDelegated

Vertex documents an OpenAI-compatible endpoint for Gemini, and serves third-party models (Claude through Anthropic’s Messages API, Llama) alongside it. Consumption through those interfaces is a managed service behind a multi-vendor standard, so the integration opinions lift to any compatible endpoint and the serving interface is Delegated. The model itself remains Google’s; what is portable is the interface, not Gemini.

Customer-Created and Managed Tools (Function Calling / Client-Side)Retained

GA through the OpenAI-compatible Vertex endpoint and Gemini function calling. The enterprise owns and operates the business logic, APIs, commands and deterministic validators invoked through tool calls: the seam where control passes from instructing the model to executing code outside it. The opinions are the enterprise's and run against a multi-vendor interface, so they lift to another provider with an adapter and re-evaluation, not a rebuild. What the enterprise cedes at this seam is the decision to act: the model chooses whether, when and with which arguments to call, and that judgment is inherited under the serving component above. Added instrument-wide on September 2, 2026 (M4): the seam exists on every runtime that hands control to portable customer code, and a real Retained position is scored, not narrated.

The House Model AdvantageCeded

Gemini on TPU: silicon designed for this model, networking designed for its topology, distributed runtime (Pathways) built for its coordination, inference optimized for its architecture, governance (Knowledge Catalog) powered by it, orchestration defaults to it. No other vendor achieves this degree of vertical optimization. Third-party models (Claude, Llama) run on NVIDIA GPUs — supported but not co-optimized. The playing field is structurally tilted. The standard-interface consumption path is scored separately; what remains here is the platform surface beyond it, where the vertical co-optimization and the governance defaults accumulate.

Pathways (Distributed Runtime)Ceded

Google’s distributed runtime for superpod-scale training and inference coordination. Proprietary to Google Cloud with no implementation elsewhere, so coordination and scaling opinions written against it have nowhere to go. This is the part of the framework stack that actually captures.

Open Frameworks on TPU (JAX, TorchTPU, vLLM, llm-d)Delegated

JAX is Apache 2.0 and runs on GPU and CPU through XLA. TorchTPU brings full PyTorch to TPU, and PyTorch code is hardware-agnostic. vLLM is optimized across GPU and TPU, and llm-d is an open-source Kubernetes-native inference serving project with multi-vendor participation. Each is open source with real alternatives, and the consumed interface is the framework API, which independent vendors implement. Model and serving code written against them lifts to other hardware without rebuilding. TPU optimization is a performance fact, not an authority fact.

NVIDIA-Provided

NVIDIA GPU Instances

A3/A3 Mega/A3 Ultra/A5X for third-party model inference. CUDA ecosystem required for non-Gemini models.

◆ Gap Analysis

Layer 2B is the center of gravity for the model-integrated stack. The model provider, runtime provider, and infrastructure provider are the same company. When the enterprise runs Gemini on Agent Platform on TPU, it borrows Google’s judgment at the model layer, runtime layer, framework layer, and silicon layer simultaneously. A single entity’s priorities shape the entire execution path. The Agent Platform collapses Layers 2B, 2C, and 3 into a single product surface: Agent Runtime (2B infrastructure), Agent Identity/Gateway/Registry/Orchestration/Observability (2C governance), Agent Studio/ADK/Antigravity (Layer 3 development). The product boundary does not align with the architectural boundary. AWS separates these: Bedrock (model access) is distinct from AgentCore Runtime (agent execution) is distinct from AgentCore Policy (governance). AWS’s separation preserves architectural boundaries the enterprise can independently govern. Google’s collapse optimizes integration but prevents swapping the governance layer (2C) while keeping the runtime (2B). The NVIDIA dependency at 2B is optional for Gemini (TPU-native) but required for third-party models. The enterprise using Claude on Google Cloud pays a structural performance tax — Claude runs on NVIDIA GPUs through a runtime designed for Gemini. Model Garden’s 200+ models are API-equal but not silicon-equal.

◆ Borrowed Judgment

The model-integrated runtime: Gemini on TPU inherits Google’s judgment at model layer (training data, alignment, safety), runtime layer (scheduling, scaling, session management), framework layer (JAX/Pathways optimization), and silicon layer (TPU architecture). Most concentrated borrowed judgment in the assessment series. The open-source hedge: llm-d, TorchTPU, vLLM, ADK provide genuine open alternatives. Google opens components that reduce adoption friction (frameworks, SDKs) while keeping authority-concentrating components closed (Agent Runtime infrastructure, Agent Gateway, Pathways). Production deployment pulls open tools into Google’s managed surface where authority shifts from Retained to Ceded.

◆ Working Notes

The 2B/2C collapse prevents the enterprise from independently governing the orchestration layer. An enterprise that wants Google’s Agent Registry and Agent Identity but AWS’s Bedrock for model access and its own governance engine for policy enforcement cannot compose that architecture. The components are bundled. Lab-measured (July 2026, reconcile v1.6): the gemma4-tpu-inference lab measured the bring-your-own serving path on the self-serve TPU lane, using a model that is itself TPU-native. The constraint it hit is common to mixture-of-experts models sharded across accelerators on any substrate: two global key-value heads cap pure tensor parallelism at 2, and at that width bf16 does not fit v5e's per-chip memory budget. The single-device contrast in the lab program, 128GB of unified memory on the GB10 serving the same model whole, is a memory-topology difference, not a runtime-openness difference. The lane ruling stands on its own terms: on reachable self-serve silicon the deployable paths were quantizing off-box or adopting Google's serving stack, and that adoption cost is this cell's Ceded reading in practice. Whether Google's own managed stack serves this model well was not the question and was not measured.

Layer 2C · ReasoningAgentic Infrastructure — The Reasoning Plane2 LABSComplete Intelligence-2C Plane, No Placementdecides: model · Ceded▼

Policy-driven placement and resource coordination — the Autonomy Layer

Vendor-Provided

Agent Identity (GA)Ceded

Agents as identity principals with authentication, authorization, audit. Control plane function: determines which agents exist as governed entities.

Agent GatewayCeded

Protocol-level governance for MCP and A2A communications. Security partner integrations (Broadcom, Check Point, Cisco, CrowdStrike, F5, Netskope, Okta, Palo Alto, Zscaler). Spans Layer 0 (networking), 2B (runtime), and 2C (orchestration).

Agent-to-Agent OrchestrationCeded

Deterministic multi-agent workflow routing. Control plane function: determines which agent handles which subtask.

Agent RegistryCeded

Catalog of agents with ownership, capabilities, protocols, invocation details. Administrator-controlled discoverability. Control plane function: determines which agents are available and who can use them.

Agent ObservabilityCeded

Monitoring, tracing, debugging across production agent populations. Feedback loop for detecting faulty reasoning and intervening.

NVIDIA-Provided

No NVIDIA Layer 2C Dependency

All Layer 2C components are Google IP. NVIDIA does not control governance, policy, or reasoning in Google’s stack.

◆ Gap Analysis

Google’s Intelligence Layer 2C is the most complete productized offering in the assessment series: Agent Identity + Gateway + Registry + Orchestration + Observability + Memory Bank. Together they constitute a genuine control plane for agent governance. Infrastructure Layer 2C — the autonomous placement engine — is NOT built as a customer-configurable product. The capacity primitives (Fluid Compute, CUDs, DWS) are building blocks, but they don’t compose into a policy-driven placement engine querying Knowledge Catalog governance metadata. Same gap as AWS and every other vendor. Google’s implicit Layer 2C is the most sophisticated in the assessment: managed services make autonomous placement, scaling, routing, and capacity decisions invisibly. The enterprise cannot see, configure, audit, or override these decisions. The model-integrated Reasoning Plane: if Google’s ‘self-driving cloud’ uses Gemini for infrastructure decisions, then the model powering the enterprise’s agents (Layer 3) is the same model governing agent orchestration (Intelligence 2C) is the same model deciding where agents run (Infrastructure 2C). One model’s judgment pervades every decision surface. Cross-cloud orchestration gap: Google federates the data plane (Cross-Cloud Lakehouse) but NOT the control plane. Agent Platform governs GCP agents only. Enterprise running agents across multiple clouds has no cross-platform agent governance surface — unless all agents route through Google’s Agent Gateway, which cedes cross-cloud governance to Google. The captive-but-best dilemma: this is evidence the control plane CAN be built as a coherent capability. The enterprise architect who wants it has one option: adopt Google Cloud. The federated alternative does not exist. Calibration (September 2, 2026): strong on the completeness criterion, not on placement. All five legs of an Intelligence-2C plane are productized and GA here: first-class agent identity, a request-time gateway, a registry, cross-agent orchestration, and observability. Salesforce lacks the identity leg, VMware the registry and orchestration legs, Nutanix everything but the gateway, and each reads moderate for that stated reason. No reasoning mechanism ships; placement stays the universal gap.

◆ Borrowed Judgment

The captive control plane: enterprise inherits Google’s orchestration model — deterministic routing, Google-managed identity, Google-governed protocols. Well-engineered but unchallengeable — cannot substitute alternative orchestration logic within the Agent Platform boundary. The model-powered control plane: if the Reasoning Plane uses Gemini for infrastructure decisions, a model judgment error at the control plane layer is invisible to the enterprise, with no fallback to human decision-making. Intelligence 2C: Low borrowed judgment in the sense that the components are productized and configurable. High borrowed judgment in the sense that the governance logic itself (Agent Gateway protocol decisions, Agent Identity authentication model, Orchestration routing patterns) is Google’s, not the enterprise’s.

◆ Working Notes

Google’s 2C proves the Control Plane Working Notes thesis: the control plane can be built. The question is whether it can be liberated from the vendor boundary — and whether the model-integrated dimension (control plane powered by the same model it governs) is a pattern to replicate or to avoid. The asymmetry: data plane federates (Cross-Cloud Lakehouse), control plane does not. This serves Google’s strategic interest — data accessibility without orchestration portability draws workloads toward GCP as governance center. Instrument note: live Infrastructure-2C — a policy engine deciding at request time where inference runs relative to the data, which model serves the request, and how cost, latency, and compliance are arbitrated — does not ship from any vendor on this instrument. It is recorded here as a universal finding about the market, not as a charge against this vendor. The sibling finding applies here as well: you cannot prompt your way to deterministic output. The controls credited in this cell answer whether an action is legal — permissions, schema and policy checks, rate and token limits, approval gates — and none of them validates whether the outcome was right under a graduated escalation policy. Controls that work by instructing the model shift the output distribution without pinning it. That deterministic outcome-validator is missing instrument-wide, and like live placement it is noted as universal rather than charged to one vendor.

Layer 3 (+1) · ApplicationsAI Application Layer — The Value PlaneOpen Model Layer, Captive Platformdecides: vendor · Ceded▼

AI-powered business capabilities — business logic, workflow automation

Vendor-Provided

Gemini Model FamilyCeded

Gemini 3.1 Pro, Gemini 3.5, Gemini Flash. The model the entire stack was designed around. Gemma open-weight models for self-hosting (the one offering where enterprise can Retain model authority).

Model Garden (200+ Models)Delegated

Gemini, Claude Opus/Sonnet/Haiku, Meta Llama, Gemma, open-source models. Model Evaluation service. Broadest model catalog of any cloud provider. Model-agnostic claim genuine at Layer 3 — more so than any other layer.

Application SurfacesRetained

ADK (code-first, open-source, model-agnostic) — agents built on the open SDK lift out; the substrate is self-hostable.

Application Surfaces (Managed Studios)Ceded

Agent Studio (low-code) and Agent Designer (no-code in the Gemini Enterprise app), plus Gemini Enterprise agent discovery and the Deep Research agent. Managed authoring surfaces — low/no-code agent definitions are captive to the platform; scored separately from the open ADK.

Consumer-Enterprise Feedback LoopCeded

Gemini powers Google Search, Gmail, Docs, Photos, Android, Chrome. Workspace Intelligence uses Gemini for agentic work. Model improvements from billions of consumer interactions directly benefit enterprise workloads. But: consumer-driven alignment and safety tuning may not align with enterprise needs.

Google Antigravity 2.0 (Agent-First Development Platform)Ceded

Announced I/O 2026. Standalone desktop app + CLI + SDK — a full developer platform built around agent orchestration. Multi-agent parallel execution: orchestrate multiple agents and execute tasks simultaneously. Dynamic subagent workflows and scheduled background automation. Antigravity CLI (Go-based, replacing Gemini CLI — deprecated June 18, 2026) for terminal-native multi-agent workflows. Antigravity SDK for building custom agents with templates in AI Studio. Powered by Gemini 3.5 Flash (co-developed using Antigravity). Native voice command support. Ecosystem integrations: Google AI Studio, Android, Firebase. Export tool for AI Studio → local development. Search integration: real-time custom UI generation within Google Search answers. AI Ultra plan ($100/month, 5x usage limits). Google's most aggressive move in the agentic coding market — positioned as the hub for multi-agent development workflow orchestration, not just code assistance.

NVIDIA-Provided

NVIDIA Models via Model Garden

NVIDIA Nemotron and other NVIDIA models available alongside all other providers.

◆ Gap Analysis

No meaningful capability gap. Broadest model catalog. Most portable agent development framework (ADK). Application surfaces from no-code through code-first. Google deliberately keeps Layer 3 more open than any other layer — while ensuring every Layer 3 application is gravitationally pulled toward Agent Platform (2B/2C). By keeping Layer 3 open, Google maximizes platform adoption: enterprises wanting Claude on Google Cloud still consume Agent Platform’s runtime, identity, gateway, registry, observability. The model is portable; the platform is captive. Consistent with the 4+1 model’s prediction that vendor lock-in concentrates at Layer 2B/2C, not Layer 3. Google has understood this prediction and built strategy accordingly. Code portability vs operational portability: ADK is open-source and model-agnostic — agent code CAN run on AWS or on-prem K8s. But Agent Registry, Memory Bank, Agent Identity, Agent Gateway, Agent Observability are Google Cloud services with no portable equivalents. Agent code is an asset the enterprise owns. Agent operations are an asset it rents. Antigravity 2.0 deepens the Layer 3 gravitational pull toward Google's platform. The desktop app + CLI + SDK creates a development surface that integrates directly with Agent Platform (2B/2C): agents built in Antigravity inherit Agent Platform's identity, gateway, registry, and observability. The Gemini CLI deprecation (June 18, 2026) forces migration to Antigravity CLI — consolidating Google's developer AI surface into one opinionated platform. The SDK enabling custom agent templates in AI Studio means Antigravity is not just a coding tool but an agent construction platform that feeds directly into the Gemini Enterprise Agent Platform. Compare to AWS Kiro (spec-driven, methodology-opinionated, Bedrock-native) and GitHub Copilot (IDE-embedded, multi-model, GitHub-native). Google's differentiator is multi-agent parallel orchestration — Antigravity coordinates multiple agents simultaneously rather than single-agent sequential interaction. This maps to the 4+1 model's Layer 2C vision: orchestrating multiple agents is a control plane function that Antigravity surfaces through a developer tool. The consumer-enterprise feedback loop extends to Antigravity: Google is using Antigravity's capabilities in consumer Search to generate real-time custom UIs as part of search answers. Developer tool innovations flow to consumer products and back — a flywheel no other vendor in the assessment possesses.

◆ Borrowed Judgment

Gemini as borrowed Layer 3 judgment: alignment changes affect agents (Layer 3), governance enrichment (Layer 1A via Knowledge Catalog), semantic models (Layer 1B via LookML Agent), and potentially infrastructure operations (Layer 2C via self-driving cloud). A single alignment decision propagates across the entire model-integrated stack. Platform defaults: Agent Studio and Agent Designer default to Gemini. Enterprise that adopts without explicitly selecting alternatives inherits Google’s model preference as a default rather than a decision. Strategic openness as borrowed judgment about lock-in location: Google’s decision to keep Layer 3 open and concentrate lock-in at 2B/2C is itself borrowed judgment the enterprise inherits. Evaluating Google on model diversity without evaluating platform captivity accepts Google’s framing of where portability matters.

◆ Working Notes

The consumer-enterprise feedback loop has no parallel in the assessment. Model improvements from billions of consumer interactions benefit enterprise workloads — but consumer-driven alignment may constrain enterprise use cases. If Google tightens content policies for consumer safety, enterprise agents inherit that tightening. The Gemini CLI → Antigravity CLI forced migration is a significant authority move. Over 100,000 GitHub stars on Gemini CLI — all those developers must migrate to Antigravity by June 18, 2026. This concentrates Google's developer AI surface into one platform and one billing model (AI Ultra at $100/month). The deprecation timeline is aggressive but consistent with Google's pattern of consolidating developer tools around Gemini. Antigravity 2.0's scheduled tasks capability (agents running automatically in the background) converts the developer tool from a single-turn interaction to a persistent automation pipeline. This blurs the boundary between Layer 3 (application) and Layer 2C (orchestration) — when Antigravity schedules background agents to perform tasks autonomously, who governs those agents? The answer is Agent Platform — reinforcing the Layer 2C gravitational pull.

Sources & revision history · v1.19 - 4+1 v2: Layer 1A Re-read on the GA Gate

Google Cloud Next 2026 (Apr 22–24), GTC 2026, NVIDIA partnership, Forrester, SiliconANGLE, The New Stack, analyst coverage. v1.2 (instrument reconciliation): 2A GKE Retained→Delegated — managed K8s behind a standard interface is Delegated, not Retained. Cloud Storage remains Ceded (GCS-native API). v1.3 (lab-validated against the TFD corpus / vCTOA RAG build): added the retrieval surface to Layer 1B - Vertex AI Search (Discovery Engine, the Bedrock Knowledge Bases analog), BigQuery Vector Search, and Firestore Vector Search, all Ceded (single-vendor APIs); demoted the preview LookML Agent from a scored component to a notes watch-line per the GA-gate; and added the self-orchestrate-versus-managed Retain/Delegate fork to 1B/1C, mirroring the AWS row. Inference-interface reconciliation (July 23, 2026): the completions-interface ruling (OpenAI v1.0) and its Messages-API extension (Anthropic v1.0) applied — model access consumed through a genuine multi-vendor standard inference interface is Delegated; proprietary surfaces beyond the interface remain Ceded. Note added at 2B (Vertex OpenAI-compatible endpoint); no chips moved — the scored components are platform-altitude. /reconcile (September 1, 2026): inference-interface rule applied as a scored facet. Model access consumed through a multi-vendor standard interface (OpenAI-compatible completions, Anthropic Messages) split out as Delegated; the platform surface beyond the interface stays Ceded. Matches the Azure v1.4 treatment. /reconcile (September 1, 2026): the Frameworks component bundled open-source JAX, TorchTPU, vLLM, and llm-d with Google-proprietary Pathways under a single Ceded call. Split: Pathways stays Ceded, the open frameworks score Delegated on the open-substrate rule. Dataflow/Dataproc moved Ceded to Delegated: managed OSS behind a standard interface, matching aws 1C Glue ETL, databricks 1C Spark, and ibm 1C ML Pipeline Stack. /reconcile (September 1, 2026): shared 2C findings stated explicitly (live-placement universality, outcome-validation finding). /reconcile (September 1, 2026): rule 7 (availability, not occupancy) applied at 2A — the grade states deployable capability on this vendor's paper and no longer reads as though the vendor's layer is automatically the customer's. /reconcile (October 5, 2026): Layer 0 runtime rule (ruled September 28 and October 4, 2026). TPUs split by adopted runtime (JAX and PyTorch/XLA Retained, TPU-specific code Ceded); the ownership reasoning retired. NVIDIA GPUs on Google Cloud split: standard instances Ceded to Retained from the customer's seat, NVL72 capacity Ceded on the integrated-system rule. Grade unchanged. /reconcile (October 5, 2026, same pass): TPUs re-read on the generally available generations (TPU7x, v6e, v5p); TPU 8t and 8i descored to named-not-scored under the GA-gate. Single-host slices split by runtime; pod-scale capacity Ceded on the integrated-system rule, matching AWS Trainium3 UltraServers. /reconcile (October 5, 2026): Layer 1C rebuilt on GA evidence (docs.cloud.google.com BigQuery, Datastream, Pub/Sub, Composer, and Knowledge Catalog docs and release notes). Strong holds on BigQuery-native pipelines, which carry the grade alone (Keith, October 5, 2026). The five prior chips couldn't carry it: Managed Lustre moved to Layer 0 and BigLake to 1A under the M2 storage ruling, Rapid Tier and Smart Storage were already at 1A, Cross-Cloud Lakehouse (Preview) and Smart Storage auto-annotate (undocumented) named, not scored. 1C authority vendor / Ceded to code / Retained. Layer 0, 1A, and 1C grades unchanged. /reconcile (October 5, 2026): authority re-read against product docs under the judgment test, override rule, and Layer 3 ruling: layer1b to vendor / Delegated. Grades and components unchanged. /reconcile (October 5, 2026): product names updated to Google's April 2026 renames (docs.cloud.google.com Gemini Enterprise Agent Platform name changes): Vertex AI Search to Agent Search, Vertex AI Feature Store and Vector Search to Gemini Enterprise Agent Platform, Dataproc to Managed Service for Apache Spark; the first use in each chip keeps the former name. Earlier entries in this log keep the names in use at the time. Grades, DAPM, and authority unchanged. /reconcile (October 5, 2026): authority read against product docs: layer2a visible true; layer0 visible false to true (physicalHostTopology documents cluster, block, sub-block, and host placement, the AWS instance-topology reading). Grades and components unchanged. Layer 1A re-read (October 5, 2026): public-source check found the Gemini context-graph features in preview (Knowledge Catalog release notes) and Smart Storage automated annotation in preview. Grade holds strong on the GA catalog core under rule 9; Smart Storage rescoped to object contexts (GA); Knowledge Catalog credited on discovery, lineage, profiling, quality, data products, and structured-data classification, with the Sensitive Data Protection aspect integration's Cloud Storage limit stated; authority vendor / not overridable / Ceded to vendor / visible / overridable / Delegated on per-object override of aspects, object contexts, and sensitivity tags; BigLake split into the Lakehouse runtime catalog (Ceded) and the Iceberg REST catalog interface (Delegated). ChatGPT peer check (reviews/gcp-1a-chatgpt.json): 8 findings upheld, 1 rejected under the watch-list GA-date rule.

4+1 Layer AI Infrastructure Model · Vendor Assessment Series · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com