{
  "id": "cloudflare",
  "name": "Cloudflare (Workers AI + AI Gateway + Agents SDK + Developer Platform)",
  "subtitle": "Mapped to the 4+1 Layer AI Infrastructure Model",
  "version": "v1.0 - 4+1 v2: Authority Split",
  "date": "September 5, 2026",
  "source": "Cloudflare developer documentation (Workers AI overview, models catalog, pricing, fine-tunes; AI Gateway overview, REST API, guardrails usage considerations and supported model types, custom providers; Agents SDK, human-in-the-loop, remote MCP server guide, MCP authorization; Workflows; Durable Objects SQLite storage and data location; Containers pricing and scaling and routing; Workers Placement; Workers for Platforms custom limits; Vectorize limits; AI Search limits and pricing; R2, R2 Data Catalog, R2 SQL, Pipelines, Super Slurper, Sippy; D1 import and export; Data Localization Suite; Queues; Workers VPC; AI Security for Apps; Cloudflare One AI controls and MCP server portals; AI Crawl Control and pay per crawl) and changelogs (Workers AI and AI Gateway unified billing, August 7, 2026; MCP detection and AI Security dashboard, August 12, 2026; Vectorize 20M vectors, August 4, 2026; VPC Networks and Mesh public beta, April 14, 2026; Durable Objects us jurisdiction, June 26, 2026; Kimi K2.6 on Workers AI, April 20, 2026); Cloudflare blog (Agents Week April 13 to 17 and August 3 to 7, 2026 posts; AI platform post, April 16, 2026; AI Security for Apps GA, March 11, 2026; MCP Server Portals; Infire and extra-large model serving; NVIDIA GPUs at the edge; Dynamic Workflows); the cloudflare/agents (MIT) and cloudflare/workerd (Apache 2.0) repositories; Q2 2026 results (August 6, 2026: revenue $696.1 million, 7.4 million Workers developers). Peer-reviewed cell by cell through the labs claims ledger (cloudflare-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (cloudflare-row-agy); totals and escalated items in labs/reviews/cloudflare-judgment.md.",
  "status": "complete",
  "summary": {
    "title": "Summary Finding",
    "paragraphs": [
      "Cloudflare is the network that became a compute platform and then an agent platform, and the map reads it as strong where the agents run and moderate everywhere they store, move, and get governed. One layer strong (2B: open-weights inference on Cloudflare's own GPUs in 180-plus cities, one gateway to 70-plus models across 12-plus providers, and a generally available durable agent runtime with documented approval gates), six moderate (Layer 0 for a network and CPU compute you can buy with no accelerator you can touch, 1A for an S3-compatible store and edge databases without a catalog, 1B for a vector index and open embedding models with the managed search in beta, 1C for queues and durable workflows without pipelines, 2A for manual containers and multi-tenant dispatch, 2C for a gateway over model calls and a firewall for large language model (LLM) endpoints, with the tool gateway and agent identity in beta), and one gap: no first-party application that runs a business process.",
      "The capture is coupled and visible, and it sits in the substrate rather than the model. The models are open weights behind standard-shaped endpoints, the gateway is a broker the enterprise can replace, and the Workers code the enterprise writes runs on the Apache 2.0 workerd runtime anywhere; those read Delegated and Retained. Everything stateful reads Ceded: Durable Objects, Workflows, Vectorize, Containers, Queues, the Agents SDK's runtime (MIT code on a substrate that exists only on Cloudflare), the gateway's routing rules and guardrail settings, and the jurisdictions and policies written into Cloudflare's dashboards. Twenty-four components read Retained 1, Delegated 6, Ceded 17. The pattern is the hyperscalers' with the accelerator subtracted: Cloudflare owns its Layer 0 and rents none of it.",
      "The buyer's trade: inference and agents in the same 330-city network their traffic already crosses, with no cluster to run and a model layer that lifts, in exchange for an agent substrate that doesn't, a data plane sized for applications rather than estates, and a governance plane whose tool-call gateway, agent identity, and traces are all in beta. Nearly everything that would move a cell has a date and a badge: AI Search, Pipelines, R2 Data Catalog, R2 SQL, Dynamic Workers, MCP Server Portals, WriteGuard, Agent Traces, Smart Placement, container autoscaling, pay per crawl, Wallets. Cloudflare ships to beta faster than any row on the map, and the instrument grades what has left it."
    ]
  },
  "layers": [
    {
      "id": "layer0",
      "label": "Layer 0",
      "shortName": "Compute",
      "title": "Compute & Network Fabric",
      "purpose": "Raw compute, networking, and acceleration fabric",
      "status": "moderate",
      "statusLabel": "A Global Network You Can Buy, CPU Compute You Can Size, GPUs You Can't Touch",
      "authority": {
        "decides": "vendor",
        "visible": true,
        "overridable": false,
        "boundary": "vendor",
        "direction": "Ceded"
      },
      "nvidia": [
        {
          "component": "NVIDIA GPUs in 180+ Cities Behind Workers AI, Invisible to the Customer",
          "detail": "Cloudflare's inference fleet is NVIDIA (the edge deployment was announced with NVIDIA; the Infire engine runs Kimi K2.5 across eight H100s with headroom for KV cache). The customer never sees, sizes, or schedules a GPU; there is no GPU instance or container class to buy. The fleet-wide GPU generation mix isn't public."
        }
      ],
      "gap": "Cloudflare owns its Layer 0 and sells parts of it. The network is the product: an anycast fabric in more than 330 cities that the enterprise buys as connectivity (Magic WAN, Cloudflare Mesh with post-quantum tunnels), and compute the enterprise sizes without seeing a host: Workers (V8 isolates, 7.4 million developers as of the second quarter of 2026), Durable Objects, and Containers (generally available on the Workers Paid plan, six instance types from a sixteenth of a vCPU and 256 MiB to four vCPUs and 12 GiB, billed for every 10 milliseconds of running time at per-second resource rates). What the enterprise can't buy is acceleration. Workers AI runs on Cloudflare-owned NVIDIA GPUs in more than 180 cities, and frontier open models run across multiple H100s through the Infire engine, but no GPU is exposed as a product: no instance, no container class, no reservation. The buyer gets a global substrate for CPU workloads and connectivity, with inference as a service on top.\n\nApplying the exposure test: the network and the CPU compute are customer-administered surfaces (instance types, bindings, jurisdictions), so the layer isn't invisible; the accelerator fleet is, and the layer's purpose names acceleration. Half a Layer 0.\n\nCalibration: CoreWeave and the hyperscalers read strong on GPU compute the customer provisions; Mistral reads gap with authority Ceded on a fleet behind an API; Elastic and Kamiwaza gap with authority Absent because they sell no substrate at all. Cloudflare sells a substrate with no accelerator in it. Moderate, with the boundary named: gap on the accelerator test alone is arguable, and the item is escalated.",
      "borrowedJudgment": "Total for placement, bounded for location. Cloudflare decides which city and host a Worker, Durable Object, or container runs on; the enterprise sets jurisdictions (EU, FedRAMP, and since June 2026 US) and location hints, which are configuration. The inference fleet's hardware, capacity, and routing are inherited invisibly. Vendor decides, visible (the instance type, the jurisdiction, the city in analytics), not overridable, Ceded.",
      "notes": "Escalated: the grade (moderate written; gap argued on the accelerator test). Watch-list, notes only: Workers VPC Networks binding Workers to Tunnels, Mesh, and WAN on-ramps (public beta since April 14, 2026). Public evidence that moves the cell: a GPU instance, container class, or reservation the customer administers.",
      "components": [
        {
          "component": "Cloudflare Network + Magic WAN + Cloudflare Mesh (Connectivity the Enterprise Administers)",
          "detail": "The anycast fabric sold as connectivity: Magic WAN and Cloudflare Mesh (encrypted private networking for users, nodes, and agents). Cloudflare's fabric: Ceded. Workers VPC Networks is public beta and unscored.",
          "dapm": "Ceded"
        },
        {
          "component": "Containers + Workers Compute (CPU Instances the Enterprise Sizes)",
          "detail": "Six container instance types from lite to standard-4, started from Durable Objects and billed in 10 millisecond increments, plus the Workers isolate runtime. Compute on Cloudflare's hosts, placed by Cloudflare: Ceded.",
          "dapm": "Ceded"
        }
      ]
    },
    {
      "id": "layer1a",
      "label": "Layer 1A",
      "shortName": "Storage",
      "title": "Data Storage & Governance",
      "purpose": "Durable, governed data foundation — the Governance Catalog that Layer 2C queries",
      "status": "moderate",
      "statusLabel": "S3-Compatible Object Store + Edge Databases + Localization Controls; Catalog in Beta",
      "authority": {
        "decides": "vendor",
        "visible": true,
        "overridable": true,
        "boundary": "vendor",
        "direction": "Delegated"
      },
      "nvidia": [
        {
          "component": "No NVIDIA Layer 1A Dependency",
          "detail": "Nothing at this layer runs on accelerators."
        }
      ],
      "gap": "Cloudflare's data plane is built for applications and agents, not for the enterprise's data estate. R2 is S3-compatible object storage with zero egress fees, the store under AI Search, Pipelines, and the Iceberg catalog; D1 is a serverless database with SQLite semantics (SQLite3-compatible data, SQL export through the command line); Workers KV is a global key-value store; SQLite-backed Durable Objects carry their own database, which is what agent state persists into; Hyperdrive fronts the enterprise's existing PostgreSQL and MySQL with connection pooling and caching at the edge. Governance is location and access, and it's narrower than it sounds: the Data Localization Suite's Regional Services restricts where Workers process requests for a configured hostname (code and secrets deploy globally, and subrequests aren't covered), R2 jurisdiction is a per-bucket setting, and Durable Object jurisdictions (EU, FedRAMP, US since June 26, 2026) confine an object's compute and storage; Cloudflare Access and Zero Trust policies gate who reaches what, and Logpush exports the record. The buyer gets durable, globally distributed storage with an open interface at the bottom (S3) and residency controls on top, applied product by product.\n\nWhat's missing is the catalog. R2 Data Catalog is a managed Apache Iceberg REST catalog inside an R2 bucket, connectable from Spark, Snowflake, and PyIceberg, and R2 SQL queries it; both are in open beta, so neither is scored, and a catalog of table metadata still wouldn't be lineage, classification, or a policy engine over data as an estate. There's no system of record at enterprise scale (D1 is per-application SQLite).\n\nCalibration: AWS reads strong on S3 plus Glue and Lake Formation; Snowflake strong on a governed lakehouse; Elastic and MongoDB moderate on a durable store with access governance and no catalog; CoreWeave moderate on open default storage. Cloudflare is the Elastic and MongoDB case with an S3 interface and residency controls. Moderate.",
      "borrowedJudgment": "Low at the open interfaces, higher above them. R2 is consumed through the S3 API, a genuine multi-vendor standard: Delegated. D1 is a managed service over the open SQLite engine with SQL export, the Hyperdrive and Atlas reading: Delegated. Workers KV and Durable Object storage are Cloudflare-bound: Ceded. Hyperdrive accelerates a database the enterprise owns and can be removed by connecting directly: Delegated. Regional Services and jurisdictions are Cloudflare controls: Ceded. The runtime tradeoff is Cloudflare's storage executing the schemas and access rules the enterprise wrote: vendor decides, visible, overridable, Delegated, the Elastic reading.",
      "notes": "Watch-list, notes only: R2 Data Catalog (managed Iceberg REST catalog, open beta) and R2 SQL (open beta). Public evidence that moves the cell: a GA catalog carrying governance functions (lineage, classification, policy, or queryable governance metadata), not a GA badge on table metadata alone.",
      "components": [
        {
          "component": "R2 Object Storage (S3-Compatible, Zero Egress)",
          "detail": "Object storage consumed through the S3 API with no egress fees, the substrate under AI Search, Pipelines, and the Iceberg catalog. A multi-vendor standard interface: Delegated.",
          "dapm": "Delegated"
        },
        {
          "component": "D1 (Serverless SQLite Database; SQL Export)",
          "detail": "A managed database with SQLite semantics and SQLite3-compatible data, exportable as SQL. A managed service over an open engine: Delegated.",
          "dapm": "Delegated"
        },
        {
          "component": "Workers KV + SQLite-Backed Durable Object Storage (Edge State and Agent Memory)",
          "detail": "A global key-value store and per-object SQLite storage where agent state and memory persist. Cloudflare-bound managed services: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "Hyperdrive (Pooling and Caching for the Enterprise's Own PostgreSQL and MySQL)",
          "detail": "An edge accelerator in front of a database the enterprise owns elsewhere; removable by connecting directly. Delegated.",
          "dapm": "Delegated"
        },
        {
          "component": "Regional Services + Durable Object and R2 Jurisdictions (EU, FedRAMP, US)",
          "detail": "Hostname-scoped regional request processing for Workers, per-bucket R2 jurisdiction, and region-bound Durable Objects. Cloudflare's controls, applied product by product: Ceded.",
          "dapm": "Ceded"
        }
      ]
    },
    {
      "id": "layer1b",
      "label": "Layer 1B",
      "shortName": "Retrieval",
      "title": "Context Management & Retrieval",
      "purpose": "Low-latency retrieval for RAG — vector/hybrid search, context windows",
      "status": "moderate",
      "statusLabel": "Vectorize + Open Embedding and Rerank Models; Managed RAG in Beta",
      "authority": {
        "decides": "vendor",
        "visible": true,
        "overridable": true,
        "boundary": "vendor",
        "direction": "Delegated"
      },
      "nvidia": [
        {
          "component": "Cloudflare's NVIDIA Fleet Under Embedding and Rerank Inference",
          "detail": "The embedding and reranker models run on Workers AI's GPUs in Cloudflare's network; no customer-side dependency."
        }
      ],
      "gap": "Retrieval on Cloudflare is assembled from primitives. Vectorize (generally available) is a managed vector index queried from Workers with metadata filtering and namespaces, up to 20 million vectors per index (doubled on August 4, 2026) at up to 1,536 dimensions, and Workers AI serves the open embedding models (BGE base, large, small, and m3; EmbeddingGemma; Qwen3 embeddings) and a BGE reranker that fill it, so the enterprise builds the pipeline: chunk in a Worker, embed on Workers AI, store in Vectorize, filter and query at request time, rerank in a second call. AI Search is the managed alternative: point it at an R2 bucket, a website, or uploaded files and it chunks, embeds, continuously re-indexes, and serves hybrid semantic-plus-keyword search through a Workers binding, a REST API, or a built-in Model Context Protocol (MCP) endpoint, with namespaces, metadata filters, custom domains, and Access protection; it's free within limits during its open beta, so it isn't scored. The buyer gets a vector store and embedding models at the edge, and a managed retrieval-augmented generation (RAG) service they can try but not yet rely on.\n\nThe architect's concern is coverage, not the vector count. Vectorize is an index, not a retrieval estate: no lexical search, no hybrid fusion, no reranking stage, no permission-aware retrieval inside it; hybrid search exists only in the beta AI Search. Elastic and MongoDB carry no per-index cap and ship the whole pipeline. Embeddings made on Workers AI are made by open models the enterprise could serve elsewhere, so the space isn't captive to Cloudflare, only to the model.\n\nCalibration: Elastic, MongoDB, and Snowflake read strong on native hybrid retrieval estates; Mistral moderate on hosted libraries; Cohere moderate on frontier models plus a fixed-function bundle; AWS strong on a managed portfolio. Cloudflare is a vector index plus open models with the managed layer in beta. Moderate.",
      "borrowedJudgment": "Low. Vectorize is Cloudflare's index and Ceded; the embedding and reranker models are open weights Cloudflare serves, Delegated on the open-weights precedent (the same models run anywhere, and the vectors recompute). The enterprise writes the pipeline, chooses the model, and composes the query: vendor decides at the index, visible, overridable, Delegated, the Elastic reading for a retrieval surface the enterprise configures end to end.",
      "notes": "Watch-list, notes only, and this is what moves the cell: AI Search (open beta; hybrid search, automated indexing, MCP endpoint, namespaces; the August 6, 2026 post added sitemap-free discovery, custom domains, and Access protection). Public evidence that moves the cell: a GA badge on AI Search, or hybrid search and reranking inside Vectorize.",
      "components": [
        {
          "component": "Vectorize (Managed Vector Index; Metadata Filtering and Namespaces; 1,536 Dimensions, 20M Vectors per Index)",
          "detail": "The generally available vector database queried from Workers. Cloudflare's index: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "Open Embedding and Rerank Models on Workers AI (BGE, EmbeddingGemma, Qwen3 Embedding, BGE Reranker)",
          "detail": "Open-weights models served in Cloudflare's network through the Workers AI binding and REST API. The enterprise can serve the same weights elsewhere and the vectors recompute: Delegated on the open-weights precedent.",
          "dapm": "Delegated"
        }
      ]
    },
    {
      "id": "layer1c",
      "label": "Layer 1C",
      "shortName": "Pipelines",
      "title": "Data Movement & Pipelines",
      "purpose": "Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering",
      "status": "moderate",
      "statusLabel": "Queues, Workflows, and Object Migration; Streaming Ingest and Iceberg in Beta",
      "authority": {
        "decides": "code",
        "visible": true,
        "overridable": true,
        "boundary": "vendor",
        "direction": "Retained"
      },
      "nvidia": [
        {
          "component": "No NVIDIA Layer 1C Dependency",
          "detail": "Nothing at this layer runs on accelerators."
        }
      ],
      "gap": "Cloudflare moves data as events and code, not as pipelines. Queues (generally available) is a message queue between Workers with batching, retries, and dead-letter queues; Workflows (generally available) is durable execution for multi-step processes with retries, sleeps, and waitForApproval gates, the orchestration primitive under agents; Super Slurper migrates objects in bulk from S3, Google Cloud Storage, and any S3-compatible store into R2, and Sippy migrates on first read from S3, Google Cloud Storage, and Azure Blob, with zero egress on the Cloudflare side; Logpush streams logs to the enterprise's destinations. The data platform proper is arriving: Pipelines ingests from HTTP and Workers bindings, transforms with SQL, and lands Iceberg tables, Parquet, or JSON in R2 with exactly-once delivery; R2 SQL queries the catalog; Workers VPC Networks and Cloudflare Mesh reach private databases and APIs. All of those are beta, and none is scored. The buyer gets movement primitives they compose in code, and a lakehouse ingest path they can preview.\n\nApplying rule 4: what's generally available is general in kind (queues and durable workflows move anything) but carries no pipeline product: no ETL, no change data capture (CDC), no lineage, no cost-aware movement beyond zero egress, no KV-cache tiering. The migration tools are one-way into R2.\n\nCalibration: Snowflake, Databricks, and AWS read strong on managed pipeline portfolios; Elastic moderate on ingest fleets without CDC or lineage; MongoDB moderate on streams and CDC with fixed destinations. Cloudflare is below MongoDB (no CDC) and above gap (queues, workflows, and migration are shipped movement). Moderate.",
      "borrowedJudgment": "Low, and split by component. Workflows executes the enterprise's step code and Queues delivers to the enterprise's consumers with no judgment of Cloudflare's own in the transformation, the override rule's third clause: code decides, visible, overridable, Retained, the Elastic and MongoDB 1C reading. Super Slurper's multipart choices and Logpush's fixed cadence are vendor-run mechanics inside those tools, named here and not the layer's dominant function. Every component is Cloudflare's to run, Ceded on portability.",
      "notes": "Watch-list, notes only, and this is what moves the cell: Pipelines (open beta; SQL transformations, exactly-once to R2, Iceberg sinks), R2 SQL (open beta), R2 Data Catalog (open beta), Workers VPC Networks and Mesh (public beta since April 14, 2026), Dynamic Workflows (MIT library on Dynamic Workers, itself open beta). Reviewed for awareness: whether the layer authority should read vendor / Delegated on the Snowflake and Databricks calibration rather than code / Retained on Elastic and MongoDB; the row keeps code / Retained because the scored engines run enterprise code without curated connectors doing judgment. Public evidence that moves the cell: GA badges on Pipelines and the catalog, or lineage.",
      "components": [
        {
          "component": "Queues (Message Queue Between Workers; Batching, Retries, Dead-Letter Queues)",
          "detail": "Generally available queueing for event movement inside the platform. Cloudflare's service: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "Workflows (Durable Execution: Steps, Retries, Sleeps, waitForApproval)",
          "detail": "Generally available durable execution engine running the enterprise's multi-step code, the orchestration primitive under Agents. Cloudflare's engine: Ceded. Dynamic Workflows is beta-dependent and unscored.",
          "dapm": "Ceded"
        },
        {
          "component": "Super Slurper (Bulk from S3, GCS, S3-Compatible) + Sippy (On First Read from S3, GCS, Azure Blob) into R2",
          "detail": "Bulk and on-demand migration into R2. One-way tools on Cloudflare's paper: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "Logpush (Log Streaming to Enterprise Destinations)",
          "detail": "Delivery of platform and Zero Trust logs to the enterprise's storage and SIEM at a cadence Cloudflare sets. Cloudflare's exporter: Ceded.",
          "dapm": "Ceded"
        }
      ]
    },
    {
      "id": "layer2a",
      "label": "Layer 2A",
      "shortName": "Orchestration",
      "title": "Infrastructure Orchestration",
      "purpose": "GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization",
      "status": "moderate",
      "statusLabel": "Manual Containers from Durable Objects, Multi-Tenant Dispatch, and Jurisdictions; No Autoscaling, No GPU Plane",
      "authority": {
        "decides": "vendor",
        "visible": true,
        "overridable": false,
        "boundary": "vendor",
        "direction": "Ceded"
      },
      "nvidia": [
        {
          "component": "No Customer-Side NVIDIA Dependency",
          "detail": "The customer never schedules a GPU; Workers AI's fleet is scheduled by Cloudflare."
        }
      ],
      "gap": "Orchestration on Cloudflare is the platform's, with a handful of levers the enterprise holds, and fewer than the marketing suggests. Workers scale to zero and to millions of isolates with no cluster to run; Durable Objects are placed near their first request, pinned by location hints, and confined by jurisdictions; Containers (generally available) are started, stopped, and addressed from a Durable Object in JavaScript, which is the whole orchestration API (the docs say it outright: instead of writing Kubernetes operators, you write JavaScript), but today they scale manually, by getting a container with a unique ID and starting it, and stateless routing is a fixed-count random choice that ignores location, with built-in autoscaling and routing on the roadmap; Workers for Platforms (generally available) dispatches tenant code into isolated namespaces, with the enterprise's own dispatch Worker choosing the target Worker and setting per-tenant CPU and subrequest limits in code on every request; Dynamic Workers (open beta) instantiate sandboxed isolates at runtime for agent-generated code; Smart Placement (beta) moves a Worker near its backends. The buyer gets a scheduler they never operate, can bound, and can't yet make elastic for containers.\n\nThe architect's concern: none of this is a GPU plane, and the container plane isn't a scheduler yet. There are no accelerator quotas, no fair-share across a shared fleet, no capacity the enterprise reserves or observes for inference; Workers AI's scheduling is Cloudflare's alone. Placement across the network is Cloudflare's judgment with the enterprise's bounds on it.\n\nCalibration: Elastic reads moderate with authority Ceded on platform-scoped orchestrators; Snowflake moderate on managed compute with GA GPU pools; MongoDB moderate on cluster autoscaling and operators; AWS strong on EKS and capacity management; Nutanix moderate. Cloudflare is the platform-scoped case with an unusually broad isolate plane and a container plane still without autoscaling. Moderate, written pending Keith's read on whether manual containers and dispatch clear the bar.",
      "borrowedJudgment": "Bounded, with enterprise code choosing what runs but not where. Jurisdictions, location hints, instance types, and namespace limits are configuration Cloudflare honors within its own placement and scaling judgment; the Workers for Platforms dispatch Worker is the enterprise's code choosing which tenant Worker runs and under what quotas on each request, which decides the workload, not its placement. Vendor decides placement and scaling, visible, not overridable, Ceded, the Elastic platform-scoped reading. The Container lifecycle code the enterprise writes is a control surface over Cloudflare's scheduler, not a scheduler of its own.",
      "notes": "Escalated: the grade (moderate written; gap argued because Containers scale manually and the isolate plane is invisible). Watch-list, notes only: Smart Placement (beta), Dynamic Workers (open beta for paid Workers accounts), built-in container autoscaling and routing (roadmap). Public evidence that moves the cell: container autoscaling reaching GA, or a GPU or shared-fleet scheduling surface the customer administers.",
      "components": [
        {
          "component": "Containers Orchestrated from Durable Objects (Explicit-ID Lifecycle, Instance Types, Fixed-Count Random Routing)",
          "detail": "Generally available containers whose lifecycle is driven from a Durable Object in JavaScript, scaled manually today, with six instance sizes. Cloudflare's scheduler behind a JavaScript control surface: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "Workers for Platforms (Dispatch Namespaces; Enterprise-Written Dispatch Worker with Per-Tenant Limits)",
          "detail": "Generally available dispatch of tenant code into isolated namespaces, routed and rate-bounded by the enterprise's dispatch Worker. Cloudflare-only platform: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "Durable Object Location Hints and Jurisdictions",
          "detail": "Placement of Durable Objects by hint or jurisdiction (EU, FedRAMP, US). Configuration over Cloudflare's placement: Ceded.",
          "dapm": "Ceded"
        }
      ]
    },
    {
      "id": "layer2b",
      "label": "Layer 2B",
      "shortName": "Runtime",
      "title": "Application Runtime & Execution",
      "purpose": "Model serving, agent execution, inference APIs, distributed inference",
      "status": "strong",
      "statusLabel": "Serverless Open-Model Inference + One Gateway to 70+ Models + Durable Agents on Every Edge",
      "authority": {
        "decides": "model",
        "visible": true,
        "overridable": true,
        "boundary": "model",
        "direction": "Delegated"
      },
      "nvidia": [
        {
          "component": "Cloudflare's NVIDIA Fleet Under Workers AI; Multi-H100 Serving for Frontier MoE Models",
          "detail": "Workers AI runs on Cloudflare-owned NVIDIA GPUs in 180+ cities; the Infire engine (Rust) serves trillion-parameter mixture-of-experts models such as Kimi K2.5 across eight H100s. No customer-side dependency; nothing the customer serves runs on their own GPUs."
        }
      ],
      "gap": "Both halves of the layer are generally available and both live at the edge. Serving: Workers AI runs more than fifty open models on Cloudflare's GPUs, including frontier open weights (Kimi K2.6, a trillion-parameter mixture-of-experts model with 32 billion active parameters and a 262K context, added April 20, 2026; GLM-5.3; DeepSeek V4; Llama 4 Scout; gpt-oss; Gemma 4, the frontier tier behind a paid billing method) through the Workers AI binding and a REST API, with bring-your-own LoRA adapters for fine-tuned inference. AI Gateway (generally available on all plans) fronts Workers AI and more than seventy models across twelve-plus providers (OpenAI, Anthropic, Google, Groq, xAI, Bedrock, Azure OpenAI, Replicate, Alibaba, Bytedance, and more) through standard-shaped endpoints: an OpenAI-compatible chat-completions path for supported LLMs, an Anthropic-compatible messages path, an OpenAI Responses path, and Cloudflare's own run envelope for every modality, with one prepaid credit balance across Cloudflare's models and the third parties' since August 7, 2026. Agents: the Agents SDK (MIT) makes each agent a Durable Object with its own SQLite state, WebSocket sessions, scheduling, long-running sessions that survive hibernation (April 2026), an MCP client to any server, and remote MCP hosting through createMcpHandler for new stateless servers (the stateful McpAgent path is deprecated) with OAuth through Cloudflare Access or the Workers OAuth Provider library, plus two documented approval gates: needsApproval on a tool (a boolean or a predicate over the arguments) that parks the turn until a human answers, and Workflows' waitForApproval for gates that hold for days. Around it: Workflows (generally available), Containers and the Sandbox SDK, Browser Run, Dynamic Workers (open beta) for agent-written code, and Cloudflare OS (open-source, August 2026). The buyer gets inference and agents in the same 330-city network their traffic already crosses, with no cluster and no cold start worth naming.\n\nThe architect's concern splits. The models are open and the interfaces are standard-shaped, so the model layer lifts; the agent layer doesn't. An agent written to the Agents SDK is a Durable Object, and Durable Objects, Workflows, and Vectorize exist only on Cloudflare; the code is MIT and the Workers runtime (workerd) is Apache 2.0 and self-hostable as an application server, but the stateful substrate the agents depend on isn't reproducible off the network. What the enterprise writes in plain Workers code (tools, handlers, MCP servers) is web-standard JavaScript that runs on workerd anywhere.\n\nCalibration: AWS, Google Cloud, and Azure read strong on serving plus agent runtimes; Mistral strong on serving plus a weights exit; Elastic moderate on brokered LLMs with no serving. Cloudflare has serving of open weights, a broker across twelve-plus providers (Azure's Foundry catalog and Google's Model Garden list more models; Cloudflare's breadth is providers behind one gateway), and a GA durable agent runtime with documented gates. Strong.",
      "borrowedJudgment": "Low for the model, high for the harness. Open weights served through Workers AI read Delegated on the inference-interface ruling, and so does raw model access through AI Gateway's standard-shaped endpoints, a broker the enterprise can replace (the Elastic and ServiceNow managed-LLM precedent); the gateway's routing rules, guardrails, keys, and billing are its control plane, scored Ceded at 2C. The Agents SDK, Durable Objects, Workflows, Containers, and Browser Run are Cloudflare's runtime, Ceded, MIT license notwithstanding: the code lifts, the substrate doesn't. Customer-authored Workers code is web-standard JavaScript that runs on the Apache 2.0 workerd runtime outside Cloudflare, written Retained on the open-source seam and escalated against the narrowed customer-tools ruling. The model decides which tool to call; needsApproval and waitForApproval are persisted gates before the effect, and AI Gateway guardrails block in line for non-streaming text-generation traffic (streaming responses on the REST path are evaluated and logged, not enforced): model decides, visible, overridable, Delegated.",
      "notes": "Escalated: whether customer-authored Workers code reads Retained (runs on workerd anywhere) or Ceded (platform-bound like Atlas Functions) under the narrowed customer-tools ruling. Watch-list, notes only: Dynamic Workers (open beta), the next-edition Agents SDK preview (Agents Week, August 2026), bring-your-own-model on Workers AI (roadmap), Browser Run's Live View and Human in the Loop (beta), Cloudflare OS (open-source core, August 5, 2026), full Guardrails support for streaming (roadmap). Inference flagged: the fleet-wide GPU mix. Public evidence that moves the cell: nothing upward from strong.",
      "components": [
        {
          "component": "Workers AI Model Serving (50+ Open Models on Cloudflare GPUs; Binding and REST API; Bring-Your-Own LoRA)",
          "detail": "Serverless inference of open weights including Kimi K2.6, GLM-5.3, DeepSeek V4, Llama 4, gpt-oss, and Gemma 4 in 180+ cities, with LoRA adapters the enterprise trained elsewhere applied at runtime. Open models behind a standard-shaped interface the enterprise can serve elsewhere: Delegated.",
          "dapm": "Delegated"
        },
        {
          "component": "Raw Model Access via AI Gateway's Standard-Shaped Endpoints (OpenAI Chat Completions and Responses, Anthropic Messages; 70+ Models, 12+ Providers)",
          "detail": "Third-party and Cloudflare models reached through OpenAI- and Anthropic-compatible paths at one base URL. A broker the enterprise can replace with a base URL swap: Delegated, the managed-LLM precedent. Routing, guardrails, BYOK, logging, and unified billing are the gateway's control plane, scored at 2C.",
          "dapm": "Delegated"
        },
        {
          "component": "Agents SDK + Durable Objects + Workflows (Durable Agent Runtime; MCP Client; createMcpHandler Servers; needsApproval and waitForApproval)",
          "detail": "MIT-licensed SDK whose agents are Durable Objects with SQLite state, hibernation, scheduling, long-running sessions, stateless remote MCP hosting with Access or OAuth Provider authorization, and documented approval gates. The code lifts, the stateful substrate exists only on Cloudflare: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "Containers + Sandbox SDK + Browser Run (Agent Execution Substrates)",
          "detail": "Containers driven from Durable Objects, the Sandbox SDK for agent code execution, and the Browser Run headless browser (Live View and Human in the Loop in beta). Cloudflare's substrates: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "Customer-Authored Workers Code (Tools, Handlers, MCP Servers; Web-Standard APIs on the Apache 2.0 workerd Runtime)",
          "detail": "The tool logic, handlers, and MCP servers the enterprise writes as standard JavaScript and TypeScript, runnable on the open-source workerd runtime outside Cloudflare (its README names self-hosting as a use). Written Retained on the open-source seam; escalated against the narrowed customer-tools ruling.",
          "dapm": "Retained"
        }
      ]
    },
    {
      "id": "layer2c",
      "label": "Layer 2C",
      "shortName": "Reasoning",
      "title": "Agentic Infrastructure — The Reasoning Plane",
      "purpose": "Policy-driven placement and resource coordination — the Autonomy Layer",
      "status": "moderate",
      "statusLabel": "Gateway Over Model Calls + Firewall for LLM Endpoints + Access for MCP Servers; Tool Gateway and Identity in Beta",
      "authority": {
        "decides": "vendor",
        "visible": true,
        "overridable": true,
        "boundary": "vendor",
        "direction": "Delegated"
      },
      "nvidia": [
        {
          "component": "Cloudflare's GPUs Under the Guardrail Models",
          "detail": "AI Gateway guardrails run Llama Guard 3 8B and Prompt Guard 2 86M on Workers AI; the firewall's detections run on Cloudflare's edge. No customer-side dependency."
        }
      ],
      "gap": "Cloudflare's reasoning plane is a request-time gateway, which is the leg it was always going to ship first. AI Gateway (generally available) sits between every application and every model with rate limits, caching, logging of prompts and completions, cost tracking, bring-your-own keys in the Secret Store, one credit balance across providers, dynamic routing that segments users, enforces quotas, and picks models with fallbacks from rules the enterprise draws or writes as JSON, and guardrails that run Llama Guard 3 and Prompt Guard 2 over prompts and responses for supported non-streaming text-generation traffic (embeddings get prompt evaluation only, unknown model types skip response evaluation, and streaming responses on the REST path are logged rather than enforced); identity-aware analytics that attribute requests to users and flag rogue behavior are in open beta. AI Security for Apps (formerly Firewall for AI, generally available March 11, 2026, an Enterprise add-on) runs incoming JSON requests to endpoints labeled as LLM endpoints through detection modules for prompt injection, personally identifiable information, and unsafe or custom topics, and attaches scores the enterprise's Web Application Firewall (WAF) rules act on. Cloudflare Access is an OAuth provider for remote MCP servers and portals, so who may reach a server is a Zero Trust policy. What's coming is the rest of the plane: MCP Server Portals (open beta) put every MCP server behind one endpoint with Access policies, per-tool aliases and visibility, and per-request logging; Gateway now detects MCP traffic (the policy selector is beta) and an AI Security dashboard reports it; WriteGuard (private beta) governs write operations through MCP; Agent Traces (beta) trace agents; the Agent Access Model (a published architecture, not a product) proposes task-scoped agent identity. The buyer gets policy over model calls and LLM endpoints today, and policy over tool calls and agent identity on a visible roadmap.\n\nApplying the test: Intelligence-2C has a gateway leg (model calls, GA), an enforcement leg (LLM-endpoint firewall, GA), and an access leg for servers (Access as MCP OAuth, GA); no registry treating agents as governed assets, no agent principal (the Access Model is a proposal), no per-tool policy outside the beta portals, no cross-agent orchestration beyond Workflows inside one application, and agent observability in beta. Infrastructure-2C, placement across model, cost, and compliance tiers, shows up as dynamic routing: rules the enterprise writes, executed deterministically. Routing is not reasoning.\n\nCalibration: AWS reads moderate on Guardrails and AgentCore Policy; Azure strong on Entra Agent ID plus a gateway; Google Cloud strong on a complete plane; Nutanix moderate on gateway governance without placement; Databricks moderate on a gateway plus catalog governance; Elastic moderate on registry, orchestration, and observability legs. Cloudflare is Nutanix's shape with a real LLM firewall and the tool gateway in beta. Moderate.",
      "borrowedJudgment": "Split between the enterprise's rules and Cloudflare's classifiers. Routing flows, rate limits, and Access policies are the enterprise's decisions executed deterministically; guardrails and the firewall are Cloudflare's models (Llama Guard 3, Prompt Guard 2, the PII and injection detectors) producing scores, with the enterprise's thresholds as configuration and its WAF rules as persisted gates before the effect. Vendor decides at the classifier, visible, overridable through the rules, Delegated, the Nutanix and Databricks gateway reading. On portability every component is Cloudflare's and Ceded: the rules are written in Cloudflare's dashboards and Terraform provider and rebuild elsewhere.",
      "notes": "Watch-list, notes only, and these are the legs that move the cell: MCP Server Portals (open beta since August 2025; per-tool aliases and visibility live here), MCP traffic detection in Gateway (beta selector) and the AI Security dashboard (August 12, 2026), WriteGuard for MCP (private beta), identity-aware AI Gateway analytics (open beta), Agent Traces (beta), the Agent Access Model (proposal, August 5, 2026), full streaming enforcement for guardrails (roadmap). Public evidence that moves the cell: GA badges on the portals and an agent principal.",
      "components": [
        {
          "component": "AI Gateway Control Plane (Rate Limits, Caching, Logging, Dynamic Routing Rules, Guardrails, BYOK, Unified Billing, Cost Tracking)",
          "detail": "The generally available policy and observability layer over every model call, with routing flows the enterprise draws or writes as JSON, Cloudflare-held provider keys, and one credit balance. Cloudflare's control plane; the rules rebuild elsewhere: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "AI Security for Apps (Prompt Injection, PII, and Topic Detection on JSON Requests to LLM Endpoints; WAF Rules on Scores)",
          "detail": "Generally available since March 11, 2026 as an Enterprise add-on: detection modules score incoming JSON prompts to labeled LLM endpoints and the enterprise's WAF rules block or challenge. Cloudflare's detections: Ceded.",
          "dapm": "Ceded"
        },
        {
          "component": "Cloudflare Access as OAuth for Remote MCP Servers and Portals (Zero Trust Policy over Server Access)",
          "detail": "Access issues and enforces the OAuth grants that gate remote MCP servers hosted on Workers and the portals in front of them; per-tool policy belongs to the beta portals. Cloudflare's identity plane: Ceded.",
          "dapm": "Ceded"
        }
      ]
    },
    {
      "id": "layer3",
      "label": "Layer 3 (+1)",
      "shortName": "Applications",
      "title": "AI Application Layer — The Value Plane",
      "purpose": "AI-powered business capabilities — business logic, workflow automation",
      "status": "gap",
      "statusLabel": "A Console for AI Traffic, Not an AI Application; the Agentic-Internet Surfaces Are in Beta",
      "authority": {
        "decides": "absent",
        "visible": null,
        "overridable": null,
        "boundary": "vendor",
        "direction": "Absent"
      },
      "nvidia": [
        {
          "component": "No NVIDIA Layer 3 Dependency",
          "detail": "Nothing at this layer exists to carry a dependency."
        }
      ],
      "gap": "Cloudflare's own applications are about the agentic Internet rather than about a business process, and the one that's generally available is a policy console. AI Crawl Control (available on all plans since August 2025) gives every site owner visibility into AI crawlers and rules to allow or block them by bot and purpose, with the enterprise's WAF rules layered over it; charging for access through pay per crawl, where a crawler either presents payment intent or receives an HTTP 402 with a price, is in private beta and limited to Enterprise plans with Bot Management. The Agent Readiness diagnostics tell a site how agents read it; Answer Engine Optimization is early access; Radar Researcher (beta) answers questions about Internet traffic in plain language. The commerce surfaces are announced rather than shipped: the Monetization Gateway (waitlist since July 1, 2026) charges for any asset behind Cloudflare through x402, and Cloudflare Wallets and cloudflare.pay (announced August 4, 2026) give agents stablecoin wallets and identity handles, rolling out over the following months. Cloudflare OS (open-source, August 5, 2026) is a platform for building agents, apps, and work, used internally first. The buyer gets a control surface for the AI traffic hitting their properties, and a roadmap for getting paid by it.\n\nApplying the test: the layer asks for AI-powered business capabilities, and what's generally available is bot management for AI traffic, a console of the enterprise's own rules that runs no business process and carries no AI inside it. Everything that would be an application here is beta, waitlist, or announced.\n\nCalibration: Elastic and Qlik read strong on first-party security, observability, and analytics applications; AWS strong on the broadest ecosystem; CoreWeave moderate on a developer plane; Nutanix moderate on platform-enabled applications; NetApp gap as a foundation beneath others' apps; Cohere strong on a workspace. Cloudflare is the platform others build applications on. Gap, Absent, written pending Keith's read on whether crawler control counts as an application surface.",
      "borrowedJudgment": "None to borrow at the layer's function: AI Crawl Control's rules are the enterprise's settings on a console, and the applications built on Workers are the enterprise's or an ISV's, scored on their rows. Absent: nothing offered at the layer's function, nothing inherited.",
      "notes": "Escalated: the grade (gap written; moderate argued on AI Crawl Control as a GA first-party surface). Watch-list, notes only: pay per crawl (private beta; Enterprise with Bot Management), Monetization Gateway (waitlist), Cloudflare Wallets and cloudflare.pay (rolling out), Radar Researcher (beta), Answer Engine Optimization (early access), Cloudflare OS (open-source core). Public evidence that moves the cell: a first-party application that runs a business process, or the commerce surfaces reaching GA.",
      "components": []
    }
  ]
}
