# Cerebras Systems (CS-3 and CS-4 Systems + Cerebras Cloud Inference and Training + the Model Zoo) — 4+1 Layer AI Infrastructure Assessment

> Mapped to the 4+1 Layer AI Infrastructure Model  
> Version: v1.0 - 4+1 v2: Authority Split · Date: September 10, 2026  
> Source: Cerebras Inference documentation at inference-docs.cerebras.ai (API reference: chat completions, completions, batch, files, models, metrics, customer management API, versions; capabilities: tool use, structured outputs, reasoning, streaming, prompt caching, image inputs, payload optimization, service tiers, batch, metrics; console: overview, API keys, projects, usage monitoring, account and billing, playground; dedicated endpoints overview; models overview and model pages; support: rate limits, pricing, deprecations, preview releases, change log; resources: OpenAI compatibility, designing for Cerebras; integrations and cookbooks; llms.txt); Cerebras training documentation at training-docs.cerebras.ai (release 2.5.0: getting started, Model Zoo overview, the Wafer-Scale Cluster and CS-3 systems); GitHub license API for Cerebras/modelzoo (Apache 2.0); cerebras.ai product pages (Inference, Cloud, System, Pricing, Partners) and posts (AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution, July 23, 2026; Accelerating GPT-5.6 Sol Ultrafast, August 13, 2026; In the News, including the May 14, 2026 Nasdaq listing). Peer-reviewed cell by cell through the labs claims ledger (cerebras-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (cerebras-row-agy); totals and escalated items in reviews/cerebras-judgment.md.  
> Published by: The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com  
> Author: Keith Townsend

[Full interactive assessment](https://layer2c.com/assessment/cerebras) · [Methodology](https://layer2c.com/methodology) · [What Is Layer 2C?](https://layer2c.com/what-is-layer-2c)

## Executive Summary

Cerebras is the other AI silicon that reached production, and since May 14, 2026 a public company (Nasdaq: CBRS). The map reads it strong at two layers, moderate at one, and gap at five, the narrowest profile in the batch. Layer 0 is strong on the NVIDIA reading: the Wafer-Scale Engine in CS-3 systems the enterprise can own and cluster, CS-4 shipping this quarter, and Cerebras-operated data centers (165 MW in Mikkeli; AMD Helios racks in the second half of 2026) sold as a Training Cloud by the hour and an Inference Cloud by the token. Layer 2B is strong: the fastest production inference of open-weight models behind an OpenAI-compatible API with tool use, constrained-decoding structured outputs, reasoning controls, and zero-retention prompt caching, on a shared tier of two models or on dedicated wafers with a long catalogue (the enterprise's own fine-tuned weights are a Private Preview); OpenAI's Ultrafast Mode for GPT-5.6 Sol runs here in a limited preview. Layer 1C is moderate on the Model Zoo's open data-preparation pipeline, a fixed-function slice whose only destination is the Cerebras trainer. Everything else is a gap by design: no data foundation, no embeddings or retrieval, no movement product, no customer-facing orchestration plane (limits and budgets on a meter; the scheduling is Cerebras's, so that gap reads Ceded), no governance plane, no application.

The capture is at the wafer and nowhere above it. The API is OpenAI-compatible and the models are other labs' open weights, so the client and the model choice lift; tools are the enterprise's; trained weights export from the Apache 2.0 Model Zoo to Hugging Face and GPU formats. What is Cerebras's and stays Cerebras's: the silicon, the systems, the cloud's capacity, the compiler path, and the dedicated endpoints, which run nowhere else. Six components read Ceded, Ceded, Retained, Delegated, Retained, and Ceded across Layer 0, 1C, and 2B; the row has no NVIDIA dependency at all, which on this map is itself the finding.

The buyer's trade: speed nobody else sells, on open models it could serve elsewhere, through an interface it already uses, in exchange for a second silicon vendor with a two-model public tier, preview badges on batch, custom weights, and service tiers, and no agent, data, or governance surface of its own. The decision-authority readings follow: vendor / Ceded at Layer 0 and 2A (Cerebras places and schedules work on its wafers, on its floor or the enterprise's), model / Delegated at 2B (the loop is in the enterprise's process), Absent everywhere else. What would move cells: the custom-weights API reaching general availability (a Delegated chip for the enterprise's own models), a hosted agent runtime (2B's escalated grade), documented capacity controls on dedicated endpoints (2A), and CS-4 shipping (nothing moves; watch-listed).

## Layer Status

| Layer | Status | Classification |
|---|---|---|
| Layer 0 · Compute | ● The Wafer-Scale Engine: CS-3 Systems You Can Own and Cluster, a Cerebras Cloud You Rent by the Hour or the Token, and Cerebras-Operated Data Centers; CS-4 Ships This Quarter; No Networking or Storage of Its Own | Compute & Network Fabric |
| Layer 1A · Storage | ○ No Data Foundation: a Files API for Batch Jobs (Private Preview) and the Enterprise's Own S3 Bucket for Uploaded Weights | Data Storage & Governance |
| Layer 1B · Retrieval | ○ No Retrieval and No Embedding Models: Chat Models Only on the Public Tier; the Documented Vector-Store Integration Is Milvus, Pinecone Appears in an Example | Context Management & Retrieval |
| Layer 1C · Pipelines | ◑ The Model Zoo's Data-Preparation Pipeline (Apache 2.0: Preprocessing, Four-Stage Deduplication, Read Hooks, Token Generators, Dataloaders) Run by the Enterprise or on the Training Cloud, for Cerebras Training Only; a Batch API in Private Preview; No Movement Product | Data Movement & Pipelines |
| Layer 2A · Orchestration | ○ A Metering Edge and a Provisioned Endpoint: Two-Level Rate Limits, Projects, and Roles; Dedicated Endpoints With No Documented Capacity Controls; Service Tiers in Private Preview; No Customer Plane | Infrastructure Orchestration |
| Layer 2B · Runtime | ● The Fastest Production Inference of Open Models Behind an OpenAI-Compatible API, Shared or on Dedicated Wafers, With Tool Use, Structured Outputs, Reasoning, and Prompt Caching; Your Own Fine-Tuned Weights in Private Preview; No Agent Runtime | Application Runtime & Execution |
| Layer 2C · Reasoning | ○ Account Controls, Not a Plane: Organizations, Projects, Roles, Rate Limits, Audit Logs; Gateways Are Partners' (Cloudflare AI Gateway, Portkey, Kong, Operant) | Agentic Infrastructure — The Reasoning Plane |
| Layer 3 (+1) · Applications | ○ No Application: a Playground, Cookbooks, and Customer Stories; the Value Plane Is the Customer's (Cognition, Lovable, Tavus, GSK) or OpenAI's | AI Application Layer — The Value Plane |

## DAPM Portability Profile (components)

| Classification | Count | Meaning |
|---|---|---|
| Retained | 2 | I possess the capability and can operate it independently of this provider |
| Delegated | 1 | Someone else provides the capability, but I can substitute that provider without reconstructing my accumulated opinions |
| Ceded | 3 | Changing providers requires reconstructing those opinions |

**Decision authority (per layer, gaps included)**

| Reading | Layers | Meaning |
|---|---|---|
| Retained | 1 | The enterprise, or code it writes or controls, decides |
| Delegated | 1 | Vendor or model decides; the enterprise can see and override |
| Ceded | 2 | Vendor or model decides; no override, often invisible |
| Absent | 4 | Nothing offered, nothing inherited |

## Strongest Layers

- **Layer 0** (Compute & Network Fabric) — The Wafer-Scale Engine: CS-3 Systems You Can Own and Cluster, a Cerebras Cloud You Rent by the Hour or the Token, and Cerebras-Operated Data Centers; CS-4 Ships This Quarter; No Networking or Storage of Its Own
- **Layer 2B** (Application Runtime & Execution) — The Fastest Production Inference of Open Models Behind an OpenAI-Compatible API, Shared or on Dedicated Wafers, With Tool Use, Structured Outputs, Reasoning, and Prompt Caching; Your Own Fine-Tuned Weights in Private Preview; No Agent Runtime

## Gap Areas

- **Layer 1A** (Data Storage & Governance) — No Data Foundation: a Files API for Batch Jobs (Private Preview) and the Enterprise's Own S3 Bucket for Uploaded Weights
- **Layer 1B** (Context Management & Retrieval) — No Retrieval and No Embedding Models: Chat Models Only on the Public Tier; the Documented Vector-Store Integration Is Milvus, Pinecone Appears in an Example
- **Layer 2A** (Infrastructure Orchestration) — A Metering Edge and a Provisioned Endpoint: Two-Level Rate Limits, Projects, and Roles; Dedicated Endpoints With No Documented Capacity Controls; Service Tiers in Private Preview; No Customer Plane
- **Layer 2C** (Agentic Infrastructure — The Reasoning Plane) — Account Controls, Not a Plane: Organizations, Projects, Roles, Rate Limits, Audit Logs; Gateways Are Partners' (Cloudflare AI Gateway, Portkey, Kong, Operant)
- **Layer 3 (+1)** (AI Application Layer — The Value Plane) — No Application: a Playground, Cookbooks, and Customer Stories; the Value Plane Is the Customer's (Cognition, Lovable, Tavus, GSK) or OpenAI's

## Layer-by-Layer Detail

### ● Layer 0 · Compute: Compute & Network Fabric

*Raw compute, networking, and acceleration fabric*  
**Status:** The Wafer-Scale Engine: CS-3 Systems You Can Own and Cluster, a Cerebras Cloud You Rent by the Hour or the Token, and Cerebras-Operated Data Centers; CS-4 Ships This Quarter; No Networking or Storage of Its Own

**Decision authority:** Ceded (decides: vendor; visible: true; overridable: false; boundary: vendor)

**CS-3 Systems and the Wafer-Scale Cluster (WSE-3; Installed and Administered by the Enterprise; Model Zoo Compilation; S3 Checkpointing)** [DAPM: Ceded]  
Proprietary silicon in Cerebras's own systems, on the enterprise's floor. Ceded, the NVIDIA DGX and silicon reading. CS-4 ('first shipments begin this quarter') is watch-listed, not scored.

**Cerebras Cloud (Cerebras-Operated Data Centers; the Training Cloud by the Hour or the Model; the Inference Cloud by the Token, Shared or as Dedicated Endpoints)** [DAPM: Ceded]  
Cerebras's capacity, sold as jobs, endpoints, and tokens rather than as instances the enterprise administers. Ceded, the NVIDIA DGX Cloud reading. AMD Helios racks (second half of 2026) are watch-listed, not scored.

**Gap Analysis:** Cerebras is a silicon company that became a cloud. The CS-3 system (WSE-3 inside) is sold as hardware the enterprise racks and clusters as a Wafer-Scale Cluster, with the training documentation covering installation, administration, S3 checkpointing, and job restarts; CS-4, 'a revolutionary rack-scale solution' on the Nexus platform with three WSE-3 Turbo wafers per system and two-microsecond wafer-to-wafer links, has 'first CS-4 shipments begin this quarter' and is watch-listed. The same systems fill Cerebras's own data centers (a 165 MW site in Mikkeli, Finland with Compute Nordic; AMD Helios racks arriving in the second half of 2026), sold two ways: the Training Cloud by the hour or by the model ('Pay Per Hour ... we'll determine the time needed to train, fine-tune, and deploy your model'; 'Pay Per Model ... Let our AI experts design, train, and fine-tune'), and the Inference Cloud by the token on a shared tier or as dedicated endpoints ('a private, provisioned instance of the Cerebras Inference service reserved exclusively for your organization'). The buyer gets a second kind of AI silicon, on its own floor or on Cerebras's.

The architect's concern is scope and the ownership line. Cerebras sells no fabric or storage; the cloud exposes no virtual machine, GPU, or wafer the enterprise administers, only jobs, endpoints, and tokens; the on-premises path is a cluster of a single vendor's systems with a Model Zoo that compiles for that silicon; and the developer documentation says nothing about the hardware, data residency, or compliance.

Calibration: NVIDIA reads strong on silicon plus DGX systems plus a hosted cloud; CoreWeave strong on GPU compute the enterprise rents by the instance; Dell and HPE strong on systems; Groq is the closest shape and is scored later in this batch; Cohere and Mistral gap on rented capacity with no customer surface. Cerebras is NVIDIA's shape at a fraction of the scale: proprietary silicon in systems the enterprise can own, or a cloud where it rents the output. Strong: Layer 0 reads by the vendor's shape (the September 15, 2026 ruling), and Cerebras sells systems the enterprise buys and clusters, which puts it in the OEM column with NVIDIA and Dell; the reviewers' rule 4 argument (no fabric, no storage) counted the purpose line's functions, and the OEM line has never required all three.

**Borrowed Judgment:** Ceded at the wafer, on both paths. The Wafer-Scale Engine, the CS systems, the Cerebras Cloud, and the Model Zoo's compilation path are Cerebras's with no second implementer: Ceded, the NVIDIA silicon reading. The runtime call at this layer has two homes and one reading: on Cerebras Cloud, Cerebras places and schedules work across its wafers and the enterprise sees jobs, endpoints, and tokens; on a customer-owned CS-3 cluster, the enterprise administers a single-vendor system whose silicon, compiler, and fabric make the placement, the NVIDIA DGX and Dell reading. Vendor decides, visible (usage, metrics, the cluster's own management), not overridable, Ceded.

### ○ Layer 1A · Storage: Data Storage & Governance

*Durable, governed data foundation — the Governance Catalog that Layer 2C queries*  
**Status:** No Data Foundation: a Files API for Batch Jobs (Private Preview) and the Enterprise's Own S3 Bucket for Uploaded Weights

**Decision authority:** Absent (decides: absent; visible: n/a; overridable: n/a; boundary: vendor)

**Gap Analysis:** Cerebras sells no durable, governed data foundation; what it holds are service and runtime artifacts, some preview-only and time-limited. The Files API ('This feature is in Private Preview') stores batch inputs and outputs for seven days by default with a configurable expiration, and batch results are retained on the same terms (configurable for enterprise customers); uploaded fine-tuned weights (Private Preview) are staged from 'an S3 bucket with cross-account access to Cerebras' that the enterprise owns, after which Cerebras creates and tracks a model version for deployment; the Training Cloud's S3 checkpointing writes to 'any S3-compatible storage'; prompt caches live in Cerebras's memory for their time to live, 'ephemeral in memory and never persisted'; 'Your playground and API requests are never used to train models.' The buyer gets nothing to store here, and that's the design.

The architect's concern is nil at this layer.

Calibration: OpenAI, Anthropic, Mistral, and Cohere read gap on trust apparatus without a data foundation; CoreWeave moderate on object storage it sells. Cerebras sells no storage. Gap, authority Absent.

**Borrowed Judgment:** Nothing offered, nothing inherited. Absent.

### ○ Layer 1B · Retrieval: Context Management & Retrieval

*Low-latency retrieval for RAG — vector/hybrid search, context windows*  
**Status:** No Retrieval and No Embedding Models: Chat Models Only on the Public Tier; the Documented Vector-Store Integration Is Milvus, Pinecone Appears in an Example

**Decision authority:** Absent (decides: absent; visible: n/a; overridable: n/a; boundary: vendor)

**Gap Analysis:** Cerebras serves language models and nothing that indexes. The public tier lists two chat models (gpt-oss-120b and qwen-3.8-27b) and no embedding, reranking, or search model; the documentation's retrieval content is the Milvus integration guide (Pinecone appears in an example, 'RAG with Pinecone + Docker'), where the vector store is the partner's and Cerebras supplies the generation step; the 'build your own Perplexity' and grounded-research cookbooks use Exa and Parallel for search, and the Gist Memory cookbook (summarizing and searching long documents with a read agent) is a customer-code pattern, not a Cerebras retrieval service. The buyer brings its own index.

The architect's concern is nil at this layer.

Calibration: OpenAI moderate on a hosted vector store and embeddings; Mistral moderate on embeddings and libraries; Cohere moderate on frontier embed and rerank models; Anthropic moderate on federated search; Cerebras offers none of those. Gap, authority Absent.

**Borrowed Judgment:** Nothing offered, nothing inherited. Absent.

### ◑ Layer 1C · Pipelines: Data Movement & Pipelines

*Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering*  
**Status:** The Model Zoo's Data-Preparation Pipeline (Apache 2.0: Preprocessing, Four-Stage Deduplication, Read Hooks, Token Generators, Dataloaders) Run by the Enterprise or on the Training Cloud, for Cerebras Training Only; a Batch API in Private Preview; No Movement Product

**Decision authority:** Retained (decides: code; visible: true; overridable: true; boundary: vendor)

**Model Zoo Data-Preparation Pipeline (Apache 2.0; Preprocessing, Four-Stage Deduplication, Read Hooks, Token Generators, On-the-Fly Processing, Dataloaders; Run on the Enterprise's Machines or the Training Cloud; Cerebras Trainer as the Only Destination)** [DAPM: Retained]  
Open code the enterprise runs and modifies; the hosted run on the Training Cloud is Cerebras's and is named with the 2B training chip. Retained, the library reading.

**Gap Analysis:** Cerebras shapes training data and moves nothing else. The Model Zoo's data preparation section (a preprocessing quick start, a four-stage deduplication pipeline, read hooks, token generators, on-the-fly processing, and dataloaders, Apache 2.0) is a library the enterprise runs on its own machines or on the Training Cloud to turn its corpus into the tokenized, packed, deduplicated inputs the Cerebras trainer takes; the Batch API ('This feature is in Private Preview'; a 24-hour window, 50,000 requests or 200 MB per file, ten concurrent batches, chat completions only, 'SDK support for Batch is not yet available during Private Preview') is bulk inference, not a pipeline, and it's pre-GA. The buyer gets a documented, open toolkit for one destination and bulk inference in preview.

The architect's concern is the destination: every stage ends in Cerebras's trainer, there's no connector, movement, lineage, or cache-tiering product, and the second requirement (moving the enterprise's data anywhere else) brings another tool.

Calibration: NetApp reads moderate on AIDE, a fixed pipeline ending in its own retrieval surface; Hugging Face moderate on a hosted runner plus open pipeline libraries; Nebius moderate on a transfer service beside Data Lab; OpenAI, Anthropic, Mistral, Cohere, and Groq gap on intake without transformation or movement. A captive-destination transform is rule 4's own example of a fixed-function slice, and the September 11, 2026 ruling reads trainer-input preparation on that line rather than carving it out. Moderate.

**Borrowed Judgment:** Retained at the library, Ceded at the hosted run. The Model Zoo pipeline is Apache 2.0 code the enterprise runs and modifies anywhere: Retained, the Distilabel and FAISS library reading. The Training Cloud run of it is Cerebras's: Ceded, named with the 2B training chip. The runtime call is the pipeline executing the stages the enterprise configured over its own corpus, with no judgment of the engine's own: code decides, visible, overridable, Retained, the Hugging Face and Elastic 1C reading.

### ○ Layer 2A · Orchestration: Infrastructure Orchestration

*GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization*  
**Status:** A Metering Edge and a Provisioned Endpoint: Two-Level Rate Limits, Projects, and Roles; Dedicated Endpoints With No Documented Capacity Controls; Service Tiers in Private Preview; No Customer Plane

**Decision authority:** Ceded (decides: vendor; visible: false; overridable: false; boundary: vendor)

**Gap Analysis:** Cerebras orchestrates its own wafers and shows the enterprise a meter. The console gives organizations and projects (Org Admin, Project Admin, Project Member; 'On Free Trial accounts, projects aren't available'), two-level rate limits in a dual-bucket model ('every organization has both an uncached token limit and a total token limit'), usage and cost analytics ('may be delayed by up to 10 minutes'), request and audit logs, and a Metrics API in Prometheus format 'designed for customers on dedicated endpoints' and enabled 'on an opt-in basis'. Dedicated endpoints are 'a private, provisioned instance ... reserved exclusively for your organization' whose sizing, replicas, and scaling aren't documented; service tiers (priority, default, auto, flex) are 'in Private Preview' and dedicated-only; the Training Cloud promises 'No DevOps Required' with preconfigured environments Cerebras runs. The buyer gets quotas and a bill.

The architect's concern is that there's no plane: no instance to size, no scaling rule to set, no scheduler to see. Cerebras decides where and when work runs on its wafers, and the enterprise's controls are limits and budgets.

Calibration: Mistral and OpenAI read gap on a metering edge with no customer plane; Anthropic gap on token quotas; Cohere moderate on Model Vault's documented capacity controls (fixed and flex capacity, replica bounds, autoscaling, pause and resume); CoreWeave strong on a real orchestrator. Cerebras is Mistral's shape: a provisioned instance without the controls that made Cohere moderate. Gap, authority Ceded: Cerebras absorbs the scheduling beneath its service, the Mistral, OpenAI, and Anthropic reading.

**Borrowed Judgment:** Nothing offered as a plane, and the scheduling inherited. Cerebras places and schedules work across its wafers invisibly; the enterprise sets limits and budgets on a meter, bounds rather than overrides: vendor decides, not visible, not overridable, Ceded, the Mistral, OpenAI, and Anthropic 2A reading.

### ● Layer 2B · Runtime: Application Runtime & Execution

*Model serving, agent execution, inference APIs, distributed inference*  
**Status:** The Fastest Production Inference of Open Models Behind an OpenAI-Compatible API, Shared or on Dedicated Wafers, With Tool Use, Structured Outputs, Reasoning, and Prompt Caching; Your Own Fine-Tuned Weights in Private Preview; No Agent Runtime

**Decision authority:** Delegated (decides: model; visible: true; overridable: true; boundary: model)

**Model Serving via the OpenAI-Compatible Inference API (Chat Completions and Completions; gpt-oss-120b and qwen-3.8-27b on the Shared Tier; Tool Use With Strict Mode and tool_choice, Parallel Calls Where the Model's Metadata Allows; Structured Outputs With Constrained Decoding; Reasoning Controls; Streaming; Automatic Zero-Retention Prompt Caching; API Version 2)** [DAPM: Delegated]  
Open-weight models their owners publish, behind a multi-vendor interface the enterprise's client already speaks. Delegated, the inference-interface reading.

**Customer-Created and Managed Tools (Function Calling, Client-Side; Any Framework)** [DAPM: Retained]  
Tool logic the enterprise writes and runs in its own process. Retained, the customer-tools ruling.

**Dedicated Endpoints and the Training Cloud on Cerebras Wafers (Provisioned Private Instances; the Dedicated Model Catalogue; Training and Fine-Tuning Through the Apache 2.0 Model Zoo With Checkpoint Conversion to Hugging Face and GPU Formats)** [DAPM: Ceded]  
Cerebras's capacity and compiler path, which run nowhere else; the weights that come out are the enterprise's. Ceded, the NVIDIA DGX Cloud reading, with the artifact exit named. Custom weights upload to dedicated endpoints is Private Preview and carried in notes.

**Gap Analysis:** Cerebras's runtime is the reason the company exists. The Inference API serves open-weight models from other labs ('We host a variety of open-source models from the community... All of our public models are unpruned'): gpt-oss-120b and qwen-3.8-27b on the public shared tier (gemma-4-31b left the public endpoints on September 3, 2026 and stays on dedicated ones; kimi-k2.7-code is customer-trial only), and a long dedicated-endpoint catalogue (Qwen3 up to 235B, Llama 4, Mistral Large 3, GLM 5.1, Kimi K2.7, DeepSeek V3.2, MiniMax, StepFun) 'as well as your own custom weights'. The API is OpenAI-compatible ('changing the API key, base URL, and model ID'), with tool use (strict mode, tool_choice, parallel calls on the models whose metadata allows it; gpt-oss-120b's doesn't), structured outputs with constrained decoding ('strict: true ... Recommended for production use'), reasoning effort controls, streaming (multiple tokens per event), automatic prompt caching at no fee with zero data retention, image inputs (Public Preview), payload optimization, and API version 2 (default since July 22, 2026). Dedicated endpoints are provisioned Cerebras capacity; the Customer Management API ('This feature is in Private Preview') uploads 'custom variants of Cerebras-supported model architectures' from the enterprise's S3 bucket and deploys them 'to a dedicated endpoint running a model with the same underlying architecture'; the Training Cloud trains and fine-tunes 'models of any size' with the Apache 2.0 Model Zoo and converts checkpoints to and from Hugging Face formats and for GPUs. Around it: Hugging Face Inference Providers, OpenRouter, AWS Marketplace, Vercel, and forty-odd integration guides (LangChain, LlamaIndex, CrewAI, LiteLLM, Portkey, Kong, Cloudflare AI Gateway); and OpenAI's own API offers 'Ultrafast Mode ... powered by Cerebras' for GPT-5.6 Sol in a limited preview since August 13, 2026 (named, not scored). The buyer gets open models answering faster than anywhere else, through a client it already has.

The architect's concern is breadth. Cerebras owns no model; the public tier carries two; the agent layer is other people's frameworks (no hosted agent runtime, no Model Context Protocol (MCP) in any cookbook); the enterprise's own weights are a Private Preview on a fixed list of architectures; batch is Private Preview; and the fastest path to the fastest frontier model runs through OpenAI's paper, not Cerebras's.

Calibration: OpenAI, Anthropic, Mistral, and Cohere read strong on frontier serving plus open-source harnesses; CoreWeave strong on open serving plus sandboxes; Cloudflare strong on serverless open-model inference plus a gateway plus agents; Hugging Face moderate on open serving under maturity gates with an experimental agent library. Cerebras is serving alone, at the frontier of speed (rule 6, on Cerebras's own published benchmarks and third-party comparisons). Strong. Agent execution is judged by whether the runtime carries the enterprise's own agent loop (tool calling, structured outputs, reasoning controls, latency over the serial turns an agent takes), not by whether the vendor hosts one (the September 11, 2026 ruling under rule 4), and Cerebras carries it on its own engine at a documented frontier; both reviewer passes voted moderate for calibration with Hugging Face, whose moderate rests on rule 5, not on a missing loop.

**Borrowed Judgment:** Delegated at the interface, Retained at the tools, Ceded at the wafer. Model access through the OpenAI-compatible API is the inference-interface reading: the enterprise's client code and prompts point anywhere, and the models are open weights their owners publish: Delegated. Tools the enterprise writes and calls client-side are its own, the customer-tools reading: Retained. Dedicated endpoints, the Training Cloud, and the Model Zoo's compiled artifacts run only on Cerebras silicon: Ceded, though trained weights export to Hugging Face and GPU formats and are the enterprise's. The decision to call a tool is the model's, but the loop runs in the enterprise's own process, where a deterministic gate before any effect is the enterprise's to write: model decides, visible, overridable, Delegated, the OpenAI and Anthropic reading.

### ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane

*Policy-driven placement and resource coordination — the Autonomy Layer*  
**Status:** Account Controls, Not a Plane: Organizations, Projects, Roles, Rate Limits, Audit Logs; Gateways Are Partners' (Cloudflare AI Gateway, Portkey, Kong, Operant)

**Decision authority:** Absent (decides: absent; visible: n/a; overridable: n/a; boundary: vendor)

**Gap Analysis:** Read against the five legs, Cerebras has none. Identity: organization and project roles and API keys identify callers to the API; no agent principal. Gateway: none of Cerebras's own in production; the integration guides route through partners' gateways (Cloudflare AI Gateway, Portkey, Kong, TrueFoundry AI Gateway with observability, cost tracking, access control, and rate limiting, Operant as a security gateway, among others), and the Cerebras Code MCP Server, an open-source server for coding assistants, is 'in research preview'. Registry and orchestration: none. Observability: request and audit logs and a Metrics API for dedicated endpoints watch the API, not agents. The buyer brings the plane.

The architect's concern is nil at this layer: there's nothing to be captured by.

Calibration: OpenAI, Anthropic, Mistral, and Cohere read gap; Cloudflare moderate on a gateway of its own. Cerebras is the OpenAI shape without the frontier-pre-GA caveat. Gap, authority Absent.

**Borrowed Judgment:** Nothing offered as a plane. Absent.

### ○ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane

*AI-powered business capabilities — business logic, workflow automation*  
**Status:** No Application: a Playground, Cookbooks, and Customer Stories; the Value Plane Is the Customer's (Cognition, Lovable, Tavus, GSK) or OpenAI's

**Decision authority:** Absent (decides: absent; visible: n/a; overridable: n/a; boundary: vendor)

**Gap Analysis:** Cerebras ships no application. The console's Playground is a developer's chat surface ('never used to train models'); the cookbooks are recipes for agents built on partners' frameworks; the case studies (Cognition, Lovable, Tavus, GSK, DeepLearning.AI) are the customers' products; and the most visible application of Cerebras inference, OpenAI's Ultrafast Mode, is OpenAI's. The buyer gets a console (Playground, projects, usage) and no first-party business application or workflow surface.

The architect's concern is nil at this layer.

Calibration: OpenAI, Anthropic, Mistral, and Cohere read strong on first-party applications; CoreWeave moderate on a developer plane; Cerebras has neither. Gap, authority Absent.

**Borrowed Judgment:** Nothing offered, nothing inherited. Absent.

---
*Layer2C · AI Infrastructure Decision Intelligence · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com*
