Cerebras is the other AI silicon that reached production, and since May 14, 2026 a public company (Nasdaq: CBRS). The map reads it strong at two layers, moderate at one, and gap at five, the narrowest profile in the batch. Layer 0 is strong on the NVIDIA reading: the Wafer-Scale Engine in CS-3 systems the enterprise can own and cluster, CS-4 shipping this quarter, and Cerebras-operated data centers (165 MW in Mikkeli; AMD Helios racks in the second half of 2026) sold as a Training Cloud by the hour and an Inference Cloud by the token. Layer 2B is strong: the fastest production inference of open-weight models behind an OpenAI-compatible API with tool use, constrained-decoding structured outputs, reasoning controls, and zero-retention prompt caching, on a shared tier of two models or on dedicated wafers with a long catalogue (the enterprise's own fine-tuned weights are a Private Preview); OpenAI's Ultrafast Mode for GPT-5.6 Sol runs here in a limited preview. Layer 1C is moderate on the Model Zoo's open data-preparation pipeline, a fixed-function slice whose only destination is the Cerebras trainer. Everything else is a gap by design: no data foundation, no embeddings or retrieval, no movement product, no customer-facing orchestration plane (limits and budgets on a meter; the scheduling is Cerebras's, so that gap reads Ceded), no governance plane, no application.
The capture is at the wafer and nowhere above it. The API is OpenAI-compatible and the models are other labs' open weights, so the client and the model choice lift; tools are the enterprise's; trained weights export from the Apache 2.0 Model Zoo to Hugging Face and GPU formats. What is Cerebras's and stays Cerebras's: the silicon, the systems, the cloud's capacity, the compiler path, and the dedicated endpoints, which run nowhere else. Six components read Ceded, Ceded, Retained, Delegated, Retained, and Ceded across Layer 0, 1C, and 2B; the row has no NVIDIA dependency at all, which on this map is itself the finding.
The buyer's trade: speed nobody else sells, on open models it could serve elsewhere, through an interface it already uses, in exchange for a second silicon vendor with a two-model public tier, preview badges on batch, custom weights, and service tiers, and no agent, data, or governance surface of its own. The decision-authority readings follow: vendor / Ceded at Layer 0 and 2A (Cerebras places and schedules work on its wafers, on its floor or the enterprise's), model / Delegated at 2B (the loop is in the enterprise's process), Absent everywhere else. What would move cells: the custom-weights API reaching general availability (a Delegated chip for the enterprise's own models), a hosted agent runtime (2B's escalated grade), documented capacity controls on dedicated endpoints (2A), and CS-4 shipping (nothing moves; watch-listed).
Layer-by-layer status: Layer 0 (The Wafer-Scale Engine: CS-3 Systems You Can Own and Cluster, a Cerebras Cloud You Rent by the Hour or the Token, and Cerebras-Operated Data Centers; CS-4 Ships This Quarter; No Networking or Storage of Its Own), Layer 1A (No Data Foundation: a Files API for Batch Jobs (Private Preview) and the Enterprise's Own S3 Bucket for Uploaded Weights), Layer 1B (No Retrieval and No Embedding Models: Chat Models Only on the Public Tier; the Documented Vector-Store Integration Is Milvus, Pinecone Appears in an Example), Layer 1C (The Model Zoo's Data-Preparation Pipeline (Apache 2.0: Preprocessing, Four-Stage Deduplication, Read Hooks, Token Generators, Dataloaders) Run by the Enterprise or on the Training Cloud, for Cerebras Training Only; a Batch API in Private Preview; No Movement Product), Layer 2A (A Metering Edge and a Provisioned Endpoint: Two-Level Rate Limits, Projects, and Roles; Dedicated Endpoints With No Documented Capacity Controls; Service Tiers in Private Preview; No Customer Plane), Layer 2B (The Fastest Production Inference of Open Models Behind an OpenAI-Compatible API, Shared or on Dedicated Wafers, With Tool Use, Structured Outputs, Reasoning, and Prompt Caching; Your Own Fine-Tuned Weights in Private Preview; No Agent Runtime), Layer 2C (Account Controls, Not a Plane: Organizations, Projects, Roles, Rate Limits, Audit Logs; Gateways Are Partners' (Cloudflare AI Gateway, Portkey, Kong, Operant)), Layer 3 (+1) (No Application: a Playground, Cookbooks, and Customer Stories; the Value Plane Is the Customer's (Cognition, Lovable, Tavus, GSK) or OpenAI's).
Assessment framework: 4+1 Layer AI Infrastructure Model. Scoring model: Decision Authority Placement Model (DAPM) — Retained, Delegated, or Ceded. Published by The CTO Advisor LLC (DBA The Advisor Bench). Author: Keith Townsend. Date assessed: September 10, 2026. Version: v1.0 - 4+1 v2: Authority Split.
Cerebras is the other AI silicon that reached production, and since May 14, 2026 a public company (Nasdaq: CBRS). The map reads it strong at two layers, moderate at one, and gap at five, the narrowest profile in the batch. Layer 0 is strong on the NVIDIA reading: the Wafer-Scale Engine in CS-3 systems the enterprise can own and cluster, CS-4 shipping this quarter, and Cerebras-operated data centers (165 MW in Mikkeli; AMD Helios racks in the second half of 2026) sold as a Training Cloud by the hour and an Inference Cloud by the token. Layer 2B is strong: the fastest production inference of open-weight models behind an OpenAI-compatible API with tool use, constrained-decoding structured outputs, reasoning controls, and zero-retention prompt caching, on a shared tier of two models or on dedicated wafers with a long catalogue (the enterprise's own fine-tuned weights are a Private Preview); OpenAI's Ultrafast Mode for GPT-5.6 Sol runs here in a limited preview. Layer 1C is moderate on the Model Zoo's open data-preparation pipeline, a fixed-function slice whose only destination is the Cerebras trainer. Everything else is a gap by design: no data foundation, no embeddings or retrieval, no movement product, no customer-facing orchestration plane (limits and budgets on a meter; the scheduling is Cerebras's, so that gap reads Ceded), no governance plane, no application.
The capture is at the wafer and nowhere above it. The API is OpenAI-compatible and the models are other labs' open weights, so the client and the model choice lift; tools are the enterprise's; trained weights export from the Apache 2.0 Model Zoo to Hugging Face and GPU formats. What is Cerebras's and stays Cerebras's: the silicon, the systems, the cloud's capacity, the compiler path, and the dedicated endpoints, which run nowhere else. Six components read Ceded, Ceded, Retained, Delegated, Retained, and Ceded across Layer 0, 1C, and 2B; the row has no NVIDIA dependency at all, which on this map is itself the finding.
The buyer's trade: speed nobody else sells, on open models it could serve elsewhere, through an interface it already uses, in exchange for a second silicon vendor with a two-model public tier, preview badges on batch, custom weights, and service tiers, and no agent, data, or governance surface of its own. The decision-authority readings follow: vendor / Ceded at Layer 0 and 2A (Cerebras places and schedules work on its wafers, on its floor or the enterprise's), model / Delegated at 2B (the loop is in the enterprise's process), Absent everywhere else. What would move cells: the custom-weights API reaching general availability (a Delegated chip for the enterprise's own models), a hosted agent runtime (2B's escalated grade), documented capacity controls on dedicated endpoints (2A), and CS-4 shipping (nothing moves; watch-listed).
Raw compute, networking, and acceleration fabric
Proprietary silicon in Cerebras's own systems, on the enterprise's floor. Ceded, the NVIDIA DGX and silicon reading. CS-4 ('first shipments begin this quarter') is watch-listed, not scored.
Cerebras's capacity, sold as jobs, endpoints, and tokens rather than as instances the enterprise administers. Ceded, the NVIDIA DGX Cloud reading. AMD Helios racks (second half of 2026) are watch-listed, not scored.
The Wafer-Scale Engine is Cerebras's own processor (WSE-3: four trillion transistors, 900,000 cores; WSE-3 Turbo in CS-4). The one accelerator partnership is AMD's: 'AMD Helios and the Cerebras Wafer-Scale Engine will operate as a single disaggregated inference workflow', with Helios in Cerebras's data centers and the joint solution 'expected to be available first through Cerebras Cloud in the second half of 2026'. NVIDIA GPUs appear only as the comparison in Cerebras's benchmarks and as the target of Model Zoo checkpoint conversion.
Cerebras is a silicon company that became a cloud. The CS-3 system (WSE-3 inside) is sold as hardware the enterprise racks and clusters as a Wafer-Scale Cluster, with the training documentation covering installation, administration, S3 checkpointing, and job restarts; CS-4, 'a revolutionary rack-scale solution' on the Nexus platform with three WSE-3 Turbo wafers per system and two-microsecond wafer-to-wafer links, has 'first CS-4 shipments begin this quarter' and is watch-listed. The same systems fill Cerebras's own data centers (a 165 MW site in Mikkeli, Finland with Compute Nordic; AMD Helios racks arriving in the second half of 2026), sold two ways: the Training Cloud by the hour or by the model ('Pay Per Hour ... we'll determine the time needed to train, fine-tune, and deploy your model'; 'Pay Per Model ... Let our AI experts design, train, and fine-tune'), and the Inference Cloud by the token on a shared tier or as dedicated endpoints ('a private, provisioned instance of the Cerebras Inference service reserved exclusively for your organization'). The buyer gets a second kind of AI silicon, on its own floor or on Cerebras's. The architect's concern is scope and the ownership line. Cerebras sells no fabric or storage; the cloud exposes no virtual machine, GPU, or wafer the enterprise administers, only jobs, endpoints, and tokens; the on-premises path is a cluster of a single vendor's systems with a Model Zoo that compiles for that silicon; and the developer documentation says nothing about the hardware, data residency, or compliance. Calibration: NVIDIA reads strong on silicon plus DGX systems plus a hosted cloud; CoreWeave strong on GPU compute the enterprise rents by the instance; Dell and HPE strong on systems; Groq is the closest shape and is scored later in this batch; Cohere and Mistral gap on rented capacity with no customer surface. Cerebras is NVIDIA's shape at a fraction of the scale: proprietary silicon in systems the enterprise can own, or a cloud where it rents the output. Strong: Layer 0 reads by the vendor's shape (the September 15, 2026 ruling), and Cerebras sells systems the enterprise buys and clusters, which puts it in the OEM column with NVIDIA and Dell; the reviewers' rule 4 argument (no fabric, no storage) counted the purpose line's functions, and the OEM line has never required all three.
Ceded at the wafer, on both paths. The Wafer-Scale Engine, the CS systems, the Cerebras Cloud, and the Model Zoo's compilation path are Cerebras's with no second implementer: Ceded, the NVIDIA silicon reading. The runtime call at this layer has two homes and one reading: on Cerebras Cloud, Cerebras places and schedules work across its wafers and the enterprise sees jobs, endpoints, and tokens; on a customer-owned CS-3 cluster, the enterprise administers a single-vendor system whose silicon, compiler, and fabric make the placement, the NVIDIA DGX and Dell reading. Vendor decides, visible (usage, metrics, the cluster's own management), not overridable, Ceded.
Ruled September 15, 2026: strong stands on the shape line (an OEM that sells the systems reads on what it sells); moderate had been argued on the purpose line's three functions. Nasdaq: CBRS since May 14, 2026 (a procurement fact, not a scoring one). Watch-listed with dates: CS-4 ('first CS-4 shipments begin this quarter', product page read September 10, 2026, Q3 2026); the AMD Helios joint solution on Cerebras Cloud (second half of 2026). Named, not scored: the Mikkeli data center (announced, no in-service date in the source set); Condor Galaxy (no mention in the current documentation). The hardware facts come from the product page and the training documentation; the inference documentation carries no hardware, residency, or compliance language. Public evidence that moves the cell: nothing upward from strong; a customer-administered cloud instance would change the authority reading.
Durable, governed data foundation — the Governance Catalog that Layer 2C queries
Nothing at this layer is a scored capability.
Cerebras sells no durable, governed data foundation; what it holds are service and runtime artifacts, some preview-only and time-limited. The Files API ('This feature is in Private Preview') stores batch inputs and outputs for seven days by default with a configurable expiration, and batch results are retained on the same terms (configurable for enterprise customers); uploaded fine-tuned weights (Private Preview) are staged from 'an S3 bucket with cross-account access to Cerebras' that the enterprise owns, after which Cerebras creates and tracks a model version for deployment; the Training Cloud's S3 checkpointing writes to 'any S3-compatible storage'; prompt caches live in Cerebras's memory for their time to live, 'ephemeral in memory and never persisted'; 'Your playground and API requests are never used to train models.' The buyer gets nothing to store here, and that's the design. The architect's concern is nil at this layer. Calibration: OpenAI, Anthropic, Mistral, and Cohere read gap on trust apparatus without a data foundation; CoreWeave moderate on object storage it sells. Cerebras sells no storage. Gap, authority Absent.
Nothing offered, nothing inherited. Absent.
Named, not scored: the Files API (Private Preview; seven-day retention). Public evidence that moves the cell: a storage or governance product, which nothing describes.
Low-latency retrieval for RAG — vector/hybrid search, context windows
Nothing at this layer is a scored capability.
Cerebras serves language models and nothing that indexes. The public tier lists two chat models (gpt-oss-120b and qwen-3.8-27b) and no embedding, reranking, or search model; the documentation's retrieval content is the Milvus integration guide (Pinecone appears in an example, 'RAG with Pinecone + Docker'), where the vector store is the partner's and Cerebras supplies the generation step; the 'build your own Perplexity' and grounded-research cookbooks use Exa and Parallel for search, and the Gist Memory cookbook (summarizing and searching long documents with a read agent) is a customer-code pattern, not a Cerebras retrieval service. The buyer brings its own index. The architect's concern is nil at this layer. Calibration: OpenAI moderate on a hosted vector store and embeddings; Mistral moderate on embeddings and libraries; Cohere moderate on frontier embed and rerank models; Anthropic moderate on federated search; Cerebras offers none of those. Gap, authority Absent.
Nothing offered, nothing inherited. Absent.
Public evidence that moves the cell: an embedding or reranking model on the API, or a retrieval service; nothing in the documentation describes one.
Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering
Open code the enterprise runs and modifies; the hosted run on the Training Cloud is Cerebras's and is named with the 2B training chip. Retained, the library reading.
The Model Zoo pipeline runs on the enterprise's own CPUs or on Cerebras's Training Cloud.
Cerebras shapes training data and moves nothing else. The Model Zoo's data preparation section (a preprocessing quick start, a four-stage deduplication pipeline, read hooks, token generators, on-the-fly processing, and dataloaders, Apache 2.0) is a library the enterprise runs on its own machines or on the Training Cloud to turn its corpus into the tokenized, packed, deduplicated inputs the Cerebras trainer takes; the Batch API ('This feature is in Private Preview'; a 24-hour window, 50,000 requests or 200 MB per file, ten concurrent batches, chat completions only, 'SDK support for Batch is not yet available during Private Preview') is bulk inference, not a pipeline, and it's pre-GA. The buyer gets a documented, open toolkit for one destination and bulk inference in preview. The architect's concern is the destination: every stage ends in Cerebras's trainer, there's no connector, movement, lineage, or cache-tiering product, and the second requirement (moving the enterprise's data anywhere else) brings another tool. Calibration: NetApp reads moderate on AIDE, a fixed pipeline ending in its own retrieval surface; Hugging Face moderate on a hosted runner plus open pipeline libraries; Nebius moderate on a transfer service beside Data Lab; OpenAI, Anthropic, Mistral, Cohere, and Groq gap on intake without transformation or movement. A captive-destination transform is rule 4's own example of a fixed-function slice, and the September 11, 2026 ruling reads trainer-input preparation on that line rather than carving it out. Moderate.
Retained at the library, Ceded at the hosted run. The Model Zoo pipeline is Apache 2.0 code the enterprise runs and modifies anywhere: Retained, the Distilabel and FAISS library reading. The Training Cloud run of it is Cerebras's: Ceded, named with the 2B training chip. The runtime call is the pipeline executing the stages the enterprise configured over its own corpus, with no judgment of the engine's own: code decides, visible, overridable, Retained, the Hugging Face and Elastic 1C reading.
Ruled September 11, 2026: gap to moderate on rule 4 as written (a captive-destination transform is a fixed-function slice, the NetApp AIDE precedent); both reviewer passes had voted moderate and the first draft's stricter line (trainer-input preparation as 2B's business) isn't in the methodology. Named, not scored: the Batch API (Private Preview). Public evidence that moves the cell: a movement, connector, or lineage product.
GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization
Nothing at this layer is a scored capability.
Cerebras orchestrates its own wafers and shows the enterprise a meter. The console gives organizations and projects (Org Admin, Project Admin, Project Member; 'On Free Trial accounts, projects aren't available'), two-level rate limits in a dual-bucket model ('every organization has both an uncached token limit and a total token limit'), usage and cost analytics ('may be delayed by up to 10 minutes'), request and audit logs, and a Metrics API in Prometheus format 'designed for customers on dedicated endpoints' and enabled 'on an opt-in basis'. Dedicated endpoints are 'a private, provisioned instance ... reserved exclusively for your organization' whose sizing, replicas, and scaling aren't documented; service tiers (priority, default, auto, flex) are 'in Private Preview' and dedicated-only; the Training Cloud promises 'No DevOps Required' with preconfigured environments Cerebras runs. The buyer gets quotas and a bill. The architect's concern is that there's no plane: no instance to size, no scaling rule to set, no scheduler to see. Cerebras decides where and when work runs on its wafers, and the enterprise's controls are limits and budgets. Calibration: Mistral and OpenAI read gap on a metering edge with no customer plane; Anthropic gap on token quotas; Cohere moderate on Model Vault's documented capacity controls (fixed and flex capacity, replica bounds, autoscaling, pause and resume); CoreWeave strong on a real orchestrator. Cerebras is Mistral's shape: a provisioned instance without the controls that made Cohere moderate. Gap, authority Ceded: Cerebras absorbs the scheduling beneath its service, the Mistral, OpenAI, and Anthropic reading.
Nothing offered as a plane, and the scheduling inherited. Cerebras places and schedules work across its wafers invisibly; the enterprise sets limits and budgets on a meter, bounds rather than overrides: vendor decides, not visible, not overridable, Ceded, the Mistral, OpenAI, and Anthropic 2A reading.
Named, not scored: service tiers (Private Preview), the Metrics API (opt-in, dedicated only). Public evidence that moves the cell: documented capacity controls on dedicated endpoints (replicas, autoscaling, pause) would argue moderate on the Cohere reading.
Model serving, agent execution, inference APIs, distributed inference
Open-weight models their owners publish, behind a multi-vendor interface the enterprise's client already speaks. Delegated, the inference-interface reading.
Tool logic the enterprise writes and runs in its own process. Retained, the customer-tools ruling.
Cerebras's capacity and compiler path, which run nowhere else; the weights that come out are the enterprise's. Ceded, the NVIDIA DGX Cloud reading, with the artifact exit named. Custom weights upload to dedicated endpoints is Private Preview and carried in notes.
Inference runs on Cerebras's Wafer-Scale Engine with 44 GB of static random-access memory (SRAM) per wafer holding weights on-chip; the Model Zoo converts checkpoints 'for GPUs' and from Hugging Face formats, so trained weights leave for NVIDIA hardware, but nothing here runs on it.
Cerebras's runtime is the reason the company exists. The Inference API serves open-weight models from other labs ('We host a variety of open-source models from the community... All of our public models are unpruned'): gpt-oss-120b and qwen-3.8-27b on the public shared tier (gemma-4-31b left the public endpoints on September 3, 2026 and stays on dedicated ones; kimi-k2.7-code is customer-trial only), and a long dedicated-endpoint catalogue (Qwen3 up to 235B, Llama 4, Mistral Large 3, GLM 5.1, Kimi K2.7, DeepSeek V3.2, MiniMax, StepFun) 'as well as your own custom weights'. The API is OpenAI-compatible ('changing the API key, base URL, and model ID'), with tool use (strict mode, tool_choice, parallel calls on the models whose metadata allows it; gpt-oss-120b's doesn't), structured outputs with constrained decoding ('strict: true ... Recommended for production use'), reasoning effort controls, streaming (multiple tokens per event), automatic prompt caching at no fee with zero data retention, image inputs (Public Preview), payload optimization, and API version 2 (default since July 22, 2026). Dedicated endpoints are provisioned Cerebras capacity; the Customer Management API ('This feature is in Private Preview') uploads 'custom variants of Cerebras-supported model architectures' from the enterprise's S3 bucket and deploys them 'to a dedicated endpoint running a model with the same underlying architecture'; the Training Cloud trains and fine-tunes 'models of any size' with the Apache 2.0 Model Zoo and converts checkpoints to and from Hugging Face formats and for GPUs. Around it: Hugging Face Inference Providers, OpenRouter, AWS Marketplace, Vercel, and forty-odd integration guides (LangChain, LlamaIndex, CrewAI, LiteLLM, Portkey, Kong, Cloudflare AI Gateway); and OpenAI's own API offers 'Ultrafast Mode ... powered by Cerebras' for GPT-5.6 Sol in a limited preview since August 13, 2026 (named, not scored). The buyer gets open models answering faster than anywhere else, through a client it already has. The architect's concern is breadth. Cerebras owns no model; the public tier carries two; the agent layer is other people's frameworks (no hosted agent runtime, no Model Context Protocol (MCP) in any cookbook); the enterprise's own weights are a Private Preview on a fixed list of architectures; batch is Private Preview; and the fastest path to the fastest frontier model runs through OpenAI's paper, not Cerebras's. Calibration: OpenAI, Anthropic, Mistral, and Cohere read strong on frontier serving plus open-source harnesses; CoreWeave strong on open serving plus sandboxes; Cloudflare strong on serverless open-model inference plus a gateway plus agents; Hugging Face moderate on open serving under maturity gates with an experimental agent library. Cerebras is serving alone, at the frontier of speed (rule 6, on Cerebras's own published benchmarks and third-party comparisons). Strong. Agent execution is judged by whether the runtime carries the enterprise's own agent loop (tool calling, structured outputs, reasoning controls, latency over the serial turns an agent takes), not by whether the vendor hosts one (the September 11, 2026 ruling under rule 4), and Cerebras carries it on its own engine at a documented frontier; both reviewer passes voted moderate for calibration with Hugging Face, whose moderate rests on rule 5, not on a missing loop.
Delegated at the interface, Retained at the tools, Ceded at the wafer. Model access through the OpenAI-compatible API is the inference-interface reading: the enterprise's client code and prompts point anywhere, and the models are open weights their owners publish: Delegated. Tools the enterprise writes and calls client-side are its own, the customer-tools reading: Retained. Dedicated endpoints, the Training Cloud, and the Model Zoo's compiled artifacts run only on Cerebras silicon: Ceded, though trained weights export to Hugging Face and GPU formats and are the enterprise's. The decision to call a tool is the model's, but the loop runs in the enterprise's own process, where a deterministic gate before any effect is the enterprise's to write: model decides, visible, overridable, Delegated, the OpenAI and Anthropic reading.
Ruled September 11, 2026: strong stands (agent execution at 2B is the runtime carrying the enterprise's own loop, not a hosted loop; a hosted runtime is a component, never the price of strong). Both reviewer passes had voted moderate. Badges: Batch API, Files API, Customer Management API (custom weights), service tiers, and the queue_threshold header are 'Private Preview' ('Invite only... No SLA'); image inputs and predicted outputs are 'Public Preview'; 'Preview models are intended for evaluation purposes only'; all held out of scored chips. OpenAI's Ultrafast Mode ('available in a limited preview today to a select group of customers', August 13, 2026) is on OpenAI's paper and is named, not scored, and removed from the grade rationale after the ChatGPT pass. Public Models API metadata lists parallel_tool_calls false for gpt-oss-120b. Trained weights: the Model Zoo is Apache 2.0 and converts checkpoints to Hugging Face and GPU formats, so fine-tunes on the Training Cloud are the enterprise's artifact (the Retained reading) once exported; the Training Cloud itself has no customer-facing documentation beyond the product page and is named inside the dedicated chip. Public evidence that moves the cell: a hosted agent runtime would settle the grade; the custom-weights API reaching general availability would add a Delegated chip for the enterprise's own models on dedicated endpoints.
Policy-driven placement and resource coordination — the Autonomy Layer
Nothing at this layer is a scored capability.
Read against the five legs, Cerebras has none. Identity: organization and project roles and API keys identify callers to the API; no agent principal. Gateway: none of Cerebras's own in production; the integration guides route through partners' gateways (Cloudflare AI Gateway, Portkey, Kong, TrueFoundry AI Gateway with observability, cost tracking, access control, and rate limiting, Operant as a security gateway, among others), and the Cerebras Code MCP Server, an open-source server for coding assistants, is 'in research preview'. Registry and orchestration: none. Observability: request and audit logs and a Metrics API for dedicated endpoints watch the API, not agents. The buyer brings the plane. The architect's concern is nil at this layer: there's nothing to be captured by. Calibration: OpenAI, Anthropic, Mistral, and Cohere read gap; Cloudflare moderate on a gateway of its own. Cerebras is the OpenAI shape without the frontier-pre-GA caveat. Gap, authority Absent.
Nothing offered as a plane. Absent.
Named, not scored: the Cerebras Code MCP Server (research preview). Public evidence that moves the cell: a Cerebras gateway or registry over agents; nothing in the documentation describes one.
AI-powered business capabilities — business logic, workflow automation
Nothing at this layer is a scored capability.
Cerebras ships no application. The console's Playground is a developer's chat surface ('never used to train models'); the cookbooks are recipes for agents built on partners' frameworks; the case studies (Cognition, Lovable, Tavus, GSK, DeepLearning.AI) are the customers' products; and the most visible application of Cerebras inference, OpenAI's Ultrafast Mode, is OpenAI's. The buyer gets a console (Playground, projects, usage) and no first-party business application or workflow surface. The architect's concern is nil at this layer. Calibration: OpenAI, Anthropic, Mistral, and Cohere read strong on first-party applications; CoreWeave moderate on a developer plane; Cerebras has neither. Gap, authority Absent.
Nothing offered, nothing inherited. Absent.
Public evidence that moves the cell: a first-party application, which nothing describes.
Cerebras Inference documentation at inference-docs.cerebras.ai (API reference: chat completions, completions, batch, files, models, metrics, customer management API, versions; capabilities: tool use, structured outputs, reasoning, streaming, prompt caching, image inputs, payload optimization, service tiers, batch, metrics; console: overview, API keys, projects, usage monitoring, account and billing, playground; dedicated endpoints overview; models overview and model pages; support: rate limits, pricing, deprecations, preview releases, change log; resources: OpenAI compatibility, designing for Cerebras; integrations and cookbooks; llms.txt); Cerebras training documentation at training-docs.cerebras.ai (release 2.5.0: getting started, Model Zoo overview, the Wafer-Scale Cluster and CS-3 systems); GitHub license API for Cerebras/modelzoo (Apache 2.0); cerebras.ai product pages (Inference, Cloud, System, Pricing, Partners) and posts (AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution, July 23, 2026; Accelerating GPT-5.6 Sol Ultrafast, August 13, 2026; In the News, including the May 14, 2026 Nasdaq listing). Peer-reviewed cell by cell through the labs claims ledger (cerebras-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (cerebras-row-agy); totals and escalated items in reviews/cerebras-judgment.md.