# Groq (GroqCloud on the LPU + Compound + Server-Side Web Retrieval; Independent Under a Non-Exclusive NVIDIA Technology License) — 4+1 Layer AI Infrastructure Assessment

> Mapped to the 4+1 Layer AI Infrastructure Model  
> Version: v1.0 - 4+1 v2: Authority Split · Date: September 11, 2026  
> Source: GroqCloud documentation at console.groq.com (models; OpenAI compatibility; Responses API; rate limits; text generation; speech to text; text to speech and Orpheus; vision; reasoning; content moderation; structured outputs; prompt caching; tool use overview, built-in tools (web search, visit website, code execution, Wolfram Alpha, browser search, browser automation), remote MCP and connectors, local tool calling; integrations catalogue; coding with Groq; Compound overview, built-in tools, systems; service tiers, performance tier, flex processing, batch processing; LoRA inference; production readiness; security onboarding and Google Cloud Private Service Connect; your data and the feedback policy; examples; Prometheus metrics; spend limits, projects, model permissions, billing FAQs, your data; SDK libraries; changelog through April 18, 2026; legacy changelog; deprecations; llms.txt); groq.com pages (platform, pricing, GroqRack, about us, newsroom index) and posts (Groq and Nvidia Enter Non-Exclusive Inference Technology Licensing Agreement, December 24, 2025; GroqCloud: Expanding to Meet Demand, February 16, 2026; Groq Raises $650M, June 22, 2026; Groq Becomes an NVIDIA Cloud Partner, August 12, 2026; Groq Closes $350 million Series A, August 17, 2026; Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market, August 24, 2026). Peer-reviewed cell by cell through the labs claims ledger (groq-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (groq-row-agy); totals and escalated items in reviews/groq-judgment.md.  
> Published by: The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com  
> Author: Keith Townsend

[Full interactive assessment](https://layer2c.com/assessment/groq) · [Methodology](https://layer2c.com/methodology) · [What Is Layer 2C?](https://layer2c.com/what-is-layer-2c)

## Executive Summary

Groq is the Language Processing Unit (LPU) inference cloud that licensed its technology to NVIDIA (non-exclusive, December 24, 2025) and kept operating: its founder and president left for NVIDIA, Simon Edwards became chief executive, GroqCloud 'will continue to operate without interruption', and the company has since raised $650 million and a $350 million round at a $3.5 billion valuation, become an NVIDIA Cloud Partner, and announced it will run NVIDIA's Groq 3 LPX with Vera Rubin NVL72 beside its own silicon. The map reads it strong at one layer, moderate at one, and gap at six. Layer 2B is the strength: open models (gpt-oss, Llama, Whisper) behind an OpenAI-compatible API with constrained-decoding structured outputs, reasoning, speech to text, and prompt caching, plus Compound, a production agentic system over a fixed harness of server-side web search, browsing, code execution, and Wolfram Alpha that takes no custom tools, and Low-Rank Adaptation (LoRA) inference of the enterprise's own adapters; text to speech and vision are preview, and the Responses API and every Model Context Protocol (MCP) path are beta. Layer 1B is moderate on the Perplexity reading: Compound's web retrieval on Tavily and Exa returns search results and citations per request, over the web and not the enterprise's data. Layer 0 is a gap by shape: Groq's own silicon in thirteen partner-operated data centers is sold only as tokens, sovereign endpoints, and a dedicated instance for LoRA serving, with GroqMetal marketed and GroqRack undocumented, so the enterprise decides nothing there. Layer 2A is a gap on the OpenAI reading: service tiers, provisioned throughput, per-project limits, spend caps, and model permissions are a metering edge, not a plane. Everything else is a gap: no data foundation, no pipelines, no governance plane, no business application.

The capture is at the LPU, which the enterprise never touches, and at Compound. The API is OpenAI-compatible and the models are other labs' open weights, so the client and the model choice lift; local tools are the enterprise's; LoRA adapters are the enterprise's standard files served behind the same door. What is Groq's and stays Groq's: the silicon and its cloud, the service tiers and account controls, Compound and the built-in tools that execute on Groq's side, and the web retrieval that is Tavily's and Exa's through Groq's paper. Five components read Ceded, Delegated, Retained, Ceded, Delegated across 1B and 2B; today's runtime has no NVIDIA dependency, and the row's NVIDIA note is a roadmap and a license rather than a chip.

The buyer's trade: LPU speed on open models through an interface it already uses, with a hosted agent that searches the web and runs code, in exchange for a silicon vendor it can't administer, a production catalogue of six models with speech out and vision still in preview, beta badges on the Responses API and every MCP path, and an agent whose tool calls run on Groq's side behind a per-request allow-list and nothing else, and don't run in the sovereign regions. The decision-authority readings follow: vendor / Ceded at Layer 0 and 2A (Groq absorbs both beneath its service; a tier is a menu, a limit is a bound; nothing below the endpoint is visible), model / Ceded at 1B (Compound decides when to search), model / Delegated at 2B on the primary path (the enterprise's loop with local tools; Compound's path is Ceded), Absent elsewhere. What would move cells: documentation for GroqMetal, a GroqRack datasheet, or a sized dedicated instance (Layer 0), a production text-to-speech or vision model or custom tools in Compound (2B chips), remote MCP leaving beta with its approval flow (a 2B chip and a 2C fragment), an instance the enterprise sizes (2A), and the NVIDIA LPX capacity coming online (an NVIDIA dependency at Layer 0).

## Layer Status

| Layer | Status | Classification |
|---|---|---|
| Layer 0 · Compute | ○ Not a Customer Surface: Groq's Own LPU Silicon in Thirteen Partner-Operated Data Centers, Sold Only as Tokens and Regional Sovereign Endpoints; a Dedicated Instance Documented Only as a LoRA Deployment Mode; GroqMetal and GroqRack Marketed Without Documentation | Compute & Network Fabric |
| Layer 1A · Storage | ○ No Data Foundation: No Retention by Default, Zero Data Retention on Request, Batch Files Kept Thirty Days in US Google Cloud Buckets; Feedback Submissions Kept up to Three Years | Data Storage & Governance |
| Layer 1B · Retrieval | ◑ Server-Side Web Retrieval in Compound and the Built-In Tools: Web Search on Tavily, Browser Search on Exa, and Visit Website, With Search Results and Citations Returned per Request; No Embedding or Reranking Model, No Index Over the Enterprise's Data | Context Management & Retrieval |
| Layer 1C · Pipelines | ○ No Pipelines: a Batch API Queues Bulk Requests at Half Price Inside a Processing Window of a Day to a Week | Data Movement & Pipelines |
| Layer 2A · Orchestration | ○ A Metering Edge, Not a Plane: On-Demand, Flex, Performance (Provisioned Throughput With a 99.9 Percent SLA), Auto, and Batch Tiers; Per-Project Limits, Spend Caps, Model Permissions; No Instance You Size | Infrastructure Orchestration |
| Layer 2B · Runtime | ● The LPU Inference Cloud: Open Models Behind an OpenAI-Compatible API With Strict Structured Outputs, Reasoning, Whisper Speech to Text, and Prompt Caching, Plus Compound, a Production Agentic System With a Fixed Set of Server-Side Tools; Text to Speech and Vision in Preview; Remote MCP and Connectors in Beta; LoRA Inference on Shared or Dedicated Hardware | Application Runtime & Execution |
| Layer 2C · Reasoning | ○ Model Permissions, Project Limits, and Roles Govern Accounts, Not Agents; Compound's Server-Side Tools Take a Per-Request Allow-List and Nothing Else; Remote MCP Approvals and Tool Filters Are Beta | Agentic Infrastructure — The Reasoning Plane |
| Layer 3 (+1) · Applications | ○ No Business Application: Groq Chat Is a Demonstration Surface, the Examples Page Holds Groq-Authored Templates and Demos, and the Coding Tools (Factory Droid, OpenCode, Kilo Code, Roo Code, Cline) Are Third-Party and Open-Source With a Groq Key | AI Application Layer — The Value Plane |

## DAPM Portability Profile (components)

| Classification | Count | Meaning |
|---|---|---|
| Retained | 1 | I possess the capability and can operate it independently of this provider |
| Delegated | 2 | Someone else provides the capability, but I can substitute that provider without reconstructing my accumulated opinions |
| Ceded | 2 | Changing providers requires reconstructing those opinions |

**Decision authority (per layer, gaps included)**

| Reading | Layers | Meaning |
|---|---|---|
| Retained | 0 | The enterprise, or code it writes or controls, decides |
| Delegated | 1 | Vendor or model decides; the enterprise can see and override |
| Ceded | 3 | Vendor or model decides; no override, often invisible |
| Absent | 4 | Nothing offered, nothing inherited |

## Strongest Layers

- **Layer 2B** (Application Runtime & Execution) — The LPU Inference Cloud: Open Models Behind an OpenAI-Compatible API With Strict Structured Outputs, Reasoning, Whisper Speech to Text, and Prompt Caching, Plus Compound, a Production Agentic System With a Fixed Set of Server-Side Tools; Text to Speech and Vision in Preview; Remote MCP and Connectors in Beta; LoRA Inference on Shared or Dedicated Hardware

## Gap Areas

- **Layer 0** (Compute & Network Fabric) — Not a Customer Surface: Groq's Own LPU Silicon in Thirteen Partner-Operated Data Centers, Sold Only as Tokens and Regional Sovereign Endpoints; a Dedicated Instance Documented Only as a LoRA Deployment Mode; GroqMetal and GroqRack Marketed Without Documentation
- **Layer 1A** (Data Storage & Governance) — No Data Foundation: No Retention by Default, Zero Data Retention on Request, Batch Files Kept Thirty Days in US Google Cloud Buckets; Feedback Submissions Kept up to Three Years
- **Layer 1C** (Data Movement & Pipelines) — No Pipelines: a Batch API Queues Bulk Requests at Half Price Inside a Processing Window of a Day to a Week
- **Layer 2A** (Infrastructure Orchestration) — A Metering Edge, Not a Plane: On-Demand, Flex, Performance (Provisioned Throughput With a 99.9 Percent SLA), Auto, and Batch Tiers; Per-Project Limits, Spend Caps, Model Permissions; No Instance You Size
- **Layer 2C** (Agentic Infrastructure — The Reasoning Plane) — Model Permissions, Project Limits, and Roles Govern Accounts, Not Agents; Compound's Server-Side Tools Take a Per-Request Allow-List and Nothing Else; Remote MCP Approvals and Tool Filters Are Beta
- **Layer 3 (+1)** (AI Application Layer — The Value Plane) — No Business Application: Groq Chat Is a Demonstration Surface, the Examples Page Holds Groq-Authored Templates and Demos, and the Coding Tools (Factory Droid, OpenCode, Kilo Code, Roo Code, Cline) Are Third-Party and Open-Source With a Groq Key

## Layer-by-Layer Detail

### ○ Layer 0 · Compute: Compute & Network Fabric

*Raw compute, networking, and acceleration fabric*  
**Status:** Not a Customer Surface: Groq's Own LPU Silicon in Thirteen Partner-Operated Data Centers, Sold Only as Tokens and Regional Sovereign Endpoints; a Dedicated Instance Documented Only as a LoRA Deployment Mode; GroqMetal and GroqRack Marketed Without Documentation

**Decision authority:** Ceded (decides: vendor; visible: false; overridable: false; boundary: vendor)

**Gap Analysis:** Groq owns the chip, operates capacity in partner data centers, and sells tokens. GroqCloud runs on Groq's Language Processing Units in thirteen data centers 'across North America, Europe, the Middle East and APAC' (Helsinki, a UK site with Equinix, Sydney, Saudi Arabia with HUMAIN, Bell Canada's network, and the United States), 'scaling toward 200 MW by 2027' after a $650 million raise (June 2026) and a $350 million round led by Disruptive with planned NVIDIA participation (August 2026, at a $3.5 billion valuation); the enterprise reaches it through api.groq.com, or, on the Enterprise plan, through regional and sovereign endpoints exposed as Google Cloud Private Service Connect published services (me-central2 and us-central1; api.me-central-1.groqcloud.com, api.us.groqcloud.com) and, for LoRA inference, on 'dedicated Groq hardware instances purchased by the customer'. GroqRack, the on-premises rack, has no datasheet or documentation in the public set; the page for it carries only the company's LPX message. The platform page markets three tiers, 'GroqMetal provides infrastructure, GroqCore adds inference, and GroqAssured adds enterprise controls', with GroqMetal as 'Dedicated bare-metal infrastructure ... with full control in your hands', and no documentation, specification, or price behind any of them. The buyer gets a second inference silicon by the token, close to its region.

The architect's concern is that nothing here is administered by the enterprise. No instance it sizes (the LoRA dedicated instance is bought through sales with no documented sizing or scheduling surface), no bare metal it can document (GroqMetal is a marketing tier), no rack it can document, no fabric; the sovereign endpoints are private doors into Groq's capacity, Compound and the built-in tools are 'not available currently for use with regional / sovereign endpoints', and the next generation of the silicon (Groq 3 LPX) is an NVIDIA product Groq will rent like anyone else.

Calibration: Layer 0 reads by the vendor's shape (the September 15, 2026 ruling under the exposure test). Groq is a service: it sells tokens and sovereign endpoints, and the LPU behind them is Groq's supply chain, which is Anthropic's cell ('compute supply chain, not customer surface') and Perplexity's ('serving fleet, not a customer surface'); Cerebras and NVIDIA read strong because they sell systems the enterprise buys, not because they own the silicon, and owning it behind a service changes nothing, as Google's TPUs behind the Gemini API change nothing. There's nothing for the enterprise to decide at this layer. Gap, the service reading; a Stack Builder wouldn't build a Layer 0 on it.

**Borrowed Judgment:** Ceded and invisible: Groq absorbs the layer beneath its service. The runtime call at this layer is Groq placing and scheduling work across its LPUs; the enterprise chooses a region through the endpoint it calls and sees nothing below it (usage and Prometheus metrics are consumption, not placement): vendor decides, not visible, not overridable, Ceded, the OpenAI, Anthropic, and Perplexity reading.

### ○ Layer 1A · Storage: Data Storage & Governance

*Durable, governed data foundation — the Governance Catalog that Layer 2C queries*  
**Status:** No Data Foundation: No Retention by Default, Zero Data Retention on Request, Batch Files Kept Thirty Days in US Google Cloud Buckets; Feedback Submissions Kept up to Three Years

**Decision authority:** Absent (decides: absent; visible: n/a; overridable: n/a; boundary: vendor)

**Gap Analysis:** Groq keeps as little as it can on the API path. 'By default, Groq does not retain customer data for inference requests'; 'All customers may enable Zero Data Retention (ZDR) in Data Controls settings'; what the API retains (batch input and output files for thirty days unless deleted earlier, LoRA adapters and training datasets until deleted, reliability and abuse logs up to thirty days) sits in 'Google Cloud Platform (GCP) buckets located in the United States' with standard contractual clauses for transfers; voluntary feedback submissions are governed separately, and 'reviewed feedback, conversation snippets, and related metadata are stored for up to 3 years'. The buyer gets nothing to store here.

The architect has no Groq-managed data foundation to evaluate; retention and US residency of what is retained are trust controls to account for, not Layer 1A capability.

Calibration: OpenAI, Anthropic, Mistral, Cohere, and Cerebras read gap on trust apparatus without a data foundation. Gap, authority Absent.

**Borrowed Judgment:** Nothing offered, nothing inherited. Absent.

### ◑ Layer 1B · Retrieval: Context Management & Retrieval

*Low-latency retrieval for RAG — vector/hybrid search, context windows*  
**Status:** Server-Side Web Retrieval in Compound and the Built-In Tools: Web Search on Tavily, Browser Search on Exa, and Visit Website, With Search Results and Citations Returned per Request; No Embedding or Reranking Model, No Index Over the Enterprise's Data

**Decision authority:** Ceded (decides: model; visible: true; overridable: false; boundary: model)

**Server-Side Web Retrieval (Web Search on Tavily, Browser Search on Exa, Visit Website; in Compound and Compound Mini and as Built-In Tools on gpt-oss Models; Search Results, Relevance Scores, and Citations in executed_tools; Domain and Country Settings)** [DAPM: Ceded]  
Partners' search engines through Groq's harness on Groq's servers. Ceded, the channel ruling and the Perplexity Search API reading.

**Gap Analysis:** Groq retrieves from the web, not from the enterprise. Compound and Compound Mini ('Production Systems') and the built-in tools on gpt-oss models run web search ('powered by Tavily'), browser search on Exa, and visit website on Groq's side, 'automatically enabled by default', with domain and country settings, and return what they found in executed_tools (search results with relevance scores, page content, citations); the catalogue has no embedding or reranking model; the Google Workspace connectors (Gmail, Calendar, Drive search and fetch) are remote MCP connectors 'currently in beta', and the Hugging Face integration is a remote MCP server Hugging Face hosts. The buyer gets live web retrieval inside an inference call, and brings its own index for everything else.

The architect's concern is that the retrieval is partners' engines through Groq's paper, scoped to the public web, on Groq's tool-selection judgment, and 'not available currently for use with regional / sovereign endpoints'.

Calibration: Perplexity reads moderate on a web Search API exposed raw; Anthropic moderate on federated search over connected systems; OpenAI moderate on file search and vector stores; Cerebras gap with partner vector stores in integration guides only. Groq is Perplexity's shape inside the model call, on someone else's index. Moderate, on the Perplexity reading.

**Borrowed Judgment:** Ceded at the tool. The engines are Tavily's and Exa's and the harness is Groq's; nothing lifts but the query: Ceded, the channel ruling (partners' engines through the vendor's paper, the NIM reading) and the Perplexity Search API reading. The runtime call is Compound deciding whether and what to search ('intelligently decides when to use each tool'), visible in executed_tools, with enabled_tools as an allow-list, a bound rather than an override: model decides, visible, not overridable, Ceded, the Perplexity reading.

### ○ Layer 1C · Pipelines: Data Movement & Pipelines

*Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering*  
**Status:** No Pipelines: a Batch API Queues Bulk Requests at Half Price Inside a Processing Window of a Day to a Week

**Decision authority:** Absent (decides: absent; visible: n/a; overridable: n/a; boundary: vendor)

**Gap Analysis:** Groq moves tokens, not data. Batch processing ('50% lower cost, no impact to your standard rate limits, and 24-hour to 7 day processing window', Developer tier and up; 'we'll process as many requests as our capacity allows', and a batch that doesn't finish expires) is bulk inference, not a pipeline. The buyer gets nothing here.

The architect's concern is nil at this layer.

Calibration: OpenAI, Anthropic, Mistral, Cohere, and Cerebras read gap. Gap, authority Absent.

**Borrowed Judgment:** Nothing offered, nothing inherited. Absent.

### ○ Layer 2A · Orchestration: Infrastructure Orchestration

*GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization*  
**Status:** A Metering Edge, Not a Plane: On-Demand, Flex, Performance (Provisioned Throughput With a 99.9 Percent SLA), Auto, and Batch Tiers; Per-Project Limits, Spend Caps, Model Permissions; No Instance You Size

**Decision authority:** Ceded (decides: vendor; visible: false; overridable: false; boundary: vendor)

**Gap Analysis:** Groq gives the enterprise a menu of how its requests get scheduled, and bounds around them. The service_tier parameter picks on-demand (the default), flex ('available for all models to paid customers only with 10x higher rate limits'; 'If flex capacity is unavailable, requests will fail quickly with status 498'), performance ('only available on enterprise plans'; 'a 99.9% availability SLA and a 99% latency guarantee'; 'delivered as provisioned throughput: you purchase input and output capacity bundles'), or auto ('leverage the best tier available to you at any given moment'); batch runs asynchronously outside the tiers. Around them: rate limits enforced 'at the organization level' by requests, tokens, and audio seconds, with cached tokens exempt; projects with their own API keys and limits 'must be ≤ org limit'; an organization-wide spend cap on paid tiers; model permissions ('Only Allow' or 'Only Block' at organization and project level); Owner, Developer, and Reader roles; Prometheus metrics for enterprise customers. The LoRA dedicated instance is 'purchased by the customer' with 'No LoRA-specific rate limiting' and no documented sizing, replica, or scheduling surface. The buyer gets priority it can pay for and budgets it can set.

The architect's concern is that it's a queue, not a plane: no instance the enterprise sizes, no replica count, no scheduler it sees; the performance tier is provisioned throughput on Groq's terms, and the rest is bounds.

Calibration: OpenAI reads gap on the same shape (priority and flex service tiers, Scale Tier and provisioned-throughput commitments, rate limits, per-project quota); Mistral and Cerebras gap on a metering edge; Cohere moderate on Model Vault, where the customer sets instance counts and replica bounds; Databricks and Snowflake moderate on managed compute with bounds. Groq's tiers are generally available and contracted where Cerebras's are preview, and that changes nothing at this layer: GA of a metering edge is still a metering edge. Gap, on the OpenAI reading.

**Borrowed Judgment:** Groq absorbs the layer beneath its service: its scheduler places every request inside the tier the enterprise named and the limits it set, and exposes no placement decision. Vendor decides, not visible, not overridable, Ceded, the OpenAI and Cerebras absorbed-layer reading; a tier is a menu and a limit is a bound.

### ● Layer 2B · Runtime: Application Runtime & Execution

*Model serving, agent execution, inference APIs, distributed inference*  
**Status:** The LPU Inference Cloud: Open Models Behind an OpenAI-Compatible API With Strict Structured Outputs, Reasoning, Whisper Speech to Text, and Prompt Caching, Plus Compound, a Production Agentic System With a Fixed Set of Server-Side Tools; Text to Speech and Vision in Preview; Remote MCP and Connectors in Beta; LoRA Inference on Shared or Dedicated Hardware

**Decision authority:** Delegated (decides: model; visible: true; overridable: true; boundary: model)

**Model Serving via the OpenAI-Compatible API (Chat Completions; gpt-oss 120B and 20B, Llama 3.1 8B and 3.3 70B on Enterprise, Whisper Large V3 and Turbo Speech to Text; Structured Outputs With Constrained Decoding; Reasoning Controls; Automatic Prompt Caching)** [DAPM: Delegated]  
Open-weight models their owners publish, behind a multi-vendor interface on Groq's silicon. Delegated, the inference-interface ruling; text to speech and vision are preview and unscored.

**Customer-Created and Managed Tools (Local Tool Calling in the Enterprise's Process; Parallel Calls on Supported Models)** [DAPM: Retained]  
Tool logic the enterprise writes and runs. Retained, the customer-tools ruling.

**Compound and Compound Mini (Production Agentic Systems Over gpt-oss-120b, Llama 4 Scout, and Llama 3.3 70B; a Fixed Harness of Web Search on Tavily, Visit Website, Code Execution in E2B, and Wolfram Alpha; Up to Ten Server-Side Tool Calls per Request for Compound, One for Compound Mini; No Custom Tools; Per-Request enabled_tools Allow-List)** [DAPM: Ceded]  
Groq's agent loop over Groq's tools on Groq's servers; the underlying models are their owners'. Ceded, the Perplexity Computer and Cohere North reading.

**LoRA Inference (Enterprise; Customer-Trained Adapters in Standard Parameter-Efficient Fine-Tuning (PEFT) Files Served Through the OpenAI-Compatible Chat Completions Path, on the Shared Cloud per Token or on a Dedicated Groq Hardware Instance the Customer Purchased; Groq Does No Fine-Tuning)** [DAPM: Delegated]  
The adapter is the enterprise's artifact and lifts to any PEFT-capable server; Groq serves it behind a multi-vendor interface. Delegated, the inference-interface ruling with the customer fine-tunes reading; the dedicated instance is a Layer 0 fact.

**Gap Analysis:** Groq's runtime is fast inference of other labs' models, and one agent of its own. The API is 'mostly compatible with OpenAI's client libraries' at api.groq.com/openai/v1 (no logprobs, logit_bias, or n>1): chat completions, a Responses API ('currently in beta'), speech to text (Whisper Large V3 and Turbo), structured outputs with constrained decoding on gpt-oss-20b, gpt-oss-120b, and qwen3.8-27b ('Streaming and tool use are not currently supported with Structured Outputs'), reasoning controls, and automatic no-fee prompt caching on the gpt-oss models. Production models are Llama 3.1 8B and 3.3 70B (Enterprise), gpt-oss 120B and 20B, and the Whisper pair; preview models (Orpheus text to speech, Prompt Guard, MiniMax M2.7, Qwen 3.6 and 3.8 with vision, the gpt-oss safeguard) are 'intended for evaluation purposes only', so speech out and vision are unscored; MiniMax M2.5 and Qwen3-VL 32B arrived for Enterprise customers in April 2026. Tool use runs three ways: local (the enterprise's functions, parallel calls on some models), remote MCP ('currently in beta'; 'Groq's servers will connect to the MCP server, discover the available tools, pass them to the model, and execute any tools'), and built-in server-side tools (web search on Tavily, browser search on Exa, visit website, Python code execution in E2B sandboxes, Wolfram Alpha with the enterprise's key; browser automation retired). Compound and Compound Mini are 'Production Systems': Groq-built agentic systems over gpt-oss-120b, Llama 4 Scout, and Llama 3.3 70B that 'solve problems by taking action', up to ten server-side tool calls per request (Compound Mini one), priced per tool, over a fixed set of four tools: 'Custom user-provided tools are not supported at this time.' LoRA inference (Enterprise only; 'We do not provide LoRA fine-tuning services') runs adapters on the shared cloud or on dedicated Groq hardware the customer purchased. The buyer gets open models at LPU speed, a hosted agent that browses and runs code, and its own adapters on Groq's silicon.

The architect's concern is what the speed is attached to. Groq owns no foundation model, the production list is six models plus two systems, the Responses API and every MCP path are beta, Compound and the built-in tools don't run on regional or sovereign endpoints, and Compound's tool calls execute on Groq's side with no documented gate before the effect.

Calibration: Cloudflare reads strong on serverless open-model inference plus a gateway plus agents; OpenAI and Anthropic strong on frontier serving plus hosted tools; Cerebras strong on open serving alone at the frontier of speed; Mistral and Cohere strong on frontier serving plus harnesses; Hugging Face moderate on open serving under maturity gates. Groq has serving on its own engine at the frontier of speed plus Compound, a production agentic system over a fixed harness of four server-side tools that takes no custom tools; local tool calling is the general path and it runs in the enterprise's process. Strong. Agent execution is judged by whether the runtime carries the enterprise's own agent loop (tool calling, structured outputs, reasoning controls, latency over the serial turns an agent takes), not by whether the vendor hosts one (the September 11, 2026 ruling under rule 4); Groq carries it on its own engine at the frontier of speed, and Compound is a hosted option that reads Ceded when used, never the price of the grade.

**Borrowed Judgment:** Delegated at the interface, Retained at the enterprise's tools, Ceded at Compound. Model access through the OpenAI-compatible API is the inference-interface reading: the client points anywhere and the models are their owners' open weights: Delegated. Local tools the enterprise writes and executes are its own: Retained, the customer-tools ruling. Compound and the built-in tools are Groq's surfaces on Groq's silicon: Ceded (Compound's underlying models are the owners'). LoRA adapters are the enterprise's standard files served behind the same OpenAI-compatible door: Delegated; the dedicated instance they may run on is a Layer 0 fact. On the primary path, chat completions with local tools, the model decides to call and the enterprise executes in its own process, where a deterministic gate is its to write: model decides, visible, overridable, Delegated, the OpenAI and Anthropic reading. On the Compound path the tool executes on Groq's side with no gate documented, the Perplexity reading, Ceded; the primary path governs the cell.

### ○ Layer 2C · Reasoning: Agentic Infrastructure — The Reasoning Plane

*Policy-driven placement and resource coordination — the Autonomy Layer*  
**Status:** Model Permissions, Project Limits, and Roles Govern Accounts, Not Agents; Compound's Server-Side Tools Take a Per-Request Allow-List and Nothing Else; Remote MCP Approvals and Tool Filters Are Beta

**Decision authority:** Absent (decides: absent; visible: n/a; overridable: n/a; boundary: vendor)

**Gap Analysis:** Read against the five legs, Groq has none as a plane. Identity: API keys, projects, and Owner, Developer, and Reader roles identify callers; no agent principal. Gateway: model permissions ('Only Allow' or 'Only Block' per organization and project, a 403 on a restricted model) are a policy on which models a key may call, and Compound takes a per-request allow-list (compound_custom.tools.enabled_tools 'restrict[s] or specif[ies] exactly which tools should be available for a particular request'); both are fragments, request-scoped and key-scoped, with no central policy, registry, or agent identity behind them. Remote MCP documents allowed_tools and require_approval ('never', 'always') with an approval flow, and remote MCP 'is currently in beta' ('MCP servers have access to all data in your AI model's context'). Registry and orchestration: none. Observability: executed_tools on a Compound response lists the tools called, their arguments, and results, per request; usage, spend tracking with a ten-to-fifteen-minute lag, and Prometheus metrics on Enterprise watch the API; nothing watches an estate of agents. The buyer brings the plane.

The architect's concern is Compound: an agent that searches, fetches pages, and runs code on Groq's side, governed by the caller's key and a per-request tool list and nothing else the documentation names.

Calibration: OpenAI, Anthropic, Cohere, Mistral, and Cerebras read gap; Cloudflare moderate on a gateway of its own. Groq is the OpenAI shape with a model-permission list and a tool allow-list. Gap, authority Absent.

**Borrowed Judgment:** Nothing offered as a plane. Absent.

### ○ Layer 3 (+1) · Applications: AI Application Layer — The Value Plane

*AI-powered business capabilities — business logic, workflow automation*  
**Status:** No Business Application: Groq Chat Is a Demonstration Surface, the Examples Page Holds Groq-Authored Templates and Demos, and the Coding Tools (Factory Droid, OpenCode, Kilo Code, Roo Code, Cline) Are Third-Party and Open-Source With a Groq Key

**Decision authority:** Absent (decides: absent; visible: n/a; overridable: n/a; boundary: vendor)

**Gap Analysis:** Groq ships no scored business application. Groq Chat is a demonstration surface (it 'is now using Orpheus' for voices); the examples page offers 'templates and community-driven solutions' with Groq-authored entries (Stock Bot, Blog Generator from Audio, Groq App Generator, Groq Desktop, a Compound CLI and MCP server) as repositories and live demos; the 'Coding with Groq' pages list Factory Droid, OpenCode ('an open-source AI coding agent'), Kilo Code, Roo Code, and Cline, third-party and open-source tools that take a Groq key ('keeping your Groq API key on-device via Bring Your Own Key'); the value the newsroom names (Paytm, HUMAIN One, Meta's Llama API, McLaren) is customers' and partners'. The buyer gets nothing to log into for its business.

The architect's concern is nil at this layer.

Calibration: OpenAI, Anthropic, Mistral, and Cohere read strong on first-party applications; Cerebras gap on a Playground and cookbooks. Groq is Cerebras's shape with a demo shelf. Gap, authority Absent.

**Borrowed Judgment:** Nothing offered, nothing inherited. Absent.

---
*Layer2C · AI Infrastructure Decision Intelligence · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com*
