Executive Summary: Groq (GroqCloud on the LPU + Compound + Server-Side Web Retrieval; Independent Under a Non-Exclusive NVIDIA Technology License)

Groq is the Language Processing Unit (LPU) inference cloud that licensed its technology to NVIDIA (non-exclusive, December 24, 2025) and kept operating: its founder and president left for NVIDIA, Simon Edwards became chief executive, GroqCloud 'will continue to operate without interruption', and the company has since raised $650 million and a $350 million round at a $3.5 billion valuation, become an NVIDIA Cloud Partner, and announced it will run NVIDIA's Groq 3 LPX with Vera Rubin NVL72 beside its own silicon. The map reads it strong at one layer, moderate at one, and gap at six. Layer 2B is the strength: open models (gpt-oss, Llama, Whisper) behind an OpenAI-compatible API with constrained-decoding structured outputs, reasoning, speech to text, and prompt caching, plus Compound, a production agentic system over a fixed harness of server-side web search, browsing, code execution, and Wolfram Alpha that takes no custom tools, and Low-Rank Adaptation (LoRA) inference of the enterprise's own adapters; text to speech and vision are preview, and the Responses API and every Model Context Protocol (MCP) path are beta. Layer 1B is moderate on the Perplexity reading: Compound's web retrieval on Tavily and Exa returns search results and citations per request, over the web and not the enterprise's data. Layer 0 is a gap by shape: Groq's own silicon in thirteen partner-operated data centers is sold only as tokens, sovereign endpoints, and a dedicated instance for LoRA serving, with GroqMetal marketed and GroqRack undocumented, so the enterprise decides nothing there. Layer 2A is a gap on the OpenAI reading: service tiers, provisioned throughput, per-project limits, spend caps, and model permissions are a metering edge, not a plane. Everything else is a gap: no data foundation, no pipelines, no governance plane, no business application.

The capture is at the LPU, which the enterprise never touches, and at Compound. The API is OpenAI-compatible and the models are other labs' open weights, so the client and the model choice lift; local tools are the enterprise's; LoRA adapters are the enterprise's standard files served behind the same door. What is Groq's and stays Groq's: the silicon and its cloud, the service tiers and account controls, Compound and the built-in tools that execute on Groq's side, and the web retrieval that is Tavily's and Exa's through Groq's paper. Five components read Ceded, Delegated, Retained, Ceded, Delegated across 1B and 2B; today's runtime has no NVIDIA dependency, and the row's NVIDIA note is a roadmap and a license rather than a chip.

The buyer's trade: LPU speed on open models through an interface it already uses, with a hosted agent that searches the web and runs code, in exchange for a silicon vendor it can't administer, a production catalogue of six models with speech out and vision still in preview, beta badges on the Responses API and every MCP path, and an agent whose tool calls run on Groq's side behind a per-request allow-list and nothing else, and don't run in the sovereign regions. The decision-authority readings follow: vendor / Ceded at Layer 0 and 2A (Groq absorbs both beneath its service; a tier is a menu, a limit is a bound; nothing below the endpoint is visible), model / Ceded at 1B (Compound decides when to search), model / Delegated at 2B on the primary path (the enterprise's loop with local tools; Compound's path is Ceded), Absent elsewhere. What would move cells: documentation for GroqMetal, a GroqRack datasheet, or a sized dedicated instance (Layer 0), a production text-to-speech or vision model or custom tools in Compound (2B chips), remote MCP leaving beta with its approval flow (a 2B chip and a 2C fragment), an instance the enterprise sizes (2A), and the NVIDIA LPX capacity coming online (an NVIDIA dependency at Layer 0).

Layer-by-layer status: Layer 0 (Not a Customer Surface: Groq's Own LPU Silicon in Thirteen Partner-Operated Data Centers, Sold Only as Tokens and Regional Sovereign Endpoints; a Dedicated Instance Documented Only as a LoRA Deployment Mode; GroqMetal and GroqRack Marketed Without Documentation), Layer 1A (No Data Foundation: No Retention by Default, Zero Data Retention on Request, Batch Files Kept Thirty Days in US Google Cloud Buckets; Feedback Submissions Kept up to Three Years), Layer 1B (Server-Side Web Retrieval in Compound and the Built-In Tools: Web Search on Tavily, Browser Search on Exa, and Visit Website, With Search Results and Citations Returned per Request; No Embedding or Reranking Model, No Index Over the Enterprise's Data), Layer 1C (No Pipelines: a Batch API Queues Bulk Requests at Half Price Inside a Processing Window of a Day to a Week), Layer 2A (A Metering Edge, Not a Plane: On-Demand, Flex, Performance (Provisioned Throughput With a 99.9 Percent SLA), Auto, and Batch Tiers; Per-Project Limits, Spend Caps, Model Permissions; No Instance You Size), Layer 2B (The LPU Inference Cloud: Open Models Behind an OpenAI-Compatible API With Strict Structured Outputs, Reasoning, Whisper Speech to Text, and Prompt Caching, Plus Compound, a Production Agentic System With a Fixed Set of Server-Side Tools; Text to Speech and Vision in Preview; Remote MCP and Connectors in Beta; LoRA Inference on Shared or Dedicated Hardware), Layer 2C (Model Permissions, Project Limits, and Roles Govern Accounts, Not Agents; Compound's Server-Side Tools Take a Per-Request Allow-List and Nothing Else; Remote MCP Approvals and Tool Filters Are Beta), Layer 3 (+1) (No Business Application: Groq Chat Is a Demonstration Surface, the Examples Page Holds Groq-Authored Templates and Demos, and the Coding Tools (Factory Droid, OpenCode, Kilo Code, Roo Code, Cline) Are Third-Party and Open-Source With a Groq Key).

Assessment framework: 4+1 Layer AI Infrastructure Model. Scoring model: Decision Authority Placement Model (DAPM) — Retained, Delegated, or Ceded. Published by The CTO Advisor LLC (DBA The Advisor Bench). Author: Keith Townsend. Date assessed: September 11, 2026. Version: v1.0 - 4+1 v2: Authority Split.

Groq (GroqCloud on the LPU + Compound + Server-Side Web Retrieval; Independent Under a Non-Exclusive NVIDIA Technology License)

Mapped to the 4+1 Layer AI Infrastructure Model

v1.0 - 4+1 v2: Authority SplitAssessed September 11, 2026Sources & revision history
ACTIVE ASSESSMENT

Summary Finding

Groq is the Language Processing Unit (LPU) inference cloud that licensed its technology to NVIDIA (non-exclusive, December 24, 2025) and kept operating: its founder and president left for NVIDIA, Simon Edwards became chief executive, GroqCloud 'will continue to operate without interruption', and the company has since raised $650 million and a $350 million round at a $3.5 billion valuation, become an NVIDIA Cloud Partner, and announced it will run NVIDIA's Groq 3 LPX with Vera Rubin NVL72 beside its own silicon. The map reads it strong at one layer, moderate at one, and gap at six. Layer 2B is the strength: open models (gpt-oss, Llama, Whisper) behind an OpenAI-compatible API with constrained-decoding structured outputs, reasoning, speech to text, and prompt caching, plus Compound, a production agentic system over a fixed harness of server-side web search, browsing, code execution, and Wolfram Alpha that takes no custom tools, and Low-Rank Adaptation (LoRA) inference of the enterprise's own adapters; text to speech and vision are preview, and the Responses API and every Model Context Protocol (MCP) path are beta. Layer 1B is moderate on the Perplexity reading: Compound's web retrieval on Tavily and Exa returns search results and citations per request, over the web and not the enterprise's data. Layer 0 is a gap by shape: Groq's own silicon in thirteen partner-operated data centers is sold only as tokens, sovereign endpoints, and a dedicated instance for LoRA serving, with GroqMetal marketed and GroqRack undocumented, so the enterprise decides nothing there. Layer 2A is a gap on the OpenAI reading: service tiers, provisioned throughput, per-project limits, spend caps, and model permissions are a metering edge, not a plane. Everything else is a gap: no data foundation, no pipelines, no governance plane, no business application.

The capture is at the LPU, which the enterprise never touches, and at Compound. The API is OpenAI-compatible and the models are other labs' open weights, so the client and the model choice lift; local tools are the enterprise's; LoRA adapters are the enterprise's standard files served behind the same door. What is Groq's and stays Groq's: the silicon and its cloud, the service tiers and account controls, Compound and the built-in tools that execute on Groq's side, and the web retrieval that is Tavily's and Exa's through Groq's paper. Five components read Ceded, Delegated, Retained, Ceded, Delegated across 1B and 2B; today's runtime has no NVIDIA dependency, and the row's NVIDIA note is a roadmap and a license rather than a chip.

The buyer's trade: LPU speed on open models through an interface it already uses, with a hosted agent that searches the web and runs code, in exchange for a silicon vendor it can't administer, a production catalogue of six models with speech out and vision still in preview, beta badges on the Responses API and every MCP path, and an agent whose tool calls run on Groq's side behind a per-request allow-list and nothing else, and don't run in the sovereign regions. The decision-authority readings follow: vendor / Ceded at Layer 0 and 2A (Groq absorbs both beneath its service; a tier is a menu, a limit is a bound; nothing below the endpoint is visible), model / Ceded at 1B (Compound decides when to search), model / Delegated at 2B on the primary path (the enterprise's loop with local tools; Compound's path is Ceded), Absent elsewhere. What would move cells: documentation for GroqMetal, a GroqRack datasheet, or a sized dedicated instance (Layer 0), a production text-to-speech or vision model or custom tools in Compound (2B chips), remote MCP leaving beta with its approval flow (a 2B chip and a 2C fragment), an instance the enterprise sizes (2A), and the NVIDIA LPX capacity coming online (an NVIDIA dependency at Layer 0).

Strength
Moderate
Gap
Partner
Layer 0 · ComputeCompute & Network FabricNot a Customer Surface: Groq's Own LPU Silicon in Thirteen Partner-Operated Data Centers, Sold Only as Tokens and Regional Sovereign Endpoints; a Dedicated Instance Documented Only as a LoRA Deployment Mode; GroqMetal and GroqRack Marketed Without Documentationdecides: vendor · Ceded▼

Raw compute, networking, and acceleration fabric

Vendor-Provided

NVIDIA-Provided

No NVIDIA Dependency Today

GroqCloud runs on Groq's own LPUs. The non-exclusive license to NVIDIA (December 24, 2025), NVIDIA Cloud Partner status (August 12, 2026), and the announced NVIDIA Groq 3 LPX and Vera Rubin NVL72 capacity (August 24, 2026; not online) are in the notes and the summary as roadmap, not as a chip.

◆ Gap Analysis

Groq owns the chip, operates capacity in partner data centers, and sells tokens. GroqCloud runs on Groq's Language Processing Units in thirteen data centers 'across North America, Europe, the Middle East and APAC' (Helsinki, a UK site with Equinix, Sydney, Saudi Arabia with HUMAIN, Bell Canada's network, and the United States), 'scaling toward 200 MW by 2027' after a $650 million raise (June 2026) and a $350 million round led by Disruptive with planned NVIDIA participation (August 2026, at a $3.5 billion valuation); the enterprise reaches it through api.groq.com, or, on the Enterprise plan, through regional and sovereign endpoints exposed as Google Cloud Private Service Connect published services (me-central2 and us-central1; api.me-central-1.groqcloud.com, api.us.groqcloud.com) and, for LoRA inference, on 'dedicated Groq hardware instances purchased by the customer'. GroqRack, the on-premises rack, has no datasheet or documentation in the public set; the page for it carries only the company's LPX message. The platform page markets three tiers, 'GroqMetal provides infrastructure, GroqCore adds inference, and GroqAssured adds enterprise controls', with GroqMetal as 'Dedicated bare-metal infrastructure ... with full control in your hands', and no documentation, specification, or price behind any of them. The buyer gets a second inference silicon by the token, close to its region. The architect's concern is that nothing here is administered by the enterprise. No instance it sizes (the LoRA dedicated instance is bought through sales with no documented sizing or scheduling surface), no bare metal it can document (GroqMetal is a marketing tier), no rack it can document, no fabric; the sovereign endpoints are private doors into Groq's capacity, Compound and the built-in tools are 'not available currently for use with regional / sovereign endpoints', and the next generation of the silicon (Groq 3 LPX) is an NVIDIA product Groq will rent like anyone else. Calibration: Layer 0 reads by the vendor's shape (the September 15, 2026 ruling under the exposure test). Groq is a service: it sells tokens and sovereign endpoints, and the LPU behind them is Groq's supply chain, which is Anthropic's cell ('compute supply chain, not customer surface') and Perplexity's ('serving fleet, not a customer surface'); Cerebras and NVIDIA read strong because they sell systems the enterprise buys, not because they own the silicon, and owning it behind a service changes nothing, as Google's TPUs behind the Gemini API change nothing. There's nothing for the enterprise to decide at this layer. Gap, the service reading; a Stack Builder wouldn't build a Layer 0 on it.

◆ Borrowed Judgment

Ceded and invisible: Groq absorbs the layer beneath its service. The runtime call at this layer is Groq placing and scheduling work across its LPUs; the enterprise chooses a region through the endpoint it calls and sees nothing below it (usage and Prometheus metrics are consumption, not placement): vendor decides, not visible, not overridable, Ceded, the OpenAI, Anthropic, and Perplexity reading.

◆ Working Notes

Ruled September 15, 2026: moderate to gap on the shape line (a service vendor's own silicon isn't a Layer 0 surface until it's sold as one); strong had been argued on silicon authority and GroqMetal, gap on the no-customer-surface reading. Named, not scored: GroqMetal, GroqCore, and GroqAssured (platform page only; the vendor-availability-without-documentation ruling), GroqRack (no documentation), the LoRA dedicated instance ('dedicated Groq hardware instances purchased by the customer'; documented as a LoRA deployment mode, not as a sized instance), NVIDIA Groq 3 LPX and Vera Rubin NVL72 capacity ('When Groq brings NVIDIA Groq 3 LPX capacity online'; no date), the data-center count and megawatt targets (press releases). Public evidence that moves the cell: documentation for GroqMetal, a GroqRack datasheet, or a sized dedicated instance, any of which would make Groq sell infrastructure and move the cell to the OEM or IaaS line.

Layer 1A · StorageData Storage & GovernanceNo Data Foundation: No Retention by Default, Zero Data Retention on Request, Batch Files Kept Thirty Days in US Google Cloud Buckets; Feedback Submissions Kept up to Three Yearsdecides: absent · Absent▼

Durable, governed data foundation — the Governance Catalog that Layer 2C queries

Vendor-Provided

NVIDIA-Provided

No NVIDIA Dependency

Nothing at this layer is a scored capability.

◆ Gap Analysis

Groq keeps as little as it can on the API path. 'By default, Groq does not retain customer data for inference requests'; 'All customers may enable Zero Data Retention (ZDR) in Data Controls settings'; what the API retains (batch input and output files for thirty days unless deleted earlier, LoRA adapters and training datasets until deleted, reliability and abuse logs up to thirty days) sits in 'Google Cloud Platform (GCP) buckets located in the United States' with standard contractual clauses for transfers; voluntary feedback submissions are governed separately, and 'reviewed feedback, conversation snippets, and related metadata are stored for up to 3 years'. The buyer gets nothing to store here. The architect has no Groq-managed data foundation to evaluate; retention and US residency of what is retained are trust controls to account for, not Layer 1A capability. Calibration: OpenAI, Anthropic, Mistral, Cohere, and Cerebras read gap on trust apparatus without a data foundation. Gap, authority Absent.

◆ Borrowed Judgment

Nothing offered, nothing inherited. Absent.

◆ Working Notes

Named, not scored: the feedback policy's three-year retention of submitted snippets (a support channel, not the API). Public evidence that moves the cell: a storage or governance product, which nothing describes.

Layer 1B · RetrievalContext Management & RetrievalServer-Side Web Retrieval in Compound and the Built-In Tools: Web Search on Tavily, Browser Search on Exa, and Visit Website, With Search Results and Citations Returned per Request; No Embedding or Reranking Model, No Index Over the Enterprise's Datadecides: model · Ceded▼

Low-latency retrieval for RAG — vector/hybrid search, context windows

Vendor-Provided

Server-Side Web Retrieval (Web Search on Tavily, Browser Search on Exa, Visit Website; in Compound and Compound Mini and as Built-In Tools on gpt-oss Models; Search Results, Relevance Scores, and Citations in executed_tools; Domain and Country Settings)Ceded

Partners' search engines through Groq's harness on Groq's servers. Ceded, the channel ruling and the Perplexity Search API reading.

NVIDIA-Provided

No NVIDIA Dependency at This Layer

The search engines are Tavily's and Exa's, reached from Groq's servers.

◆ Gap Analysis

Groq retrieves from the web, not from the enterprise. Compound and Compound Mini ('Production Systems') and the built-in tools on gpt-oss models run web search ('powered by Tavily'), browser search on Exa, and visit website on Groq's side, 'automatically enabled by default', with domain and country settings, and return what they found in executed_tools (search results with relevance scores, page content, citations); the catalogue has no embedding or reranking model; the Google Workspace connectors (Gmail, Calendar, Drive search and fetch) are remote MCP connectors 'currently in beta', and the Hugging Face integration is a remote MCP server Hugging Face hosts. The buyer gets live web retrieval inside an inference call, and brings its own index for everything else. The architect's concern is that the retrieval is partners' engines through Groq's paper, scoped to the public web, on Groq's tool-selection judgment, and 'not available currently for use with regional / sovereign endpoints'. Calibration: Perplexity reads moderate on a web Search API exposed raw; Anthropic moderate on federated search over connected systems; OpenAI moderate on file search and vector stores; Cerebras gap with partner vector stores in integration guides only. Groq is Perplexity's shape inside the model call, on someone else's index. Moderate, on the Perplexity reading.

◆ Borrowed Judgment

Ceded at the tool. The engines are Tavily's and Exa's and the harness is Groq's; nothing lifts but the query: Ceded, the channel ruling (partners' engines through the vendor's paper, the NIM reading) and the Perplexity Search API reading. The runtime call is Compound deciding whether and what to search ('intelligently decides when to use each tool'), visible in executed_tools, with enabled_tools as an allow-list, a bound rather than an override: model decides, visible, not overridable, Ceded, the Perplexity reading.

◆ Working Notes

Badges: the Google Workspace MCP connectors (Gmail read and search, Calendar, Drive search and fetch) are 'currently in beta', unscored, no GA date; browser automation 'has been retired'. Compound and the built-in tools are 'not available currently for use with regional / sovereign endpoints'. Public evidence that moves the cell: an embedding model on the API or a retrieval service over the enterprise's data.

Layer 1C · PipelinesData Movement & PipelinesNo Pipelines: a Batch API Queues Bulk Requests at Half Price Inside a Processing Window of a Day to a Weekdecides: absent · Absent▼

Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering

Vendor-Provided

NVIDIA-Provided

No NVIDIA Dependency

Nothing at this layer is a scored capability.

◆ Gap Analysis

Groq moves tokens, not data. Batch processing ('50% lower cost, no impact to your standard rate limits, and 24-hour to 7 day processing window', Developer tier and up; 'we'll process as many requests as our capacity allows', and a batch that doesn't finish expires) is bulk inference, not a pipeline. The buyer gets nothing here. The architect's concern is nil at this layer. Calibration: OpenAI, Anthropic, Mistral, Cohere, and Cerebras read gap. Gap, authority Absent.

◆ Borrowed Judgment

Nothing offered, nothing inherited. Absent.

◆ Working Notes

Named, not scored: optical character recognition (OCR) and structured JSON extraction on the Qwen 3.6 and 3.8 vision models (Preview; 'intended for evaluation purposes only'). Public evidence that moves the cell: nothing described.

Layer 2A · OrchestrationInfrastructure OrchestrationA Metering Edge, Not a Plane: On-Demand, Flex, Performance (Provisioned Throughput With a 99.9 Percent SLA), Auto, and Batch Tiers; Per-Project Limits, Spend Caps, Model Permissions; No Instance You Sizedecides: vendor · Ceded▼

GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization

Vendor-Provided

NVIDIA-Provided

No NVIDIA Dependency

Nothing at this layer is a scored capability.

◆ Gap Analysis

Groq gives the enterprise a menu of how its requests get scheduled, and bounds around them. The service_tier parameter picks on-demand (the default), flex ('available for all models to paid customers only with 10x higher rate limits'; 'If flex capacity is unavailable, requests will fail quickly with status 498'), performance ('only available on enterprise plans'; 'a 99.9% availability SLA and a 99% latency guarantee'; 'delivered as provisioned throughput: you purchase input and output capacity bundles'), or auto ('leverage the best tier available to you at any given moment'); batch runs asynchronously outside the tiers. Around them: rate limits enforced 'at the organization level' by requests, tokens, and audio seconds, with cached tokens exempt; projects with their own API keys and limits 'must be ≤ org limit'; an organization-wide spend cap on paid tiers; model permissions ('Only Allow' or 'Only Block' at organization and project level); Owner, Developer, and Reader roles; Prometheus metrics for enterprise customers. The LoRA dedicated instance is 'purchased by the customer' with 'No LoRA-specific rate limiting' and no documented sizing, replica, or scheduling surface. The buyer gets priority it can pay for and budgets it can set. The architect's concern is that it's a queue, not a plane: no instance the enterprise sizes, no replica count, no scheduler it sees; the performance tier is provisioned throughput on Groq's terms, and the rest is bounds. Calibration: OpenAI reads gap on the same shape (priority and flex service tiers, Scale Tier and provisioned-throughput commitments, rate limits, per-project quota); Mistral and Cerebras gap on a metering edge; Cohere moderate on Model Vault, where the customer sets instance counts and replica bounds; Databricks and Snowflake moderate on managed compute with bounds. Groq's tiers are generally available and contracted where Cerebras's are preview, and that changes nothing at this layer: GA of a metering edge is still a metering edge. Gap, on the OpenAI reading.

◆ Borrowed Judgment

Groq absorbs the layer beneath its service: its scheduler places every request inside the tier the enterprise named and the limits it set, and exposes no placement decision. Vendor decides, not visible, not overridable, Ceded, the OpenAI and Cerebras absorbed-layer reading; a tier is a menu and a limit is a bound.

◆ Working Notes

Written gap on the OpenAI reading after the ChatGPT pass; the first draft read moderate on Cohere. Named, not scored: the LoRA dedicated instance (a documented dedicated-capacity fact with no customer-facing capacity object). Public evidence that moves the cell: an instance or replica surface the enterprise sizes, which would argue the Cohere reading.

Layer 2B · RuntimeApplication Runtime & ExecutionThe LPU Inference Cloud: Open Models Behind an OpenAI-Compatible API With Strict Structured Outputs, Reasoning, Whisper Speech to Text, and Prompt Caching, Plus Compound, a Production Agentic System With a Fixed Set of Server-Side Tools; Text to Speech and Vision in Preview; Remote MCP and Connectors in Beta; LoRA Inference on Shared or Dedicated Hardwaredecides: model · Delegated▼

Model serving, agent execution, inference APIs, distributed inference

Vendor-Provided

Model Serving via the OpenAI-Compatible API (Chat Completions; gpt-oss 120B and 20B, Llama 3.1 8B and 3.3 70B on Enterprise, Whisper Large V3 and Turbo Speech to Text; Structured Outputs With Constrained Decoding; Reasoning Controls; Automatic Prompt Caching)Delegated

Open-weight models their owners publish, behind a multi-vendor interface on Groq's silicon. Delegated, the inference-interface ruling; text to speech and vision are preview and unscored.

Customer-Created and Managed Tools (Local Tool Calling in the Enterprise's Process; Parallel Calls on Supported Models)Retained

Tool logic the enterprise writes and runs. Retained, the customer-tools ruling.

Compound and Compound Mini (Production Agentic Systems Over gpt-oss-120b, Llama 4 Scout, and Llama 3.3 70B; a Fixed Harness of Web Search on Tavily, Visit Website, Code Execution in E2B, and Wolfram Alpha; Up to Ten Server-Side Tool Calls per Request for Compound, One for Compound Mini; No Custom Tools; Per-Request enabled_tools Allow-List)Ceded

Groq's agent loop over Groq's tools on Groq's servers; the underlying models are their owners'. Ceded, the Perplexity Computer and Cohere North reading.

LoRA Inference (Enterprise; Customer-Trained Adapters in Standard Parameter-Efficient Fine-Tuning (PEFT) Files Served Through the OpenAI-Compatible Chat Completions Path, on the Shared Cloud per Token or on a Dedicated Groq Hardware Instance the Customer Purchased; Groq Does No Fine-Tuning)Delegated

The adapter is the enterprise's artifact and lifts to any PEFT-capable server; Groq serves it behind a multi-vendor interface. Delegated, the inference-interface ruling with the customer fine-tunes reading; the dedicated instance is a Layer 0 fact.

NVIDIA-Provided

No NVIDIA Dependency in the Runtime Today

Inference runs on Groq's LPUs. NVIDIA Groq 3 LPX and Vera Rubin NVL72 capacity is announced and not online; NVIDIA holds a non-exclusive license to Groq's inference technology.

◆ Gap Analysis

Groq's runtime is fast inference of other labs' models, and one agent of its own. The API is 'mostly compatible with OpenAI's client libraries' at api.groq.com/openai/v1 (no logprobs, logit_bias, or n>1): chat completions, a Responses API ('currently in beta'), speech to text (Whisper Large V3 and Turbo), structured outputs with constrained decoding on gpt-oss-20b, gpt-oss-120b, and qwen3.8-27b ('Streaming and tool use are not currently supported with Structured Outputs'), reasoning controls, and automatic no-fee prompt caching on the gpt-oss models. Production models are Llama 3.1 8B and 3.3 70B (Enterprise), gpt-oss 120B and 20B, and the Whisper pair; preview models (Orpheus text to speech, Prompt Guard, MiniMax M2.7, Qwen 3.6 and 3.8 with vision, the gpt-oss safeguard) are 'intended for evaluation purposes only', so speech out and vision are unscored; MiniMax M2.5 and Qwen3-VL 32B arrived for Enterprise customers in April 2026. Tool use runs three ways: local (the enterprise's functions, parallel calls on some models), remote MCP ('currently in beta'; 'Groq's servers will connect to the MCP server, discover the available tools, pass them to the model, and execute any tools'), and built-in server-side tools (web search on Tavily, browser search on Exa, visit website, Python code execution in E2B sandboxes, Wolfram Alpha with the enterprise's key; browser automation retired). Compound and Compound Mini are 'Production Systems': Groq-built agentic systems over gpt-oss-120b, Llama 4 Scout, and Llama 3.3 70B that 'solve problems by taking action', up to ten server-side tool calls per request (Compound Mini one), priced per tool, over a fixed set of four tools: 'Custom user-provided tools are not supported at this time.' LoRA inference (Enterprise only; 'We do not provide LoRA fine-tuning services') runs adapters on the shared cloud or on dedicated Groq hardware the customer purchased. The buyer gets open models at LPU speed, a hosted agent that browses and runs code, and its own adapters on Groq's silicon. The architect's concern is what the speed is attached to. Groq owns no foundation model, the production list is six models plus two systems, the Responses API and every MCP path are beta, Compound and the built-in tools don't run on regional or sovereign endpoints, and Compound's tool calls execute on Groq's side with no documented gate before the effect. Calibration: Cloudflare reads strong on serverless open-model inference plus a gateway plus agents; OpenAI and Anthropic strong on frontier serving plus hosted tools; Cerebras strong on open serving alone at the frontier of speed; Mistral and Cohere strong on frontier serving plus harnesses; Hugging Face moderate on open serving under maturity gates. Groq has serving on its own engine at the frontier of speed plus Compound, a production agentic system over a fixed harness of four server-side tools that takes no custom tools; local tool calling is the general path and it runs in the enterprise's process. Strong. Agent execution is judged by whether the runtime carries the enterprise's own agent loop (tool calling, structured outputs, reasoning controls, latency over the serial turns an agent takes), not by whether the vendor hosts one (the September 11, 2026 ruling under rule 4); Groq carries it on its own engine at the frontier of speed, and Compound is a hosted option that reads Ceded when used, never the price of the grade.

◆ Borrowed Judgment

Delegated at the interface, Retained at the enterprise's tools, Ceded at Compound. Model access through the OpenAI-compatible API is the inference-interface reading: the client points anywhere and the models are their owners' open weights: Delegated. Local tools the enterprise writes and executes are its own: Retained, the customer-tools ruling. Compound and the built-in tools are Groq's surfaces on Groq's silicon: Ceded (Compound's underlying models are the owners'). LoRA adapters are the enterprise's standard files served behind the same OpenAI-compatible door: Delegated; the dedicated instance they may run on is a Layer 0 fact. On the primary path, chat completions with local tools, the model decides to call and the enterprise executes in its own process, where a deterministic gate is its to write: model decides, visible, overridable, Delegated, the OpenAI and Anthropic reading. On the Compound path the tool executes on Groq's side with no gate documented, the Perplexity reading, Ceded; the primary path governs the cell.

◆ Working Notes

Ruled September 11, 2026: strong stands (agent execution at 2B is the runtime carrying the enterprise's own loop; Compound is a Ceded component, not the price of strong). Both reviewer passes had voted moderate. Badges: the Responses API, remote MCP (with allowed_tools and require_approval), and MCP Connectors (Google Workspace) are 'currently in beta' and unscored; preview models (Orpheus text to speech, Qwen 3.6 and 3.8 vision and OCR, Prompt Guard, MiniMax M2.7, the gpt-oss safeguard) are 'intended for evaluation purposes only'; Qwen3-VL 32B for Enterprise appears in the April 2026 changelog and not in the models table; browser automation 'has been retired'. Compound and the built-in tools are 'not available currently for use with regional / sovereign endpoints'. No Anthropic-compatible endpoint appears in the documentation. Public evidence that moves the cell: a production text-to-speech or vision model, custom tools in Compound, or remote MCP leaving beta.

Layer 2C · ReasoningAgentic Infrastructure — The Reasoning PlaneModel Permissions, Project Limits, and Roles Govern Accounts, Not Agents; Compound's Server-Side Tools Take a Per-Request Allow-List and Nothing Else; Remote MCP Approvals and Tool Filters Are Betadecides: absent · Absent▼

Policy-driven placement and resource coordination — the Autonomy Layer

Vendor-Provided

NVIDIA-Provided

No NVIDIA Dependency

Nothing at this layer is a scored capability.

◆ Gap Analysis

Read against the five legs, Groq has none as a plane. Identity: API keys, projects, and Owner, Developer, and Reader roles identify callers; no agent principal. Gateway: model permissions ('Only Allow' or 'Only Block' per organization and project, a 403 on a restricted model) are a policy on which models a key may call, and Compound takes a per-request allow-list (compound_custom.tools.enabled_tools 'restrict[s] or specif[ies] exactly which tools should be available for a particular request'); both are fragments, request-scoped and key-scoped, with no central policy, registry, or agent identity behind them. Remote MCP documents allowed_tools and require_approval ('never', 'always') with an approval flow, and remote MCP 'is currently in beta' ('MCP servers have access to all data in your AI model's context'). Registry and orchestration: none. Observability: executed_tools on a Compound response lists the tools called, their arguments, and results, per request; usage, spend tracking with a ten-to-fifteen-minute lag, and Prometheus metrics on Enterprise watch the API; nothing watches an estate of agents. The buyer brings the plane. The architect's concern is Compound: an agent that searches, fetches pages, and runs code on Groq's side, governed by the caller's key and a per-request tool list and nothing else the documentation names. Calibration: OpenAI, Anthropic, Cohere, Mistral, and Cerebras read gap; Cloudflare moderate on a gateway of its own. Groq is the OpenAI shape with a model-permission list and a tool allow-list. Gap, authority Absent.

◆ Borrowed Judgment

Nothing offered as a plane. Absent.

◆ Working Notes

Named, not scored: model permissions and the Compound enabled_tools allow-list (request- and key-scoped fragments of a gateway leg); executed_tools (per-request tool visibility); remote MCP allowed_tools and require_approval (beta, unscored). Public evidence that moves the cell: remote MCP leaving beta with its approval flow, or a central policy, registry, or agent identity over Compound's and MCP tool calls.

Layer 3 (+1) · ApplicationsAI Application Layer — The Value PlaneNo Business Application: Groq Chat Is a Demonstration Surface, the Examples Page Holds Groq-Authored Templates and Demos, and the Coding Tools (Factory Droid, OpenCode, Kilo Code, Roo Code, Cline) Are Third-Party and Open-Source With a Groq Keydecides: absent · Absent▼

AI-powered business capabilities — business logic, workflow automation

Vendor-Provided

NVIDIA-Provided

No NVIDIA Dependency

Nothing at this layer is a scored capability.

◆ Gap Analysis

Groq ships no scored business application. Groq Chat is a demonstration surface (it 'is now using Orpheus' for voices); the examples page offers 'templates and community-driven solutions' with Groq-authored entries (Stock Bot, Blog Generator from Audio, Groq App Generator, Groq Desktop, a Compound CLI and MCP server) as repositories and live demos; the 'Coding with Groq' pages list Factory Droid, OpenCode ('an open-source AI coding agent'), Kilo Code, Roo Code, and Cline, third-party and open-source tools that take a Groq key ('keeping your Groq API key on-device via Bring Your Own Key'); the value the newsroom names (Paytm, HUMAIN One, Meta's Llama API, McLaren) is customers' and partners'. The buyer gets nothing to log into for its business. The architect's concern is nil at this layer. Calibration: OpenAI, Anthropic, Mistral, and Cohere read strong on first-party applications; Cerebras gap on a Playground and cookbooks. Groq is Cerebras's shape with a demo shelf. Gap, authority Absent.

◆ Borrowed Judgment

Nothing offered, nothing inherited. Absent.

◆ Working Notes

Named, not scored: Groq Chat and the examples page's Groq-authored templates and live demos (demonstration surfaces, not products). Public evidence that moves the cell: a first-party application, which nothing describes.

Sources & revision history · v1.0 - 4+1 v2: Authority Split

GroqCloud documentation at console.groq.com (models; OpenAI compatibility; Responses API; rate limits; text generation; speech to text; text to speech and Orpheus; vision; reasoning; content moderation; structured outputs; prompt caching; tool use overview, built-in tools (web search, visit website, code execution, Wolfram Alpha, browser search, browser automation), remote MCP and connectors, local tool calling; integrations catalogue; coding with Groq; Compound overview, built-in tools, systems; service tiers, performance tier, flex processing, batch processing; LoRA inference; production readiness; security onboarding and Google Cloud Private Service Connect; your data and the feedback policy; examples; Prometheus metrics; spend limits, projects, model permissions, billing FAQs, your data; SDK libraries; changelog through April 18, 2026; legacy changelog; deprecations; llms.txt); groq.com pages (platform, pricing, GroqRack, about us, newsroom index) and posts (Groq and Nvidia Enter Non-Exclusive Inference Technology Licensing Agreement, December 24, 2025; GroqCloud: Expanding to Meet Demand, February 16, 2026; Groq Raises $650M, June 22, 2026; Groq Becomes an NVIDIA Cloud Partner, August 12, 2026; Groq Closes $350 million Series A, August 17, 2026; Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market, August 24, 2026). Peer-reviewed cell by cell through the labs claims ledger (groq-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (groq-row-agy); totals and escalated items in reviews/groq-judgment.md.

4+1 Layer AI Infrastructure Model · Vendor Assessment Series · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com