Executive Summary: Mistral AI

Mistral AI is a frontier lab that sells inference on the industry's standard interface, publishes frontier-class weights under Apache 2.0, and runs a first-party application estate in Vibe with an embedded-engineer applied-AI lane behind it. It is strong at Layer 2B and Layer 3 and a gap everywhere else. It is headquartered in Paris and it sells in the United States. A US inference region since August 2026. Distribution through Microsoft Foundry and Copilot Studio, Bedrock, Vertex, SageMaker and Snowflake. A Palo Alto office, a US general manager who now runs global revenue, and a Seattle-based CMO hired in June. The infrastructure story (Mistral Compute) and the platform story (Studio's Agents, Workflows, Registry and Observability) are both real and both pre-GA. Every cell rests on public documentation.

The capture splits by altitude, and the split is the finding. At 2B it is decoupled and mostly reversible. The OpenAI-compatible interface keeps integration opinions portable. Large 3, Small 4, Ministral 3 and Devstral Small 2 give the enterprise an owned-weights exit no other frontier lab matches at this class. The capture that does exist there is quiet. Medium 3.5 and Devstral 2, the coding flagships, carry a Modified MIT license that stops at $20 million of monthly revenue, which makes them proprietary for exactly the buyer this instrument serves. At Layer 3 the capture is coupled and visible. Vibe's projects, libraries, approval grants and delegation habits accumulate per user per day. Work can't be driven from outside. A co-developed platform is inseparable from the people and stack that built it. On-premises deployment changes where that runs. It does not change who holds it.

Mistral's distinctive offer is jurisdiction as a selectable property, and the choice is real in both directions. Pin inference to the EU or the US. Deploy Vibe inside your own perimeter. Serve the Apache weights without a Mistral contract. That is local-first, and it is delivered. Control-first is not. There is no governance plane, no retrieval estate the enterprise operates, no orchestration artifact on the customer's paper, and a rule surface that stops at platform administration. The data plane moves where you put it; the control plane does not come with it, in Frankfurt or in Virginia. The trade is a real exit at the model in exchange for a thinner platform above it, a license gate on the best coding weights, and an application estate as captive as anyone's. Buy the residency and the weights. Do not mistake either for authority.

Layer-by-layer status: Layer 0 (Sovereign Fleet, Not Yet a Customer Surface), Layer 1A (Sovereignty Apparatus, Not a Governance Catalog), Layer 1B (Hosted Libraries, Consumed In-Agent), Layer 1C (Transform Primitive, No Pipeline), Layer 2A (Metering Edge, No Customer Plane (Compute Pending)), Layer 2B (Frontier Serving + Owned-Weights Exit, Agent API Pending), Layer 2C (Guardrails and Approvals, Not a Governance Plane), Layer 3 (+1) (First-Party Value Plane: Vibe Work, Vibe Code, Applied AI).

Assessment framework: 4+1 Layer AI Infrastructure Model. Scoring model: Decision Authority Placement Model (DAPM) — Retained, Delegated, or Ceded. Published by The CTO Advisor LLC (DBA The Advisor Bench). Author: Keith Townsend. Date assessed: September 2, 2026. Version: v1.3 - 4+1 v2: Gap Prose Reads Its Authority.

Mistral AI

Mapped to the 4+1 Layer AI Infrastructure Model

v1.3 - 4+1 v2: Gap Prose Reads Its AuthorityAssessed September 2, 2026Source: Mistral product documentation (models overview and model cards, deployment options, regional inference, Libraries and Document Library, Search Toolkit, Agentic Search, Agents and Conversations API reference, Workflows, Document AI and annotations, moderation and Custom Guardrails, admin panel, API keys, audit logs, Vibe Work and Vibe Code including Skills, connectors, scheduled tasks, safety and approvals, offline models and provider configuration), Mistral announcements (AI Studio, October 2025; Forge, March 17, 2026; Mistral Medium 3.5, April 29, 2026; remote agents in Vibe, May 22, 2026; Vibe Work and Code, May 28, 2026; Search Toolkit, May 28, 2026; OCR 4, June 23, 2026; connector controls, June 24, 2026; prompts and skills in Studio, July 9, 2026; Shieldstral, August 4, 2026; regional inference, open models and European compute, August 11, 2026; Agentic Search, August 20, 2026; OCR 4.1 changelog, August 31, 2026), Mistral Compute and AI Cloud product pages, Hugging Face model cards and license texts (Mistral Medium 3.5, Devstral 2, Devstral Small 2), the mistral-vibe repository (Apache 2.0), Microsoft and Mistral partnership release (July 21, 2026), CMA CGM releases (April 2025; May 27, 2026), Accenture release (February 26, 2026), Mistral customer stories (HTX, DSO), Artificial Analysis evaluation of Mistral Medium 3.5, Layer2C Labs 013, 014, 018 and 019, published 4+1 model. Rulings ratified on this row: early access is not GA (Mistral Compute watch-listed at Layer 0 and 2A); no credit budget across layers (a function is tested independently against every layer's purpose line); revenue-capped open weights are Ceded regardless of licensing practice; open source is continuity, not authority, unless the harness is documented against other vendors' models; the instrument scores only on public documentation, posts or published lab results, never on briefings; a lab validates only the surface it put under test, and a lab's own acceptance test is input to the canon, not a replacement for it.
ACTIVE ASSESSMENT
Strength
Moderate
Gap
Partner
Layer 0 · ComputeCompute & Network FabricSovereign Fleet, Not Yet a Customer Surfacedecides: vendor · Ceded

Raw compute, networking, and acceleration fabric

Vendor-Provided

NVIDIA-Provided

NVIDIA-Only Fleet, Upstream of the Customer

Mistral's own serving and training fleet is GB200 and GB300 on NVIDIA reference architectures, with the 1.4 GW Paris campus joint venture (MGX, Bpifrance, NVIDIA) behind it and no own-silicon hedge shipping. None of it is a customer surface today. When Mistral Compute reaches general availability this column fills end to end; until then the dependency sits behind the API boundary, where the customer can't see it.

Gap Analysis

The buyer never thinks about silicon, and for the API buyer that is the pitch. Behind the endpoint sits a European fleet the enterprise reads about in the press. 13,800 GB300s at Bruyeres-le-Chatel. A 10 MW inference site at Les Ulis, a Swedish build with EcoDataCenter, a joint venture for a 1.4 GW Paris campus. A multibillion-dollar July 2026 commitment from Microsoft to draw on that capacity. Mistral markets the fleet as Mistral Compute: GB200, GB300 and B300 nodes with Grace and x86 hosts, Kubernetes-native orchestration on bare metal, VAST underneath, and form factors from bare-metal servers to managed platform. The substrate-surface rule (ratified on the OpenAI row) says Layer 0 credit requires the substrate itself to be the purchasable thing: silicon choice, instance types, fabric, placement. Mistral Compute would clear that rule in kind. It fails the GA-gate in status. The product page calls it early access. The timeline reads "first external customers onboarded" in March 2026. There is no customer-facing documentation, no price list, and no order form, and the only named customer is a hyperscaler under a negotiated agreement. Can an architect buy it today? Not on the docs. Limited availability is a disqualifying status under the GA-gate, the same reading that watch-listed OpenAI Frontier. An architect building a production system today can't buy, rent or administer Mistral compute on the vendor's paper. So the cell reads as OpenAI and Anthropic read: compute is a supply chain, not a customer surface. The counterweight is the same channel-choice escape the peer rows carry, with one Mistral-specific addition. Consuming Mistral through Foundry, Bedrock, Vertex or SageMaker resolves the Layer 0 question in that cloud's row. Serving the Apache 2.0 weights on your own hardware resolves it in your own data center, with no Mistral relationship at all. That second path is scored at 2B; it is named here because it is the only way a Mistral buyer chooses their silicon today.

Borrowed Judgment

Total for the API channel, and invisible. Mistral chooses the accelerator, the facility and the placement and exposes only the model and the service contract. The one thing the customer can see that OpenAI and Anthropic customers can't: since August 2026 they can pin the request to a jurisdiction. That is a placement fact about geography, not a compute surface. No scored component exists; the decision-authority reading is Ceded and invisible, because Mistral decides placement beneath the service and the enterprise can't watch.

Working Notes

Watch-list, dateless (no GA date has been published, so this stays in notes rather than the GA schedule): Mistral Compute, early access, first external customers March 2026, Microsoft European capacity agreement July 21, 2026. On GA, re-test here and at 2A against the CoreWeave calibration: a rented NVIDIA-only fleet is Ceded at Layer 0, and Kubernetes plus Slurm shipped as the product reads Delegated on open substrate at 2A. Watch-list, dateless: European Compute Units (August 11, 2026), a coalition converting multi-year enterprise commitments into capacity access, not a purchasable product. Custom-silicon exploration reported in May 2026 is not a product. Public evidence that moves this cell: Mistral Compute documentation, a price list, and the early-access language gone.

Layer 1A · StorageData Storage & GovernanceSovereignty Apparatus, Not a Governance Catalogdecides: absent · Absent

Durable, governed data foundation — the Governance Catalog that Layer 2C queries

Vendor-Provided

NVIDIA-Provided

No NVIDIA Layer 1A Dependency

No storage product exists for NVIDIA to accelerate. The VAST Data selection is Mistral's own supply chain for Mistral Compute, invisible to customers and not scoreable under the exposure test. The near-empty column holds.

Gap Analysis

The buyer gets the answer that unblocks a regulated deal: where the data sits and under whose law. Mistral ships EU-resident infrastructure, SAML single sign-on (SSO), and role-based access control (RBAC) over members and groups. Audit logs on Enterprise plans record actor, event, target and metadata for both users and API keys. AES-256 at rest with bring-your-own-key, and an admin control plane for workspaces, billing and access policy. Vibe Enterprise deploys on-premises or in a private cloud with full data residency. That is a real trust posture. OpenAI offers more self-serve regions (ten); Mistral offers fewer regions and a perimeter the enterprise can own, which is the more jurisdiction-explicit of the two. None of it is a data foundation. Tested on function against this layer's purpose line, every candidate fails. Prompt and skill version control (generally available July 9, 2026) governs instruction artifacts with immutable versions, ownership, classification labels and audit. Does that govern the enterprise's data? No. The corpus under it is what the model is told, not what the enterprise knows. Libraries hold customer documents in Mistral's European cloud with library-level access control. But the corpus is a derived, rebuildable copy whose system of record stays in SharePoint or the file share. No classification engine, no cross-system lineage, no authorization authority past its own walls. It is scored at 1B where it performs retrieval, the same reading OpenAI's files and vector stores received. Connectors honor the source system's permissions, so Mistral is a permission consumer, not a catalog. The admin plane governs Mistral's own tenancy, not the enterprise's data estate. So the gap holds on function, not on routing: nothing Mistral ships today performs 1A's job for the enterprise's authoritative data. Calibration is the OpenAI and Anthropic cells exactly. Mistral's apparatus is thicker on residency and thinner on compliance tooling: there is no Compliance API, no published integration ecosystem, and the admin docs say SCIM provisioning arrives "when it becomes generally available." Neither direction moves a gap.

Borrowed Judgment

None at the layer's architectural function. The enterprise's authoritative data, classifications, lineage and access policies stay in source systems and existing governance platforms. Mistral administers the storage and lifecycle mechanics for Mistral-held content, captive administration that stays below the capability threshold, the same treatment the metering apparatus receives at 2A. With on-premises deployment available, the enterprise's data foundation need never leave its own perimeter at all, which earns no capability credit but makes the Absent reading cleaner here than on any other model-provider row: nothing offered, nothing inherited.

Working Notes

Watch-list, dateless: the AI Registry announced with AI Studio (private beta) is a system of record for agents, models, datasets, judges, tools and workflows with lineage, ownership, versioning, access control, moderation policies and promotion gates. It is the item to re-test at this layer on GA, since lineage and access control over registered datasets is the nearest a model provider has come to a governance catalog; today it has no docs section and does not score. SCIM: docs state GA pending. Inference flag: no doc-confirmed region list or ISO 27001 / SOC 2 statement was found; the sovereignty claim rests on product positioning and shapes narration, not grade. The trust apparatus is named here as sub-threshold platform administration, not scored capability.

Layer 1B · RetrievalContext Management & RetrievalHosted Libraries, Consumed In-Agentdecides: vendor · Ceded

Low-latency retrieval for RAG — vector/hybrid search, context windows

Vendor-Provided

Libraries + Agentic Search (Vibe and Studio)Ceded

GA as a product surface in Vibe and Studio. Managed ingest, embedding and search over customer documents held in Mistral's European cloud, with user, workspace and organization sharing, and an agentic multi-step retrieval loop returning chunks and citations. Consumed only inside Vibe or a Mistral agent; no standalone search endpoint. The programmatic path (the Libraries management API and the Agents API's document_library tool) is beta and is the scope gate, not the capability. Single-vendor API on Mistral's embeddings: lift-to-leave is re-ingest and rebuild elsewhere.

Embeddings API (mistral-embed, Codestral Embed)Ceded

GA, Premier (API-only). The embeddings carve-out ratified on the OpenAI row governs: vectors are useless without the same model at query time, no second vendor serves the space, and Mistral publishes no open-weight embedder to self-host it. An embedded corpus does not lift out.

MCP Connectors (Grounding Interface)Delegated

GA (June 24, 2026 controls). Live-fetch grounding against Drive, SharePoint, Slack, GitHub, Atlassian, Box and others, with data "fetched in real time for the current request and not stored." A genuine multi-vendor standard: connector integrations lift to any MCP-capable host, and the source systems stay authoritative. Matches the Anthropic 1B call.

NVIDIA-Provided

No Customer-Facing NVIDIA Surface at Layer 1B

Whatever silicon serves embedding and retrieval is invisible behind the API, and Search Toolkit runs on whatever the customer has. The near-empty column holds.

Gap Analysis

The buyer gets a persistent, shareable, EU-resident knowledge base with almost no setup. Libraries accept most enterprise formats (Office, PDF, images, code, email), process in minutes, and share at user, workspace or organization scope with Viewer and Editor roles. Since August 20, 2026 they run Agentic Search: a multi-step loop (search, open, navigate, read, grep) that inspects sources rather than returning the first hit. It works from Vibe today, and, through the beta Agents API, from an agent carrying the document_library tool, returning chunks with tool_reference citations. Beside it sit mistral-embed (1024 dimensions) and Codestral Embed, and live-fetch connectors over the Model Context Protocol (MCP) for grounding against Drive, SharePoint, Slack, GitHub and the rest. Three findings hold the cell at moderate, in descending weight. First, the retrieval answers only inside Mistral's agents. The API reference documents create, upload, status, extracted-text and access endpoints for Libraries, and no search endpoint. A non-Mistral application can't call the index. That is the external-consumption gate that separated Salesforce (strong) from Qlik (moderate), and that OpenAI clears with a standalone vector-store search endpoint. Mistral doesn't clear it; Perplexity's Internal Knowledge Search was read the same way, a permission-honoring index that lives inside the product and answers only there. Second, the pipeline is fixed. Mistral parses, chunks, embeds with its own models and searches; the customer gets a chunk-size knob. No bring-your-own embeddings, no exposed hybrid or rerank configuration, no insertable stages. Rule 4: a slice, not the general function. Third, the management surface is beta. Every Libraries endpoint is labeled "(beta) Libraries API" and the SDK path is client.beta.libraries, while the product itself (Vibe Libraries, the agent tool) carries no such label. The cell scores the deployable capability and names the beta management API as a scope gate. Calibration puts Mistral between its two closest peers. OpenAI sits at moderate's upper boundary with a hosted estate plus a standalone search endpoint; Mistral has the estate and not the endpoint, so it reads below OpenAI. Anthropic sits at the lower boundary with a search experience and no estate; Mistral has a real persistent estate, so it reads above Anthropic. The frontier (rule 6) is the governed retrieval estate the enterprise operates, from AWS, Azure, Google, Databricks, VAST and Salesforce, and Mistral does not hand one over today.

Borrowed Judgment

The retrieval intelligence is Mistral's and opaque: parsing, chunking defaults, the embedding model, the agentic search loop's stopping rules. The capture is the decoupled kind: documents stay in SharePoint and Drive, feeling portable, while the library configurations, the embedded corpus and the agent bindings accumulate in Mistral's namespace. The embedding space is model-captive by the instrument's own carve-out, and no open-weight Mistral embedder exists to self-host it. The one authority reversal on this cell is structural: Mistral removed its indexed estate in August 2026 rather than extending it, so less of the enterprise's grounding sits in Mistral than a year ago. Less capability, less capture.

Working Notes

Watch-list, dateless: Search Toolkit (announced May 28, 2026, Apache 2.0, PyPI 0.0.13 on August 31, 2026) is the most general retrieval asset a model provider on this instrument has shipped. A customer-run framework where every stage is swappable (parsers, chunkers, embedders, query rewriters, rerankers, hybrid BM25 plus vector), over a customer-owned index (Vespa, pgvector, custom), with an evaluation harness. No Mistral-hosted anything required. On GA it is a Retained component on an open-source substrate and the grade is re-read toward strong. Mistral's own announcement calls it public preview, so it is a note, not a component. Watch-list: the API reference lists beta RAG ingestion pipeline configuration and search index endpoints with no docs page yet; at GA they are the standalone-retrieval fact this cell is missing. Structural finding: Knowledge Connectors (admin-driven indexes of Drive and SharePoint into EU data centers with replicated permissions) were scheduled for removal at the end of August 2026. The replacement is MCP connectors that fetch live and store nothing, with no automatic migration and indexed data deleted. Secondary-source plus absence from current docs. Residency finding for US buyers: Libraries are stored "in our European cloud infrastructure" regardless of endpoint. Regional endpoints (GA August 11, 2026) exclude Agents, Batch and Files, so the document_library tool is not reachable through api.us.mistral.ai or api.eu.mistral.ai. Inference pins to a jurisdiction. The retrieval estate does not (inferred from the Agents exclusion; docs do not name Libraries). Public evidence that moves the cell: a documented standalone library search endpoint returning chunks and scores.

Layer 1C · PipelinesData Movement & PipelinesTransform Primitive, No Pipelinedecides: absent · Absent

Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering

Vendor-Provided

NVIDIA-Provided

No NVIDIA Layer 1C Dependency

No movement platform exists to accelerate, and OCR serving silicon is invisible behind the API. The near-empty column continues.

Gap Analysis

The buyer gets one genuinely useful stage of a pipeline without building the pipeline. Document AI (OCR 4.1, generally available August 31, 2026) turns any document into schema-defined JSON. Pass a Pydantic, Zod or JSON schema with nested objects, arrays and enums, and document_annotation returns the whole document structured to it. bbox_annotation does the same for figures and charts. Document QnA answers over the page. It runs at $4 per 1,000 pages, $2 through the Batch API, and lands on Studio, SageMaker and Microsoft Foundry. Beside it: a Batch API for asynchronous bulk inference and prompt caching billed as cached reads. There is no movement platform here. No extract-transform-load (ETL) authoring, no schema mapping between systems, no change data capture, no lineage, no cost-aware movement, and no administrable KV-cache tier (caching is a billing discount, per the OpenAI ruling). Connectors fetch live and store nothing; the one sync product Mistral had, the indexed Knowledge Connectors, was removed at the end of August 2026. Movement between the enterprise's systems stays with the enterprise's integration stack. Document AI is the cell that pushes against three gaps in a row, so the reasoning is stated in full. It is a real transform, not intake: arbitrary schema (the customer's problem defines the output shape) and a choosable destination (JSON returns to the caller, who lands it anywhere). That is materially more than OpenAI's upload-parse-chunk-embed intake, which terminates in OpenAI's own retrieval surface, and it is the opposite of the Perplexity topology finding: this stage has a downstream edge a pipeline can attach to. It is also one stage. No sources, no scheduling, no destination management, no lineage. Day two, the second requirement arrives: move records between two systems. What runs it? Airflow, Glue or MuleSoft, and the layer's responsibility never transferred. Calibration decides it: Dell's moderate anchor carries a full orchestration engine (Dataloop), and Salesforce's strong runs on two owned movement platforms. A single transform primitive sits below Dell's floor. Gap, with Document AI named as the strongest 1C signal on the model-provider side of the instrument.

Borrowed Judgment

None at the layer's architectural function. Pipeline definitions, transformations, scheduling, lineage and destination choices stay in the enterprise's integration stack. What a Document AI customer does inherit is extraction judgment: a bank running know-your-customer extraction against custom schemas re-validates every schema if it swaps OCR vendors. That is switching cost, not accumulated pipeline authority, and it does not trigger the real-dependence guardrail.

Working Notes

Watch-list, dateless, and this is the item that moves the cell: Workflows (Public Preview) is Temporal-backed durable execution with retries, human-in-the-loop and OpenTelemetry. It has real scheduling: calendars, ISO 8601 intervals, five-field cron, six overlap policies, a 365-day catch-up window. Connectors-in-workflows is also in preview. That is an Airflow-shaped orchestration surface for AI steps; at GA it reads moderate here and scores at 2B independently. The overview page carries the preview banner and the scheduling subpage does not; the parent governs. Watch-list: Vibe Scheduled Tasks (Public Preview; once, daily, weekly, monthly, yearly; runs Work with inline connectors under pre-authorized permissions; no declared scope; built on Workflows). Even at GA that is the Perplexity shape, agent labor rather than a pipeline artifact. Watch-list: Search Toolkit ingestion (public preview; customer-run loaders for S3, Azure Blob and GCS into a customer-owned index), which strengthens 1B more than 1C. Sub-threshold, named not scored: Batch API, prompt caching. Fact question carried from the Anthropic row and answered here from the docs: no generally available Mistral service persistently synchronizes data between enterprise systems independently of a query; the only one that did was removed. Self-managed OCR deployment is "available to enterprise customers" by contact, not a documented product.

Layer 2A · OrchestrationInfrastructure OrchestrationMetering Edge, No Customer Plane (Compute Pending)decides: vendor · Ceded

GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization

Vendor-Provided

NVIDIA-Provided

No Customer-Facing GPU Plane

Layer 2A is where NVIDIA dependency concentrates for most of the map (Run:ai, GPU operators, schedulers), and Mistral exposes no plane in which to be dependent today. The Compute plane, when it ships, is NVIDIA reference architecture end to end and this column fills with it.

Gap Analysis

There is nothing to operate, and for the API buyer that is the offer. Mistral provisions, schedules and scales the serving fleet invisibly. What the customer touches is the commercial metering edge. Rate limits, workspaces that partition "teams, products, environments, or budgets," RBAC over members and groups, the Batch API queue. Regional endpoints (GA August 11, 2026) that pin inference to the EU or US at 1.1 times list. A Priority Tier with custom rate limits and a service-level agreement, in public preview. Every control is denominated in tokens, dollars and geography. None is denominated in compute, scheduling or placement. The Salesforce precedent decides the grade. Its 2A moderate required a deployable orchestration artifact (Runtime Fabric, run on the customer's own Kubernetes), with consumption metering ruled to weigh on authority rather than lift capability. Mistral has no such artifact on the customer's paper. Studio's self-hosted, hybrid and dedicated deployment models and Vibe's on-premises option are contact-sales custom deployments with no install guide, chart, operator or admin documentation; under the GA-gate a sales-gated engagement is not a doc-confirmed product. Self-deployment of the open-weight models through vLLM, TensorRT-LLM, TGI, SkyPilot or NIM is real and well documented. Whose orchestration is it? The customer's Kubernetes or Slurm, or the cloud's 2A for the Bedrock, Vertex and Foundry channels. Same reading as gpt-oss on the OpenAI row: the layer's function stays with whoever hosts the workload. Mistral supplies a model artifact and a deployment guide, not a plane. Regional endpoints deserve one honest sentence. They are the first placement control on any model-provider row that is geographic rather than commercial, and a genuine jurisdiction lever. They are still a consumption control over Mistral's fleet, not a scheduling surface the enterprise operates; Anthropic's inference-geography controls were read the same way.

Borrowed Judgment

Total for the serving substrate, and unauditable. Mistral decides accelerator selection, placement, failover and overload policy; the customer administers the token bucket, the workspace budget and, since August 2026, the region. The one thing a Mistral customer can audit that OpenAI and Anthropic customers can't is which legal jurisdiction served the request. No orchestration artifact accumulates customer-operable opinions. The decision-authority reading is Ceded and invisible: Mistral decides placement beneath its service, and the enterprise can't watch.

Working Notes

Watch-list, dateless: Mistral Compute (early access; first external customers March 2026; Microsoft European capacity agreement July 2026) advertises exactly the artifact this layer scores, Kubernetes-native orchestration on bare metal, topology-aware across cluster sizes, with RBAC mapped to Slurm accounts. That is the CoreWeave shape, open schedulers shipped as the product, which CoreWeave scored strong and Delegated on open substrate rather than Ceded. On GA this cell re-reads from gap toward that calibration and the decision-authority reading moves from Ceded-invisible to a genuine Delegated reading. Watch-list: Priority Tier (public preview, August 11, 2026; 99.5% uptime SLA at 1.75 times list); metering texture even at GA. Public evidence that reopens moderate: an install guide, chart or operator for a customer-operated Studio or Vibe deployment, or published enterprise terms allocating committed capacity by workspace or tenant in compute units rather than tokens. Read of the docs today: neither exists.

Layer 2B · RuntimeApplication Runtime & ExecutionFrontier Serving + Owned-Weights Exit, Agent API Pendingdecides: model · Delegated

Model serving, agent execution, inference APIs, distributed inference

Vendor-Provided

Model Serving via the OpenAI-Compatible Chat Completions InterfaceDelegated

GA. Frontier models consumed through the interface independent vendors implement, so integration opinions lift and providers hot-swap without rebuilding: the S3 of inference. Regional endpoints (EU and US, GA August 11, 2026), third-party open models under the same controls, batch, caching and structured outputs are texture inside the interface. The residual switching cost is behavioral recalibration, graduated by how much application logic lives in deterministic code outside the model.

Open-Weight Models, Apache 2.0 (Large 3, Small 4, Ministral 3, Devstral Small 2, Voxtral, Leanstral, Shieldstral)Retained

GA, self-managed, with Mistral-authored deployment docs for vLLM, TensorRT-LLM, TGI, SkyPilot and NIM. The enterprise owns the artifact and runs it on any serving product; prompts and evaluations calibrated to Large 3 run wherever Large 3 runs. The gpt-oss precedent at a higher capability class, validated on owned hardware in Labs 013 and 014.

Revenue-Capped Weights, Modified MIT (Mistral Medium 3.5, Devstral 2 123B)Ceded

GA as open weights on Hugging Face, with a license that bars use above $20 million of consolidated monthly revenue absent a discretionary commercial license from Mistral. The coding flagship line at both sizes. For the enterprise buyer this is not an open substrate: opinions built on these models cannot be taken to another product without Mistral's consent, so the component is scored as proprietary and named separately so the license is legible.

Vibe Work + Remote Code Runtime (Hosted)Ceded

GA (Work May 28, 2026; remote agents May 22, 2026). Hosted agent execution with connectors, skills, approvals and session persistence, on Mistral Cloud or deployed on-premises and in private cloud by contact-sales engagement. Connector configurations, approval grants, projects and scheduled tasks are Vibe's; only skills (open spec) port. Self-deployable is not Retained. Symmetric with Claude Code and Codex.

Vibe CLI (Open-Source, Apache 2.0)Retained

GA. Open harness the enterprise operates itself, documented against any OpenAI-compatible endpoint (a providers block with api_style "openai" and backend "generic", an OpenRouter example, and local vLLM, llama.cpp, LM Studio and Ollama servers), with MCP, Agent Skills and subagents. Retained on its own facts, the same fact the peer SDK calls rest on: the enterprise can leave Mistral's models and keep operating its agent code.

Customer-Created and Managed Tools (Function Calling / Client-Side)Retained

GA, and the only tool class supported on regional endpoints. The enterprise owns and operates the business logic, APIs, commands and deterministic validators invoked through function calls: the seam where control passes from instructing the model to executing code outside it. Replacing Mistral requires an adapter and re-evaluation, not rebuilding the capability. The opinions are the enterprise's and run against a multi-vendor interface, so they lift to another provider with an adapter and re-evaluation, not a rebuild. What the enterprise cedes at this seam is the decision to act: the model chooses whether, when and with which arguments to call, and that judgment is inherited under the serving component above.

MCP (Consumed by Vibe CLI and Vibe Work)Delegated

GA in Vibe. The multi-vendor protocol for connecting tools and data to agent runtimes, under neutral governance; MCP servers and their integrations serve any compatible host. The Agents API's MCP-connector registration is beta and not credited here.

NVIDIA-Provided

NVIDIA Is the Customer's Choice Here, Not Mistral's Imposition

Mistral's own fleet is GB200 and GB300; its models ship as NVIDIA NIM containers (Medium 3 and 3.5 on build.nvidia.com); its self-hosting guidance is sized in H100 and H200; Shieldstral launched inside the Open Secure AI Alliance alongside NVIDIA. The API boundary transmits none of it, and the open weights run through vLLM on any CUDA-class or ROCm substrate. This is the one row where the self-hosted path is where the enterprise meets NVIDIA, and can decline to.

Gap Analysis

Frontier inference on the interface the industry standardized on, with a real exit underneath it. Chat completions is OpenAI-compatible (a base_url change), with function calling, structured outputs, fill-in-the-middle, batch at half price, prompt caching, audio through Voxtral, OCR, and generally available regional endpoints that pin inference to the EU or the US. The model family is frontier-class. Medium 3.5 is 128B dense with 256K context, a vendor-stated 77.6% on SWE-bench Verified, and an independently measured Artificial Analysis Intelligence Index of 30, third of 63 open-weight models. Large 3 is a 675B mixture of experts. Small 4, Ministral 3, Codestral and Devstral fill out the line, and since August 11 third-party open models (Z.ai's GLM-5.2) are served under the same regional controls. It lands on Azure AI Foundry, Bedrock, Vertex, SageMaker, Snowflake Cortex, watsonx and Outscale. Then the thing no other frontier lab on the instrument does at this class. Large 3, Small 4, Ministral 3 and Devstral Small 2 ship under Apache 2.0, with Mistral-authored self-deployment docs for vLLM, TensorRT-LLM, TGI, SkyPilot and NVIDIA NIM. What does that buy? An enterprise can run 675B-parameter Mistral-lineage intelligence on its own hardware with no Mistral relationship at all. gpt-oss is o4-mini-class; Anthropic ships nothing. The owned-weights exit is not theoretical. Lab 013 served Devstral Small 2 through vLLM on one DGX Spark and cleared 15 of 22 gate-verified repairs. Lab 014 clustered Devstral 2 123B across two Sparks without incident, at 5.8 tokens per second. Lab 018 named Devstral as the tool-trained counter-evidence that off-the-shelf open weights can win self-hosting with no fine-tune at all. Agent execution is generally available at product altitude. Vibe Work (May 28, 2026) runs long, multi-step tasks across connectors with skills, per-function approvals and reasoning transparency. Vibe Code remote agents (May 22) run parallel coding sessions on a cloud runtime that persists when the laptop is closed, with VS Code, JetBrains and Zed integrations over the Agent Client Protocol. And the Vibe command-line interface (CLI) is open source under Apache 2.0, with MCP, skills and subagents. Per the official configuration docs, a providers block points it at any OpenAI-compatible endpoint, OpenRouter included, or at a local vLLM, llama.cpp, LM Studio or Ollama server. Two findings keep the cell honest. First, the developer-facing agent runtime is beta. The Agents, Conversations and Connectors APIs, with their built-in tool bindings (web search, code interpreter, image generation, the document_library tool, server-side handoffs), all sit under "(beta)" in the API reference, and Workflows (Temporal-backed durable execution) is Public Preview. The product surfaces those APIs manage (Vibe Libraries, Skills, connectors) are generally available and are scored at 1B and Layer 3 on the product; the API is the beta gate, not the feature. Studio's Agent Runtime pillar is, in shipping terms, those two things. So the hosted platform surface that OpenAI scores as the Responses API and Anthropic as server tools is not scoreable here today; rule 5's first-release gate, named in the cell. Second, model customization is a services lane, not a product. The fine-tuning API is deprecated ("no longer actively supported"). Forge (March 17, 2026) covers pre-training, supervised fine-tuning, direct preference optimization, reinforcement learning and synthetic data on Mistral, customer or cloud infrastructure. It is by sign-up, with no GA statement and no self-serve path. Strong is frontier-pegged (rule 6). Serving stands with the best; owned-weights distributed inference is the best on the instrument; agent execution is GA at product altitude and beta at API altitude. The alternative reading is moderate on rule 5. Strong holds because each of the layer's four named functions is GA in at least one deployable form, and the beta gate is stated rather than hidden. The tidy-story risk runs the other way here: the open-weights narrative is attractive enough to over-credit, which is why the license component and the beta gate sit in the cell rather than the notes.

Borrowed Judgment

The model-judgment concentration every model-vendor row carries, alignment, refusal behavior, safety tuning, lifecycle, with the widest set of mitigations on the instrument: version pinning, seven cloud channels, two jurisdictions, and weights you can serve yourself. The switching cost is architectural, not fixed (Lab 006: put the judgment in the constraints, not the weights), and here the exit runs all the way down to ownership. Except where it doesn't. Medium 3.5 and Devstral 2, the coding flagships, are published under a Modified MIT license. Its rights end at $20 million of consolidated monthly revenue unless Mistral grants a commercial license "at its sole discretion." For the enterprise this instrument serves, the best coding weights are proprietary. The open-weight story has to be read one model at a time. Whether the commercial license is routinely granted is a procurement fact; the litmus is whether the enterprise can take opinions built on the model to another product without Mistral's say-so, and above the cap it can't.

Working Notes

Watch-list, dateless. Agents, Conversations and Connectors (beta): on GA, a Ceded hosted-runtime component and the removal of the rule-5 gate. Workflows (Public Preview): on GA, a durable-execution component here and the 1C promotion noted there. Forge (no GA): customer stories say the trained model belongs to the enterprise and HTX operates its Forge-built models on its own infrastructure, so on GA it likely reads as a Retained owned-weights customization component, the inverse of the OpenAI fine-tuning call. Scheduled Tasks and Priority Tier (preview). Observability (private preview, Enterprise-only). Deprecated: the fine-tuning API; the Continue-based "Mistral Code Enterprise" JetBrains plugin (docs: "should not be used; the supported path is Vibe through ACP"), a wrapper-turnover finding of the OpenAI-row kind. Third-party evidence: Artificial Analysis measures Medium 3.5 at 136.6 tokens per second and $0.41 per task; the SWE-bench figure stays vendor-stated. Instrument follow-up for /reconcile: OpenAI's Codex CLI is itself Apache 2.0 and the OpenAI row scores the Codex bundle Ceded without a split. The Vibe CLI call here rests on documented operation against other vendors' models. Check the same fact there before any symmetric split. Public evidence that moves the cell: the beta label leaving Agents and Conversations; Forge product docs stating weight delivery and a purchasable path.

Layer 2C · ReasoningAgentic Infrastructure — The Reasoning Plane1 LABGuardrails and Approvals, Not a Governance Planedecides: absent · Absent

Policy-driven placement and resource coordination — the Autonomy Layer

Vendor-Provided

NVIDIA-Provided

No NVIDIA Layer 2C Dependency

NVIDIA controls no governance surface here. The only touchpoint is Shieldstral launching inside the Open Secure AI Alliance alongside NVIDIA, which is provenance, not dependency. The near-empty column continues.

Gap Analysis

What ships is real and deliberately placed. Custom Guardrails let the buyer declare moderation thresholds per category directly in a chat-completions request (and on beta Conversations and Agents) and get a hard 403 block, backed by mistral-moderation-2603 across ten categories including jailbreaking. Vibe Work approvals stop and ask before every interactive tool call (send, post, create, delete), with per-function "Always allow" pre-authorization that is per user, and organization-wide connector disable and skill force-enable for admins. The June 24, 2026 connectors release added connector-scoped API keys so an automated workload can't impersonate a person, plus per-workspace, per-tool access controls. Prompt and skill versioning (GA July 9) gives instruction assets ownership, classification labels, immutable versions and audit. Regional endpoints (GA August 11) pin inference to the EU or US. And Shieldstral 1.0 (August 4, Apache 2.0, one 16 GB GPU) takes a plain-language policy at inference time and returns a calibrated safety score, which the enterprise can self-host as its own guard model. None of it is a reasoning plane, and the calibration is crisp. The moderate cohort (Salesforce, AWS, Databricks, IBM, Cisco, VMware) earned it on a productized, generally available governance plane across an agent estate: gateway plus registry plus governed agent identity, enforcing policy at request time across runtimes. The gap cohort is inherited permissions plus at most a thin gateway, and Mistral's GA surface is that description almost exactly. No agent registry and no cross-runtime agent identity: agent versioning and aliases exist in the Agents API, in beta, and connector-scoped keys are user-bound workload credentials, not agent identities with declared intent. No gateway or broker across agents: Custom Guardrails act at one API boundary on content categories, and server-side handoffs (beta) are dispatch inside Mistral's runtime, where routing is not reasoning. Vibe approvals are runtime permissions inside one harness, the Anthropic precedent word for word: deterministic gates on the legality of an action, living inside the vendor's execution surface, strengthening 2B and not establishing 2C. What exists is platform administration of Mistral's own products, not a policy object that binds a subject to a capability with a scope and a duration and survives the session. The two universal findings are logged as universal rather than charged here. No live inference placement: regional endpoints are a static partition the customer chooses, the Salesforce GDPR reading, where you architect the boundary and the platform serves inside it. No deterministic outcome validation: Custom Guardrails and Shieldstral are probabilistic classifiers with thresholds. They score. They don't validate. You can't prompt your way to deterministic output, and you can't classify your way there either.

Borrowed Judgment

No scored Mistral 2C component exists, so the decision-authority reading is Absent: nothing offered, nothing inherited. In practice the enterprise discharges the function with its own code at the approval and gate boundary, or with another vendor's plane (Agent 365, Agent Fabric, AgentMinder registering Mistral-powered agents). Who governs a Mistral-powered agent estate today? Someone other than Mistral. That is the most consequential architecture decision a Mistral buyer makes, and it is made outside the Mistral relationship. The one Mistral-specific twist: Shieldstral hands the enterprise an open-weight guard it can run inside its own plane, a Retained artifact scored at 2B and a seam here. On-premises Vibe does not change this reading. Local-first moves the data plane inside the perimeter; the control plane is the same administrative surface wherever it runs.

Working Notes

Watch-list, dateless: the AI Registry (private beta) is the closest thing on any model-provider row to the moderate cohort's registry leg, with lineage, ownership, access controls, moderation policies and promotion gates over agents, models, datasets, judges, tools and workflows. On GA it is a real registry; it would still lack the gateway and the identity model, so the read is the gap/moderate boundary, the Nutanix Agent Gateway question in a different shape. Watch-list: Workflows human-in-the-loop gates and Judges (preview) are state gates and probabilistic evaluators, sub-threshold even at GA; Agents API versioning, aliases and server-side handoffs (beta). Sub-threshold GA signals named in prose: Custom Guardrails, Vibe approvals, connector-scoped keys, prompt and skill versioning, regional endpoints, Enterprise audit logs (actor, event, target, metadata; no export; tool calls and approval decisions not documented as logged). Lab 019, 'Local-first is not control-first' (August 26, 2026, https://labs.layer2c.com/labs/local-not-control), is cited as lab-derived evidence and mapped to the canon's requirements; the cell is graded on the instrument's own rules. Its finding that keys and consent are user-bound with no agent or workload identity maps to the canon's governed-identity leg, absent here. Its finding that no surface lets a ruling persist past the session maps to the canon's request-time policy engine enforcing authored policy, absent here; what exists is consent, per user, per session. Its finding that local-first is not control-first maps to the canon's self-deployable-is-not-Retained rule at application altitude, applied to on-premises Vibe and to co-developed platforms. Its finding that the application is always the caller is carried to Layer 3 for Vibe Work. Lab 013 (August 5, 2026) tags this cell in its own assessment_sources; it exercised a Mistral model inside a customer-built harness and is evidence for 2B serving, not for this plane, in either direction. Instrument note: live Infrastructure-2C, a policy engine deciding at request time where inference runs relative to the data, which model serves the request, and how cost, latency and compliance are arbitrated, does not ship from any vendor on this instrument. It is recorded here as a universal finding about the market, not as a charge against this vendor. Public evidence that moves the cell: a docs section for the AI Registry with promotion gates GA; a gateway or policy engine evaluating tool access across agents at request time; agent identity across runtimes.

Layer 3 (+1) · ApplicationsAI Application Layer — The Value PlaneFirst-Party Value Plane: Vibe Work, Vibe Code, Applied AIdecides: vendor · Ceded

AI-powered business capabilities — business logic, workflow automation

Vendor-Provided

Vibe Work (Agentic Knowledge-Work Application)Ceded

GA May 28, 2026, web and mobile, with Team and Enterprise tiers. Multi-step task execution across connected work systems with Libraries, Projects, Skills, approvals and transparency, deployable on Mistral Cloud, in a private cloud or on-premises. The proprietary application and workspace estate; caller-only, with no outbound interface; on-premises deployment does not change who holds it.

Vibe Code (Developer Application)Ceded

GA across CLI, VS Code, JetBrains, Zed and Vibe Code Web, with remote agents since May 22, 2026. Remote sessions, IDE integrations and delegation patterns are the application relationship, and it is Mistral's. The open-source CLI is the runtime harness scored Retained at 2B and is not double-counted here. Symmetric with Codex and Claude Code.

Applied AI: Co-Developed Custom PlatformsCeded

Shipping in named engagements (CMA CGM's MAIA from June 1, 2026; Stellantis; ASML; HTX; DSO). Mistral engineers embedded in the customer, building on Mistral's stack under multi-year agreements. The Palantir Forward Deployed Engineering precedent: value-accelerating, and it deepens the vendor-operated dependency, because the delivered platform does not lift.

Customer-Authored Skills (Agent Skills Specification)Retained

GA in Vibe Work and Vibe Code. Per Mistral's docs, "Skills follow the open Agent Skills standard, originally developed by Anthropic and now adopted across multiple agentic clients." A SKILL.md folder with metadata, instructions and supporting files, shareable workspace-wide and force-enabled by admins. Plain-file procedural knowledge the enterprise owns; matches the Anthropic call.

MCP Connector EcosystemDelegated

GA. Featured and custom connectors over the Model Context Protocol, with admin controls per workspace and per tool. Menu-altitude Delegated matching the OpenAI and Anthropic calls: connector implementations and source-system authority stay outside Mistral, while Vibe-specific installation and configuration must be recreated on another host.

NVIDIA-Provided

No NVIDIA Layer 3 Dependency

NVIDIA supplies nothing at application altitude. The near-empty column continues.

Gap Analysis

Two first-party application relationships and a third lane no peer runs the same way. Vibe Work (GA May 28, 2026, web and mobile) is the agentic knowledge-work application. It takes a multi-step task and executes it across Google Workspace, Outlook, SharePoint, Slack, GitHub and the connector estate. Libraries and Projects supply context, Skills supply repeatable method, and reasoning summaries, tool-call transparency and stop-and-ask approvals sit on every interactive action. Chat mode is being folded into it. Vibe Code is the developer application: remote agents (GA May 22) running parallel cloud sessions that persist when the laptop closes, Vibe Code Web, VS Code, JetBrains and Zed over the Agent Client Protocol, and the open-source CLI. The Enterprise tier adds SAML SSO, white-label deployment on a custom hostname, audit logs, and deployment on-premises, in a private cloud, or on Mistral Cloud with full data residency, the deployment choice no US peer offers at this altitude. The third lane is applied AI, co-developed, with Mistral's own engineers embedded in the customer. CMA CGM's MAIA, "Powered by Mistral," is an agentic platform orchestrating agents against business knowledge and internal applications. It is rolling out from June 1, 2026 to nearly 80,000 employees across CMA CGM, CEVA Logistics and CMA Media, with around twenty Mistral engineers seated in Marseille under a 100 million euro, five-year agreement. Stellantis (a company-wide custom industrial model), ASML (co-creation on manufacturing data), HTX and DSO (sovereign custom models, on-premises) are the same motion. Calibration is strong, frontier-pegged, and the peers set it. OpenAI is strong at assistant altitude with Frontier pre-GA; Anthropic is strong because Cowork moved agentic knowledge work to GA; Perplexity is strong on the widest estate and a browser. Mistral stands with all three: Work is Cowork-class and GA, Code is Codex- and Claude-Code-class. It lacks a browser and any vertical product, and it adds sovereign deployment plus an embedded-engineer lane. The assistant-altitude caveat travels as the peers carry it: systems-of-record execution under commit-boundary governance is still Salesforce's story. MAIA-type builds do execute in systems of record, but as custom engagements, not product.

Borrowed Judgment

Capture here is coupled and visible, as on every model-provider row. The workforce's projects, libraries, delegation habits, per-user approval grants and workspace configuration accumulate in Vibe and lift nowhere, per user, per day, and they survive model interchangeability entirely: swap the model underneath and none of it moves. Three Mistral-specific sharpenings. Vibe Work is always the caller. No API, webhook or CLI drives a Work run from outside; the Slack trigger is Vibe Code only and "in progressive rollout," and Workflows' webhook plugin inherits its parent's Public Preview. So the workflow can't be composed into anything the enterprise already runs, the reading the Perplexity row gives Portable Computer. Local-first is not control-first: on-premises Vibe and a co-developed platform move the data plane inside the customer's perimeter and do not move who holds the application. And the applied-AI lane is Palantir-shaped, not IBM-shaped. The IBM Consulting Delegated call rests on a substitutable integrator delivering the same outcome. Mistral's embedded engineers are Mistral's people building on Mistral's stack, and the delivered platform is inseparable from it. Could another integrator deliver MAIA? Not this MAIA. A systems integrator can implement Mistral products (Capgemini and Accenture are both on the customer list), which is a different thing from this engagement and is not a scored component.

Working Notes

Watch-list, dateless: Slack trigger for Vibe Code ("in progressive rollout"); Scheduled Tasks (Public Preview); Chat mode sunset into Work. Wrapper turnover: the Continue-based "Mistral Code Enterprise" JetBrains plugin is deprecated in favor of Vibe through ACP, the OpenAI-row finding at smaller scale. Inference flag: the solutions page presents verticals (financial services, healthcare, manufacturing) as packaged offerings; the linked material reads as reference architectures and engagements, not stock-keeping units, and MAIA is a customer's platform rather than a Mistral product. Skills portability rests on the Agent Skills specification being adopted across implementations, which Mistral's own docs state and which is doc-supported rather than doc-confirmed; the Retained call survives either way because the artifact is a plain-file folder in the enterprise's possession. Public evidence that would settle the open questions: a doc stating who operates an on-premises Vibe deployment; a Vibe Work API or trigger doc; a data-export doc reconstituting projects, libraries, skills and instructions in a vendor-neutral format.

Summary Finding

Mistral AI is a frontier lab that sells inference on the industry's standard interface, publishes frontier-class weights under Apache 2.0, and runs a first-party application estate in Vibe with an embedded-engineer applied-AI lane behind it. It is strong at Layer 2B and Layer 3 and a gap everywhere else. It is headquartered in Paris and it sells in the United States. A US inference region since August 2026. Distribution through Microsoft Foundry and Copilot Studio, Bedrock, Vertex, SageMaker and Snowflake. A Palo Alto office, a US general manager who now runs global revenue, and a Seattle-based CMO hired in June. The infrastructure story (Mistral Compute) and the platform story (Studio's Agents, Workflows, Registry and Observability) are both real and both pre-GA. Every cell rests on public documentation.

The capture splits by altitude, and the split is the finding. At 2B it is decoupled and mostly reversible. The OpenAI-compatible interface keeps integration opinions portable. Large 3, Small 4, Ministral 3 and Devstral Small 2 give the enterprise an owned-weights exit no other frontier lab matches at this class. The capture that does exist there is quiet. Medium 3.5 and Devstral 2, the coding flagships, carry a Modified MIT license that stops at $20 million of monthly revenue, which makes them proprietary for exactly the buyer this instrument serves. At Layer 3 the capture is coupled and visible. Vibe's projects, libraries, approval grants and delegation habits accumulate per user per day. Work can't be driven from outside. A co-developed platform is inseparable from the people and stack that built it. On-premises deployment changes where that runs. It does not change who holds it.

Mistral's distinctive offer is jurisdiction as a selectable property, and the choice is real in both directions. Pin inference to the EU or the US. Deploy Vibe inside your own perimeter. Serve the Apache weights without a Mistral contract. That is local-first, and it is delivered. Control-first is not. There is no governance plane, no retrieval estate the enterprise operates, no orchestration artifact on the customer's paper, and a rule surface that stops at platform administration. The data plane moves where you put it; the control plane does not come with it, in Frankfurt or in Virginia. The trade is a real exit at the model in exchange for a thinner platform above it, a license gate on the best coding weights, and an application estate as captive as anyone's. Buy the residency and the weights. Do not mistake either for authority.

4+1 Layer AI Infrastructure Model · Vendor Assessment Series · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com