Executive Summary: Hugging Face (Hub + Team and Enterprise Plans + Storage Buckets + Inference Endpoints + Inference Providers + Jobs + Spaces + the Open-Source Libraries)

Hugging Face is the distribution point for open models and, since September 3, 2026, the subject of an announced NVIDIA acquisition for $12.93 billion that the NVIDIA post says will leave it 'an open platform' with 'NVIDIA compute ... not required'; the agreement hasn't closed and nothing in the documentation has changed, so the row reads Hugging Face as it ships today. The map reads it moderate at four layers and gap at three, with no strong cell: an open-serving platform whose agent, pipeline, and application surfaces are libraries, a runner, and a host rather than products. Layer 2B is the heartland: Inference Endpoints deploy any Hub model on a chosen instance in AWS, Azure, or Google Cloud Platform (GCP) under vLLM, SGLang, or TEI (TGI is in maintenance mode), Inference Providers route chat completions through an OpenAI-compatible API to eighteen partners with policies, failover, and billing at the provider's rates, and the libraries are Apache 2.0 code the enterprise runs; what's missing is a managed agent runtime, and the one full agent harness (smolagents) is experimental by its own documentation, so the cell reads moderate on rule 5 with strong argued. Layer 0 is a gap by shape: GPU, CPU, and TPU time by the minute (Jobs, Spaces, ZeroGPU) is a setting on Hugging Face's products, on capacity it rents in the hyperscalers, with nothing the enterprise administers or decides. Layer 1A is moderate on the Hub as a governed artifact registry (Git repositories on Xet, S3-compatible buckets, resource groups, gating, audit logs, SSO and SCIM, storage regions) that governs repositories, not tables. Layer 1B is moderate on the open embedders, the engine that serves them (TEI, Sentence Transformers), and the datasets library's indexes, with no managed retrieval estate. Layer 1C is moderate on Jobs as a hosted, event-triggered runner for the enterprise's data code plus Distilabel, with no connectors, CDC, or lineage. Layer 2A is moderate on endpoint autoscaling with replica bounds, instance choices, Jobs flavors, and resource-group spend caps. Layer 3 is a gap under the marketplace ruling: Spaces host a million of other people's Gradio and Docker apps and the Playground and Data Studio are developer tools, with no first-party business application. Layer 2C is a gap: a routing and budget gateway over model calls, a trace viewer, and a catalog of agent artifacts are legs, not a plane.

The capture is the lowest on the map for a vendor at this altitude, because the artifacts and the engines are open. Permissively licensed open weights the enterprise pulls are its own; transformers, tiny-agents, huggingface_hub, TEI, TGI, Distilabel, datasets, Gradio, the Deep Learning Containers, and the Xet protocol are Apache 2.0; Git repositories clone out whole; buckets speak a subset of S3; Inference Endpoints and the Providers chat router sit behind interfaces the enterprise could point elsewhere. That is why 11 of the 21 components read Retained or Delegated at the engine and the artifact. What is Hugging Face's and stays Hugging Face's: the hosted capacity and its flavor catalog, the organization model and Hub workflow (resource groups, gating, audit, tokens, service accounts, network security, storage regions, pull requests and discussions), the Inference Endpoints and Jobs control planes, the Providers routing and billing surface and its non-chat client, the MCP server over the Hub, Spaces as a host, and the Playground, Data Studio, and catalogs; and weights under their owners' restrictive terms are the owners', not the enterprise's, wherever they're served.

The buyer's trade: every open model on any hardware behind interfaces it already uses, and a governed registry for its weights and data, in exchange for a hosted plane whose placement, catalog, and organization model are Hugging Face's, with no agent runtime, no pipeline product, and no application above the developer's desk. The decision-authority readings follow: vendor / Ceded at Layer 0 (invisible; Hugging Face absorbs the substrate beneath its products) and 2A (a flavor is a menu and a replica bound is a bound, neither an override of Hugging Face's placement); vendor / Delegated at 1A and 1B on per-object controls (repository roles, per-request model choice); code / Retained at 1C, where the runner executes what the enterprise wrote; model / Delegated at 2B because the agent loop runs in the enterprise's own process, where the gate before an effect is the enterprise's to write; Absent at 2C and Layer 3. What would move cells: a first-party business application (Layer 3), the engines and tool support leaving their maturity gates (2B), smolagents leaving its experimental label (a 2C leg), a Hugging Face vector index (1B), durable-execution semantics on Jobs (1C), capacity the enterprise administers (Layer 0), and whatever the NVIDIA acquisition does to the compute offer once it closes.

Layer-by-layer status: Layer 0 (Not a Customer Surface: GPU, CPU, and TPU Time by the Minute Attached to Jobs, Spaces, ZeroGPU, and Endpoints, on Capacity Hugging Face Rents in the Hyperscalers; No Fabric, No Bare Metal, Nothing You Administer), Layer 1A (The Hub as an Artifact Registry: Git Repositories on Xet, S3-Compatible Storage Buckets, Resource Groups, Gating, Audit Logs, SSO, and Storage Regions; No Table, No Policy Below the Repository), Layer 1B (Open Embedding and Reranker Models Served Anywhere (Text Embeddings Inference, Sentence Transformers, Inference Endpoints, Inference Providers) and the datasets Library's FAISS and Elasticsearch Indexes; No Managed Retrieval Estate), Layer 1C (Jobs as a Hosted Data-Processing Runner (Schedules, Webhooks, Parallel Runs, Streamed and Mounted Datasets and Buckets) Plus Distilabel and the datasets Library; No Connectors, No CDC, No Lineage), Layer 2A (Inference Endpoints Autoscaling With Replica Bounds and Scale-to-Zero, Instance and Region Pins, Jobs Flavors and Timeouts, Resource-Group Spend Caps; No Scheduler Over a Fleet You Own), Layer 2B (Serve Any Open Model Anywhere: Inference Endpoints on Three Clouds (vLLM, SGLang, TEI; TGI in Maintenance) and an OpenAI-Compatible Chat Router Across Eighteen Providers; the Agent Libraries Are Experimental or Thin, and There's No Managed Agent Runtime), Layer 2C (Organization Governance for People and Tokens, a Tool Server, and a Place to Upload Agent Traces; Not a Plane Over Agents), Layer 3 (+1) (No Business Application: Spaces Host a Million of Other People's Gradio and Docker Apps, the Inference Playground and Data Studio Are Developer Tools, and Gradio Is an Application Framework Rather Than an Application).

Assessment framework: 4+1 Layer AI Infrastructure Model. Scoring model: Decision Authority Placement Model (DAPM) — Retained, Delegated, or Ceded. Published by The CTO Advisor LLC (DBA The Advisor Bench). Author: Keith Townsend. Date assessed: September 10, 2026. Version: v1.0 - 4+1 v2: Authority Split.

Hugging Face (Hub + Team and Enterprise Plans + Storage Buckets + Inference Endpoints + Inference Providers + Jobs + Spaces + the Open-Source Libraries)

Mapped to the 4+1 Layer AI Infrastructure Model

v1.0 - 4+1 v2: Authority SplitAssessed September 10, 2026Sources & revision history
ACTIVE ASSESSMENT

Summary Finding

Hugging Face is the distribution point for open models and, since September 3, 2026, the subject of an announced NVIDIA acquisition for $12.93 billion that the NVIDIA post says will leave it 'an open platform' with 'NVIDIA compute ... not required'; the agreement hasn't closed and nothing in the documentation has changed, so the row reads Hugging Face as it ships today. The map reads it moderate at four layers and gap at three, with no strong cell: an open-serving platform whose agent, pipeline, and application surfaces are libraries, a runner, and a host rather than products. Layer 2B is the heartland: Inference Endpoints deploy any Hub model on a chosen instance in AWS, Azure, or Google Cloud Platform (GCP) under vLLM, SGLang, or TEI (TGI is in maintenance mode), Inference Providers route chat completions through an OpenAI-compatible API to eighteen partners with policies, failover, and billing at the provider's rates, and the libraries are Apache 2.0 code the enterprise runs; what's missing is a managed agent runtime, and the one full agent harness (smolagents) is experimental by its own documentation, so the cell reads moderate on rule 5 with strong argued. Layer 0 is a gap by shape: GPU, CPU, and TPU time by the minute (Jobs, Spaces, ZeroGPU) is a setting on Hugging Face's products, on capacity it rents in the hyperscalers, with nothing the enterprise administers or decides. Layer 1A is moderate on the Hub as a governed artifact registry (Git repositories on Xet, S3-compatible buckets, resource groups, gating, audit logs, SSO and SCIM, storage regions) that governs repositories, not tables. Layer 1B is moderate on the open embedders, the engine that serves them (TEI, Sentence Transformers), and the datasets library's indexes, with no managed retrieval estate. Layer 1C is moderate on Jobs as a hosted, event-triggered runner for the enterprise's data code plus Distilabel, with no connectors, CDC, or lineage. Layer 2A is moderate on endpoint autoscaling with replica bounds, instance choices, Jobs flavors, and resource-group spend caps. Layer 3 is a gap under the marketplace ruling: Spaces host a million of other people's Gradio and Docker apps and the Playground and Data Studio are developer tools, with no first-party business application. Layer 2C is a gap: a routing and budget gateway over model calls, a trace viewer, and a catalog of agent artifacts are legs, not a plane.

The capture is the lowest on the map for a vendor at this altitude, because the artifacts and the engines are open. Permissively licensed open weights the enterprise pulls are its own; transformers, tiny-agents, huggingface_hub, TEI, TGI, Distilabel, datasets, Gradio, the Deep Learning Containers, and the Xet protocol are Apache 2.0; Git repositories clone out whole; buckets speak a subset of S3; Inference Endpoints and the Providers chat router sit behind interfaces the enterprise could point elsewhere. That is why 11 of the 21 components read Retained or Delegated at the engine and the artifact. What is Hugging Face's and stays Hugging Face's: the hosted capacity and its flavor catalog, the organization model and Hub workflow (resource groups, gating, audit, tokens, service accounts, network security, storage regions, pull requests and discussions), the Inference Endpoints and Jobs control planes, the Providers routing and billing surface and its non-chat client, the MCP server over the Hub, Spaces as a host, and the Playground, Data Studio, and catalogs; and weights under their owners' restrictive terms are the owners', not the enterprise's, wherever they're served.

The buyer's trade: every open model on any hardware behind interfaces it already uses, and a governed registry for its weights and data, in exchange for a hosted plane whose placement, catalog, and organization model are Hugging Face's, with no agent runtime, no pipeline product, and no application above the developer's desk. The decision-authority readings follow: vendor / Ceded at Layer 0 (invisible; Hugging Face absorbs the substrate beneath its products) and 2A (a flavor is a menu and a replica bound is a bound, neither an override of Hugging Face's placement); vendor / Delegated at 1A and 1B on per-object controls (repository roles, per-request model choice); code / Retained at 1C, where the runner executes what the enterprise wrote; model / Delegated at 2B because the agent loop runs in the enterprise's own process, where the gate before an effect is the enterprise's to write; Absent at 2C and Layer 3. What would move cells: a first-party business application (Layer 3), the engines and tool support leaving their maturity gates (2B), smolagents leaving its experimental label (a 2C leg), a Hugging Face vector index (1B), durable-execution semantics on Jobs (1C), capacity the enterprise administers (Layer 0), and whatever the NVIDIA acquisition does to the compute offer once it closes.

Strength
Moderate
Gap
Partner
Layer 0 · ComputeCompute & Network FabricNot a Customer Surface: GPU, CPU, and TPU Time by the Minute Attached to Jobs, Spaces, ZeroGPU, and Endpoints, on Capacity Hugging Face Rents in the Hyperscalers; No Fabric, No Bare Metal, Nothing You Administerdecides: vendor · Ceded▼

Raw compute, networking, and acceleration fabric

Vendor-Provided

NVIDIA-Provided

NVIDIA Dominates the Hosted GPU Menu; CPU, TPU, and Inferentia and Trainium Paths Are the Alternatives

Spaces and Jobs GPU flavors are NVIDIA (T4, L4, L40S, A10G, A100; H100 and 8x H100 removed from Spaces December 2025); ZeroGPU runs on NVIDIA RTX Pro 6000 Blackwell slices. Jobs also list CPU flavors (Basic through Performance) and TPU flavors, Inference Endpoints list Intel Sapphire Rapids CPU instances and non-NVIDIA accelerators, and the cloud partner documentation covers AWS Trainium and Inferentia and Google TPUs. On September 3, 2026 NVIDIA announced an agreement to acquire Hugging Face for $12.93 billion (definitive agreement dated September 2 per NVIDIA's 8-K); the post says 'NVIDIA compute will not be required to build on or deploy through Hugging Face.' Announced, not closed; nothing in the documentation has changed.

◆ Gap Analysis

Hugging Face sells compute the way a cloud does, one flavor at a time, on capacity it operates in the hyperscalers. Jobs run any command in any Docker image on a chosen hardware flavor (CPU, GPU, or TPU; a100-large in the quick start), billed by the minute with a default 30-minute timeout, from the hf CLI or Python; Spaces run Gradio or Docker apps on a price list from a free 2 vCPU tier through T4, L4, L40S, and A100 instances by the hour; ZeroGPU hands Spaces dynamic slices of RTX Pro 6000 Blackwell GPUs against a daily quota (five minutes free, forty on PRO, sixty for Enterprise members, then $1 per ten minutes of credit); Inference Endpoints place a model on an instance type in AWS us-east-1 or eu-west-1, Azure eastus, or Google Cloud Platform (GCP) us-east4, all in Hugging Face's accounts. Storage is the Hub's own (Git repositories on the Xet backend and S3-compatible buckets, US and EU regions, Asia-Pacific and GCC 'coming soon'). The buyer gets GPU time without a cloud account and a container runner that takes its own images, as settings on Hugging Face's products rather than as infrastructure it holds. The architect's concern is that there's no substrate in any of it. No fabric, no bare metal, no virtual private cloud of the enterprise's own: a Private Endpoint shows up in the enterprise's VPC as one elastic network interface over AWS PrivateLink and nothing else (the security page also names Azure Private Link; the guide and FAQ document AWS, and no GCP private path is documented), the regions are four, the underlying instance is Hugging Face's choice, and Hugging Face Generative AI Services (HUGS), the one product that ran in the enterprise's own cloud account, is deprecated (September 2025). Calibration: Layer 0 reads by the vendor's shape (the September 15, 2026 ruling under the exposure test). Hugging Face is a service, not an IaaS: the flavors are settings on Jobs, Spaces, and Endpoints, on capacity Hugging Face rents from the hyperscalers, which is Cohere's and Perplexity's shape ('serving fleet, not a customer surface') and Dataiku's (a compute quota inside its product), not Cloudflare's compute the enterprise sizes as infrastructure. There's nothing for the enterprise to decide at this layer. Gap, the service reading; a Stack Builder wouldn't build a Layer 0 on it.

◆ Borrowed Judgment

Ceded and invisible: Hugging Face absorbs the layer beneath its products. The enterprise picks a flavor, a region for an endpoint, and a timeout as product settings; Hugging Face picks the cloud account, the instance, the placement, and the retirement schedule (H100 removed from Spaces in December 2025). Vendor decides, not visible, not overridable, Ceded, the OpenAI, Anthropic, and Perplexity reading. The Deep Learning Containers the enterprise runs on its own SageMaker, Vertex AI, GKE, or Azure ML are scored at Layer 2B; the substrate there belongs to the cloud's row.

◆ Working Notes

Ruled September 15, 2026: moderate to gap on the shape line (a service vendor's hosted capacity isn't a Layer 0 surface, however the flavors are exposed); gap had been argued on the Cohere and Dataiku reading and moderate on Cloudflare's. Named, not scored: the Jobs, Spaces, ZeroGPU, and Endpoint flavor menus (product settings, scored with their products at 2A and 2B), HUGS (deprecated September 2025), Spaces H100 flavors (removed December 2025). Inference Endpoints regions per the FAQ: AWS us-east-1 and eu-west-1, Azure eastus, GCP us-east4. Public evidence that moves the cell: Hugging Face selling capacity the enterprise administers as its own (a dedicated cluster, a bare-metal tier), or the NVIDIA acquisition closing with a compute offer that changes the documentation.

Layer 1A · StorageData Storage & GovernanceThe Hub as an Artifact Registry: Git Repositories on Xet, S3-Compatible Storage Buckets, Resource Groups, Gating, Audit Logs, SSO, and Storage Regions; No Table, No Policy Below the Repositorydecides: vendor · Delegated▼

Durable, governed data foundation — the Governance Catalog that Layer 2C queries

Vendor-Provided

Hub Repositories for Models, Datasets, and Spaces (Git With Full History; Xet Storage Backend, Apache 2.0; Git LFS Legacy)Delegated

Versioned artifacts behind Git and an open transfer protocol; a repository clones out whole with its history. Delegated, the standard-interface reading. The Hub's pull requests and discussions (no forks, refs/pr/N over Git, pin, lock, hide) are Hub workflow and sit in the Ceded chip below.

Storage Buckets Through the S3-Compatible API at s3.hf.co (Mutable, Non-Versioned Object Storage on Xet; US and EU Regions; AES-256 at Rest)Delegated

Delegated on the S3 door alone: S3 tooling works against it and data copies out with rclone or s5cmd, though the gateway is single-region and supports no ACLs, bucket policies, versioning, lifecycle rules, server-side encryption, notifications, user metadata, or cross-namespace copy. The hf:// paths, the hf CLI, and Jobs and Spaces volume mounts are Hugging Face's own access conventions and don't carry the reading. The Cloudflare R2 reading, with the gaps named.

Organization Governance and Hub Workflow (Resource Groups With Per-Repository Roles on Team and Enterprise; Per-Feature Access, Cost Attribution, and Spend Caps on Enterprise and Above; Gating and Group Collections; Audit Logs; Fine-Grained Tokens; Service Accounts on Enterprise; Network Security and Managed SSO on Enterprise Plus; Storage Regions; Pull Requests and Discussions)Ceded

Hugging Face's organization model and collaboration workflow over Hugging Face's repositories, with no second implementer. Ceded, the CoreWeave dedicated-tier and Confluent Stream Governance reading.

SSO and SCIM (SAML 2.0 and OIDC; Basic SSO on Enterprise, Managed SSO on Enterprise Plus)Delegated

Standard federation protocols against the enterprise's own identity provider; what they provision into is the Hub. Delegated at the protocol.

Dataset Viewer, Data Studio (Text-to-SQL Over Datasets), Model Cards and License Metadata, Partner Scanning (Protect AI Over Public Repository Files, JFrog Over Model Files)Ceded

The Hub's own inspection and metadata surfaces; the scanners are partners' results shown on the Hub. Ceded.

NVIDIA-Provided

No NVIDIA Dependency at This Layer

Repositories, buckets, and the governance surfaces run on CPUs and object storage.

◆ Gap Analysis

Hugging Face's data foundation is a registry of artifacts, not a lakehouse. Models, datasets, and Spaces are Git repositories with full history, pull requests, and discussions, stored on the Xet backend (chunk-level deduplication; Git Large File Storage (LFS) remains supported as legacy); Storage Buckets are the mutable, non-versioned counterpart for checkpoints, logs, and training data, reachable as hf:// paths, as volume mounts in Jobs and Spaces, and through an S3-compatible API at s3.hf.co (AWS CLI, boto3, s5cmd) that's 'currently single-region' and doesn't support access control lists (ACLs), bucket policies, tagging, versioning, lifecycle rules, server-side encryption, bucket notifications, user metadata, ListObjectsV1, cross-namespace copy, or GetObject preconditions (a partial list from the documentation); everything is AES-256 at rest and TLS in transit, SOC 2 Type 2, General Data Protection Regulation (GDPR) compliant. Storage is priced per terabyte with egress and the content delivery network (CDN) included (Team 12 TB base plus 1 TB per seat, Enterprise 200 TB, Enterprise Plus 500 TB). Governance is tiered by plan: resource groups with no_access, read, contributor, write, and admin roles per repository (Team and Enterprise) and, since August 12, 2026, per-feature access, cost attribution, and monthly spend caps per group (Enterprise and above); gating with group collections and, on Enterprise Plus, location-based blocking; audit logs downloadable as JSON (new settings-change events since June 16, 2026); single sign-on (SSO) over Security Assertion Markup Language (SAML) 2.0 and OpenID Connect (OIDC) with System for Cross-domain Identity Management (SCIM) provisioning (Basic SSO on Enterprise invites existing Hugging Face users, Managed SSO on Enterprise Plus replaces the Hugging Face login and owns the user lifecycle); fine-grained tokens with administrator approval (revocation on Enterprise and above); service accounts (Enterprise and Enterprise Plus); network security by allowlisted outbound IP ranges (Enterprise Plus, after a manual verification with Hugging Face engineers); storage regions (US and EU, Asia-Pacific and GCC 'coming soon'); model cards with a license field the Hub recognizes (Apache 2.0, MIT, the OpenRAIL family, the Llama and Gemma terms); and partner scanners (Protect AI over public repository files, JFrog over model files). The buyer gets a governed place for weights, datasets, and checkpoints, with an S3 door. The architect's concern is altitude. The unit of governance is the repository and the bucket; there's no table, no column, no row-level policy, no catalog of anything outside the Hub, and the S3 door is a gateway with a short feature list. Calibration: Databricks, Snowflake, and Cloudera read strong on lakehouses whose governance owns the tables' grants; Cloudflare moderate on an S3-compatible store plus edge databases; CoreWeave moderate on open default storage with a captive governance tier; Kamiwaza moderate on a derived-artifact layer; Confluent moderate on a governed log. Hugging Face is CoreWeave's and Kamiwaza's shape: open storage of AI artifacts plus a captive governance layer scoped to them. Moderate.

◆ Borrowed Judgment

Low at the storage, captive at the plan. A Git repository clones out with its history, and the Xet protocol is Apache 2.0 (xet-core): Delegated for the repositories, the standard-interface reading. Buckets speak a subset of S3, so tooling and data lift: Delegated, the Cloudflare R2 reading, with the missing features named. Resource groups, gating, audit logs, tokens, service accounts, network security, and storage regions are Hugging Face's organization model with no second implementer: Ceded, the CoreWeave dedicated-tier reading; SSO and SCIM are standard protocols, but what they provision into is the Hub. The runtime call at this layer is the Hub enforcing visibility, gating, and resource-group roles the enterprise set per repository, with override per object: vendor decides, visible (audit logs), overridable, Delegated, the Databricks and Cloudflare 1A reading.

◆ Working Notes

Storage regions 'coming soon' for Asia-Pacific and GCC carry no date and aren't watch-listed. The S3 gateway's unsupported list (ACLs, bucket policies, tagging, versioning, lifecycle, server-side encryption) is from the documentation and is why the bucket chip is Delegated with a caveat rather than a peer of R2. Publisher Analytics and the unique-downloader export (Enterprise Plus add-on) are a publisher's tool, not the enterprise's governance. Public evidence that moves the cell: a catalog or policy surface over data the enterprise holds outside the Hub, or the bucket gateway reaching multi-region with policies.

Layer 1B · RetrievalContext Management & RetrievalOpen Embedding and Reranker Models Served Anywhere (Text Embeddings Inference, Sentence Transformers, Inference Endpoints, Inference Providers) and the datasets Library's FAISS and Elasticsearch Indexes; No Managed Retrieval Estatedecides: vendor · Delegated▼

Low-latency retrieval for RAG — vector/hybrid search, context windows

Vendor-Provided

Text Embeddings Inference (TEI, Apache 2.0; Embedding, Reranking, and Classification Serving on CPU and GPU) + Sentence Transformers (Apache 2.0)Retained

Open-source serving and training software the enterprise runs anywhere. Retained on the open-source seam.

datasets Library Search Indexes (Apache 2.0; FAISS and Elasticsearch Indexes Over a Dataset Column; Save and Load)Retained

A library the enterprise runs in its own process, with the index it built. Retained on the open-source seam, the NVIDIA cuVS precedent for scoring a library here.

Open Embedding and Reranking Models Served on Inference Endpoints or Routed Through Inference Providers (Feature Extraction Through the Hugging Face InferenceClient)Delegated

Open weights the enterprise could pull and serve itself under TEI, hosted on Hugging Face's or a provider's hardware. Delegated on the open-weights embeddings ruling, not on the interface (the Providers path is Hugging Face's own client); the carve-out reads Ceded only where the weights are the host's.

Hub Search Through the Hugging Face MCP Server (hf_fs: Semantic Search Over Documentation and Spaces; Search Over Models, Datasets, and Papers)Ceded

Hugging Face's search over Hugging Face's catalog, exposed over MCP. Ceded, the vendor-hosted MCP reading; retrieval over the enterprise's own data isn't offered.

NVIDIA-Provided

Text Embeddings Inference Targets NVIDIA GPUs First

TEI ships CUDA images for Ampere, Ada, and Hopper alongside CPU builds; Inference Endpoints run embedding models on the same NVIDIA flavors as everything else. Not an NVIDIA dependency the enterprise carries: TEI runs on CPU and the models are open weights.

◆ Gap Analysis

Hugging Face is where the embedding models live and one of the places they run. The Hub hosts the open embedding and reranking models the industry uses (BGE, Qwen3 Embedding, EmbeddingGemma, nomic, the Sentence Transformers family, with multi-vector and multilingual encoders arriving through 2026); Text Embeddings Inference (TEI, Apache 2.0) serves them on CPU or GPU as a standalone container or as an Inference Endpoints engine; Sentence Transformers (Apache 2.0) trains and runs them; Inference Providers route feature-extraction requests to third-party providers behind one API; the datasets library (Apache 2.0) adds FAISS and Elasticsearch indexes to a dataset in the enterprise's own process (add_faiss_index, add_elasticsearch_index, get_nearest_examples, save and load of the index); and the hf_fs tool in the Hugging Face Model Context Protocol (MCP) server gives an assistant semantic search over documentation and Spaces and general search over models, datasets, and papers. The buyer gets the embedders and the serving engine, on its own hardware or Hugging Face's. The architect's concern is that there's no managed retrieval estate. Hugging Face operates no vector store, no hybrid search, and no reranking service of its own beyond serving the open models (the datasets indexes are a library the enterprise runs), and the only search it operates is over the Hub's own catalog; every vector belongs to whichever open model made it and lives in whichever store the enterprise chose. Calibration: NVIDIA reads moderate on retrieval acceleration and model enablement (cuVS, NeMo Retriever, NIM embeddings); Cloudflare moderate on a managed index plus open embedding and rerank models; Cohere moderate on frontier embed and rerank models without an index; Elastic and MongoDB strong on native hybrid engines. Hugging Face is NVIDIA's and Cloudflare's shape minus the index: open embedders plus an engine that serves them. Moderate.

◆ Borrowed Judgment

Low, because everything here is open. TEI and Sentence Transformers are Apache 2.0 software the enterprise runs: Retained. TEI and Sentence Transformers are Apache 2.0 software the enterprise runs, and the datasets library's FAISS and Elasticsearch indexes are the same: Retained. Open embedding and reranking models served on Inference Endpoints or through Inference Providers are Delegated on the open-weights embeddings ruling alone (the Arctic Embed precedent): the enterprise can pull the same weights and serve them under TEI, so the vectors aren't captive to a host. The interface isn't the argument; the Inference Providers feature-extraction path goes through Hugging Face's own client, and the OpenAI-compatible endpoint is chat-only. Hub search through the MCP server is Hugging Face's index of Hugging Face's catalog: Ceded. The runtime call is an endpoint embedding the text the enterprise sent, with per-request override of model and provider: vendor decides, visible, overridable, Delegated, the Elastic reading.

◆ Working Notes

Embeddings here are the open-weights case, not the carve-out: the enterprise can pull the same weights and serve them under TEI, which is why the chip is Delegated where Cohere's and OpenAI's hosted embeddings read Ceded. Inference Providers' feature-extraction task is routed to whichever provider hosts the model; the provider list is in the Layer 2B facts. Public evidence that moves the cell: a Hugging Face vector index or a managed retrieval service; nothing in the documentation describes one.

Layer 1C · PipelinesData Movement & PipelinesJobs as a Hosted Data-Processing Runner (Schedules, Webhooks, Parallel Runs, Streamed and Mounted Datasets and Buckets) Plus Distilabel and the datasets Library; No Connectors, No CDC, No Lineagedecides: code · Retained▼

Move/transform data — ETL/ELT, lineage, cost-aware movement, KV cache tiering

Vendor-Provided

Jobs as a Data-Processing Runner (Commands in Docker Images on a Flavor; Parallel Runs; Cron Schedules and Repository Webhooks; Streamed, Queried, and Mounted Datasets and Buckets; Results Back to the Hub)Ceded

A hosted runner and its triggers, mounts, and billing. Ceded, the Cloudflare Workflows hosted-runner reading; the code it runs is the enterprise's.

Distilabel (Apache 2.0 Pipelines of Steps, Tasks, and LLMs) + the datasets Library (Map, Filter, Stream, Push) + the hf CLI and S3 Gateway for Moving FilesRetained

Open-source code the enterprise runs in its own process against the Hub or anything else. Retained on the open-source seam.

NVIDIA-Provided

No NVIDIA Dependency Beyond the Job Flavor

Data-processing Jobs run on the CPU flavors or, when the enterprise picks one, an NVIDIA GPU flavor scored at Layer 0.

◆ Gap Analysis

Hugging Face moves and transforms data with a runner and libraries, not a pipeline product. Jobs run a command in a Docker image on a chosen flavor, 'ideal ... for data ingestion and processing as well', many in parallel, triggered on a cron schedule or by a webhook when a repository changes (the payload reaches the job); a job streams a dataset or bucket from the Hub without a local copy, queries it over hf:// with Polars or DuckDB, or mounts it as a filesystem, and writes results back to a dataset or bucket; hf buckets sync, hf upload, and rclone against the S3 gateway move files in and out; Distilabel (Apache 2.0) is 'the framework for synthetic data and AI feedback' with pipelines of steps, tasks, and large language models (LLMs); the datasets library maps, filters, streams, and pushes; Argilla labels and AutoTrain trains. The buyer gets a hosted place to run its data code on a trigger, with the Hub as source and sink. The architect's concern is what a pipeline product would add and this doesn't: no connector catalog, no change data capture, no orchestration graph with retries and approvals, no lineage, and the runner's triggers are a cron and the Hub's own repository events. The transforms are the enterprise's scripts. Calibration: Cloudflare reads moderate on Queues and Workflows as durable execution with steps, retries, and approvals; Mistral gap on a transform primitive with no pipeline; CoreWeave gap on fast transport; Confluent and Cloudera strong on movement platforms. Hugging Face sits between Cloudflare and Mistral: a hosted, event-triggered runner for data code plus an open pipeline framework, without durable-execution semantics. Moderate: the runner is general in kind (the enterprise's own pipeline code runs there and can reach any destination), which rule 4 reads above a captive slice, and the missing pipeline product holds it below strong; gap had been argued on the Mistral and CoreWeave reading and was closed by the September 11, 2026 ruling on captive data preparation.

◆ Borrowed Judgment

Low at the code, hosted at the runner. Distilabel, the datasets library, and the hf CLI are Apache 2.0 code the enterprise runs anywhere: Retained. Jobs as a runner, with its flavors, schedules, webhooks, and Hub mounts, is Hugging Face's hosted service: Ceded, the Cloudflare Workflows hosted-runner reading. The runtime call is the runner executing the command the enterprise wrote, with the engine exercising no judgment of its own (placement is a Layer 0 and 2A call): code decides, visible (logs and metrics), overridable, Retained, the Elastic, MongoDB, Cloudflare, and Confluent 1C reading.

◆ Working Notes

Ruled September 11, 2026: moderate stands (a general hosted runner for the enterprise's data code reads above the captive-slice line that Cerebras's Model Zoo and Nebius's Data Lab sit on); gap had been argued on the Mistral and CoreWeave reading. Named, not scored: Argilla, AutoTrain. Public evidence that moves the cell: a connector or movement product, durable-execution semantics on Jobs, or lineage across repositories and buckets.

Layer 2A · OrchestrationInfrastructure OrchestrationInference Endpoints Autoscaling With Replica Bounds and Scale-to-Zero, Instance and Region Pins, Jobs Flavors and Timeouts, Resource-Group Spend Caps; No Scheduler Over a Fleet You Owndecides: vendor · Ceded▼

GPU scheduling, quotas, RBAC, fair-share scheduling, utilization optimization

Vendor-Provided

Inference Endpoints Lifecycle (Cloud, Region, and Instance Pins; Min and Max Replicas; Utilization Autoscaling; Scale to Zero; Enterprise-Plan Control of the Autoscaling Definitions)Ceded

Hugging Face's control plane sizing Hugging Face's endpoints inside bounds the enterprise set. Ceded, the Cohere Model Vault reading.

Jobs Runner and Scheduler (Flavors, Timeouts, Cron Schedules, Secrets, Labels, Per-Minute Billing) + Spaces Hardware Assignment, Sleep, and ZeroGPU AllocationCeded

A hosted container runner and app host with placement Hugging Face makes. Ceded.

Resource-Group Compute Controls (Cost Attribution, Per-Product Spend Caps, and Per-Feature Access on Enterprise and Above; Resource Groups Themselves on Team and Enterprise)Ceded

Budget and permission bounds in Hugging Face's organization model. Ceded.

NVIDIA-Provided

No GPU Scheduling Beyond the Flavor

The enterprise picks an NVIDIA flavor per endpoint, job, or Space; Hugging Face places it. There's no fleet, no fair share, and no Multi-Instance GPU (MIG) partitioning.

◆ Gap Analysis

Hugging Face orchestrates its own hosted products, one object at a time. Inference Endpoints: the enterprise picks cloud, region, and instance type per endpoint, sets minimum and maximum replicas, and gets utilization-based autoscaling (a replica added when average GPU utilization over one minute reaches 80 percent; scale to zero after a configurable idle period, one hour by default per the configuration guide, where an older autoscaling page says fifteen minutes) with 'full control over the autoscaling definitions' on the Enterprise plan; scaling on pending requests (more than 1.5 per replica over twenty seconds) is a beta feature the documentation calls experimental. Jobs: a flavor, a timeout (default thirty minutes), a schedule, environment and secrets, labels, and per-minute billing; on Enterprise plans admins restrict who can run Jobs and attribute costs to resource groups. Spaces: a hardware assignment per Space, sleep settings, and ZeroGPU's dynamic allocation against a quota. Across the organization, on Enterprise and above: resource groups cap monthly compute spend in total or per product (Inference Providers, Spaces, Jobs, Inference Endpoints), attribute costs, and, since August 12, 2026, restrict which features each group may use. The buyer gets bounded autoscaling for endpoints and a container runner with a budget. The architect's concern is that there's no plane. Nothing schedules across the enterprise's own fleet (HUGS is deprecated; the Deep Learning Containers run under SageMaker, Vertex AI, GKE, or Azure ML, whose orchestration is scored on those rows), no priority or fair share exists between endpoints, and the utilization rule is Hugging Face's on non-Enterprise plans. Calibration: Cohere reads moderate on Model Vault capacity controls (replica bounds, autoscaling, pause and resume); Databricks and Snowflake moderate on managed compute with bounds; Cloudflare moderate on manual containers with no autoscaling; Mistral and OpenAI gap on a metering edge. Hugging Face is Cohere's shape with a container runner attached. Moderate.

◆ Borrowed Judgment

Ceded on the loop, and on the menu. The enterprise picks cloud, region, and instance type per endpoint and a flavor per job, sets replica bounds and a scale-to-zero window, and caps spend; Hugging Face places the endpoint inside its own capacity and scales it, and on non-Enterprise plans its thresholds are its own. A flavor is a menu and a replica bound is a bound, the same controls Cohere's Model Vault and Cloudflare's Containers carry, and the row's own Layer 0 reading says so; Dataiku's Delegated rested on a pin to a named customer cluster, which nothing here offers. Vendor decides, visible (replica and utilization metrics), not overridable, Ceded, the Cohere, Cloudflare, Snowflake, and Databricks reading.

◆ Working Notes

Cross-row: the 2A column under the override rule (September 5 sheet) covers this cell; written Ceded on the Cohere and Cloudflare reading after the ChatGPT pass argued the per-endpoint instance choice is configuration, not an override of Hugging Face's placement. Unscored on its badge: pending-request autoscaling ('beta feature', 'currently an experimental feature'). Scale-to-zero: the configuration guide says one hour by default and configurable; the older autoscaling page says fifteen minutes; the more specific guide governs. Granular feature access per resource group is dated August 12, 2026 in the changelog. Public evidence that moves the cell: a scheduler across endpoints or the enterprise's own hardware, or a documented per-endpoint placement pin the enterprise can reverse.

Layer 2B · RuntimeApplication Runtime & ExecutionServe Any Open Model Anywhere: Inference Endpoints on Three Clouds (vLLM, SGLang, TEI; TGI in Maintenance) and an OpenAI-Compatible Chat Router Across Eighteen Providers; the Agent Libraries Are Experimental or Thin, and There's No Managed Agent Runtimedecides: model · Delegated▼

Model serving, agent execution, inference APIs, distributed inference

Vendor-Provided

Permissively Licensed Open Weights (Apache 2.0, MIT) Pulled From the Hub and Run Under transformers, vLLM, TEI, or the Deep Learning Containers on the Enterprise's Own Hardware or Cloud (SageMaker, Azure ML and Foundry, Vertex AI, GKE)Retained

Weights the enterprise possesses under permissive terms and open engines it runs. Retained, the Cohere Command A+, Mistral Apache 2.0, and OpenAI gpt-oss reading; the cloud's orchestration is scored on the cloud's row.

Open Weights Under Their Owners' Restrictive Terms (Llama and Gemma Community Licenses, Revenue-Capped and Non-Commercial Releases) Distributed Through the HubCeded

The Hub records the license; the owner grants it and can bound it. Ceded to the owner, the Mistral revenue-capped reading, whatever host serves them.

Inference Endpoints (Dedicated Serving of Any Hub Model on a Chosen Instance in AWS, Azure, or GCP; vLLM, SGLang, TEI; TGI in Maintenance Mode; Public, Protected, and Private Over AWS PrivateLink; Enterprise Plan With 24/7 SLAs)Delegated

An operated implementation behind the engines' standard interfaces, on an instance the enterprise picked. Delegated, the Dataiku local-serving and Cloudflare Workers AI reading.

Inference Providers Chat Completions (OpenAI-Compatible Router to Eighteen Partners; the Enterprise's Own Provider Keys Optional)Delegated

Model access behind a multi-vendor interface the enterprise could point at any provider directly. Delegated, the inference-interface ruling; the models are their owners'.

Inference Providers Routing Control Surface (Proxy; :fastest, :cheapest, and :preferred Policies; Automatic Failover; Billing at Provider Rates; X-HF-Bill-To Attribution; Per-Product Spend Caps and Feature Access) + Non-Chat Tasks Through the Hugging Face InferenceClient (Image, Video, Speech; Embeddings Scored at 1B)Ceded

The gateway's decisions and the single-vendor client for everything that isn't chat. Ceded, the Cloudflare AI Gateway split and the customer-tools SDK-binding reading.

tiny-agents and the huggingface_hub MCP Client (Apache 2.0; Agents Loaded From the Hub by Id or agent.json) + Customer-Authored ToolsRetained

Thin agent clients and tools the enterprise runs in its own process against any model. Retained, the OpenAI Agents SDK and customer-tools reading; smolagents, the fuller harness, is experimental and named in notes.

Hugging Face MCP Server (hf_fs Over Models, Datasets, Spaces, Documentation, Papers; Jobs; Spaces as Tools) + the hf CLI Skill for Coding AgentsCeded

Vendor-hosted MCP over Hugging Face's own resources, and a skill that teaches coding agents the Hub. Ceded, the Snowflake and Confluent managed-MCP reading.

NVIDIA-Provided

Engines and Endpoints Run on NVIDIA GPUs by Default, With TPU, Trainium, and CPU Paths

Inference Endpoints instance flavors are NVIDIA; TGI and TEI ship CUDA images and TPU and Inferentia builds through the cloud partner documentation; vLLM and SGLang are the recommended engines and run wherever they run. The open weights themselves carry no accelerator dependency.

◆ Gap Analysis

Hugging Face is the distribution point for open models and a place to run them. The Hub hosts more than three million models under licenses the Hub records; transformers (Apache 2.0) loads them; Inference Endpoints (dedicated) deploy any Hub model on a chosen instance in AWS, Azure, or Google Cloud Platform (GCP) under vLLM or SGLang (recommended), Text Generation Inference (TGI, 'in maintenance mode as of 12/11/2025'), or Text Embeddings Inference (TEI), as public, protected (Hugging Face token), or private (AWS or Azure PrivateLink) endpoints, pay-as-you-go from $0.032 per CPU core-hour and $0.50 per GPU-hour, with an Enterprise plan that adds 24/7 service level agreements (SLAs) and uptime guarantees; Inference Providers route chat, vision, feature-extraction, image, video, and speech requests through Hugging Face's proxy to eighteen partners (Baseten, Cerebras, Cohere, DeepInfra, Fal, Featherless, Fireworks, Groq, HF Inference, Novita, Nscale, OVHcloud, Public AI, Replicate, Scaleway, Together, WaveSpeed, Z.ai) behind an OpenAI-compatible API that the documentation says 'is currently available for chat completion tasks only' (other tasks go through the Hugging Face InferenceClient), with provider selection policies (:fastest by default, :cheapest, :preferred) and failover, billed by Hugging Face at the provider's rates with no markup, or by the provider directly under the enterprise's own key, with X-HF-Bill-To attributing spend to an organization or resource group; the Deep Learning Containers (Apache 2.0) and the SageMaker, Azure ML, Foundry, Vertex AI, and Google Kubernetes Engine (GKE) integrations deploy Hub models into the enterprise's own cloud accounts. Agents: tiny-agents (huggingface_hub and @huggingface/tiny-agents, Model Context Protocol (MCP) clients that load an agent from the Hub by id or an agent.json) are libraries the enterprise runs; smolagents (Apache 2.0; CodeAgent and ToolCallingAgent; tools from any MCP server, LangChain, or a Space) is the fuller harness, and its reference documentation says it 'is an experimental API which is subject to change at any time', so it's named, not scored; the hf CLI ships as a Skill for Claude Code, Codex, Gemini CLI, and Cursor; the Hugging Face MCP server exposes the Hub (hf_fs search, Jobs, Sandboxes, Spaces as tools) to any MCP client; Sandboxes (July 22, 2026 changelog) give an assistant execution environments on Jobs attached to buckets and repositories. The buyer gets every open model, on its own hardware, on Hugging Face's, or on a provider's, behind interfaces it already uses. The architect's concern is what Hugging Face itself runs. It serves no frontier proprietary model; the hosted agent surface is a catalog tool server and a sandbox on a job runner, not an agent runtime with sessions, evaluation, or tracing of the enterprise's agents (agent traces on the Hub are uploaded session files, from Claude Code, Codex, and Pi or from any harness that writes the open format); the one full agent harness is experimental by its own documentation; TGI is in maintenance mode with vLLM and SGLang recommended in its place; and the Sandboxes guide exists only on the main branch of the library documentation. Calibration: Cloudflare reads strong on serverless open-model inference plus a gateway plus durable agents; CoreWeave strong on open serving under KServe; NVIDIA strong on NIM plus Dynamo; OpenAI, Anthropic, Mistral, and Cohere strong on frontier serving plus open-source harnesses; Elastic and Confluent moderate on first-release maturity. Hugging Face is where open models are distributed and served first, on engines it doesn't own and under gates it hasn't cleared. Moderate, on rule 5: agent execution at 2B is judged by whether the runtime carries the enterprise's own agent loop (the September 11, 2026 ruling), and here the engines are under maintenance and beta gates, tool support varies by provider through the router, and the one harness is experimental; strong had been argued on the Cloudflare and CoreWeave open-serving precedent and rule 6.

◆ Borrowed Judgment

The lowest capture at 2B on the map, because the artifacts and the engines are open. Open weights under permissive licenses (Apache 2.0, MIT) that the enterprise pulls from the Hub and runs under transformers, vLLM, or TEI on its own hardware are its own: Retained, the Cohere Command A+, Mistral, and OpenAI gpt-oss reading; weights under their owners' restrictive terms (the Llama and Gemma community licenses, revenue-capped releases) are Ceded to the owner, the Mistral revenue-capped reading, whatever Hugging Face does with them. tiny-agents, huggingface_hub, and the Deep Learning Containers are Apache 2.0 code the enterprise runs: Retained, the OpenAI Agents SDK reading; tools the enterprise writes are Retained on the customer-tools ruling. Inference Endpoints are an operated implementation behind the engines' interfaces on an instance the enterprise picked: Delegated, the Dataiku local-serving reading. Inference Providers split on the interface: chat completions behind the OpenAI-compatible router are Delegated, the inference-interface ruling; the other tasks go through Hugging Face's own client and are Ceded on the single-vendor binding; the routing policies, failover, proxy, and billing are Hugging Face's control surface, Ceded, the Cloudflare AI Gateway split. The Hugging Face MCP server over Hub resources is vendor-hosted MCP over captive tools: Ceded. The decision to act is the model's, but the loop runs in the enterprise's own tiny-agents or smolagents process, where a deterministic gate before any effect is the enterprise's to write: model decides, visible, overridable, Delegated, the OpenAI and Anthropic reading.

◆ Working Notes

Ruled September 11, 2026: moderate stands, on rule 5 (engines under maintenance and beta gates, tool support that varies by provider through the router, an experimental harness) rather than on a missing hosted loop; agent execution at 2B is judged by whether the runtime carries the enterprise's own agent loop, and a hosted runtime is never the price of strong. Strong had been argued on the open-serving precedent. Badges: smolagents 'is an experimental API which is subject to change at any time' (reference documentation), named, not scored; TGI 'in maintenance mode as of 12/11/2025'; pending-request autoscaling 'beta feature'; HUGS 'currently deprecated' (September 2025); Sandboxes shipped per the July 22, 2026 changelog but the huggingface_hub guide exists on the main branch only, named, not scored. Inference Providers' pricing page: Hugging Face 'charges you the same rates as the provider, with no additional fees'; Free $0.10 monthly credit, PRO $2, Team and Enterprise $2 per seat. Model licenses are the owners' (Llama and Gemma terms are recorded on the Hub, not granted by it), which is why the open-weights chip splits by license. Instrument follow-up: the standing sentence that open weights the enterprise can serve itself read Delegated conflicts with the Cohere, Mistral, and OpenAI chips that read permissive open weights Retained; this row follows the rows. Public evidence that moves the cell: smolagents losing its experimental label, or a managed agent runtime with sessions and evaluation, would argue strong; a Hugging Face-hosted frontier model would change the shape.

Layer 2C · ReasoningAgentic Infrastructure — The Reasoning PlaneOrganization Governance for People and Tokens, a Tool Server, and a Place to Upload Agent Traces; Not a Plane Over Agentsdecides: absent · Absent▼

Policy-driven placement and resource coordination — the Autonomy Layer

Vendor-Provided

NVIDIA-Provided

No NVIDIA Dependency

Nothing at this layer is a scored capability.

◆ Gap Analysis

Read against the five legs, Hugging Face has none in production as a plane over agents. Identity: single sign-on (SSO), SCIM provisioning, fine-grained tokens with administrator approval, and service accounts (Enterprise and above) identify people and automations to the Hub; a service account is a non-human principal for Hub resources, the single-resource shape the MongoDB sheet escalated, not an agent identity across runtimes. Gateway: Inference Providers' proxy selects providers by policy (:fastest, :cheapest, :preferred), fails over, attributes bills, and on Enterprise and above rejects requests from members without feature access or from a resource group past its spend cap; that's a routing and budget gateway over model calls, one leg by the Cloudflare reading, and there's no gateway in front of tool calls. Registry: the Hub hosts Skills as SKILL.md repositories, tiny-agents definitions that load by id from an agent.json, and thousands of Spaces exposed as Model Context Protocol (MCP) tools, all under repository and resource-group governance: a runnable, addressable catalog of agent artifacts on the Hub, not a registry of the enterprise's agents across runtimes. Orchestration: none offered; smolagents' managed agents run in the enterprise's process. Observability: agent traces from Claude Code, Codex, and Pi, or from any harness that writes the open Session Trace Simple Format (JSONL), can be uploaded to a dataset or bucket and rendered in the Hub's trace viewer (the documentation warns they 'can include prompts, tool inputs, command output, local paths, screenshots, secrets, private code, and personal data'), a viewer for uploaded sessions, not tracing of a runtime. What Hugging Face gives the reasoning plane is artifacts and audit: a governed registry of weights, skills, and traces with logs of who touched them. The architect's concern is the same as at 2B: the loop is wherever the enterprise runs it, and the governance that exists governs Hub resources. Calibration: Anthropic, OpenAI, and NVIDIA read gap with runtime permissions and no plane; Cloudflare moderate on a gateway over model calls plus Access for MCP servers; Databricks moderate on Unity Catalog agent governance; MongoDB gap with a single-resource principal escalated. Hugging Face is the Anthropic and NVIDIA shape with a routing gateway and a trace viewer. Gap, authority Absent (ruled September 15, 2026): the 2C floor is the Nutanix line, a GA gateway every agent and model call passes through with RBAC, rate limiting, and full audit including MCP-call audit; the Providers router is routing (fastest, cheapest, preferred) plus spend caps and feature access over Hugging Face's own front door, with no RBAC or audit over the enterprise's agent traffic, and routing isn't reasoning. The ChatGPT pass had voted a thin moderate.

◆ Borrowed Judgment

Nothing offered as a plane. The enterprise that wants agent identity, a request gateway, a registry, or orchestration over agents built with smolagents brings its own, and every piece is substitutable. Absent.

◆ Working Notes

Ruled September 15, 2026: gap stands below the Nutanix 2C floor (gateway governance with RBAC and audit over every agent and model call); a thin moderate had been argued on the Inference Providers router as one GA leg, with the trace viewer and the Hub agent catalog as partial second and third legs. Named, not scored: service accounts (Enterprise and above; a non-human principal for Hub resources, the MongoDB single-resource question), agent traces and the Session Trace Simple Format, Sandboxes, the Skills and tiny-agents catalog. Public evidence that moves the cell: a request-time gateway for agents' tool calls, a registry of the enterprise's agents across runtimes, or tracing of agents Hugging Face runs.

Layer 3 (+1) · ApplicationsAI Application Layer — The Value PlaneNo Business Application: Spaces Host a Million of Other People's Gradio and Docker Apps, the Inference Playground and Data Studio Are Developer Tools, and Gradio Is an Application Framework Rather Than an Applicationdecides: absent · Absent▼

AI-powered business capabilities — business logic, workflow automation

Vendor-Provided

NVIDIA-Provided

No NVIDIA Dependency Beyond the Spaces Hardware Scored at Layer 0

Gradio and Docker Spaces run on the flavors listed at Layer 0 or, on ZeroGPU, on RTX Pro 6000 Blackwell slices; static Spaces are served without compute.

◆ Gap Analysis

Hugging Face's application layer is the Hub as a developer plane. Spaces host more than a million Gradio, Docker, and static applications (Gradio and Docker Spaces need PRO, Team, or Enterprise to create, except that free personal accounts in good standing may host two Gradio Spaces on ZeroGPU; static Spaces are free and run on no compute; visibility public, protected, or private; Dev Mode with SSH and VS Code Web on paid plans; a custom domain on public and protected Spaces; a July 16, 2026 flow that hands an agent the command to build a Space); Gradio (Apache 2.0) is Hugging Face's own application framework, with Model Context Protocol (MCP) compatible Gradio apps serving as tools (Gradio Workflows, announced in August 2026 blog posts, is named, not scored); the Inference Playground is a chat surface across models and providers; Data Studio turns questions into SQL over datasets; leaderboards and the model, dataset, and paper catalogs are the discovery surface; the Expert Support Program is services. The buyer gets a place to ship an AI application and a community to ship it to, and pays Hugging Face for the hardware under it. The architect's concern is what it isn't. There's no business application, no knowledge-work estate, and no agentic workspace: the Playground and Data Studio are a developer's tools, the apps on Spaces are their authors', and the value the enterprise builds on Spaces is its own Gradio or Docker code. Calibration: under the Layer 3 marketplace ruling (September 11, 2026), a hosting shelf of other people's applications earns no capability grade, a developer tool doesn't either, and the grade rests on first-party business applications; Mistral, OpenAI, Anthropic, and Cohere read strong on first-party knowledge-work applications; Cloudflare gap on a console; Cerebras and Groq gap. Hugging Face has no first-party business application. Gap, authority Absent.

◆ Borrowed Judgment

Nothing offered as a business application. Absent. The relationship boundary records on DAPM where it exists, and here it doesn't: Spaces bill hosting to the enterprise's account and carry no commercial relationship for the applications themselves, whose authors are the buyer's counterparty.

◆ Working Notes

Ruled September 11, 2026 (the Layer 3 marketplace ruling): the first draft read moderate, escalated, on the CoreWeave and NVIDIA developer-plane reading; a marketplace or hosting shelf of others' applications isn't creditworthy at Layer 3. Named, not scored: Spaces hosting (hardware assignment, ZeroGPU, sleep, visibility, custom domains, Dev Mode; H100 hardware removed December 2025), Gradio and Gradio Workflows (an application framework; Gradio Workflows dated only by the August 25 and September 10, 2026 blog posts), the Inference Playground, Data Studio, leaderboards, and the catalogs. Public evidence that moves the cell: a first-party knowledge-work or agentic business application; nothing in the documentation describes one.

Sources & revision history · v1.0 - 4+1 v2: Authority Split

Hugging Face Hub documentation (index; Team and Enterprise plans; SSO; audit logs; resource groups; tokens management; network security; storage regions; gating group collections; advanced security; Protect AI and JFrog scanning; analytics; service accounts; Storage Buckets, S3-compatible API, security, limits, regions, access; Xet; Jobs overview, quick start, scheduling, serving, pricing; Spaces overview, GPUs, ZeroGPU, Dev Mode, MCP servers, agents, Docker; Agents overview, SDK, CLI, libraries, local, MCP, skills, traces; Hugging Face MCP server; model cards; repository licenses; security; datasets overview and viewer); Inference Endpoints documentation (index, access, create endpoint, autoscaling, security, PrivateLink, FAQ, engines TGI, vLLM, SGLang); Inference Providers documentation (index, pricing, Hub integration); HUGS documentation (deprecation notice); smolagents, Text Generation Inference, and Text Embeddings Inference documentation; Hugging Face on AWS, Microsoft Azure, and Google Cloud documentation; huggingface.co pricing and enterprise pages; Hugging Face blog index and changelog (Granular Feature Access August 12, 2026; MCP Server Enhancements July 22, 2026; Build Spaces with AI Agents July 16, 2026); GitHub license API for transformers, huggingface_hub, smolagents, text-generation-inference, text-embeddings-inference (Apache 2.0), hf-mcp-server (MIT), xet-core (Apache 2.0); NVIDIA blog, NVIDIA to Acquire Hugging Face (September 10, 2026). Peer-reviewed cell by cell through the labs claims ledger (huggingface-<layer>-chatgpt, ChatGPT gpt-5.5) and as a whole row by Antigravity (huggingface-row-agy); totals and escalated items in reviews/huggingface-judgment.md.

4+1 Layer AI Infrastructure Model · Vendor Assessment Series · The CTO Advisor LLC (DBA The Advisor Bench) · thectoadvisor.com