NVIDIA Details BlueField-4 Scale-In Infrastructure for Agentic AI Factories
NVIDIA introduced Scale-In as a fifth AI-networking pillar, using BlueField-4, DOCA and Spectrum-X to offload security, storage, telemetry and access services from host CPUs.
NVIDIA adds a Scale-In infrastructure layer
NVIDIA published an August 24, 2026 technical deep dive describing Scale-In network infrastructure for agentic AI factories. The architecture uses BlueField-4 data processing units, NVIDIA DOCA software and Spectrum-X Ethernet to move security, storage access, telemetry, provisioning and other infrastructure services into a dedicated host-independent processing domain.
NVIDIA positions Scale-In as a fifth networking pillar alongside the connectivity used to scale accelerators within systems, across racks and across larger facilities. The emphasis is north-south infrastructure traffic: the users, agents, applications, data sources, storage systems and cloud services surrounding the accelerated-compute domain.
Why host-independent processing matters
Agentic AI workloads can generate far more than model-to-model traffic. A single user request may trigger storage reads, retrieval, policy checks, tool calls, network connections and telemetry while many tenants share the same accelerated infrastructure. If all of those services depend on the host CPU, NVIDIA argues they can compete with workload execution and create security or performance bottlenecks.
BlueField-4 provides a separate infrastructure-processing domain where operators can run policy enforcement, storage access, observability and service state independently of the tenant host. NVIDIA says the DPU supports up to 800 Gb/s throughput for the relevant infrastructure path and works with DOCA microservices and Spectrum-X networking.
Security, storage and fleet operations
The Scale-In design is intended to provide consistent tenant isolation, runtime security, storage virtualization and fleet-wide observability as AI systems grow. NVIDIA also describes BlueField Astra as extending trusted control from north-south Scale-In traffic into the east-west Scale-Out network, with policy installation and telemetry separated from tenant software.
This kind of architecture is especially relevant to shared AI infrastructure, where the operator needs to enforce policy even if a tenant workload is compromised or consuming most of its assigned CPU resources. A hardware-isolated control plane can make infrastructure services less dependent on the behavior of the hosted application.
Relationship to storage and inference context
NVIDIA distinguishes Scale-In from CMX, its approach to pod-level storage for shared KV cache and reusable inference state. Scale-In connects external and enterprise data systems to AI compute; CMX is intended to preserve and share inference context closer to the compute pod. BlueField-4 participates in both roles as an infrastructure and data processor.
What to validate in production
The architecture is technically significant, but operators should evaluate real storage patterns, tenant isolation requirements, network topologies, telemetry overhead and compatibility with existing security controls before assuming vendor-published capacity translates directly to a particular data center. Dedicated offload can remove pressure from host CPUs, but it also adds another programmable infrastructure layer that must be securely operated and updated.
Release status
BlueField is an existing NVIDIA platform. The August 24 publication introduces and explains the Scale-In architecture around BlueField-4 for emerging agentic AI factories rather than representing the first release of DPU technology itself.
This article is built from the source material below. Open the originals for full context and the latest updates.