If your AI inference roadmap has outrun the 10-15 kW cabinets in your existing facility,
this paper shows how to close the gap without a facility upgrade. It compares six high-density
heterogeneous compute platforms that fit inside air-cooled brownfield racks, and sets out
a decision framework for on-premise AI inference deployment under strict thermal budgets.
The mean rack density reported by the Uptime Institute in 2025 was 7.6 kW [3], and even the
AI-active AFCOM respondent base averaged 16 kW in 2025 and 27 kW in 2026 [1][2] – four
to sixteen times below the 40-100 kW continuous draw of a modest H100 or H200 rack [8][9].
Upgrading a brownfield facility costs four to eight million dollars per megawatt and can stall
a year or more on transformer and switchgear lead times [57][60].
A different class of hardware has arrived in response. High-density heterogeneous compute
platforms – CPU, NPU, GPU, or RDU combinations engineered for inference throughput per watt
rather than peak training FLOPS – are designed to fit within the 10-15 kW cabinets enterprises
already own. This paper surveys six current implementations, compares them on published
performance per watt, memory, workload fit, and software maturity, and provides a capex
framework architects can apply to their fleets. The central finding: for inference workloads,
facility-upgrade-free deployment is now a credible alternative to retrofitting power and cooling.
AI inference is the center of gravity in data center compute. Omdia forecasts inference will account for roughly two-thirds
of compute in 2026, up from one-third in 2023 [55]. IDC projects AI infrastructure spending at $758 billion by 2029, with
enterprise investment rising from $307 billion in 2025 to $632 billion in 2028 [56]. Yet more than 70 percent of global data
center capacity sits in existing buildings [5], much of it underutilized for AI [6].
For regulated and data-gravity workloads – healthcare, defense, finance, pharma – on-premise AI inference is not optional.
The mismatch between where AI must run and where enterprise real estate can support it is an immediate constraint.
The thesis: high-density heterogeneous compute platforms – integrating CPUs, NPUs, and GPUs on a single
motherboard – offer a facility-upgrade-free path to inference throughput within strict thermal and power limits.
The Furiosa NXT RNGD Server, NVIDIA® L40S-based systems, SambaNova SN40L, Intel® Xeon with Gaudi 3, IBM®
Spyre on Power11, and the CAPE open architecture exemplify this approach.
Most installed enterprise and colocation cabinets were provisioned at 5-15 kW ceilings [62]. Data Center Dynamics reports
enterprise facilities typically operate at 3-5 kW per rack, with new-build space engineered for 8-10 kW [4]. The Uptime
Institute’s 2025 Global Data Center Survey places the industry-wide mean at 7.6 kW, up from 6.8 kW the year prior [3].
The 2026 AFCOM survey reports 27 kW – a 69 percent year-over-year increase, though AFCOM respondents skew toward
operators actively planning for AI [1][2]. The installed-base median remains well below 30 kW.
Conventional AI accelerator clusters sit well above that envelope. An NVIDIA DGX H100 is rated at 10.2 kW maximum per
8-GPU chassis [8]; NVIDIA’s own guidance recommends no more than four DGX H100 systems per rack for thermodynamic
reasons, pushing continuous draw past 40 kW [8]. Trade-press analysis citing Schneider Electric and Vertiv projections
places fully loaded current-generation GPU racks at ~132 kW, with next-generation deployments expected at 240 kW within
a year [9].
Figure 1: Power-constrained rack density versus AI accelerator deployment draw, 2025–2026. Sources: [1], [2], [3], [4], [8], [9].
The gap is structural. A brownfield cabinet at 5-15 kW sits four to sixteen times below a modest H100 deployment, and up to twenty-five
times below a fully loaded DGX rack. Facility upgrade is neither fast nor cheap.
A second class of server architecture sidesteps the power gap rather than bridging it. Heterogeneous compute platforms
integrate CPUs with Neural Processing Units (NPUs), GPUs, or Reconfigurable Dataflow Units (RDUs) on a single
motherboard, tuned for inference performance per watt rather than peak training FLOPS.
The tradeoff is architectural. Homogeneous GPU servers such as the DGX H100 allocate silicon to the matrix-multiplication
bandwidth and memory capacity training demands, consuming ten kilowatts or more per chassis and relying on NVLink and
NVSwitch for GPU-to-GPU bandwidth [10]. Heterogeneous platforms allocate more die area to on-chip SRAM and
lower-power interconnect, trading peak FP16 and FP8 throughput for better performance per watt on inference-shaped
workloads with lower batch sizes and tight latency [11][12].
Two efficiency metrics dominate inference-platform comparisons: tokens per second per watt and joules per token
(its reciprocal, gaining currency in 2025–2026). Either belongs in a spec-sheet evaluation.
Peer-reviewed research in Systems found NPU-based servers delivered 2.9x higher throughput than GPUs in specific
inference workloads at comparable power efficiency, and outperformed GPUs by ~58 percent on video and LLM
matrix-vector operations [11]. An arXiv analysis of multi-stage inference pipelines argues retrieval and reranking suit
large-memory CPUs, embedding suits NPUs or GPUs, and prefill or decode scales with compute-heavy NPUs or
GPUs – a case for heterogeneous pipelines over monolithic GPU stacks [12].
Rack-density math drives the case. Furiosa reports the NXT RNGD Server – a 4U, 8-card system – draws 3 kW and
delivers 4 PFLOPS of FP8, placing five servers per 15 kW cabinet for 20 PFLOPS within the existing envelope [7][15].
An 8-way NVIDIA L40S 4U chassis draws 3-4 kW with host overhead, fitting three to four per cabinet [16][17][25]. An NVIDIA
HGX H100 chassis at 10.2 kW admits only one unit per 15 kW cabinet after networking and PDU overhead [8].
Figure 2: High-density AI server rack fit: servers per 15 kW air-cooled cabinet by platform, based on vendor-reported chassis power. Sources: [7], [8], [15], [16], [17], [25].
Six implementations define the current heterogeneous inference landscape. The data below distinguishes vendor-reported
figures from independent benchmarks.
Table 1: Six heterogeneous inference platforms, compared on published specifications and vendor-reported performance-per-watt efficiency. Sources:[7], [13], [15], [16], [17], [18], [24], [25], [27], [28], [29], [30], [34], [35], [36], [42], [43], [45], [50], [51]. SambaNova per-chip TDP is not publicly disclosed – Tier 3 verification gap.
The Furiosa RNGD is a second-generation NPU on TSMC 5nm: 180 W per card, 48 GB of HBM3 at 1.5 TB/s, 256 MB of
on-chip SRAM, and 512 TFLOPS of FP8 [13][18][19]. A 4U NXT RNGD Server carries eight cards for 384 GB of HBM3 and 12
TB/s aggregate bandwidth within 3 kW [15]. Furiosa’s Hot Chips 2024 disclosure reports ~1.4 TFLOPS per watt on BF16
inference of LG’s EXAONE 3.5 32B; vendor-internal testing claims 40 percent better performance per watt than L40S on the
MLPerf GPT-J reference – all vendor-reported, with no MLPerf submission [13][14][22]. Mass production began January 2026
with a 4,000-unit TSMC/ASUS batch [19][23].
The L40S is NVIDIA’s Ada Lovelace data center GPU: 350 W per card, 48 GB of GDDR6 at 864 GB/s, and 733 dense FP8
TFLOPS via fourth-generation Tensor Cores [16][17][24]. A Dell PowerEdge XE7745 MLPerf Inference v5.0 submission showed
8-way L40S configurations placing among the top perf-per-watt results in the server category [25][26]. Red Hat’s v5.1
submission reported 1,642 tokens per second offline on a single L40S with FP8 Llama 3.1 8B via vLLM [27]. Software
ecosystem breadth is the widest of the six surveyed platforms – CUDA, cuDNN, TensorRT, Triton, vLLM, and TGI are
natively supported – and GDDR6 memory hedges against CoWoS-HBM supply constraints discussed below.
SambaNova’s SN40L RDU is a 2.5D-packaged accelerator on TSMC 5nm with 102 billion transistors per socket, published
at ISSCC 2025 [28][29]. Per-chip TDP is not publicly disclosed – a verification gap this paper flags. Three-tier memory pairs
520 MiB of on-chip SRAM with 64 GiB of HBM and up to 1.5 TiB of external DDR per socket, delivering 638 BF16
TFLOPS [28]. A 16-socket SambaRack SN40L-16 draws 7-14.5 kW (10 kW typical) and aggregates 10.2 PFLOPS of BF16
in a 19-inch air-cooled enclosure [30]. SambaNova reports 1,000-plus tokens per second on Llama 3 8B and 114 tokens per
second on Llama 3.1 405B per 16-socket node – vendor figures, not independently verified [30][31][32].
Intel’s Gaudi 3 carries 128 GB of HBM2e at 3.7 TB/s and 1.8 PFLOPS of FP8 across two compute dies on TSMC 5nm [34][35].
TDP varies by form factor: 450-900 W for the HL-325L OAM, 900 W (up to 1,200 W) for the liquid-cooled HL-335, and 600
W for the HL-338 PCIe card [34][35][36]. Intel-published benchmarks show H100-range performance, with some workloads up to
70 percent higher and others up to 10 percent lower [35]. Intel submitted to MLPerf Inference v5.0 and v5.1, though Gaudi
3-specific headline results are not highlighted in MLCommons summaries; prior-generation Gaudi 2 recorded 8,035 offline
tokens per second on Llama 2 70B [37][38]. Intel cancelled Falcon Shores in January 2025; Jaguar Shores is the 2026
rack-scale successor, placing Gaudi 3 late in its cycle [40][41].
Spyre is an inference-focused SoC on Samsung 5LPE, rated at 75 W per PCIe 5.0 card, with 128 GB of LPDDR5 at ~200
GB/s and more than 300 TOPS of FP16 [42][43][44][45]. A Power11 I/O drawer holds up to eight cards for 1 TB of combined
memory, 1.6 TB/s aggregate bandwidth, and 2.4 PFLOPS at 600 W [45]. IBM Power11 reached GA on July 25, 2025 [49];
Spyre reached GA on IBM Z and LinuxONE on October 28, 2025, and on Power11 on December 12, 2025 [45][46][48].
IBM i and AIX integration via RHEL 9.6 partitions plus Live Partition Mobility differentiates Spyre for existing Power
customers [45].
CAPE is a Horizon Europe research project (grant 101189899) running December 2024 through November 2027, with
€5,996,250 in funding, coordinated by Universität Bielefeld [50][51]. The architecture combines Composable Infrastructure via
CXL with two open COM-HPC platforms – the embedded High-Performance Server (eHPS) and the Embedded Micro Data
Center (EMDC) – composing CPU, NPU, GPU, and RISC-V accelerators through CXL-attached pools [50][52]. CAPE is a
reference architecture, not a commercial product; it signals the direction of open, sovereignty-oriented heterogeneous
compute. Omdia projects Open Rack enclosures at more than 70 percent of data center rack revenue by 2030 [67].
Heterogeneous platforms are not a replacement for dedicated GPU clusters. The match depends on workload shape.
Table 2: AI workload-to-platform fit within the 10-15 kW envelope. Sources:[10], [11], [12], [14], [20], [21], [27], [30], [42], [53], [54]
Inference-shaped workloads at enterprise concurrency – chatbots, copilots, internal RAG, recommendation scoring, classification – fit the
heterogeneous envelope. Frontier training and hyperscale MoE serving do not, and belong on dedicated GPU infrastructure.
For a tenant in a colocation 15 kW rack, the upgrade option is often unavailable – the facility operator sets the ceiling.
Brownfield retrofit runs $4-8 million per megawatt excluding hardware; brownfield projects are 30 to 50 percent cheaper
than greenfield ($5-6M vs. $8-10M per megawatt) and avoid two or more years of revenue delay [57][58]. Liquid cooling adds
$1,000-2,000 per kW [59]. Grid upgrades can exceed $2 million per megawatt and take 12-24 months [57]. Power transformer
lead times reached 120 weeks in NERC’s 2024 tracking – large transformers 210 weeks, switchgear 44 weeks – and AI
retrofits can stretch total site cost to $20 million or more [57][60][61].
Figure 3: Brownfield retrofit costs versus in-envelope deployment. Units differ: first two bars are thousands of USD per
By contrast, in-envelope deployment requires no facility work. A Furiosa NXT RNGD Server at 3 kW fits five per cabinet for
20 PFLOPS of FP8 [15]. An 8-way L40S 4U at 3-4 kW fits three to four per cabinet [16][25]. A SambaRack SN40L-16 at 10 kW
typical fits most envelopes [30]. Xeon plus Gaudi 3 PCIe at ~4.8 kW fits; liquid-cooled OAM variants do not. Spyre on
Power11 at 600 W per eight-card drawer fits comfortably [45].
Upgrading a 20-rack room from 15 to 30 kW at $5 million per megawatt adds $1.5 million in facility capex plus multi-month
procurement. Equivalent density via Furiosa or L40S deploys in weeks.
Supply chain is the dominant constraint in 2026 AI infrastructure procurement. TSMC CoWoS advanced packaging is
reported booked through mid-2027, with hyperscaler demand consuming leading-edge capacity [66]. Platforms on mature
5nm lines – Spyre, Furiosa, Gaudi 3 – face less CoWoS exposure than HBM-heavy frontier GPUs. GDDR6 memory on the
L40S is a related hedge, trading peak bandwidth for decoupling from the CoWoS-HBM queue. L40S ships ex-stock from
Dell, Supermicro, Lenovo, and ASUS [25][64][65]. Spyre has shipped since December 2025 [45][46], and Furiosa began mass
production in January 2026 [19][23].
Roadmap risk varies. Intel’s Falcon Shores cancellation and Jaguar Shores’ 2026-2027 successor positioning place Gaudi 3
late-cycle [40][41]. Furiosa carries small-company execution risk, offset by a Series D up to $500 million and the LG customer
win [14][63]. SambaNova closed a $350 million round in February 2026 with Intel Capital and expanded via SoftBank Japan [33].
IBM has committed Spyre across Power, IBM Z, and LinuxONE for a long product life [42][46].
NVIDIA L40S has the broadest software ecosystem of the six surveyed platforms, with the full CUDA, cuDNN, TensorRT,
and Triton stack plus native vLLM and TGI [16][39]. SambaNova’s SambaFlow and SambaStudio platform offer Composition of
Experts as a differentiator; independent tooling is narrower [30][32]. Intel’s SynapseAI plus PyTorch supports vLLM and TGI but
not Triton, and custom CUDA kernels must be reimplemented [39]. IBM Spyre ships with Python, RHEL tooling, and Live
Partition Mobility [45][47]; Furiosa’s SDK added vLLM-compatible APIs and HuggingFace Hub integration in 2025 [20][21]. CAPE
is pre-product.
Table 3: Operational tradeoffs across heterogeneous inference platforms, Q2 2026. Sources:[14], [16], [19], [20], [21], [23], [30], [32], [33], [39], [40], [41], [45], [46], [47], [50], [51], [63].
Once the decision framework favors in-envelope deployment, the question shifts from whether to how to configure the
hardware. Nodestream, a Blockware company, is positioned to configure and deliver heterogeneous compute platforms for
enterprises and colocation tenants running on-premise AI inference in existing 10-15 kW cabinets. The focus is intended
to translate the comparison above into rack-level configurations – server counts, power budgeting, cooling headroom,
and software stack – matched to workload mix and facility constraints.
Three principles guide the approach. Hardware selection is designed to be workload-led: RAG pipelines, embedding
services, mid-size LLM serving, and recommendation engines map onto the surveyed platforms by concurrency, latency,
and model size. Supply chain transparency is emphasized; with advanced-packaging bottlenecks pushing frontier GPU lead
times beyond typical planning windows, access to mature 5nm platforms is positioned as a practical advantage.
Configurations are designed to fit existing cabinets, avoiding facility upgrades the market is increasingly unable to absorb.
For organizations sizing AI inference capacity against cabinets provisioned for a pre-AI era, in-envelope
deployment of heterogeneous hardware is a credible alternative. Talk to Nodestream about configuring
heterogeneous compute for your existing racks.
Nodestream, a Blockware company, does not provide legal, tax, accounting, business, regulatory,
or financial advice.
A server that integrates multiple accelerator types – CPU plus NPU, CPU plus GPU, or CPU plus NPU plus
GPU – on a single motherboard, tuned for inference throughput per watt rather than peak training FLOPS.
Heterogeneous platforms fit a wider range of inference workloads within a constrained power envelope than
homogeneous GPU-only servers [11][12].
Five Furiosa NXT RNGD servers (3 kW each, 4U) fit within 15 kW for a combined 20 PFLOPS of FP8
compute [15]. An 8-way NVIDIA L40S 4U server draws 3 to 4 kW, so three to four fit in the same
envelope [16][25]. A SambaRack SN40L-16 draws 10 kW typical in a single 19-inch rack [30].
NPUs allocate more die area to on-chip SRAM and lower-power interconnect, trading peak FP16 throughput
for better performance per watt on inference-shaped workloads with lower batch sizes and tight latency.
GPUs retain advantages on large-batch matrix operations. Peer-reviewed research reports NPUs delivering
2.9 times higher throughput in specific inference tasks at comparable power efficiency [11].
Yes, on platforms designed for the envelope. Furiosa NXT RNGD (3 kW, air-cooled), 8-way NVIDIA L40S
systems (3 to 4 kW, air-cooled), SambaRack SN40L-16 (10 kW, air-cooled), Intel Gaudi 3 PCIe
configurations, and IBM Spyre on Power11 all fit in air-cooled 10 to 15 kW cabinets [15][16][25][30][45].
Yes. The HL-338 PCIe variant of Gaudi 3 is rated at 600 W per card, so an 8-card server draws roughly 4.8
kW for the accelerators plus host overhead and fits within a 15 kW air-cooled envelope. The OAM HL-325L
(450-900 W) and liquid-cooled HL-335 (up to 1,200 W) variants do not fit the same envelope without facility
changes [34][35][36].
Brownfield retrofits run $4 million to $8 million per megawatt excluding hardware, plus $1,000 to $2,000
per kW for liquid cooling where required [57][59]. Grid upgrades can exceed $2 million per megawatt alone
and take 12 to 24 months; transformer lead times averaged 120 weeks in NERC’s 2024 tracking [57][60].
Colocation tenants often cannot trigger those upgrades directly – the facility operator sets the cabinet ceiling.
Both describe inference energy efficiency. Tokens per second per watt measures useful output rate per
unit power; joules per token (or energy per token) measures the energy cost to produce each token and is
the reciprocal metric rising in 2025-2026 academic and trade literature. A platform that scores well on one
typically scores well on the other; both are worth tracking as inference-efficiency vocabulary evolves.
Platforms that do not rely on HBM and CoWoS advanced packaging are less exposed to the current TSMC
CoWoS bottleneck reported as booked through mid-2027 [66]. The NVIDIA L40S uses GDDR6 rather than
HBM; IBM Spyre uses LPDDR5; CPU-anchored heterogeneous configurations reduce HBM count per
server. These are hedges, not replacements, for the peak-bandwidth advantages HBM-based
accelerators provide.
Yes. IBM announced GA of Spyre on IBM Z and LinuxONE on October 28, 2025, and on Power11
on December 12, 2025 [45][46][48]. As of April 2026, the platform is shipping in early availability.
[2] AFCOM, “2026 State of the Data Center Report,” industry survey, March 2026 (via [1]).
[33] HPCwire, “SambaNova-Intel partnership and SoftBank Japan deployments,” cross-referenced trade press coverage,
2025-2026.