Heterogeneous Compute for Power-Constrained AI Inference

Existing data centers cannot support the power and cooling demands of conventional AI clusters. This whitepaper explores heterogeneous CPU, GPU, NPU, and RDU platforms designed for efficient inference inside 10 to 15 kW cabinets. Compare six architectures, performance-per-watt metrics, workload fit, software maturity, and deployment economics without costly facility upgrades.

Download the whitepaper now!

Executive Summary

If your AI inference roadmap has outrun the 10-15 kW cabinets in your existing facility, this paper shows how to close the gap without a facility upgrade. It compares six high-density heterogeneous compute platforms that fit inside air-cooled brownfield racks, and sets out a decision framework for on-premise AI inference deployment under strict thermal budgets.
The mean rack density reported by the Uptime Institute in 2025 was 7.6 kW [3], and even the AI-active AFCOM respondent base averaged 16 kW in 2025 and 27 kW in 2026 [1][2] – four to sixteen times below the 40-100 kW continuous draw of a modest H100 or H200 rack [8][9]. Upgrading a brownfield facility costs four to eight million dollars per megawatt and can stall a year or more on transformer and switchgear lead times [57][60].
A different class of hardware has arrived in response. High-density heterogeneous compute platforms – CPU, NPU, GPU, or RDU combinations engineered for inference throughput per watt rather than peak training FLOPS – are designed to fit within the 10-15 kW cabinets enterprises already own. This paper surveys six current implementations, compares them on published performance per watt, memory, workload fit, and software maturity, and provides a capex framework architects can apply to their fleets. The central finding: for inference workloads, facility-upgrade-free deployment is now a credible alternative to retrofitting power and cooling.

Introduction

AI inference is the center of gravity in data center compute. Omdia forecasts inference will account for roughly two-thirds of compute in 2026, up from one-third in 2023 [55]. IDC projects AI infrastructure spending at $758 billion by 2029, with enterprise investment rising from $307 billion in 2025 to $632 billion in 2028 [56]. Yet more than 70 percent of global data center capacity sits in existing buildings [5], much of it underutilized for AI [6].
For regulated and data-gravity workloads – healthcare, defense, finance, pharma – on-premise AI inference is not optional. The mismatch between where AI must run and where enterprise real estate can support it is an immediate constraint. The thesis: high-density heterogeneous compute platforms – integrating CPUs, NPUs, and GPUs on a single motherboard – offer a facility-upgrade-free path to inference throughput within strict thermal and power limits. The Furiosa NXT RNGD Server, NVIDIA® L40S-based systems, SambaNova SN40L, Intel® Xeon with Gaudi 3, IBM® Spyre on Power11, and the CAPE open architecture exemplify this approach.

The Power Gap: Why Legacy Racks Cannot Host Conventional AI Accelerators

Most installed enterprise and colocation cabinets were provisioned at 5-15 kW ceilings [62]. Data Center Dynamics reports enterprise facilities typically operate at 3-5 kW per rack, with new-build space engineered for 8-10 kW [4]. The Uptime Institute’s 2025 Global Data Center Survey places the industry-wide mean at 7.6 kW, up from 6.8 kW the year prior [3]. The 2026 AFCOM survey reports 27 kW – a 69 percent year-over-year increase, though AFCOM respondents skew toward operators actively planning for AI [1][2]. The installed-base median remains well below 30 kW.
Conventional AI accelerator clusters sit well above that envelope. An NVIDIA DGX H100 is rated at 10.2 kW maximum per 8-GPU chassis [8]; NVIDIA’s own guidance recommends no more than four DGX H100 systems per rack for thermodynamic reasons, pushing continuous draw past 40 kW [8]. Trade-press analysis citing Schneider Electric and Vertiv projections places fully loaded current-generation GPU racks at ~132 kW, with next-generation deployments expected at 240 kW within a year [9].
Figure 1: Power-constrained rack density versus AI accelerator deployment draw, 2025–2026. Sources: [1], [2], [3], [4], [8], [9].
The gap is structural. A brownfield cabinet at 5-15 kW sits four to sixteen times below a modest H100 deployment, and up to twenty-five times below a fully loaded DGX rack. Facility upgrade is neither fast nor cheap.

How Heterogeneous CPU + NPU + GPU Platforms Change the Equation

A second class of server architecture sidesteps the power gap rather than bridging it. Heterogeneous compute platforms integrate CPUs with Neural Processing Units (NPUs), GPUs, or Reconfigurable Dataflow Units (RDUs) on a single motherboard, tuned for inference performance per watt rather than peak training FLOPS.
The tradeoff is architectural. Homogeneous GPU servers such as the DGX H100 allocate silicon to the matrix-multiplication bandwidth and memory capacity training demands, consuming ten kilowatts or more per chassis and relying on NVLink and NVSwitch for GPU-to-GPU bandwidth [10]. Heterogeneous platforms allocate more die area to on-chip SRAM and lower-power interconnect, trading peak FP16 and FP8 throughput for better performance per watt on inference-shaped workloads with lower batch sizes and tight latency [11][12].
Two efficiency metrics dominate inference-platform comparisons: tokens per second per watt and joules per token (its reciprocal, gaining currency in 2025–2026). Either belongs in a spec-sheet evaluation.
Peer-reviewed research in Systems found NPU-based servers delivered 2.9x higher throughput than GPUs in specific inference workloads at comparable power efficiency, and outperformed GPUs by ~58 percent on video and LLM matrix-vector operations [11]. An arXiv analysis of multi-stage inference pipelines argues retrieval and reranking suit large-memory CPUs, embedding suits NPUs or GPUs, and prefill or decode scales with compute-heavy NPUs or GPUs – a case for heterogeneous pipelines over monolithic GPU stacks [12].
Rack-density math drives the case. Furiosa reports the NXT RNGD Server – a 4U, 8-card system – draws 3 kW and delivers 4 PFLOPS of FP8, placing five servers per 15 kW cabinet for 20 PFLOPS within the existing envelope [7][15]. An 8-way NVIDIA L40S 4U chassis draws 3-4 kW with host overhead, fitting three to four per cabinet [16][17][25]. An NVIDIA HGX H100 chassis at 10.2 kW admits only one unit per 15 kW cabinet after networking and PDU overhead [8].
Figure 2: High-density AI server rack fit: servers per 15 kW air-cooled cabinet by platform, based on vendor-reported chassis power. Sources: [7], [8], [15], [16], [17], [25].

Platform Comparison: Air-Cooled AI Server Options for Performance per Watt

Six implementations define the current heterogeneous inference landscape. The data below distinguishes vendor-reported figures from independent benchmarks.
Platform Accelerator TDP Memory Peak compute Server draw Efficiency (vendor-reported) Independent benchmark status
Furiosa NXT RNGD 180 W/card 48 GB HBM3, 1.5 TB/s 512 TFLOPS FP8 3 kW per 4U (8 cards) ~1.4 TFLOPS/W BF16 (EXAONE 3.5)[13] No MLPerf submission
NVIDIA L40S 350 W/card 48 GB GDDR6, 864 GB/s 733 TFLOPS FP8 ~3-4 kW per 8-card 4U 1,642 tok/s/L40S offline, Llama 3.1 8B[27] MLPerf Inference v5.0 and v5.1
SambaNova SN40L Not publicly disclosed 64 GiB HBM + up to 1.5 TiB DDR per socket 638 BF16 TFLOPS/socket 10 kW typical per SambaRack SN40L-16 ~10.2 PFLOPS BF16 per 10 kW rack[30] No MLPerf submission
Intel Gaudi 3 (Xeon + Gaudi) 450-900 W OAM; 600 W PCIe 128 GB HBM2e, 3.7 TB/s 1.8 PFLOPS FP8 Varies; PCIe fits 10-15 kW H100-range per workload[35] MLPerf v5.0 and v5.1
IBM Spyre on Power11 75 W/card 128 GB LPDDR5, ~200 GB/s 300+ TOPS FP16 600 W for 8 cards per drawer 4 TOPS/W target class[42] No MLPerf submission
CAPE (eHPS / EMDC) Research platform CXL-attached pools Not yet specified Not yet specified Not applicable – project runs to Nov 2027 Not applicable
Table 1: Six heterogeneous inference platforms, compared on published specifications and vendor-reported performance-per-watt efficiency. Sources:[7], [13], [15], [16], [17], [18], [24], [25], [27], [28], [29], [30], [34], [35], [36], [42], [43], [45], [50], [51]. SambaNova per-chip TDP is not publicly disclosed – Tier 3 verification gap.

Furiosa NXT RNGD Server

The Furiosa RNGD is a second-generation NPU on TSMC 5nm: 180 W per card, 48 GB of HBM3 at 1.5 TB/s, 256 MB of on-chip SRAM, and 512 TFLOPS of FP8 [13][18][19]. A 4U NXT RNGD Server carries eight cards for 384 GB of HBM3 and 12 TB/s aggregate bandwidth within 3 kW [15]. Furiosa’s Hot Chips 2024 disclosure reports ~1.4 TFLOPS per watt on BF16 inference of LG’s EXAONE 3.5 32B; vendor-internal testing claims 40 percent better performance per watt than L40S on the MLPerf GPT-J reference – all vendor-reported, with no MLPerf submission [13][14][22]. Mass production began January 2026 with a 4,000-unit TSMC/ASUS batch [19][23].

NVIDIA L40S-Based Systems

The L40S is NVIDIA’s Ada Lovelace data center GPU: 350 W per card, 48 GB of GDDR6 at 864 GB/s, and 733 dense FP8 TFLOPS via fourth-generation Tensor Cores [16][17][24]. A Dell PowerEdge XE7745 MLPerf Inference v5.0 submission showed 8-way L40S configurations placing among the top perf-per-watt results in the server category [25][26]. Red Hat’s v5.1 submission reported 1,642 tokens per second offline on a single L40S with FP8 Llama 3.1 8B via vLLM [27]. Software ecosystem breadth is the widest of the six surveyed platforms – CUDA, cuDNN, TensorRT, Triton, vLLM, and TGI are natively supported – and GDDR6 memory hedges against CoWoS-HBM supply constraints discussed below.

SambaNova SN40L Hybrid Platform

SambaNova’s SN40L RDU is a 2.5D-packaged accelerator on TSMC 5nm with 102 billion transistors per socket, published at ISSCC 2025 [28][29]. Per-chip TDP is not publicly disclosed – a verification gap this paper flags. Three-tier memory pairs 520 MiB of on-chip SRAM with 64 GiB of HBM and up to 1.5 TiB of external DDR per socket, delivering 638 BF16 TFLOPS [28]. A 16-socket SambaRack SN40L-16 draws 7-14.5 kW (10 kW typical) and aggregates 10.2 PFLOPS of BF16 in a 19-inch air-cooled enclosure [30]. SambaNova reports 1,000-plus tokens per second on Llama 3 8B and 114 tokens per second on Llama 3.1 405B per 16-socket node – vendor figures, not independently verified [30][31][32].

Intel Xeon with Gaudi 3

Intel’s Gaudi 3 carries 128 GB of HBM2e at 3.7 TB/s and 1.8 PFLOPS of FP8 across two compute dies on TSMC 5nm [34][35]. TDP varies by form factor: 450-900 W for the HL-325L OAM, 900 W (up to 1,200 W) for the liquid-cooled HL-335, and 600 W for the HL-338 PCIe card [34][35][36]. Intel-published benchmarks show H100-range performance, with some workloads up to 70 percent higher and others up to 10 percent lower [35]. Intel submitted to MLPerf Inference v5.0 and v5.1, though Gaudi 3-specific headline results are not highlighted in MLCommons summaries; prior-generation Gaudi 2 recorded 8,035 offline tokens per second on Llama 2 70B [37][38]. Intel cancelled Falcon Shores in January 2025; Jaguar Shores is the 2026 rack-scale successor, placing Gaudi 3 late in its cycle [40][41].

IBM Spyre on Power11

Spyre is an inference-focused SoC on Samsung 5LPE, rated at 75 W per PCIe 5.0 card, with 128 GB of LPDDR5 at ~200 GB/s and more than 300 TOPS of FP16 [42][43][44][45]. A Power11 I/O drawer holds up to eight cards for 1 TB of combined memory, 1.6 TB/s aggregate bandwidth, and 2.4 PFLOPS at 600 W [45]. IBM Power11 reached GA on July 25, 2025 [49]; Spyre reached GA on IBM Z and LinuxONE on October 28, 2025, and on Power11 on December 12, 2025 [45][46][48]. IBM i and AIX integration via RHEL 9.6 partitions plus Live Partition Mobility differentiates Spyre for existing Power customers [45].

CAPE (European Open Compute Architecture for Powerful Edge)

CAPE is a Horizon Europe research project (grant 101189899) running December 2024 through November 2027, with €5,996,250 in funding, coordinated by Universität Bielefeld [50][51]. The architecture combines Composable Infrastructure via CXL with two open COM-HPC platforms – the embedded High-Performance Server (eHPS) and the Embedded Micro Data Center (EMDC) – composing CPU, NPU, GPU, and RISC-V accelerators through CXL-attached pools [50][52]. CAPE is a reference architecture, not a commercial product; it signals the direction of open, sovereignty-oriented heterogeneous compute. Omdia projects Open Rack enclosures at more than 70 percent of data center rack revenue by 2030 [67].

Workload Fit: Which AI Inference Tasks Belong on Heterogeneous Platforms

Heterogeneous platforms are not a replacement for dedicated GPU clusters. The match depends on workload shape.
Workload Best fit Heterogeneous viability Source
Small-to-mid LLM inference (7B-70B) Heterogeneous Strong fit. NPU perf-per-watt lead is largest here. [14][27][42]
Retrieval-Augmented Generation pipelines Heterogeneous Strong fit. CPU + NPU stages map naturally. [12]
Embedding generation at scale Heterogeneous Strong fit. NPU acceleration of embedding. [12]
Recommendation systems Heterogeneous Strong fit. Feature stores on CPU, scoring on NPU/GPU. [11]
Small-model fine-tuning / LoRA Heterogeneous Strong fit. Vendor SDK support covers this regime. [20][21]
Latency-sensitive single-request Heterogeneous Strong fit. NPUs deliver sub-ms latency. [11]
Frontier-model training (>70B from scratch) Dedicated GPU Out of scope for 15 kW cabinets. [10]
Extreme-batch 405B inference Dedicated GPU / specialized rack SambaRack SN40L-16 is a credible exception. [30]
Mixture-of-Experts at hyperscale Dedicated GPU Blackwell delivers 10x throughput-per-MW over Hopper. [53][54]
Table 2: AI workload-to-platform fit within the 10-15 kW envelope. Sources:[10], [11], [12], [14], [20], [21], [27], [30], [42], [53], [54]
Inference-shaped workloads at enterprise concurrency – chatbots, copilots, internal RAG, recommendation scoring, classification – fit the heterogeneous envelope. Frontier training and hyperscale MoE serving do not, and belong on dedicated GPU infrastructure.

Capex Math: Facility Upgrade vs. In-Envelope Deployment in a Colocation 15 kW Rack

For a tenant in a colocation 15 kW rack, the upgrade option is often unavailable – the facility operator sets the ceiling.
Brownfield retrofit runs $4-8 million per megawatt excluding hardware; brownfield projects are 30 to 50 percent cheaper than greenfield ($5-6M vs. $8-10M per megawatt) and avoid two or more years of revenue delay [57][58]. Liquid cooling adds $1,000-2,000 per kW [59]. Grid upgrades can exceed $2 million per megawatt and take 12-24 months [57]. Power transformer lead times reached 120 weeks in NERC’s 2024 tracking – large transformers 210 weeks, switchgear 44 weeks – and AI retrofits can stretch total site cost to $20 million or more [57][60][61].
Figure 3: Brownfield retrofit costs versus in-envelope deployment. Units differ: first two bars are thousands of USD per
By contrast, in-envelope deployment requires no facility work. A Furiosa NXT RNGD Server at 3 kW fits five per cabinet for 20 PFLOPS of FP8 [15]. An 8-way L40S 4U at 3-4 kW fits three to four per cabinet [16][25]. A SambaRack SN40L-16 at 10 kW typical fits most envelopes [30]. Xeon plus Gaudi 3 PCIe at ~4.8 kW fits; liquid-cooled OAM variants do not. Spyre on Power11 at 600 W per eight-card drawer fits comfortably [45].
Upgrading a 20-rack room from 15 to 30 kW at $5 million per megawatt adds $1.5 million in facility capex plus multi-month procurement. Equivalent density via Furiosa or L40S deploys in weeks.

Supply Chain, Roadmap Risk, and Software Ecosystem

Supply chain is the dominant constraint in 2026 AI infrastructure procurement. TSMC CoWoS advanced packaging is reported booked through mid-2027, with hyperscaler demand consuming leading-edge capacity [66]. Platforms on mature 5nm lines – Spyre, Furiosa, Gaudi 3 – face less CoWoS exposure than HBM-heavy frontier GPUs. GDDR6 memory on the L40S is a related hedge, trading peak bandwidth for decoupling from the CoWoS-HBM queue. L40S ships ex-stock from Dell, Supermicro, Lenovo, and ASUS [25][64][65]. Spyre has shipped since December 2025 [45][46], and Furiosa began mass production in January 2026 [19][23].
Roadmap risk varies. Intel’s Falcon Shores cancellation and Jaguar Shores’ 2026-2027 successor positioning place Gaudi 3 late-cycle [40][41]. Furiosa carries small-company execution risk, offset by a Series D up to $500 million and the LG customer win [14][63]. SambaNova closed a $350 million round in February 2026 with Intel Capital and expanded via SoftBank Japan [33]. IBM has committed Spyre across Power, IBM Z, and LinuxONE for a long product life [42][46].
NVIDIA L40S has the broadest software ecosystem of the six surveyed platforms, with the full CUDA, cuDNN, TensorRT, and Triton stack plus native vLLM and TGI [16][39]. SambaNova’s SambaFlow and SambaStudio platform offer Composition of Experts as a differentiator; independent tooling is narrower [30][32]. Intel’s SynapseAI plus PyTorch supports vLLM and TGI but not Triton, and custom CUDA kernels must be reimplemented [39]. IBM Spyre ships with Python, RHEL tooling, and Live Partition Mobility [45][47]; Furiosa’s SDK added vLLM-compatible APIs and HuggingFace Hub integration in 2025 [20][21]. CAPE is pre-product.
Platform Software ecosystem Roadmap signal Supply in Q2 2026
NVIDIA L40S Broadest among surveyed platforms. CUDA, TensorRT, Triton, vLLM, TGI. Line continuing. No deprecation signal. Ex-stock from multiple OEMs.
SambaNova SN40L Proprietary. Llama + CoE Closed $350M round Feb 2026. Expanding. Direct and cloud channels.
Intel Gaudi 3 SynapseAI + PyTorch. No Triton. Late-cycle. Jaguar Shores 2026-2027. GA since 2H 2024.
IBM Spyre on Power11 Python, RHEL. Live Partition Mobility. Committed across Power, Z, LinuxONE. Shipping since Dec 2025.
Furiosa NXT RNGD vLLM-compatible. Narrower model zoo. Series D up to $500M. Gen-3 pending. Mass production Jan 2026.
CAPE Pre-product. Project ends Nov 2027. Not commercial.
Table 3: Operational tradeoffs across heterogeneous inference platforms, Q2 2026. Sources:[14], [16], [19], [20], [21], [23], [30], [32], [33], [39], [40], [41], [45], [46], [47], [50], [51], [63].

Solution Context: Configuring Heterogeneous Compute Inside the Existing Envelope

Once the decision framework favors in-envelope deployment, the question shifts from whether to how to configure the hardware. Nodestream, a Blockware company, is positioned to configure and deliver heterogeneous compute platforms for enterprises and colocation tenants running on-premise AI inference in existing 10-15 kW cabinets. The focus is intended to translate the comparison above into rack-level configurations – server counts, power budgeting, cooling headroom, and software stack – matched to workload mix and facility constraints.
Three principles guide the approach. Hardware selection is designed to be workload-led: RAG pipelines, embedding services, mid-size LLM serving, and recommendation engines map onto the surveyed platforms by concurrency, latency, and model size. Supply chain transparency is emphasized; with advanced-packaging bottlenecks pushing frontier GPU lead times beyond typical planning windows, access to mature 5nm platforms is positioned as a practical advantage. Configurations are designed to fit existing cabinets, avoiding facility upgrades the market is increasingly unable to absorb.

Conclusion

The gap between AI inference demand and installed data center power envelopes is structural. Brownfield
cabinets at 5-15 kW sit well below the draw of a conventional GPU rack, and retrofit costs millions per
megawatt across multi-year procurement queues. Heterogeneous compute platforms – Furiosa NXT RNGD,
NVIDIA L40S-based systems, SambaNova SN40L, Intel Xeon with Gaudi 3, IBM Spyre on Power11, and the
CAPE open architecture – offer a different path: inference throughput adequate for enterprise AI today, inside
racks organizations already own. Workload fit, performance per watt, software maturity, roadmap risk, and
supply chain availability form the framework for matching platform to deployment.

For organizations sizing AI inference capacity against cabinets provisioned for a pre-AI era, in-envelope deployment of heterogeneous hardware is a credible alternative. Talk to Nodestream about configuring heterogeneous compute for your existing racks.
Nodestream, a Blockware company, does not provide legal, tax, accounting, business, regulatory, or financial advice.

Frequently Asked Questions

What is a heterogeneous compute platform for AI inference?

A server that integrates multiple accelerator types – CPU plus NPU, CPU plus GPU, or CPU plus NPU plus GPU – on a single motherboard, tuned for inference throughput per watt rather than peak training FLOPS. Heterogeneous platforms fit a wider range of inference workloads within a constrained power envelope than homogeneous GPU-only servers [11][12].

How many AI inference servers fit in a 15 kW data center cabinet?

Five Furiosa NXT RNGD servers (3 kW each, 4U) fit within 15 kW for a combined 20 PFLOPS of FP8 compute [15]. An 8-way NVIDIA L40S 4U server draws 3 to 4 kW, so three to four fit in the same envelope [16][25]. A SambaRack SN40L-16 draws 10 kW typical in a single 19-inch rack [30].

What is the difference between an NPU and a GPU for LLM inference?

NPUs allocate more die area to on-chip SRAM and lower-power interconnect, trading peak FP16 throughput for better performance per watt on inference-shaped workloads with lower batch sizes and tight latency. GPUs retain advantages on large-batch matrix operations. Peer-reviewed research reports NPUs delivering 2.9 times higher throughput in specific inference tasks at comparable power efficiency [11].

Can you run LLM inference in an air-cooled 10 to 15 kW rack?

Yes, on platforms designed for the envelope. Furiosa NXT RNGD (3 kW, air-cooled), 8-way NVIDIA L40S systems (3 to 4 kW, air-cooled), SambaRack SN40L-16 (10 kW, air-cooled), Intel Gaudi 3 PCIe configurations, and IBM Spyre on Power11 all fit in air-cooled 10 to 15 kW cabinets [15][16][25][30][45].

Can Intel Gaudi 3 PCIe cards run in a 15 kW air-cooled cabinet?

Yes. The HL-338 PCIe variant of Gaudi 3 is rated at 600 W per card, so an 8-card server draws roughly 4.8 kW for the accelerators plus host overhead and fits within a 15 kW air-cooled envelope. The OAM HL-325L (450-900 W) and liquid-cooled HL-335 (up to 1,200 W) variants do not fit the same envelope without facility changes [34][35][36].

What does it cost to upgrade a colocation cabinet from 15 kW to 30 kW?

Brownfield retrofits run $4 million to $8 million per megawatt excluding hardware, plus $1,000 to $2,000 per kW for liquid cooling where required [57][59]. Grid upgrades can exceed $2 million per megawatt alone and take 12 to 24 months; transformer lead times averaged 120 weeks in NERC’s 2024 tracking [57][60]. Colocation tenants often cannot trigger those upgrades directly – the facility operator sets the cabinet ceiling.

What is the difference between joules per token and tokens per second per watt?

Both describe inference energy efficiency. Tokens per second per watt measures useful output rate per unit power; joules per token (or energy per token) measures the energy cost to produce each token and is the reciprocal metric rising in 2025-2026 academic and trade literature. A platform that scores well on one typically scores well on the other; both are worth tracking as inference-efficiency vocabulary evolves.

What AI inference hardware avoids HBM supply chain shortages?

Platforms that do not rely on HBM and CoWoS advanced packaging are less exposed to the current TSMC CoWoS bottleneck reported as booked through mid-2027 [66]. The NVIDIA L40S uses GDDR6 rather than HBM; IBM Spyre uses LPDDR5; CPU-anchored heterogeneous configurations reduce HBM count per server. These are hedges, not replacements, for the peak-bandwidth advantages HBM-based accelerators provide.

Is IBM Spyre on Power11 available for enterprise AI inference?

Yes. IBM announced GA of Spyre on IBM Z and LinuxONE on October 28, 2025, and on Power11 on December 12, 2025 [45][46][48]. As of April 2026, the platform is shipping in early availability.

What is the CAPE open compute architecture?

CAPE – European Open Compute Architecture for Powerful Edge – is an EU Horizon Europe research
project (grant 101189899) running December 2024 through November 2027, coordinated by Universität
Bielefeld. It specifies two open hardware platforms (eHPS and EMDC) built on COM-HPC and CXL
for composable heterogeneous compute. CAPE is a reference architecture, not a commercial
product [50][51][52].

References

[1] Data Center Knowledge, “AFCOM: Rack Density Surges as AI Overhauls Data Center Design,” industry news, March 2026.
https://www.datacenterknowledge.com/data-center-construction/afcom-rack-density-and-build-outs-surge-as-ai-overhauls -data-center-design
[2] AFCOM, “2026 State of the Data Center Report,” industry survey, March 2026 (via [1]).
[3] Uptime Institute, “Global Data Center Survey Results 2025,” 15th annual, 2025.
https://uptimeinstitute.com/resources/research-and-reports/uptime-institute-global-data-center-survey-results-2025

[4] Data Center Dynamics, “Watts up?” undated analysis, accessed 2026.
https://uptimeinstitute.com/resources/research-and-reports/uptime-institute-global-data-center-survey-results-2025

CITATION RISK: Reference [4] is undated (trade press, accessed 2026). Rule 10 requires the date or time period for cited statistics. The market-data claim derived from this source (enterprise facilities typically operate at 3-5 kW per rack) cannot be verified as within 18 months. Confirm the publication date; if unrecoverable, replace with a dated source or remove the claim.

[5] Data Center Dynamics, “Brownfield retrofits: The overlooked advantage in the AI race,” 2024-2025.
https://www.datacenterdynamics.com/en/opinions/brownfield-retrofits-the-overlooked-advantage-in-the-ai-race/
[6] Schneider Electric Blog, “Brownfield data center modernization for AI,” 2026-02-26.
https://blog.se.com/datacenter/2026/02/26/brownfield-data-center-modernization-ai/
[7] FuriosaAI, “Introducing Furiosa NXT RNGD Server,” vendor blog, September 2025.
https://furiosa.ai/blog/introducing-rngd-server-efficient-ai-inference-at-data-center-scale
[8] NVIDIA, “Electrical Specifications — DGX SuperPOD H100 Data Center Design Guide,” undated vendor documentation, accessed 2026.
https://docs.nvidia.com/dgx-superpod/design-guides/dgx-superpod-data-center-design-h100/latest/electrical.html
[9] ZincFive, “Data Center Rack Power Trends and What They Mean for Build-Outs,” 2025-11-25.
https://zincfive.com/blog/2025/11/25/data-center-rack-power-trends-and-what-they-mean-for-build-outs/
[10] NVIDIA, “Introduction to NVIDIA DGX H100/H200 Systems,” undated, accessed 2026.
https://docs.nvidia.com/dgx/dgxh100-user-guide/introduction-to-dgxh100.html
[11] MDPI Systems, “Performance and Efficiency Gains of NPU-Based Servers over GPUs for AI Model Inference,” vol. 13, no. 9, p. 797, 2025
https://www.mdpi.com/2079-8954/13/9/797
[12] arXiv 2504.09775v2, “Understanding and Optimizing Multi-Stage AI Inference Pipelines,” preprint, April 2025.
https://arxiv.org/html/2504.09775v2
[13] Chips and Cheese, “FuriosaAI’s RNGD at Hot Chips 2024,” 2024-09-11.
https://chipsandcheese.com/p/furiosaais-rngd-at-hot-chips-2024-accelerating-ai-with-a-more-flexible-primitive
[14] The Register, “How AI chip upstart FuriosaAI won over LG,” 2025-07-22.
https://www.theregister.com/2025/07/22/sk_furiosa_ai_lg/
[15] FuriosaAI, “NXT RNGD Server product details,” vendor blog, September 2025.
https://furiosa.ai/blog/introducing-rngd-server-efficient-ai-inference-at-data-center-scale
[16] NVIDIA, “L40S GPU for AI and Graphics Performance,” vendor product page, accessed 2026.
https://www.nvidia.com/en-us/data-center/l40s/
[17] NVIDIA, “L40S GPU Accelerator Product Brief,” vendor, accessed 2026.
https://gzhls.at/blob/ldb/0/5/9/0/39bfaa3357547208384aea622ed99ac414d0.pdf
[18] FuriosaAI, “RNGD Product Page,” vendor, accessed 2026.
https://furiosa.ai/rngd
[21] FuriosaAI, “Furiosa SDK 2025.3 boosts RNGD performance,” 2025.
https://furiosa.ai/blog/furiosaai-sdk-2025-3-boosts-rngd-performance-with-multichip-scaling-and-more
[22] FuriosaAI Developer Center, “Running MLPerf Inference Benchmark,” v2025.2.0, 2025.
https://developer.furiosa.ai/v2025.2.0/en/getting_started/furiosa_mlperf.html
[24] NVIDIA, “L40S product page,” vendor, accessed 2026.
https://www.nvidia.com/en-us/data-center/l40s/
[25] Dell, “PowerEdge XE7745,” vendor product page, accessed 2026.
https://www.dell.com/en-us/shop/ipovw/poweredge-xe7745
[26] MLCommons, “MLPerf Inference v5.0 Benchmark Results,” 2025-04-02.
https://mlcommons.org/2025/04/mlperf-inference-v5-0-results/
[27] Red Hat, “Efficient and reproducible LLM inference with Red Hat: MLPerf Inference v5.1 results,” September 2025.
https://www.redhat.com/en/blog/efficient-and-reproducible-llm-inference-red-hat-mlperf-inference-v51-results
[28] IEEE Xplore, “SambaNova SN40L: A 5nm 2.5D Dataflow Accelerator with Three Memory Tiers for Trillion Parameter AI,” ISSCC 2025.
https://ieeexplore.ieee.org/document/10904578/
[29] arXiv 2405.07518v1, “SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts,”preprint, May 2024.
https://arxiv.org/html/2405.07518v1
[30] SambaNova, “SambaRack SN40L-16 data sheet,” 2025-07-09.
https://sambanova.ai/hubfs/SambaRack%20data%20sheet%20template%2007%2009%2025.pdf
[31] The Register, “SambaNova Cloud serves up Llama 3.1 405B at 100+ token/s,” 2024-09-10.
https://www.theregister.com/2024/09/10/sambanovas_inference_cloud/
[32] SambaNova, “Why SambaNova’s SN40L Chip Is the Best for Inference,” vendor, undated, accessed 2026.
https://sambanova.ai/blog/sn40l-chip-best-inference-solution
[33] HPCwire, “SambaNova-Intel partnership and SoftBank Japan deployments,” cross-referenced trade press coverage, 2025-2026.
[34] Intel, “Gaudi 3 AI Accelerator White Paper,” vendor white paper, 2024.
https://cdrdv2-public.intel.com/817486/gaudi-3-ai-accelerator-white-paper.pdf
[35] Intel, “Gaudi 3 HL-325L OAM Mezzanine Card Product Brief,” vendor product brief, 2024.
https://cdrdv2-public.intel.com/817487/gaudi-3-ai-accelerator-hl-325l-oam-mezzanine-card-product-brief.pdf
[36] Intel, “Gaudi 3 PCIe Product Brief,” 2024-2025.
https://cdrdv2-public.intel.com/817488/Gaudi%203%20PCIe%20Product%20Brief_RB_1_V6.pdf
[37] MLCommons, “MLPerf Inference v5.1 Benchmark Results,” 2025-09-09.
https://mlcommons.org/2025/09/mlperf-inference-v5-1-results/
[38] Intel Newsroom, “Intel Gaudi Enables Lower Cost Alternative for AI Compute and GenAI,” press release, 2024.
https://newsroom.intel.com/artificial-intelligence/intel-gaudi-enables-lower-cost-choice-for-genai-2
[39] Introl, “Intel Gaudi 3 Deployment Guide: H100 Alternative,” 2024-2025.
https://introl.com/blog/intel-gaudi-3-deployment-guide-h100-alternative
[41] The Next Platform, “Intel Pushes Out Clearwater Forest Xeon 7, Sidelines Falcon Shores Accelerator,” 2025-01-31.
https://www.nextplatform.com/2025/01/31/intel-pushes-out-clearwater-forest-xeon-7-sidelines-falcon-shores-accelerator/
[42] IBM Research, “Introducing the IBM Spyre AI accelerator chip,” research blog, 2024.
https://research.ibm.com/blog/spyre-for-z
[43] The Next Platform, “IBM Shows Off Next-Gen AI Acceleration,” 2024-08-27.
https://www.nextplatform.com/2024/08/27/ibm-shows-off-next-gen-ai-acceleration-on-chip-dpu-for-big-iron/
[44] G. Bonshor, “IBM’s Spyre AI Accelerator Deep Dive,” More Than Moore Substack, 2025.
https://morethanmoore.substack.com/p/ibms-spyre-ai-accelerator-deep-dive
[45] IT Jungle, “A Bit More Insight Into IBM’s Spyre AI Accelerator For Power,” 2025-10-20.
https://www.itjungle.com/2025/10/20/a-bit-more-insight-into-ibms-spyre-ai-accelerator-for-power/
[46] IBM Newsroom, “IBM Introduces the Spyre Accelerator for Commercial Availability,” 2025-10-07.
https://newsroom.ibm.com/2025-10-07-ibm-introduces-the-spyre-accelerator-for-commercial-availability
[47] IBM Docs, “IBM Spyre Accelerator for Power,” IBM documentation, accessed 2026.
https://www.ibm.com/docs/en/SSV8TX/pdf/spyreonpower.pdf
[48] TechChannel, “IBM’s Spyre AI Accelerator Gets Oct. 28 GA for Mainframes, December Debut for Power11,” October 2025.
https://techchannel.com/artificial-intelligence/spyre-october-28-ga/
[49] IBM Newsroom, “IBM Power11 Raises the Bar for Enterprise IT,” 2025-07-08.
https://newsroom.ibm.com/2025-07-08-ibm-power11-raises-the-bar-for-enterprise-it
[50] European Commission / CORDIS, “CAPE — European Open Compute Architecture for Powerful Edge — Fact Sheet,” project 101189899, 2024.
https://cordis.europa.eu/project/id/101189899
[51] Innovation News Network, “Composability for powerful edge computing,” 2025-2026.
https://www.innovationnewsnetwork.com/composability-for-powerful-edge-computing/66804/
[52] CAPE Project, “Superpowering the Edge,” homepage, accessed 2026.
https://cape-project.eu/
[53] NVIDIA, “Blackwell Leads on SemiAnalysis InferenceMAX v1 Benchmarks,” October 2025.
https://developer.nvidia.com/blog/nvidia-blackwell-leads-on-new-semianalysis-inferencemax-benchmarks/
[54] SemiAnalysis, “InferenceMAX Open Source Inference Benchmarking,” October 2025.
https://newsletter.semianalysis.com/p/inferencemax-open-source-inference
[56] IDC, “Artificial Intelligence Infrastructure Spending to Reach $758Bn USD Mark by 2029,” press release, 2025.
https://my.idc.com/getdoc.jsp?containerId=prUS53894425

CITATION RISK: IDC citation from 2025 (no month specified) — publication date is incomplete. Rule 17 requires a full publication date for analyst citations; "2025" alone is non-compliant. Replace with a more precisely dated IDC source, paraphrase without citing IDC, or escalate if central to the thesis. If the exact publication date can be confirmed as within 18 months of April 2026, update the reference with the specific date.

[57] TrueLook, “Data Center Construction Costs Explained,” industry guide, 2025.
https://www.truelook.com/blog/data-center-construction-costs
[58] Data Center Dynamics, “Brownfield retrofits: The overlooked advantage in the AI race,” 2024-2025.
https://www.datacenterdynamics.com/en/opinions/brownfield-retrofits-the-overlooked-advantage-in-the-ai-race/
[59] California Energy Commission, “Demonstration of Low-Cost Data Center Liquid Cooling,” CEC-500-2024-061, June 2024.
https://www.energy.ca.gov/sites/default/files/2024-06/CEC-500-2024-061.pdf
[60] Power Magazine, “Transformers in 2026: Shortage, Scramble, or Self-Inflicted Crisis?,” 2026.
https://www.powermag.com/transformers-in-2026-shortage-scramble-or-self-inflicted-crisis/
[61] Vertiv, “The cost impact of AI data center design, build and operations,” undated vendor content, accessed 2026.
https://www.vertiv.com/en-us/about/news-and-insights/articles/educational-articles/the-cost-impact-of-ai-data-center-design-build-and-operations/

CITATION RISK: Reference [61] is undated vendor content (accessed 2026). Rule 10 requires the date or time period for cited statistics. The cost-projection figures sourced from this reference cannot be verified as current. Confirm the publication date; if unrecoverable, corroborate with a dated independent source or remove.

[61] Vertiv, “The cost impact of AI data center design, build and operations,” undated vendor content, accessed 2026.
https://www.vertiv.com/en-us/about/news-and-insights/articles/educational-articles/the-cost-impact-of-ai-data-center-design-build-and-operations/

[62] Datacenters.com, “kW per Rack Explained,” industry explainer, 2025.
https://www.datacenters.com/news/kw-per-rack-explained-optimize-your-data-center
[63] Verdict, “FuriosaAI targets up to $500m in Series D for AI chip expansion,” August 2025.
https://www.verdict.co.uk/furiosaai-series-d-ai-chips/
[64] Supermicro, “NVIDIA L40S Optimized Systems,” product catalog, accessed 2026.
https://www.supermicro.com/en/accelerators/nvidia/l40s
[65] Lenovo Press, “ThinkSystem NVIDIA L40S 48GB PCIe Gen4 Passive GPU Product Guide,” 2023-2025.
https://lenovopress.lenovo.com/lp1812-nvidia-l40s-48gb-pcie-gen4-passive-gpu
[66] Silicon Analysts, “AI Data Center Value Chain Analysis 2025,” research report, 2025.
https://siliconanalysts.com/research/ai-data-center-value-chain

CITATION RISK: Silicon Analysts citation from 2025 (no month specified) — publication date is incomplete. Rule 17 requires a full publication date for analyst citations; "2025" alone is non-compliant. Replace with a more precisely dated source, paraphrase without citing Silicon Analysts, or escalate if central to the thesis. If the exact publication date can be confirmed as within 18 months of April 2026, update the reference with the specific d

[67] Omdia, “Data Center Racks Tracker — 2025,” 2025.
https://omdia.tech.informa.com/om136685/data-center-racks-tracker–2025

CITATION RISK: Omdia citation from 2025 (no month specified) — publication date is incomplete. Rule 17 requires a full publication date for analyst citations; "2025" alone is non-compliant. Replace with a more precisely dated Omdia source, paraphrase without citing Omdia, or escalate if central to the thesis. If the exact publication date can be confirmed as within 18 months of April 2026, update the reference with the specific date.