Bristol Myers Squibb said Monday it will deploy an NVIDIA DGX SuperPOD, a prepackaged supercomputing platform that NVIDIA sells as a single standard unit, combining multiple compute racks, high-speed networking and management software into one cluster. BMS describes it as the “most powerful and energy-efficient single-owned NVIDIA infrastructure in life sciences.”
That claim lands nine months after Eli Lilly announced the industry’s most powerful AI supercomputer in October, and four months after Roche claimed the industry’s largest announced hybrid-cloud AI factory in March.
The build is an expansion rather than a first move. BMS has operated a DGX SuperPOD since 2024, and NVIDIA’s blog says that system is now saturated. The two will be combined into a single environment accessible across BMS sites.
What BMS bought
Each rack in the new BMS cluster is a Vera Rubin NVL72 system, pairing 72 Rubin GPUs with 36 Vera CPUs so the rack operates as one machine rather than 72 separate ones. NVIDIA’s blog puts the build at eight racks, or 576 GPUs, a figure absent from BMS’s own release.
Rubin is the generation that follows Blackwell, the chips behind both Lilly’s and Roche’s systems. For context, at Blackwell’s launch in March 2024, NVIDIA said its GB200 NVL72 rack, which links 72 Blackwell GPUs into a single system, delivered up to 30 times the LLM inference performance of the same number of H100 GPUs and cut cost and energy consumption by up to 25 times.. NVIDIA said Rubin entered full production earlier this year, with partner availability in the second half of 2026. NVIDIA claims roughly 10 times the inference throughput per watt of Blackwell at the rack level. Meanwhile, BMS cites up to 10 times the performance per megawatt over the system it is replacing. Both figures come from the vendor.
Reuters reported BMS as the first life-sciences company to buy a Rubin-based SuperPOD. That holds with one qualifier: NVIDIA said in January that Lilly’s $1 billion co-innovation lab in South San Francisco would be built on the Vera Rubin architecture. Lilly named the architecture. BMS named a rack count.
Three systems, three different disclosures
The table below uses theoretical dense FP8 training performance, a common low-precision AI metric, for the two systems. These are peak reference-spec estimates. Measured performance on pharmaceutical workloads remains undisclosed. Roche’s disclosure omits the Blackwell model and the network topology.
System
GPU hardware
Peak dense FP8 training*
GPU memory
Disclosed layout
BMS, planned
576 Rubin GPUs
10.1 exaflops
166 TB
Eight NVL72 racks, each with 72 GPUs and 36 Vera CPUs
LillyPod, live
1,016 B300 Blackwell Ultra GPUs
4.6 exaflops
293 TB
Eight GPUs per DGX B300 system, linked through a SuperPOD network
Roche, operating
2,176 Blackwell GPUs on premises, more than 3,500 total
Unavailable from disclosure
Unavailable from disclosure
Hybrid footprint across U.S. and European sites plus cloud, GPU model and topology undisclosed
*The estimates multiply NVIDIA’s current reference specifications across the disclosed hardware. Vera Rubin NVL72 provides 1.26 exaflops of dense FP8/FP6 training performance per rack. DGX B300 provides 72 petaflops of sparse FP8 performance per eight-GPU system, equal to 36 petaflops dense. NVIDIA labels the Rubin specifications preliminary, and separate B300 module documentation implies a higher dense FP8 rate than the DGX system figure, which would raise the Lilly estimate. Lilly’s published figure of more than 9,000 petaflops aligns with NVIDIA’s sparse FP8 specification across 127 DGX B300 systems; the table converts it to dense FP8 for comparison. Memory totals use 20.7 TB per Rubin rack and 288 GB per B300 GPU.
Filed Under: Drug Discovery
Tagged With: AI infrastructure, Blackwell Ultra, BMS, Bristol-Myers Squibb, DGX SuperPOD, drug discovery AI, Eli Lilly, life sciences computing, LillyPod, NVIDIA, pharma supercomputer, Pharmaceutical AI, Roche AI factory, Rubin GPUs, Vera Rubin