This section presents the computational requirements for training state-of-the-art AI models on the Nvidia A100 SXM 40 GB GPU. While AI training typically occurs in large-scale data center environments supporting diverse digital workloads, this analysis isolates the material requirements directly attributable to GPU-based AI training. It is important to note that the results here reflect only the GPU unit. Broader infrastructure components, such as networking, storage, and cooling systems, are excluded from this analysis. Including these hardware components would substantially increase the total material footprint, and the values reported here should therefore be interpreted as conservative lower-bound estimates of AI’s overall material demand.
GPU requirement per AI model
Using Eq. 2, we derive the computational budgets (measured in FLOPs) of eight large-scale dense transformer models developed between 2022 and 2024, trained on Nvidia A100 GPUs (see Table 1). The estimated number of FLOPs per model is derived from reported or inferred parameter and token counts. Where official data on parameters or token counts is unavailable, estimates based on industry conventions and credible leaked data sources (for more details, see supplementary materials (SM) Table S1).
Table 1 GPU demand for AI training under varying lifespan and MFU scenarios
For GPT-4, we separately estimate a range of plausible computational budget scenarios based on MoE architectural assumptions (see SM Table S2). The estimated GPT-4 MoE configurations are based on publicly discussed architectural hypotheses and disclosed MoE architectures assuming a training dataset of 13 trillion tokens. For the remainder of this study, we adopt the scenario outlined by SemiAnalysis, which estimates GPT-4 was trained on 1.76 T total parameters, 222 B active parameters, and 13 T tokens28.
For our estimations, GPU demand is expressed as equivalent hardware lifetime consumed, translating cumulative computational workloads into fractional depletion of GPU productive capacity regardless of the parallelization strategy employed. This lets us directly estimate how many GPUs are effectively used during a training run and, thus, the material demand involved. It further allows standardized comparison across models with heterogeneous training strategies. While it is also possible to estimate the proportional wear and tear across all GPUs used in parallel, the hardware lifetime consumed modeling approach more effectively demonstrates the scale and intensity of hardware turnover implicit in contemporary AI training workloads.
Using Eq. 5, we estimate the number of Nvidia A100 GPUs required to train each AI model based on five life span scenarios and applying a lower bound MFU of 20% and an upper bound estimate of 50% MFU (see Table 1). GPU requirements scale with model complexity, i.e., compute budget. Training GPT-4 (based on the MoE SemiAnalysis scenario) with an MFU of 20% and a hardware lifespan of 1 year requires 8800 A100 GPUs, decreasing to approximately 2934 GPUs with a hardware lifespan of 3 years. Improving the MFU to 50% reduces the range from 3520 GPUs (1-year lifespan) to 1174 GPUs (3-year lifespan). GPT-4 requires the highest number of GPUs, followed by Amazon Titan (2439 to 326 GPUs), Mistral large 2 (752 to 151 GPUs), LLaMa 2 (427 to 57 GPUs), DeepSeekLLM (409 to 55 GPUs), BLOOM (196 to 27 GPUs), GPT-3.5 (160 to 22 GPUs), and Falcon (122 to 17 GPUs). Pythia demonstrates the lowest requirements (11 to 2 GPUs), illustrating the substantial variation in computational demands across different model sizes.
Elemental composition of the Nvidia A100 GPU
The ICP-OES elemental analysis of the GPU identifies 32 elements from the periodic table in the Nvidia A100 SXM 40 GB GPU (see Fig. 1), spanning metals (84%), metalloids (12.5%), and one nonmetal. The five dominant elements by mass are copper, iron, tin, silicon, and nickel, with copper alone accounting for 1374 grams. At the other extreme, beryllium registers as the least abundant at 0.0000238 grams (see Table 2). Four of the eight precious metals—gold, silver, platinum, and palladium—were detected, though primarily in trace amounts; silver is the most abundant among them at 0.55 grams per GPU.
Fig. 1: Elemental composition of the NVIDIA A100 SXM GPU.
The alternative text for this image may have been generated using AI.
Proportion of elements in the Nvidia A100 SXM 40 GB GPU (author illustration).
Table 2 Elemental composition of the A100 GPU by component group
Beyond total quantities, the elemental composition differs considerably across the four GPU components: the heatsink, PCB, GPU chip, and power-on-packages (PoP) (see Table 2). The heatsink, the most substantial component by mass, is composed of 98.1% copper, reflecting its primary role in thermal management. The PCB exhibits a more heterogeneous profile, comprising 46.5% copper and 28% iron alongside smaller proportions of silicon, tin, and calcium. The internal computing components show greater elemental complexity: the PoP comprises 52.6% copper, 19% iron, and 6.6% magnesium with notable amounts of barium and zinc, while the GPU chip is characterized by 41% chromium, 29% silicon, 17% tin, and smaller portions of bismuth and aluminum. Copper and iron thus dominate the structural components, while silicon and nickel are predominantly concentrated in the functional components. The substantial amount of silicon of the GPU chip is primarily attributable to the large die size of the main processor. Our analysis of the dismantled unit reveals that approximately 1353 mm2 of silicon area is integrated within a single 55 mm × 55 mm, 12-layer packed GPU inside the Nvidia A100 SXM. In comparison, the die area of its predecessor, the Nvidia Tesla V100, is much smaller at 815 mm2 29.
The resource cost of AI training
Scaling the material level from individual units to GPU counts involved in AI model training reveals a substantial physical footprint. Training only one single round of GPT-4 at a reported MFU of 35%30 requires the computational capacity of approximately 2515 A100 GPUs under the most plausible baseline scenario of a 2-year hardware lifespan (see Fig. 2). This corresponds to the extraction of about 3750 kg of the analyzed materials.
Fig. 2: GPU and elemental requirements for training GPT-4 across lifespans.
The alternative text for this image may have been generated using AI.
Estimated hardware and elemental requirements for training GPT-4 at reported 35% MFU across varying hardware lifespan scenarios (1–3 years). Results are expressed in terms of the total number of GPUs required and the total elemental mass (kg) on a logarithmic scale. Extending the hardware’s operational lifespan to 2 years halves the GPU demand to 2515 GPUs, while a 3-year lifespan reduces requirements by approximately 67% to 1676 GPUs. (author illustration).
These figures are relevant because the GPU contains a broad suite of heavy metals, including arsenic, mercury, lead, cadmium, chromium, zinc, copper, nickel, antimony, cobalt, and beryllium, that are classified as hazardous and have well-documented toxic properties if released during mining, manufacturing or disposal at the end of their life31,32,33. The elemental analysis of the GPU indicates that 93% of the Nvidia A100 GPU consist of elements that are classified as hazardous based on their toxic properties under relevant exposure conditions. For instance, exposure to these elements through inhalation, dermal contact, or
ingestion of contaminated water in these contexts can lead to lung cancer, neurological impairment, gastrointestinal disorders, and other long-term health impacts31. Toxic metals pose a health hazard in mining at varying concentrations depending on the metal; lead, for example, is dangerous at extremely low parts-per-billion (ppb) levels33. In developing countries, untreated or inadequately treated industrial wastewater and mining are primary sources of metal pollution in freshwater systems34. In sub-Saharan Africa, for example, the rapid expansion of mining and processing has increased the concentration of toxic metals in terrestrial, aquatic, and atmospheric systems, thereby exacerbating ecological degradation and health risks for workers and surrounding communities35. More specifically, in many of those mining regions, concentrations of toxic metals in the soil and water substantially exceed WHO drinking water thresholds, posing risks to communities35 (e.g., levels are above 10 µg/L for As and Pb, 3 µg/L for Cd, 20 µg/L for Cr, and 2000 µg/L for Cu33). While the concentrations of individual metals may appear modest at the level of a single GPU unit, large-scale AI training requires thousands of GPUs, thereby magnifying the pressures of upstream extraction and the risks of downstream contamination across soil, air, and groundwater systems.
Aggregating the material consumption across the nine models listed in Table 1, the most plausible baseline scenario (MFU = 35%, lifespan = 2 years) requires 4455.45 GPUs, corresponding to a total material footprint of about 6640 kg of extracted resources, of which 6175 kg are hazardous metals with recognized toxic properties. For reference, a lower-bound scenario (20% MFU, 1-year lifespan) requires 13,315 GPUs and 19,940 kg of extracted resources, while an upper-bound scenario (50% MFU, 3-year lifespan) reduces material extraction to approximately 2740 kg, of which 2550 kg remain hazardous material with recognized toxic properties.
These findings demonstrate that the environmental impact of large-scale AI model training extends beyond operational energy and carbon emissions. The material intensity of hardware production, including the extraction, processing, and disposal of potentially hazardous elements, constitutes a critical yet frequently overlooked aspect of AI sustainability27. The environmental impact of mining and e-waste disposal is moreover concentrated in regions with limited environmental governance and capacity to mitigate associated health and environmental risks36.
The nine models considered in this study represent only a fraction of the aggregate AI industry resource consumption. Thus, comparing material quantities from individual training runs to global annual metal extraction would yield negligible ratios that risk obscuring rather than contextualizing the scale of AI’s material demand; the meaningful comparison is at the industry aggregate level, not the single-training-run level. To quantify sector-wide resource consumption and compare with other industries, annual GPU shipment data is required. However, this information is not publicly disclosed by Nvidia, and current assessments therefore remain constrained to training-run-level analyses rather than comprehensive industry-scale evaluations. Moreover, to meet the computational demands of large-scale AI training, semiconductor devices are increasingly scaling their physical dimensions, either through larger die areas or advanced packaging that integrates multiple smaller dies. These trends
lead to increased silicon consumption as AI models continue to grow. However, constraints, such as the maximum reticle area, manufacturing costs, cooling challenges, and diminishing functional die yield, impose practical limits on die scaling, a phenomenon often referred to as “area-wall” 37. Among these limitations, thermal management emerges as a particularly critical challenge as chip size increases37, underscoring the need for enhanced cooling solutions and the consideration of waste heat utilization38. Hence, we suggest that meeting the computational demands of increasingly large AI models will necessitate scaling via additional GPU units rather than fewer, enhanced individual chip capabilities. This shift is accompanied by substantial implications for overall material consumption. As AI models continue to grow, the full spectrum of resources required for manufacturing complete GPU units will be extracted, rather than primarily increasing silicon consumption by enlarging individual, more powerful chips.
Performance vs. resource consumption
The AI research community has developed a range of standardized benchmarks to systematically monitor technical progress in AI model capabilities over time39. Despite their inherent limitations, these benchmarks serve as practical tools for evaluating discrete intelligent competencies, such as image classification and multiple-choice question answering, across AI models (see Fig. 3).
Fig. 3: Model performance vs. GPU requirements.
The alternative text for this image may have been generated using AI.
Illustration of the relationship between Falcon, Llama 2, GPT-3.5, and GPT-4 model performance—measured across five standard benchmarks—and the corresponding GPU requirements (on log scale) for model training (author illustration).
Benchmark proposal in the AI field are commonly published as arXiv preprints rather than through traditional peer-reviewed venues. The rapid pace of model development and the need for the timely, open dissemination of evaluation methods render traditional peer-review cycles impractical for this purpose. Accordingly, the benchmark papers cited in this study originate from arXiv preprints.
This analysis examines the relationship between benchmark performance and GPU resource requirements, focusing on the following five widely used evaluation frameworks:
MATH (mathematical reasoning): MATH serves as a widely adopted benchmark for evaluating mathematical problem- solving skills based on 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanations40.
MMLU (multidisciplinary knowledge): the massive multitask language understanding (MMLU) benchmark encompasses 57 tasks, including elementary mathematics, US history, computer science, law, and additional domains41.
HumanEval (programming proficiency): HumanEval measures functional correctness in program synthesis from doc-strings, comprising 164 original programming problems that assess language comprehension, algorithmic thinking, and mathematical reasoning comparable to entry-level software engineering interviews42.
ARC-c (multidisciplinary knowledge): the AI2’s reasoning challenge (ARC-c) dataset presents multiple-choice questions derived from science examinations spanning grades 3–9, with the challenge partition containing complex problems requiring advanced reasoning capabilities43.
HellaSwag (commonsense understanding): HellaSwag tests a model’s commonsense reasoning via natural language inference, focusing on plausible sentence completions in everyday scenarios44.
To better understand the relationship between computational investments and AI performance, we compare models along three dimensions. First, we examine the transition from GPT-3.5 to GPT-4, both developed by OpenAI, to assess the impact of scaling large-scale models within a consistent organizational context. Second, we compare two models trained with nearly identical computational budgets to isolate differences in training efficiency and model performance. Lastly, we contrast the second-smallest with the largest model in our dataset to explore how performance scales at the extremes of model size and resource consumption. To identify the most plausible GPU requirements for each model, this analysis and the resulting Fig. 4 adopt model-specific MFU values and a 2-year hardware lifespan as the baseline scenario. Where MFU has been reported or credibly estimated by an independent technical analysis, these values are used directly: LLama 2 at 53%45, GPT-3.5 at 20%, and GPT-4 at 35%46. For Falcon, where no MFU has been reported, a value of 40% is adopted based on comparable models of similar size and architecture. The resulting GPU counts, therefore, represent the most plausible central estimates for each model under typical training conditions. However, MFU can vary substantially across different training runs depending on hardware infrastructure, parallelization strategy, and software optimization. The subsequent section quantifies how GPU requirements vary across the MFU range through a sensitivity analysis.
Fig. 4: Model performance and resource consumption.
The alternative text for this image may have been generated using AI.
Comparison of model performance between a GPT-3.5 and GPT-4, b GPT-3.5 and LLaMa 2, and c Falcon and GPT-4, across five benchmarks, alongside their respective GPU resource requirement for training (author illustration).
To examine in-house performance improvements over time, we first compare OpenAI’s GPT-3.5 with its successor, GPT-4 (see Figs. 3 and 4a). The transition from GPT-3.5 to GPT-4 illustrates substantial computational scaling accompanied by mixed performance returns. GPT-3.5 and GPT-4 MFU values are based on OpenAI’s reported 19.6% MFU for GPT-346, from which we adopt 20% for GPT-3.5 as a closely related successor, and on the reported 35% MFU for GPT-446. GPT-4 required approximately 31.5 times more GPU resources for training than GPT-3.5 (2515 vs. 80 GPUs), representing a more than 3000% increase in computational resources. While GPT-4 delivered substantial performance improvements in certain domains,
achieving +61.1% over GPT-3.5 on the MATH benchmark and +39.3% on HumanEval, other benchmarks showed only modest gains. Overall, these results suggest diminishing returns in terms of performance relative to computational investment. This raises critical questions about the efficiency and sustainability of current scaling trends and whether performance evaluation benchmarks are already saturated.
The computational demands of training large-scale AI models have increased substantially in recent years. The most straight-forward way to measure this is the number of FLOPs required to train an AI model. In 2020, only 11 models required more than 1023 FLOPs to train47. However, between January 2020 and June 2025, 432 notable AI models entered the market48. Among these, 271 models provided estimates of their training FLOPs, while 160 did not. Notably, 111 models among those with available data, roughly 41%, exceed the 1023 FLOPs threshold during training. This trend reflects a substantial escalation in the computational scale of modern AI model development. While increasing computational resources has enabled notable advancement in some domains, it also underscores the limitations of brute-forcing intelligence.
To further explore the relationship between compute efficiency and model performance, we compare models trained with similar computational budgets above the 1023 FLOPs threshold. This comparison helps isolate factors that contribute to performance increase beyond sheer scaling. Meta’s LLaMa 2 (8.4 × 1023 Flops) and OpenAI’s GPT-3.5 (3.15 × 1023 Flops) utilize nearly identical training resources yet demonstrate contrasting efficiency profiles across hardware utilization and model performance evaluations (see Fig. 4b). LLaMa 2 achieved a higher MFU of 53%45, compared to an estimated 20% MFU for GPT-3.546. This suggests that Meta’s training infrastructure and optimization strategies were more effective in extracting computational throughput from the hardware. However, despite this efficiency advantage and comparable computational budgets, GPT-3.5 consistently outperforms LLaMa 2 across all evaluated benchmarks (see Fig. 4), especially in mathematical reasoning and programming tasks. This contrast illustrates an important distinction: models trained with similar computational resources can achieve substantially different performance outcomes depending on factors such as training data quality, model architectural design choices, and other optimization strategies (e.g., post-training reinforcement learning). GPT-3.5’s stronger performance in mathematical and coding domains, despite lower MFU, may indicate that OpenAI’s training data curation and model architecture were better suited for these specific capabilities, even if their training process was less computationally efficient. The comparison between the second smallest model in this study, Falcon (2.4 × 1023 Flops), and the largest model, GPT-4 (1.73 × 1025 FLOPs), illustrates the relationship between massive resource investment and capability increase (see Fig. 4c).
GPT-4 required approximately 81 times the computational resources of Falcon during training. Again, mathematical reasoning capabilities showed the most dramatic improvement, with GPT-4 achieving more than 7 times the performance of Falcon on the math benchmark. This substantial gain suggests that mathematical reasoning represents a particularly resource-intensive capability that may justify extreme computational investments for specific applications. Conversely, commonsense understanding capabilities, as measured by HellaSwag, demonstrate only a modest improvement of 14%, suggesting that distinct cognitive abilities exhibit differential scaling efficiency in response to increased computational resources.
Resource savings via training efficiency and lifespan improvements
ML engineers, data center operators, and semiconductor manufacturers all play a crucial role in reducing the material footprint of AI training. However, it should be noted that the following analysis focuses exclusively on GPU hardware; broader infrastructure components, such as networking, storage, power delivery, and cooling systems, are not captured in these estimates, and total infrastructure-level material savings would be substantially larger.
Within the scope of GPU hardware, two primary strategies can reduce the number of GPUs required for model training: software-based improvements, such as those that optimize GPU utilization; and hardware-based measures, maximizing the lifespan of GPUs in data centers. Data center design, particularly cooling efficiency, has a substantial influence on GPU wear and longevity. These considerations begin at the semiconductor level, where chip designers work to improve thermal management. Prolonged high-energy workloads, common during AI training, can lead to considerable heat generation on silicon chips, accelerating hardware degradation, and ultimately catastrophic failure49,50.
Increasing the MFU from 20 to 60%, while keeping the GPU lifespan constant, reduces the number of GPUs required for model training by approximately 67%. Similarly, a lifespan expansion from 1 to 3 years, while keeping MFU constant, results in about a 67% reduction. A further lifespan extension to 5 years would yield an estimated 80% reduction. To illustrate the impact of these combined optimizations: training GPT-4 with a relatively low MFU of 20% over a 1-year lifespan requires 8800 GPUs. In contrast, under an optimized scenario, a five-year lifespan and a 60% MFU would require only 587 GPUs. This
represents a potential reduction of about 93% in GPU usage (see Fig. 5). Notably, the decline in GPU requirements with increasing MFU is steeper for shorter lifespans and becomes more gradual for longer lifespans. This trend indicates diminishing returns in GPU savings from MFU improvements once hardware longevity is maximized (see Fig. 5).
Fig. 5: GPU requirements vs. MFU by lifespan.
The alternative text for this image may have been generated using AI.
GPU requirements for training GPT-4 for varying MFU scenarios and varying lifespans (author illustration).