Following recent innovations in the field of environmentally adjusted sovereign credit risk analysis31,56, a random forest classification model is trained on credit ratings data describing ratings issued by S&P and macroeconomic fundamentals to predict sovereign ratings. We then integrate forward-looking macroeconomic paths derived from GTAP-InVEST via a sovereign-level aggregation and indicator-adjustment layer, linking scenario outputs to the six ratings determinants in our prediction model (Supplementary Fig. 1). This implementation entails propagating nature-adjusted GDP (from GTAP-InVEST) into fiscal/external indicators via practitioner-based mappings, and embedding the resulting nature-adjusted GDP and nature-adjusted fiscal indicators into a ratings–PD–costs of borrowing pipeline.

We use the GTAP-InVEST earth–economy framework to link ecosystem service dynamics with macroeconomic outcomes. The model constitutes an original methodological contribution that integrates the GTAP computable general-equilibrium model of the world economy with the InVEST spatially explicit biophysical models. GTAP provides a consistent representation of global production, consumption and trade across 37 aggregated regions and 17 sectors, subdivided into 18 agro-ecological zones (341 AEZ-region units). InVEST is executed globally at this resolution to simulate changes in the provision of three ecosystem services—crop pollination, timber provision and marine fisheries—under alternative land-use and policy conditions. Timber productivity is represented through changes in forest biomass, crop pollination as a reduction in the ability of wild pollinators to pollinate crops, and marine fisheries as a reduction in the total catch biomass of marine fisheries.

Biophysical outputs from InVEST are converted into percentage productivity shocks for the corresponding GTAP sectors (agriculture, forestry and fisheries). These shocks are implemented as factor-neutral changes in total-factor productivity within the GTAP model, allowing ecosystem-service changes to propagate through production, trade and income linkages. The resulting general-equilibrium adjustments yield country-level measures of GDP, welfare and related macroeconomic indicators.

The most up-to-date GTAP-InVEST estimates impose two restrictions on our modelling: (1) reliance on three ecosystem services that can be modelled endogenously, and (2) the sample of countries for which the GTAP-InVEST framework provides data that can be reliably disaggregated to the national level, in our case 23 sovereign states.

While GTAP-InVEST constitutes a step change for integrated ecological–economic modelling, a range of assumptions and limitations remain. These arise from the nature of CGE modelling in general, the specific GTAP CGE model itself, the InVEST model, and issues encountered when linking models across multiple spatial scales. First, in its current form, GTAP-InVEST can only be used for comparative static analyses: it can compare a baseline equilibrium against future equilibria but is not a dynamic model that can endogenize environmental–economic–fiscal feedback loops over time. Second, the three ecosystem services included here are only a limited subset of the full spectrum of economically valuable services generated by nature. Third, and relatedly, many of the benefits generated by nature accrue to humans outside of formal markets and would not be captured in this analysis. Fourth, at present, the GTAP-InVEST model does not include a coupled, integrated financial sector. Lastly, CGE models entail a range of assumptions for key parameter values, including total factor productivity growth, substitution elasticities, representative household behaviour (in this case, Cobb–Douglas utility functions) and production frameworks (in this case, nested constant elasticities of substitution production functions). Accordingly, our results should be considered lower-bound estimates and interpreted as scenario simulations rather than predictions.

Training the ratings model on historical data

Our parsimonious model uses six macroeconomic variables to predict sovereign credit ratings: GDP per capita, GDP growth, net general government debt/GDP, general government balance/GDP, narrow net external debt/CAR, and current account balance/GDP. We choose this six-indicator set because it aligns with the core economic (income level and growth), fiscal (debt stock and fiscal balance) and external (external leverage and external balance) pillars used in sovereign credit assessment, offers consistent cross-country coverage, and can be systematically adjusted under our biodiversity scenarios. GDP per capita and GDP growth are taken directly from the GTAP-InVEST outputs for each scenario. The four government-performance and external indicators are then updated using practitioner’s evidence57 on how GDP shortfalls relate to movements in these variables. Concretely, we fit low-order polynomial functions to the published relationships between disaster-induced GDP losses and each indicator, and we apply the nature-adjusted GDP outcomes to generate internally consistent, scenario-specific series by country. This design preserves predictive performance while avoiding ‘anchoring’ on determinants that cannot be adjusted on the basis of scientific evidence, and it keeps the scenario mechanics close to observable ratings practice.

We trained the algorithm using data on 113 countries from 2015 to 2020. Typically, in machine learning exercises it is best to include the largest possible sample. Despite longer data availability for most countries, we intentionally avoid the 2008–2009 global/Eurozone crises and the 2021–2023 pandemic years to reduce noise and maintain the highest possible predictive accuracy. We tested various periods, including 2004–2020 and 2012–2020, which are reported in Supplementary Section 3 with an extended time series. This process enables us to build statistical parameters within which we can begin to make predictions around how changes in our variables will influence credit ratings.

Random forest model

Random forest models are classification algorithms. In our case, we classify sovereigns into investment-grade versus speculative-grade as a starting point, using six explanatory features. The model chooses the split that most reduces classification error—for example, if a GDP-per-capita threshold neatly separates investment-grade from speculative-grade issuers, that variable becomes the first split (the root node; Supplementary Fig. 2).

After the initial split, observations pass to the next node where the algorithm again selects the variable that most effectively reduces error. Iterating this yields clear boundaries between different levels of credit quality. Our implementation expands the simplified example to 20 rating classes and aggregates results over 2,000 trees. The estimated objects are the thresholds/boundaries separating categories. This approach offers two advantages: (1) each tree is trained on a perturbed version of the baseline data, producing largely uncorrelated predictions that we average; and (2) it mitigates overfitting, improving out-of-sample accuracy. Each split considers only a random subset of features, and the forest prediction is the average across all trees, which is typically more stable and reliable than a single model.

As illustrated in Supplementary Fig. 3 (left hand-side box), higher ln(GDP per capita) pushes ratings over successive thresholds, with pronounced nonlinearities: at very low income levels, further declines can leave predicted ratings unchanged, indicating other variables dominate in that range. GDP growth has strong effects over a narrower band: rising growth helps to move ratings into investment grade, but the effect fades beyond a point—potentially because very high growth may reflect post-shock rebounds where lingering vulnerabilities matter more. The remaining panels show analogous relationships for fiscal and external indicators.

Variable importance

The importance of each variable for predicting ratings can be assessed by observing the percentage drop in R2 when each variable is permuted at random (Supplementary Fig. 4). This test indicates how much predictive accuracy would be affected if an individual variable from the set were replaced with a random value. While the ranking is illustrative and may vary by country, ln(GDP per capita) emerges as the dominant driver on average, followed by debt, fiscal balance and GDP growth.

To unpack country-level mechanics, Supplementary Fig. 5 decomposes predicted ratings for a range of economies by sequentially adding variables. Starting from the sample average predicted rating (11.155), incorporating additional information into the model ln(GDP per capita) lifts Canada by +5.428 notches (reflecting high income) but lowers China (given lower per-capita income). Proceeding through the remaining variables yields Canada’s final prediction of 19.672. This figure reinforces that contributions differ by country, underscoring heterogeneous variable importance.

Benefits of random forest models

Random forests offer several benefits over econometric approaches for predicting sovereign ratings. First, we implement the above-described process thousands of times with slightly modified versions of the original dataset, each making use of a varied pool of the original six variables. This means that our prediction model will perform much better when presented with new data, adding precision to our estimates that no parametric approach such as regression can offer58,59.

Second, this approach enables us to model nonlinearities with greater ease60. Ratings data are peculiar, as they are discrete in nature (alphabetical ratings are translated into numerical scale such as the one we are using AAA = 20, AA+ = 19, SD = 1, with AAA representing the highest creditworthiness and SD the lowest). Incremental shifts through the rating scale do not represent equally meaningful changes in creditworthiness.

For instance, a 1-notch change at the top of the rating scale (for example, AAA to AA+) entails only a negligible increase in risk and default probability, whereas the same 1-notch change at the lower end of the scale (for example BB− to B+) can have substantial impacts on the cost and access to debt finance. Machine learning ultimately captures these dynamics. Moreover, ratings data exhibit distributional properties that can lead to bias in econometric modelling. There are often more observations at the top end of the ratings scale than across the rest of the rating categories. These features make linear modelling of credit ratings difficult and can lead to error. Finally, ratings are not purely quantitative assessments and reflect committee judgement and other qualitative considerations. Non-parametric methods help capture complex and nonlinear relationships between fundamentals and ratings more flexibly than traditional (parametric) approaches61. Therefore, using a methodology that can handle distributional properties, nonlinearities and qualitative components is essential.

Adjusting ratings factors for environmental risk

Our ratings prediction model includes input variables that need to be adjusted to reflect natural capital and ecosystem service decline56. Shortfalls in GDP per capita and GDP growth for each scenario can be taken directly from GTAP-InVEST2. To adjust the remaining variables (net general government debt/GDP, general government balance/GDP, narrow net external debt/CAR, and current account balance/GDP), we exploit industry research conducted by S&P. S&P credit analysts were asked to assess how natural disasters such as earthquakes, tropical cyclones, floods and winter storms would affect GDP losses and government performance indicators62. Supplementary Fig. 8 shows data relating per capita GDP loss with analysts’ assessment of the impact on (the log of) net general debt to GDP. To produce the function that best describes this relationship, we fit polynomials of increasing order until no further improvement in fit could be found. We use this fitted model to relate GDP losses under each GTAP-InVEST scenario to each government performance variable.

One exception to this procedure is the production of results for Madagascar. S&P did not issue a rating for Madagascar during our 2015–2020 sample period. However, the country is included in the GTAP-InVEST results.

To avoid losing Madagascar from the sample, we used GDP per capita and growth data from the World Bank to estimate government performance variables based on the implied associations in the rest of the historical data. These were in turn used to estimate a starting point credit rating using the same random forest model described above. Incorporating the GTAP-InVEST results—as for any other country in the sample—produces biodiversity-adjusted credit ratings for Madagascar. For context, our simulated rating was 5.6, part way between B− (5) and B (6) (Supplementary Fig. 9). In April 2022, S&P issued Madagascar a B− rating. This exercise has no impact on the rest of the results and could be considered inconsequential, as Johnson et al.2 estimate losses exceeding 100% of Madagascar’s economy.

Treatment of uncertainty

Our workflow links biodiversity loss to sovereign credit risk through a series of steps. Ecosystem service shocks generate scenario-conditioned macropaths for GDP levels and growth. These paths are then used to update fiscal and external indicators using practitioner-based mappings. The six adjusted inputs are finally translated into ratings and, from there, into PD and borrowing costs. Each stage introduces uncertainty (biophysical parameters, economic elasticities, sovereign-level harmonization and statistical prediction).

We report point estimates by scenario to match market practice in sovereign credit assessment. Credit ratings are single alphanumeric opinions on future default risk. Committees and market users usually communicate central judgements rather than full probability distributions. Practitioners such as S&P’s57,62 estimate how extreme weather events would affect sovereign creditworthiness, expressed as a change in rating notches relative to a base case rather than as a full distribution of outcomes. Our presentation mirrors this practice, while acknowledging that parameter and model uncertainty is present at several stages.

Although sampling over the full uncertainty space would be ideal, the principles guiding this study emphasize staying close to how rating agencies handle uncertainty in practice. Our team includes a former Global Sovereign Chief Rating Officer at S&P who helped build the sovereign rating methodology and sat on hundreds of committees, including those that issued ratings for countries examined here. This experience informs our choice of presentation.

We address uncertainty in three ways. First, we use clearly defined scenarios that function as stress tests. These bracket plausible outcomes under large ecological shocks rather than assigning explicit probabilities. Second, when mapping nature-induced GDP shortfalls to fiscal and external indicators, we fit low-order polynomial functions to practitioner evidence. Third, in the ratings step, we report out-of-sample diagnostics such as accuracy within ±1 or ±2 notches, and we rely on the internal resampling of the random forest procedure to gauge predictive dispersion without imposing a parametric error structure.

Sampling over the full uncertainty space remains an important issue and is a priority for future work. It is outside the scope of this study.

Calculating default probabilities and additional cost of debt

Once we obtain the biodiversity-adjusted credit ratings, we can translate them into additional costs of borrowing of sovereigns and corporates. We first converted S&P’s sovereign ratings to a 20-notch scale (Supplementary Fig. 9).

To calculate the additional cost of borrowing associated with nature-driven sovereign downgrades, we estimate the additional premium over the risk-free rate that would be associated with a sovereign at its baseline rating. With the newly estimated, nature-driven sovereign credit rating, we re-establish the additional premium. The difference between these two values is the additional cost of debt. For example, China, which has a baseline estimated rating of 15, would have a spread of 1.05%. With a nature-driven rating of 9.5, the new spread would be 2.81. The increase in the cost of borrowing is therefore 1.76% (2.81 − 1.05 = 1.76). The cost of borrowing figures for the rating categories are established by taking the median option-adjusted spread for each rating category and interpolating as per Supplementary Fig. 10.

A similar process is conducted using data from S&P Global to estimate the increase in the 10-year PD. Supplementary Fig. 10 reveals the model used to estimate the spread increase in the cost of debt (Supplementary Fig. 10a) and the increase in the PD, respectively (Supplementary Fig. 10b). Our model for estimating the increase in cost of debt and the PD is given by the parameters from a cubic and quintic regression, respectively.

Calculating additional debt burden from nature loss for the median citizen

Figure 3 (right) was calculated by obtaining the median income or consumption from Our World in Data63. For each country, we take the most recent year observation to calculate the per capita increase in the cost of debt (associated with the nature-driven downgrade) using the population data from the corresponding year. We then divide this number by the median income or consumption (multiplied by 365, as the data are daily) for the same year to obtain the proportion of the nature-driven borrowing cost that represents the median person’s annual income.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.