This work complies with all relevant ethics regulations. The Cambridge South Research Ethics Committee (10/H0305/83), The Republic of Ireland Research Ethics Committee (GEN/284/12) and The East of England–Cambridge Central Research Ethics Committee (20/EE/0035) gave ethical approval for this work. ALSPAC was accessed under application number B3565, MCS under GDAC_2022_14 and the UK Biobank under 44165. This study was not pre-registered. No statistical methods were used to pre-determine sample sizes, but our sample sizes are similar to those reported in previous publications24,52.

Cohorts

We report results from the Avon Longitudinal Study of Parents and Children33,68, Millennium Cohort Study34 and the UK Biobank cohorts. Participants in ALSPAC and MCS had their sex assigned at birth noted, whereas in the UK Biobank participants were asked their sex with at least the options of ‘female’ or ‘male’, and in all cohorts participants had their sex genetically inferred via genotype microarray data (male having XY and female having XX chromosomes). Individuals with discordant self-reported and genetically inferred sex were removed.

ALSPAC

In ALSPAC, pregnant women with expected deliveries between 1 April and 31 December 1991 were recruited in the greater Bristol area (formerly Avon county), resulting in an initial sample of 14,541 pregnancies enrolled in the study, of which 13,988 resulted in live births of children surviving to age 1. Data collected after the age of 7 were available for an additional 906 pregnancies from other phases of enrollment, which resulted in an additional 913 children that survived to age 1. At initial enrollment, 14,203 unique mothers were in the study, which increased to 14,833 unique mothers enrolled after the additional phases of enrollment (G0 mothers). The partners of G0 mothers (G0 partners) were also invited to participate in the study, of which 12,113 provided data at one point in the study and 3,807 are currently enrolled. From birth to early adulthood, mother and children were followed up with questionnaires and clinical and psychometric data collection across several timepoints. Biosamples used for genotyping and exome sequencing were obtained from most children and some of the mothers and fathers.

In the current study, data from 6,495 unrelated children with European inferred genetic ancestry (G1 children; described below) were used, along with genetic and/or survey data collected on 4,968 G0 mothers and 4,563 G0 partners. Note that the study website (http://www.bristol.ac.uk/alspac/researchers/our-data/) contains details of all data available through a fully searchable data dictionary and variable search tool. Ethics approval for the study was obtained from the ALSPAC Ethics and Law Committee and the Local Research Ethics Committees. Consent for biological samples was collected in accordance with the Human Tissue Act (2004). Informed consent for the use of data collected via questionnaires and clinics was obtained from participants following the recommendations of the ALSPAC Ethics and Law Committee at the time. At age 18, study children were sent ‘fair processing’ materials describing ALSPAC’s intended use of their health and administrative records, and were given clear means to consent or object via a written form. Data were not extracted for participants who objected, or who were not sent fair processing materials.

MCS

MCS recruited 18,552 pregnant mothers (2000–2002) using a sampling scheme to ensure a nationally representative sample across the UK, as previously described34. Mothers and children were followed longitudinally and genetic data were collected when the children were age 14, including from mothers and fathers where available. Ethics approval for the collection of saliva samples from these individuals as part of the sixth sweep was obtained from London-Central Research Ethics Committee. Genotype and newly generated exome data are available for ~8,000 and ~7,000 children, respectively (~13,000 and ~7,000 parents).

UK Biobank

The UK Biobank is a prospective cohort of over 500,000 individuals sampled throughout the UK between 2006 and 2010. Individuals have extensive phenotype data and genetic data including genotype array69 and whole-exome sequence data70.

Cognitive performance measures across cohorts

In this study, we used various cognitive performance and school performance measures across the three cohorts.

In ALSPAC, children had IQ measured at ages 4 (Wechsler Preschool and Primary Scale of Intelligence), 8 (Wechsler Intelligence Scale for Children) and 16 (Wechsler Abbreviated Scale of Intelligence). Linked educational records are available, including national standardized exam scores in English, Math and Science administered at the end of Key Stage 2 (Year 6 at ~11 years old) and Key Stage 3 (Year 9 at ~14 years old). To create a composite measure of academic performance, we standardized each score and, at a given Key Stage examination point, computed a one-factor model score with the ‘factanal’ function in R using the Bartlett method for scoring, explaining 74% and 77% of the variance at Key Stages 2 and 3 exams, respectively. The factor loadings and factor variance explained for the three exams were nearly identical at Key Stages 2 and 3 (Supplementary Table 1), implying longitudinal invariance.

In MCS, children completed various cognitive performance tests at several ages. Those at later ages had a greatly reduced sample size, so we focused on those at ages 3 to 7. These included Bracken School Readiness at age 3, reading vocabulary at ages 3 and 5, pattern construction at ages 5 and 7, and word reading and progress in math at age 7. These were summarized into a single cognitive performance measure using a one-factor model score with the factanal function in R and Bartlett method for scoring, explaining 39% of the variance.

In the UK Biobank, individuals were asked 13 questions testing verbal–numerical reasoning in a limited time frame during the initial assessment visit. We used the sum of questions answered correctly as a measure of fluid intelligence (data field 20016.0), as has been used previously in genetic studies of cognitive ability71.

Imputation of IQ values in ALSPAC

Imputation of missing IQ values was carried out using SoftImpute37 with rank.max = 4.5 and lambda = 4.5. Before imputation, all variables were standardized to have a mean of 0 and a variance of 1. To assess the imputation accuracy obtained using three potential sets of variables (described below), we set 100 random individuals with measured IQ values at a given age to ‘missing’, conducted phenotype imputation and calculated correlations between the true measured values and the imputed values, repeating this procedure 100 times.

We considered the following sets of variables:

  1. 1.

    Base set: sex (kz021), birth weight (kz030), maternal and paternal socioeconomic groups (b_seg_m and b_seg_p), and IQ values at ages 4 (cf813), 8 (f8ws112) and 16 (fh6280).

  2. 2.

    Expanded set: base set plus development score at age 2 (cf783), sociability score at age 3 (kg623a/c), additional verbal and performance IQ at age 4 (cf811, cf812), non-word repetition and multisyllabic word repetition at age 5 (cf470, cf480), communication score at age 6 (kq517), Skuse social cognition score at age 7 (kr554a/b), cognitive scores at age 8 including Sky Search, Diagnostic Analysis of Nonverbal Accuracy, non-word repetition, verbal and performance IQ, and children communication checklist score (f8at062, f8dv443, f8dv444, 8dv445, f8dv446, f8sl100, f8sl101, f8sl102, f8ws110, f8ws111 and ku506a).

  3. 3.

    Auxiliary set: expanded set excluding base set variables.

When calculating the imputation accuracy using the expanded or auxiliary sets, we removed verbal and performance IQ at each age, although in practice very few children had only one measured at a given age.

For the imputation of the final IQ values, we used the expanded set and required that each individual had non-missing values for at least one IQ test and at least 20% of the variables used for imputation. Final IQ values were standardized to a mean 0 and a standard deviation of 1 separately for measured and imputed values, and the measured and imputed values were then combined into a single variable.

Genotype data preparation and imputationALSPAC

ALSPAC genotype data were generated and processed as described in ref. 68. We further removed individuals with high genotype missingness (>3%). We restricted analyses to autosomal SNPs with MAF > 0.5%, missingness rate <3%, and that had Hardy–Weinberg Equilibrium (HWE) test P > 1 × 10−5.

We used KING72 for relatedness inference and removed 152 samples who had first degree relationships but were reported as coming from different families, and another 16 samples who did match available exome sequencing data supposedly for the same individuals, resulting in 8,831 children, 9,302 mothers and 1,706 fathers. To identify a set of unrelated children, we iteratively identified the child with the greatest number of genetically inferred relatives (3rd degree or closer) and removed them, recalculated the number of inferred relatives per child, and repeated the first step until no children were inferred to have any relatives. This resulted in a total of 6,495 unrelated children with genotype array and whole-exome sequence data.

To identify individuals of genetically inferred European ancestry, we projected samples onto 1,000 Genomes Phase 3 individuals73 using the smartpca function in EIGENSOFT (v.7.2.1)74. We used linkage disequilibrium (LD)-pruned SNPs (pairwise r2 < 0.2 in batches of 50 SNPs with sliding windows of 5) with MAF > 5% and removed 24 regions with high or long-range LD, including the human leucocyte antigen region75. All ALSPAC samples projected onto European ancestry samples. We then performed principal component analysis (PCA) identically on the unrelated ALSPAC samples and projected all individuals onto the PCA space.

Before imputation, we removed palindromic SNPs, SNPs that were not in the imputation reference panel, and SNPs with mismatched alleles. We imputed the samples to the TOPMed r2 reference panel using the TOPMed imputation server76,77,78. We kept well-imputed common variants with Minimac4 R2 > 0.8 and MAF > 1%.

MCS

MCS genotyped data were generated and processed as described in ref. 79. A set of unrelated individuals with genetically inferred European ancestry was identified identically to the procedure used in ALSPAC above. Imputation was also conducted identically as in ALSPAC. We kept well-imputed common variants with Minimac4 R2 > 0.8 and MAF > 1%.

UK Biobank

UK Biobank genotype data were generated, processed and imputed by UKB as described in ref. 69. Sample quality control consisted of excluding individuals with >3% missingness, inconsistent sex, sex aneuploidy or withdrawn consent, and relatedness was similarly calculated using KING as previously described. As in the other cohorts, we kept autosomal SNPs with MAF > 1%, missingness rate <3% and that passed the HWE test (P > 1 × 10−5). To identify UKB individuals with genetically inferred European ancestry, we similarly projected samples onto the 1,000 Genomes Phase 3 individuals and assigned individuals to a genetic ancestry based on their Mahalanobis distance to the nearest continental ancestry group centroid using the top 6 PCs. Those more than 6 standard deviations from the centroid along any axis were removed. Unrelated individuals were identified using the iterative exclusion procedure as previously described, resulting in 387,531 unrelated individuals. PCs were recalculated in this subset of unrelated individuals using smartpca as previously described, and related individuals were projected onto this PCA space.

Calculating polygenic scores

SNP weights for polygenic scores were estimated using LDpred2-auto80, which does not require a tuning dataset and automatically estimates the required hyperparameters from the discovery sample81. The LD reference panel was computed from unrelated individuals from the target dataset for a set of 1,444,196 HapMap3+82 variants69. GWAS summary statistics for EA13 excluding 23andMe or 23andMe only (used for PGIEA UK Biobank samples), EA-Cog and EA-NonCog41 were matched with the list of overlapping SNPs.

Once the weights were generated, we scored individuals using the ‘–score’ function in PLINK v.1.9, which calculates the weighted sum of genotypes across a set of SNPs for each individual.

Mendelian imputation of parental genotypes

When an individual had at least one parent with genotype data in the cohort, we imputed the expected parental genotype of the other parent. To do so, we used the ‘snipar’17 package across the three cohorts. We supplied relatedness inference using the KING output and used the default options throughout. We used the previous weights derived from each GWAS using LDpred2-auto to calculate the PGIs in the full trios using the ‘pgs.py’ script provided in the snipar package.

GWAS of IQ and cognitive performance

We used the linear regression function ‘lm’ in R to conduct GWASs on the IQ measures pre- and post-imputation in ALSPAC and the cognitive performance measure in MCS on unrelated genetically inferred European ancestry individuals, controlling for 10 genetic PCs and sex. We removed variants with MAF < 1% or missingness >2%.

Heritability and genetic correlations

We used both GREML-LDMS38 and LD score regression83,84 to estimate SNP heritabilities and genetic correlations from the summary statistics of the GWASs conducted on the IQ measures and external GWASs using the previously described LD reference panel on HapMap3 SNPs.

Other exposures associated with IQ in ALSPAC and MCS

We considered the following exposures: maternal and paternal educational attainment (EA), weeks born preterm, and a composite variable representing maternal illness in ALSPAC only. For educational attainment, in ALSPAC we used the highest qualification reported by the parent and coded it as previously reported13, with a university degree equalling 20 years of education, A levels as 13, O levels/vocational as 10, and Certificate of Secondary Education/no degree as 5 (c645a and c666a for the mother and father, respectively); in MCS we used the equivalent self-reported derived variables in sweep 1 from each parent questionnaire (APLFTE00). Maternal illness was defined as a binary variable indicating whether or not a mother had at least one of the following conditions reported during the pregnancy in her obstetric clinical records: preeclampsia, anaemia (DELP_1060), any diabetes (pregnancy_diabetes), or genital herpes, gonorrhoea, syphilis, urinary tract infection, vaginal infection, or hepatitis B noted during delivery (DEL_P1050-1055). Weeks born preterm was coded as 40 minus the gestational age at birth (in weeks) assessed in a formal paediatric assessment (DEL_B4401). All variables were standardized to have a mean of zero and a standard deviation of 1.

Exome-sequencing data preparation

Whole-exome sequencing data quality control (QC) for ALSPAC and MCS cohorts was carried out by the Human Genetics Informatics team at the Sanger Institute as described in ref. 42. Briefly, GATK v.4.2 was used to call short variants (single-nucleotide variants (SNVs) and insertions/deletions (indels)) in 11,994 samples from ALSPAC and 15,050 samples from MCS. Sample QC measures were employed to remove outliers on several metrics (for example, heterozygosity, variant counts) and likely sample mismatches. To identify low-quality variants (variant QC), a random forest was trained on pre-defined truth sets in each cohort individually. Random forest filtering was then applied in combination with genotype-level and missingness filters to balance precision, recall, true and false positive rates, and synonymous transmission ratios. Specifically, SNVs were filtered (genotypes set to missing) if they had an allele depth (DP) < 5, a heterozygous allele balance ratio (AB) < 0.2, or a genotype quality (GQ) < 20 (ALSPAC) or <15 (MCS); indels were filtered using these thresholds: DP < 10, AB < 0.3, GQ < 10 (ALSPAC) or GQ < 20 (MCS). The variants were excluded if they failed the random forest filtering or if the fraction of missing genotypes (missingness) exceeded 0.5. The final dataset included 8,436 children and 3,215 parents in ALSPAC, and 7,667 children and 6,925 parents in MCS. Calling and QC of de novo mutations is described in the Supplementary Methods.

In the UK Biobank, we performed quality control for whole-exome sequencing data on 469,836 participants within the UKB research analysis platform. First, we split and left-aligned multi-allelic variants in the population-level variant call format files into separate alleles using ‘bcftools norm’85. Next, we performed genotype-level filtering using bcftools to separately filter for SNVs and indels using a missingness-based approach. Specifically, SNV genotypes with a depth lower than 7 and genotype quality lower than 20, or indel genotypes with a depth lower than 10 and genotype quality lower than 20, were set to missing. We further tested for an expected alternate allele contribution of 50% for heterozygous SNVs using a binomial test, and SNV genotypes with a binomial test P ≤ 0.0001 were set to missing. Finally, we recalculated the proportion of individuals with a missing genotype for each variant and excluded all variants with a missingness value greater than 50%.

Classification of deleterious rare variants

All variants were annotated using the MANE transcript from Ensembl86. In MCS and ALSPAC, rare pLoFs were defined as pLoFs annotated as high confidence by LOFTEE31 that had a combined annotation dependent depletion (CADD) score > 25 (ref. 87) (if SNVs), were not located in the last exon or intron, and had a gnom-AD31 V3 allele frequency of <3 × 10−5 (up to ~5 occurrences in gnomAD r3 genomes & ~10× in exomes), and an in-sample allele frequency of <0.1% among the unrelated set of children from that cohort86. Rare damaging missense variants were defined using identical allele frequency and CADD filters to pLoFs but were additionally required to have a missense badness, PolyPhen-2 and constraint (MPC) score88 ≥2. Rare synonymous variants were defined with identical allele frequency thresholds.

In UKB individuals with inferred European genetic ancestry, we defined rare variants as those with a within-sample allele frequency <0.001%. For pLoFs, we retained only those variants defined as high-confidence pLoFs by LOFTEE and had CADD > 25 as in MCS and ALSPAC. For missense variants, we defined the damaging missense variants by including variants with REVEL >0.5, AlphaMissense89 >0.56 and MPC >2.

Calculating rare variant burden

We calculated RVB using the following formula:

$${\mathrm{RVB}}_{i,c}\,=\,\mathop{\sum }\limits_{g\in G}{I}_{g,c}(i){S}_{g}$$

(1)

where \({I}_{g,c}(i)\) is an indicator function for whether an individual i has a variant of consequence c in gene g, \({S}_{g}\) is the fitness cost for heterozygous carriers of a pLoF allele in gene i estimated in ref. 44, and G is the set of all autosomal genes.

For analyses to ascertain the relative contribution of de novo versus inherited variants (Supplementary Note 4), we also calculated a separate rare variant burden metric which we call ‘constrained variant count’:

$${\mathrm{constrained}\,\mathrm{pLoF}\,\mathrm{count}}_{i}=\mathop{\sum }\limits_{g\in \mathrm{constrained}\,\mathrm{genes}}{I}_{g,\mathrm{pLoF}}(i)$$

(2)

where \({I}_{g,\mathrm{pLoF}}(i)\) is an indicator function for whether an individual i has a pLoF in gene g, and ‘constrained genes’ is the set of genes identified as constrained in ref. 44. For Extended Data Fig. 4, we calculated each RVB in three mutually exclusive gene sets from Li et al.90, as well as in genes prioritized from various GWASs6,91,92.

Sampling and non-response weights in MCS

Sampling weights were developed by MCS to adjust for the non-random sampling scheme devised for the study93. We used the full UK sampling weights. Non-response inverse probability weights were developed for each sample as done in ref. 23. Specifically, we fitted a logistic regression model predicting whether an MCS child had available genotype data, using the following variables collected at the first study sweep (where missingness was minimal, with over 96% of individuals having complete data): housing tenure, parental years of education, language spoken at home, study sampling stratum, single parent status and breastfeeding status. The study stratum variable comprised nine categories defined by country within the UK (England, Wales, Scotland, Northern Ireland), each stratified by area-level advantage or disadvantage, with an additional ethnic minority stratum in England. We fitted this model separately for two target samples: first, to predict membership in the set of unrelated individuals with genetically inferred European ancestry and genotype data (N = 5,884 of 6,036 children with complete covariate data); and second, to predict membership in the subset of these children who additionally had genotyped parents forming complete trios (N = 2,445 of 2,498 trio children with complete covariate data). In the trio model, an additional covariate indicating whether the child had any genotype data was included. These models achieved Nagelkerke R2 values of 0.51 and 0.42, respectively. For each genotyped individual, we extracted the predicted probability of being in the target sample and took its inverse as the non-response weight, such that individuals with lower predicted probabilities of inclusion who were nonetheless present in the sample received higher weights. Individuals whose weights exceeded three standard deviations above the mean were excluded as likely reflecting data errors driving implausibly low predicted probabilities; this removed 14 individuals from the full genotyped sample and 3 from the trio subset. Final combined weights were obtained by multiplying the MCS-provided sampling weights by our non-response weights, and these were incorporated into all relevant regression analyses using the weights argument in R.

Associations between genetic measures and traits

Cross-sectional associations for a genetic score and IQ/cognitive performance measure were conducted using the lm function in R, restricting to unrelated samples with genetically inferred European ancestry as follows:

$${{\rm{IQ}}}_{i}\,\sim \,{G}_{i}\,+\,{\rm{PC}}{1}_{i}\,+\,\ldots \,+{\rm{PC}}{10}_{i}\,+\,{{\rm{sex}}}_{i}$$

(3)

where Gi is the genetic score (PGI or RVB) for child i and PC1–10i are their genetic PCs.

Genetic scores, unless otherwise stated, were standardized to a mean of 0 and a standard deviation of 1. In trio-based analyses, the parental genetic values were similarly standardized and regression was conducted as follows:

$${{\rm{IQ}}}_{i}\sim {G}_{i}+{G}_{i,{\rm{m}}}+{G}_{i,{\rm{p}}}+{\rm{PC}}{1}_{i}+\ldots +{\rm{PC}}{10}_{i}+{{\rm{sex}}}_{i}$$

(4)

where Gi,m and Gi,p are the genetic scores for child i’s mother and father, respectively.

Mixed-effects linear models were conducted using the ‘lme4’ package in R94. For our primary models, in addition to Wald-test P values, we obtained heteroscedasticity-robust inference by performing a parametric bootstrap of the fixed-effect estimates (1,000 replicates) via bootMer (type = ‘parametric’) in lme4. Bootstrap standard errors and P values are reported in Supplementary Tables 4 and 6. The same covariates were used for the cross-sectional analysis with the addition of age, an age × genetic score interaction effect and a child-specific random intercept term as follows:

$${{\rm{IQ}}}_{i,t}\sim {{\rm{age}}}_{i,t}+{G}_{i}+{{\rm{age}}}_{i,t}\ast G{\rm{i}}+{\rm{PC}}{1}_{i}+\ldots +{\rm{PC}}{10}_{i}+{{\rm{sex}}}_{i}+{C}_{i}$$

(5)

where agei,t and IQi,t are child i’s age and IQ at age t, and Ci is a child-specific random effect intercept. Age was coded as age-4 such that the genetic predictor intercept estimate was equivalent to the effect at the first age at which IQ data were collected. In the trio-based analyses, an age interaction effect was additionally modelled for the parental genetic values as follows:

$$\begin{array}{l}{\mathrm{IQ}}_{{\rm{i}},{\rm{t}}}\sim {\mathrm{age}}_{{\rm{i}},{\rm{t}}}+{{\rm{G}}}_{{\rm{i}}}+{\mathrm{age}}_{{\rm{i}},{\rm{t}}}\ast {{\rm{G}}}_{{\rm{i}}}+{{\rm{G}}}_{{\rm{i}},{\rm{m}}}+{\mathrm{age}}_{{\rm{i}},{\rm{t}}}\ast {{\rm{G}}}_{{\rm{i}},{\rm{m}}}\\ \,\,\,\,\,\,\,\,+{{\rm{G}}}_{{\rm{i}},{\rm{p}}}+{\mathrm{age}}_{{\rm{i}},{\rm{t}}}\ast {{\rm{G}}}_{{\rm{i}},{\rm{p}}}+\mathrm{PC}{1}_{{\rm{i}}}+\ldots +\mathrm{PC}{10}_{{\rm{i}}}+{\mathrm{sex}}_{{\rm{i}}}+{{\rm{C}}}_{{\rm{i}}}\end{array}$$

(6)

For cross-sectional and longitudinal quantile regression modelling, we used the R package ‘quantreg’95 and ‘rqpd’96, respectively, using the same models as above. Mixed-effects model standard errors were determined using bootstrapping with 500 bootstrap replications.

For the academic performance score, the cross-sectional models were performed in the same way as described above, with the addition of age at testing (in weeks) as a covariate (ks2age_w and ks3age_w). As there were only two timepoints, a slightly different longitudinal model was used for the linear modelling (Extended Data Fig. 2) as follows:

$$({{\rm{APS}}}_{9,i}-{{\rm{APS}}}_{6,i})\sim {G}_{i}+{{\rm{age}}}_{6,i}+{{\rm{age}}}_{9,i}+{\rm{PC}}{1}_{i}+\ldots +{\rm{PC}}{10}_{i}+se{x}_{i}$$

(7)

where APS6/9,i and age6/9,i are the academic performance scores and ages of individual i at Years 6 and 9, respectively. The coefficient estimated for Gi in this regression is equivalent to G’s effect at Year 9 minus that at Year 6 within an individual, that is, the age interaction effect. To compare the effects for the quantile regression effect size estimates at Years 6 and 9, we conducted a Z-test between the effect sizes (Extended Data Fig. 7).

For the linear models, we assumed normally distributed residuals and additionally, for mixed-effects linear models, random effects. We assessed the potential impact of heteroscedasticity on our inference by re-estimating key models with bootstrap-derived standard errors (Supplementary Tables 4 and 6), which produced virtually identical results. Quantile regression does not assume normally distributed errors and is robust to heteroscedasticity.

All statistical tests were two-tailed, except for the F-test of variance differences, which was one-tailed by construction. Assumptions for statistical tests were not explicitly tested. This study was not pre-registered.

Marginal and conditional associations between genetic and other factors and IQ

To assess the associations of genetic and other factors with IQ (Fig. 4), we considered both marginal and conditional models. For marginal models, we conducted the following regressions:

$${{\rm{IQ}}}_{i,t}\sim {{\rm{age}}}_{i,t}+{E}_{i}+{{\rm{age}}}_{i,t}\ast {E}_{i}+{\rm{PC}}{1}_{i}+\ldots +{\rm{PC}}{10}_{i}+{{\rm{sex}}}_{i}+{C}_{i}$$

(8)

where Ei is the given measure of interest. In Fig. 4, the coefficient on the Ei term (that is, the main effect) is shown in the top panel, and the coefficient on the agei,t × Ei term (that is, the interaction effect) in the bottom panel.

As maternal and paternal EA are highly correlated, we modified the model for the effects of parental EA as follows:

$$\begin{array}{l}{\mathrm{IQ}}_{{\rm{i}},{\rm{t}}}\sim {\mathrm{age}}_{{\rm{i}},{\rm{t}}}+{\mathrm{EA}}_{{\rm{i}},{\rm{m}}}+{\mathrm{EA}}_{{\rm{i}},{\rm{p}}}+{\mathrm{age}}_{{\rm{i}},{\rm{t}}}\ast {\mathrm{EA}}_{{\rm{i}},{\rm{m}}}\\ \,\,\,\,\,\,\,\,+{\mathrm{age}}_{{\rm{i}},{\rm{t}}}\ast {\mathrm{EA}}_{{\rm{i}},{\rm{p}}}+\mathrm{PC}{1}_{{\rm{i}}}+\ldots +\mathrm{PC}{10}_{{\rm{i}}}+{\mathrm{sex}}_{{\rm{i}}}+{{\rm{C}}}_{{\rm{i}}}\end{array}$$

(9)

where EAi,m and EAi,p are the maternal and paternal EA, respectively. Similarly, we fit the PGIEA values for the child, mother and father jointly in a trio model.

To calculate the variance explained by a given variable, we took the square of the standardized effect size of that variable’s main effect. To calculate the total variance explained by two uncorrelated variables (for example, RVBpLoF and RVBMissense that have a correlation of −0.01, P = 0.16), we summed the squares of their standardized effect sizes. To test for differences in effect size between the combination of RVBMissense and RVBpLoF versus other factors, we used the square root of the previous variance explained estimate as the effect size estimate and determined the standard error for that effect size by summing the square of the standard error for each effect estimate and then taking the square root of the new estimate. We were then able to use a Z-test to compare the effect size estimates as done previously.

For the full joint models, we included all genetic and other measures in a single model, with a main and age interaction effect for each measure. Since only a subset of probands had information on gestational age, we fitted the full joint model both with and without the ‘weeks born preterm’ variable.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.