{"id":71923,"date":"2026-08-10T17:04:26","date_gmt":"2026-08-10T17:04:26","guid":{"rendered":"https:\/\/www.europesays.com\/japan\/71923\/"},"modified":"2026-08-10T17:04:26","modified_gmt":"2026-08-10T17:04:26","slug":"prevalence-and-chronology-of-colibactin-associated-mutational-processes-and-their-microbiome-spectra-in-japanese-colorectal-cancer","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/japan\/71923\/","title":{"rendered":"Prevalence and chronology of colibactin-associated mutational processes and their microbiome spectra in Japanese colorectal cancer"},"content":{"rendered":"<p>Patients and samples<\/p>\n<p>This study was approved by the Research Ethics Committees of the National Cancer Center, the University of Osaka and Institute of Science Tokyo as meeting the ethical guidelines for medical and health research involving human individuals (National Cancer Center Institutional Review Board, 2013-244; Research Ethics Committee, the University of Osaka, 20064-2: Human Subjects Research Ethics Review Committee, Institute of Science Tokyo, 2014018). Written informed consent was obtained from all participants before enrollment. All data were de-identified before analysis, and a double-anonymization procedure was applied to all participant identifiers.<\/p>\n<p>This study enrolled a total of 600 individuals, comprising 454 patients with histologically confirmed CRC and 146 HCs, who were recruited at the National Cancer Center Hospital (Tokyo, Japan) and the University of Osaka Hospital (Osaka, Japan) between 2014 and 2025. Among these, cohort 1 consisted of cases prospectively collected specifically for the present study, whereas cohort 2 comprised cases previously described in our published work<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Yachida, S. et al. Metagenomic and metabolomic analyses reveal distinct stage-specific phenotypes of the gut microbiota in colorectal cancer. Nat. Med. 25, 968&#x2013;976 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR8\" id=\"ref-link-section-d4054757e2915\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a> (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">1a<\/a>).<\/p>\n<p>These studies were designed to be observational and included participants for whom the biological specimens required for comprehensive analyses\u2014such as stool samples, tumor tissues and matched normal DNA\u2014were available between 2014 and 2025. Individuals with hereditary CRC syndromes, including familial adenomatous polyposis and hereditary non-polyposis CRC, or with inflammatory bowel disease, were excluded from the analysis.<\/p>\n<p>WGS data were obtained from primary CRC tissues in 200 cases from cohort 1. WGMS data derived from stool samples were available for a total of 395 patients with CRC and 146 HCs across both cohorts. Detailed information on participants\u2019 lifestyle and dietary habits was collected using a comprehensive self-administered questionnaire comprising 475 items across 25 pages. The questionnaire was developed based on the framework of the Japan Public Health Center-based Next Generation Study (JPHC-NEXT)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 46\" title=\"Tsugane, S. &amp; Sawada, N. The JPHC study: design and some findings on the typical Japanese diet. Jpn J. Clin. Oncol. 44, 777&#x2013;782 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR46\" id=\"ref-link-section-d4054757e2928\" rel=\"nofollow noopener\" target=\"_blank\">46<\/a>.<\/p>\n<p>Stool sample collection and processing<\/p>\n<p>Individuals classified as HC were those who presented with a positive fecal occult blood test during a hospital visit and subsequently underwent colonoscopy for secondary screening, with no clinically significant findings. Stool samples were obtained from HCs and CRC patients before the initiation of any treatment, including surgical resection or chemotherapy. The first stool passed at the hospital on the day of colonoscopy was collected for analysis<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Nishimoto, Y. et al. High stability of faecal microbiome composition in guanidine thiocyanate solution at room temperature and robustness during colonoscopy. Gut 65, 1574&#x2013;1575 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR47\" id=\"ref-link-section-d4054757e2940\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 48\" title=\"Erawijantari, P. P. et al. Influence of gastrectomy for gastric cancer treatment on faecal microbiome and metabolome profiles. Gut 69, 1404&#x2013;1415 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR48\" id=\"ref-link-section-d4054757e2943\" rel=\"nofollow noopener\" target=\"_blank\">48<\/a>. In most cases, the stool samples collected on the day of colonoscopy were solid. Consistent with our previous observations<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Nishimoto, Y. et al. High stability of faecal microbiome composition in guanidine thiocyanate solution at room temperature and robustness during colonoscopy. Gut 65, 1574&#x2013;1575 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR47\" id=\"ref-link-section-d4054757e2947\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a>, the microbial taxonomic composition showed a high degree of concordance with that of standard frozen stool samples obtained without bowel preparation (Pearson\u2019s correlation coefficient\u2009=\u20090.91, P\u2009&lt;\u20090.01). All stool samples were immediately frozen on dry ice and subsequently stored at \u221280\u2009\u00b0C until DNA extraction. Participants were instructed to consume a low-residue diet on the day preceding colonoscopy. On the day of examination, all participants received a bowel-cleansing agent followed by colonoscopic evaluation.<\/p>\n<p>Tumor and matched normal sample collection<\/p>\n<p>Tumor tissues were obtained from surgically resected specimens of patients with CRC. Matched non-tumor DNA was primarily extracted from peripheral blood lymphocytes; in selected cases, non-neoplastic colonic mucosa collected from regions anatomically distant from the primary tumor was also used as a normal control.<\/p>\n<p>Clinical endpoint definitions<\/p>\n<p>Overall survival was defined as the interval from the initiation of treatment to death from any cause, with censoring at the date of last confirmed follow-up for surviving patients (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>). PFS was defined as the interval between the initiation of treatment and the first documented evidence of disease progression, as determined by serial radiological examinations (for example, computed tomography or magnetic resonance imaging) or by clinical confirmation of tumor recurrence or metastasis.<\/p>\n<p>Microbial subtype classification and reproducibility assessment<\/p>\n<p>For the combined test dataset (cohorts 1\u2009+\u20092), comprising 199 CRC cases from our previous publication<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Yachida, S. et al. Metagenomic and metabolomic analyses reveal distinct stage-specific phenotypes of the gut microbiota in colorectal cancer. Nat. Med. 25, 968&#x2013;976 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR8\" id=\"ref-link-section-d4054757e2982\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a> and 196 CRC cases newly collected from the National Cancer Center Hospital and the University of Osaka Hospital, we applied a previously published microbial subtype classification algorithm<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"Rynazal, R. et al. Leveraging explainable AI for gut microbiome-based colorectal cancer classification. Genome Biol. 24, 21 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR16\" id=\"ref-link-section-d4054757e2986\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>. Cases were categorized into four microbial subtypes, and their bacterial community compositions were compared to assess reproducibility.<\/p>\n<p>Baseline clinicopathological characteristics were summarized across subtypes, and overall survival was analyzed using the Kaplan\u2013Meier method with comparisons performed with the log-rank test. These univariate comparisons assessed differences in survival that may reflect both subtype-related and background clinical variations. To determine whether microbial subtypes independently influence prognosis, we further applied multivariate Cox proportional hazards models adjusted for established prognostic covariates.<\/p>\n<p>WGMS of fecal samples<\/p>\n<p>Genomic DNA was extracted from frozen stool samples using a bead-beating protocol<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 49\" title=\"Furet, J. P. et al. Comparative assessment of human and farm animal faecal microbiota using real-time quantitative PCR. FEMS Microbiol. Ecol. 68, 351&#x2013;362 (2009).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR49\" id=\"ref-link-section-d4054757e3001\" rel=\"nofollow noopener\" target=\"_blank\">49<\/a> in conjunction with the GNOME DNA Isolation Kit (MP Biomedicals). DNA integrity and concentration were evaluated using the Agilent 4200 TapeStation system (Agilent Technologies). Following ethanol precipitation, the purified DNA was resuspended in TE buffer and stored at \u221280\u2009\u00b0C until library preparation. Sequencing libraries were constructed from fecal DNA using the Nextera XT DNA Sample Prep Kit (Illumina) according to the manufacturer\u2019s instructions. WGMS was performed on the NovaSeq 6000 platform (Illumina) with paired-end (PE) reads (2\u2009\u00d7\u2009150\u2009bp), targeting an average sequencing depth of 5.0\u2009Gb per sample.<\/p>\n<p>Metagenomic sequencing quality control and taxonomic profiling<\/p>\n<p>In total, 9,875,823,911 PE reads (51,843,293 on average) from cohort 1 and 10,156,584,408 PE reads (51,038,112 on average) from cohort 2 were generated from 150-bp sequencing and subjected to stringent quality control. Raw reads containing undetermined bases (\u2018N\u2019) were removed. Reads containing bacteriophage phiX sequences were identified and filtered by alignment to the reference genome using Bowtie2 (v.2.2.9) with the \u2018\u2013fast-local\u2019 preset. Adaptor and primer sequences were trimmed using Cutadapt (v.1.9.1) with the following parameters: for forward reads, -a CTGTCTCTTATACACATCTCCGAGCCCACGAGAC -O 33 -q 17; for reverse reads, -a CTGTCTCTTATACACATCTGACGCTGCCGACGA -O 32 -q 17. Within Cutadapt, reads with consecutive quality scores of \u226417 were trimmed at the 3\u2032 end, and reads shorter than 50\u2009bp after trimming were discarded. Reads with an average Phred quality score of \u226425 were also excluded. High-quality reads were aligned to the human reference genome (GRCh38; gi 568336000\u2013568336023) using Bowtie2 (v.2.2.9), and all reads mapping to the human genome were removed. Unpaired reads were subsequently discarded. After these filtering steps, a total of 8,777,929,480 PE reads (46,899,434 on average) from cohort 1 and 8,896,915,982 PE reads (44,708,121 on average) from cohort 2 were retained for downstream analyses (hereafter referred to as high-quality reads). Quality-controlled reads were processed using MetaPhlAn3 with default parameters to generate species-level and genus-level taxonomic profiles<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 50\" title=\"Beghini, F. et al. Integrating taxonomic, functional, and strain-level profiling of diverse microbial communities with bioBakery 3. eLife 10, e65088 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR50\" id=\"ref-link-section-d4054757e3016\" rel=\"nofollow noopener\" target=\"_blank\">50<\/a>.<\/p>\n<p>Identification of CRC subtypes based on metagenome-derived SHAP valuesConstruction of a CRC classifier using a random forest model<\/p>\n<p>A random forest model was constructed to estimate the likelihood of CRC, hereafter referred to as the \u2018microbial CRC trait score\u2019, which was calculated for each sample. The model was trained on four independent metagenomic cohorts<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Wirbel, J. et al. Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer. Nat. Med. 25, 679&#x2013;689 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR7\" id=\"ref-link-section-d4054757e3032\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Yu, J. et al. Metagenomic analysis of faecal microbiome as a tool towards targeted non-invasive biomarkers for colorectal cancer. Gut 66, 70&#x2013;78 (2017).\" href=\"#ref-CR17\" id=\"ref-link-section-d4054757e3035\">17<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Vogtmann, E. et al. Colorectal cancer and the human gut microbiome: reproducibility with whole-genome shotgun sequencing. PLoS ONE 11, e0155362 (2016).\" href=\"#ref-CR18\" id=\"ref-link-section-d4054757e3035_1\">18<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 19\" title=\"Zeller, G. et al. Potential of fecal microbiota for early-stage detection of colorectal cancer. Mol. Syst. Biol. 10, 766 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR19\" id=\"ref-link-section-d4054757e3038\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a> curated through the curatedMetagenomicData R package<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 20\" title=\"Pasolli, E. et al. Accessible, curated metagenomic data through ExperimentHub. Nat. Methods 14, 1023&#x2013;1024 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR20\" id=\"ref-link-section-d4054757e3042\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a>, applying filtering thresholds of 1\u2009\u00d7\u200910\u22125 for taxonomic abundance and 0.95 for prevalence. Model training and evaluation were performed using the scikit-learn library (v.1.2.2). Model performance was assessed by tenfold cross-validation, yielding an AUC of 0.84. The trained classifier was subsequently applied to MetaPhlAn3-derived species-level relative abundance profiles from 196 CRC samples (cohort 1) and 146 HCs previously reported<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Yachida, S. et al. Metagenomic and metabolomic analyses reveal distinct stage-specific phenotypes of the gut microbiota in colorectal cancer. Nat. Med. 25, 968&#x2013;976 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR8\" id=\"ref-link-section-d4054757e3048\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>. This evaluation yielded an AUC of 0.74, confirming consistent discriminatory performance in an independent dataset (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">1b<\/a>).<\/p>\n<p>SHAP analysis and subtype identification<\/p>\n<p>Interpretable artificial intelligence frameworks such as SHAP and LIME (local interpretable model-agnostic explanations) have been developed to elucidate the internal decision-making mechanisms of complex predictive models<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 51\" title=\"Alkhanbouli, R., Matar Abdulla Almadhaani, H., Alhosani, F. &amp; Simsekler, M. C. E. The role of explainable artificial intelligence in disease prediction: a systematic literature review and future research directions. BMC Med. Inform. Decis. Mak. 25, 110 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR51\" id=\"ref-link-section-d4054757e3063\" rel=\"nofollow noopener\" target=\"_blank\">51<\/a>. Among these, SHAP provides a mathematically rigorous and reproducible framework for feature attribution, quantifying the contribution of each input variable to the model\u2019s output. SHAP analysis is grounded in cooperative game theory, in which Shapley values represent the fair distribution of contributions among all predictors. In this study, SHAP was used to interpret the internal decision process of the random forest model and to identify distinct microbial features associated with CRC. To further characterize CRC heterogeneity, unsupervised clustering of the SHAP values was performed using the k-means algorithm. The optimal number of CRC subtypes was determined by minimizing the within-cluster sum of squares (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>). Given that SHAP does not inherently account for potential confounding factors such as age, tumor stage or primary tumor site, we systematically examined these variables across SHAP-derived clusters to evaluate potential confounding influences. The reproducibility of the SHAP-based subtype classification was validated in an independent cohort of 199 CRC cases (cohort 2) (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">2b,c<\/a>). All analyses were performed using the shapmat Python package (<a href=\"https:\/\/github.com\/ryzary\/shapmat\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/ryzary\/shapmat<\/a>).<\/p>\n<p>WGS analysis of CRC<\/p>\n<p>Genomic DNA was extracted from primary tumor tissues and matched peripheral blood samples. WGS libraries with an average insert size of 550\u2009bp were prepared from 2\u2009\u00b5g of genomic DNA using the TruSeq DNA PCR-Free Library Preparation Kit (Illumina). Sequencing was performed on the NovaSeq 6000 platform (Illumina) with 150\u2009bp PE reads. The median sequencing depth was 52\u00d7 (mean, 56\u00d7) for tumor samples and 32\u00d7 (mean, 34\u00d7) for matched normal samples. The median estimated tumor purity was 0.52 (range, 0.11\u20130.97), and all cases met the inclusion criterion of \u226510% tumor purity.<\/p>\n<p>Mutation calling<\/p>\n<p>PE reads were aligned to the human reference genome (GRCh37) using BWA-MEM<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 52\" title=\"Li, H. &amp; Durbin, R. Fast and accurate short read alignment with Burrows&#x2013;Wheeler transform. Bioinformatics 25, 1754&#x2013;1760 (2009).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR52\" id=\"ref-link-section-d4054757e3100\" rel=\"nofollow noopener\" target=\"_blank\">52<\/a>. PCR duplicates were removed by eliminating paired reads that mapped to identical genomic coordinates, and pileup files were subsequently generated using SAMtools<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 53\" title=\"Li, H. et al. The Sequence Alignment\/Map format and SAMtools. Bioinformatics 25, 2078&#x2013;2079 (2009).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR53\" id=\"ref-link-section-d4054757e3104\" rel=\"nofollow noopener\" target=\"_blank\">53<\/a>. Somatic point mutations, including SNVs and short indels, were identified according to the following eight criteria: (1) mapping quality of \u226520 and (2) base quality of \u226510. Candidate somatic mutations were further filtered by applying the following conditions: (3) in each tumor sample, variants were required to be supported by at least four reads when the tumor variant allele frequency (TVAF) was \u22650.15, or by at least eight reads when 0.15\u2009&gt;\u2009TVAF\u2009\u2265\u20090.05, with at least one supporting read having a base quality \u226530; (4) the variant allele frequency (VAF) of the matched non-tumor sample had to be &lt;0.03, with a minimum read depth of eight. To reduce context-dependent sequencing errors, reads from all non-tumor samples were aggregated to identify recurrent false-positive sites. For each genomic position with a non-tumor sequence depth of \u226510 and VAF\u2009&lt;\u20090.2, the aggregate non-tumor VAF (NVAF) was calculated. Variants were retained only if (5) NVAF\u2009&lt;\u20090.03 for sites with TVAF\u2009\u2265\u20090.15 or NVAF\u2009&lt;\u20090.01 for sites with 0.15\u2009&gt;\u2009TVAF\u2009\u2265\u20090.05, and (6) the ratio of TVAF to NVAF was \u226520. To exclude germline single-nucleotide polymorphisms, (7) the proportion of non-tumor samples with a VAF\u2009\u2265\u20090.1 was required to be &lt;0.002. Finally, (8) variants showing extreme strand bias (&gt;95% of reads derived from a single strand) were excluded.<\/p>\n<p>Copy number analysis<\/p>\n<p>Copy number ratios, tumor purity and ploidy were estimated using Battenberg (v.2.2.10)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 54\" title=\"Nik-Zainal, S. et al. The life history of 21 breast cancers. Cell 149, 994&#x2013;1007 (2012).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR54\" id=\"ref-link-section-d4054757e3117\" rel=\"nofollow noopener\" target=\"_blank\">54<\/a> with default parameters. The inferred copy number ratios were subsequently adjusted according to the estimated tumor purity for each sample to obtain purity-corrected copy number profiles.<\/p>\n<p>Significantly mutated gene analysis<\/p>\n<p>Significantly mutated genes were identified using three complementary approaches: the inactivation bias method<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Yachida, S. et al. Comprehensive genomic profiling of neuroendocrine carcinomas of the gastrointestinal system. Cancer Discov. 12, 692&#x2013;711 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR22\" id=\"ref-link-section-d4054757e3129\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>, the activation bias method and dNdScv<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 55\" title=\"Martincorena, I. et al. Universal patterns of selection in cancer and somatic tissues. Cell 171, 1029&#x2013;1041.e21 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR55\" id=\"ref-link-section-d4054757e3133\" rel=\"nofollow noopener\" target=\"_blank\">55<\/a>. For the inactivation bias test, the number of samples harboring inactivating mutations (nonsense, read-through, splice-site or frameshift variants) was compared with the number carrying alternative mutations using Fisher\u2019s exact test. Conversely, for the activation bias test, the number of activating (hotspot missense) mutations was compared with the number of other mutations using Fisher\u2019s exact test. Genomic positions exhibiting at least two identical mutations were defined as mutational hotspots. Mutations occurring within 5\u2009bp of a hotspot were also classified as hotspot mutations. As hotspot missense mutations are generally activating but may occasionally be inactivating or functionally ambiguous, genes with P\u2009&lt;\u20090.05 by the inactivation bias method were designated as putative tumor suppressor genes, and their activation bias P\u2009values were fixed at 1. Multiple hypothesis testing was corrected using the Benjamini\u2013Hochberg false discovery rate procedure, and adjusted P\u2009values are reported as q\u2009values. Genes were considered significantly mutated if any of the three tests yielded a q\u2009value of &lt;0.1.<\/p>\n<p>Mutational signature analysis<\/p>\n<p>De novo decomposition of SBS and small ID mutational signatures was performed using SigProfilerExtractor on WGS data from 1,153 CRC cases, comprising 200 Japanese CRC cases from the present study and additional non-Japanese CRC cases reported previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Diaz-Gay, M. et al. Geographic and age variations in mutational processes in colorectal cancer. Nature 643, 230&#x2013;240 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR15\" id=\"ref-link-section-d4054757e3161\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>. Given that mutational spectra differ substantially between hypermutated and non-hypermutated tumors, de novo decomposition was conducted separately for these two groups, following the approach described in a previous study<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Diaz-Gay, M. et al. Geographic and age variations in mutational processes in colorectal cancer. Nature 643, 230&#x2013;240 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR15\" id=\"ref-link-section-d4054757e3165\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>. The number of decomposed signatures was determined based on the optimal solution estimated by SigProfilerExtractor. For each decomposed signature, cosine similarity was calculated relative to the COSMIC reference signature set. Signatures with cosine similarity of \u22650.75 were assigned to known COSMIC signatures, whereas those with cosine similarity of &lt;0.75 were interpreted as novel or undecomposed composite signatures and classified as undetermined.<\/p>\n<p>For the 200 CRC cases analyzed in this study, reconstruction accuracy was evaluated by comparing the reconstructed and observed mutational spectra. SBS signatures demonstrated consistently high reconstruction fidelity (cosine similarity of \u22650.97 in all cases). By contrast, two cases exhibited ID signature cosine similarity values of &lt;0.9, both of which contained a very small number of indels (44 and 68, respectively), indicating insufficient data for reliable signature decomposition. These two cases were therefore excluded from subsequent ID signature analyses.<\/p>\n<p>Definition of colibactin-associated mutational signatures<\/p>\n<p>The presence of a colibactin-associated mutational signature was defined as a relative contribution of SBS88 of \u22655% or ID18 of \u22655%. This threshold was adopted because ten out of 82 colibactin-exposed cases exhibited only one of the two signatures (SBS88 or ID18). Moreover, a 5% cutoff has been widely applied in prior investigations, including the first report characterizing the colibactin-associated mutational process<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 12\" title=\"Pleguezuelos-Manzano, C. et al. Mutational signature in colorectal cancer caused by genotoxic pks+ E.&#x2009;coli. Nature 580, 269&#x2013;273 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR12\" id=\"ref-link-section-d4054757e3180\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>. Given that mutational signature decomposition algorithms may artifactually assign low-frequency components, contributions below 5% are generally regarded as unreliable and are typically truncated to zero, whereas contributions \u22655% are considered robust signals.<\/p>\n<p>Expected probability of mutational signatures<\/p>\n<p>The expected probability of each mutational signature for a given trinucleotide context in an individual patient was calculated using the following equations:<\/p>\n<p>$${N}_{{ijk}}={S}_{{ik}}\\times {C}_{{ij}}$$<\/p>\n<p>$${P}_{{ijk}}={N}_{{ijk}}\/\\mathop{\\sum }\\limits_{i}{N}_{{ijk}}$$<\/p>\n<p>where \\({S}_{{ik}}\\) denotes the number of mutations attributed to mutational signature i in patient k, \\({C}_{{ij}}\\) represents the proportion of trinucleotide context j within mutational signature i, \\({N}_{{ijk}}\\) is the estimated number of mutations corresponding to mutational signature i and trinucleotide context j in patient k, and \\({P}_{{ijk}}\\) is the expected probability of observing mutational signature i in trinucleotide context j for patient k (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Chen, B. et al. Contribution of pks+ E.&#x2009;coli mutations to colorectal carcinogenesis. Nat. Commun. 14, 7827 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR27\" id=\"ref-link-section-d4054757e3427\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>).<\/p>\n<p>Clonal analysis of mutational signatures<\/p>\n<p>Mutations were categorized as clonal or subclonal using MutationTimeR (v.1.00.2)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Gerstung, M. et al. The evolutionary history of 2,658 cancers. Nature 578, 122&#x2013;128 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR33\" id=\"ref-link-section-d4054757e3439\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a>. Mutations that occurred before copy number gains were annotated as clonal (early); those that arose following copy number gains were annotated as clonal (late); and mutations for which timing could not be determined were designated as clonal (NA). For each mutation group, mutational signature decomposition was performed using the deconstructSigs R package (v.1.9.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 56\" title=\"Rosenthal, R., McGranahan, N., Herrero, J., Taylor, B. S. &amp; Swanton, C. DeconstructSigs: delineating mutational processes in single tumors distinguishes DNA repair deficiencies and patterns of carcinoma evolution. Genome Biol. 17, 31 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR56\" id=\"ref-link-section-d4054757e3443\" rel=\"nofollow noopener\" target=\"_blank\">56<\/a>. The parameter signature \u2018signature.cutoff\u2019 was set to zero, and \u2018signatures.ref\u2019 was assigned to the SBS and ID signatures extracted by SigProfilerExtractor, as described above. Clonal (early), clonal (late) and clonal (NA) categories were collectively defined as clonal mutations. Cases with fewer than 50 mutations in any group were excluded from the analysis to ensure robust signature estimation.<\/p>\n<p>Timing analysis of copy number gains<\/p>\n<p>The relative timing of three classes of copy number gains was inferred from (1) CNN-LOH\u2009\/\u2009loss\u2009+\u2009gain (N:0), (2) monoallelic gains (N:1) and (3) biallelic gains (N:M). Analyses were performed using MutationTimeR (v.1.00.2)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Gerstung, M. et al. The evolutionary history of 2,658 cancers. Nature 578, 122&#x2013;128 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR33\" id=\"ref-link-section-d4054757e3456\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a>, which integrates somatic SNVs and copy number variation data derived from WGS. For each tumor, copy number timing estimates were generated using genomic segments containing at least 100 somatic mutations to ensure reliable inference. Weighted averages of mutation counts were computed across segments for each copy number gain category. Comparative analyses among the four CRC subtypes were performed after exclusion of hypermutated cases. To further reconstruct the temporal sequence of genomic events within individual tumors, PhylogicNDT SinglePatientTiming (v.1.0; <a href=\"https:\/\/doi.org\/10.1101\/508127\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.1101\/508127<\/a>) was applied. The inferred relative timing of events was scaled from zero, representing the zygotic (fertilization) stage, to one, corresponding to the time of clinical diagnosis.<\/p>\n<p>Transcriptome sequencing (RNA-seq)<\/p>\n<p>Total RNA was extracted from fresh\u2013frozen tumor tissues using the miRNeasy Mini Kit (Qiagen) according to the manufacturer\u2019s protocol. RNA integrity was assessed using the BioAnalyzer system (Agilent Technologies), and samples with an RNA integrity number of &gt;6.0 were selected for sequencing. RNA-seq libraries were prepared using the SureSelect Strand-Specific RNA Library Preparation Kit (Agilent Technologies) with 300\u2009ng of total RNA as input. Sequencing was performed on the HiSeq 2500 platform (Illumina) to generate 101\u2009bp PE reads. Reads were aligned to the human reference genome (GRCh37\/hg19) using STAR<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 57\" title=\"Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15&#x2013;21 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR57\" id=\"ref-link-section-d4054757e3475\" rel=\"nofollow noopener\" target=\"_blank\">57<\/a>, and gene-level read counts were quantified against the UCSC human transcriptome reference using HTSeq<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 58\" title=\"Anders, S., Pyl, P. T. &amp; Huber, W. HTSeq&#x2014;a Python framework to work with high-throughput sequencing data. Bioinformatics 31, 166&#x2013;169 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR58\" id=\"ref-link-section-d4054757e3479\" rel=\"nofollow noopener\" target=\"_blank\">58<\/a>. Expression levels were calculated as fragments per kilobase of exon per million mapped fragments, incorporating strand-specific information. Differential gene expression was assessed using the Wilcoxon rank-sum test.<\/p>\n<p>Estimation of non-human bacterial reads in tumor tissues<\/p>\n<p>We used the kneaddata tool (v.0.12.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 59\" title=\"McIver, L. J. et al. bioBakery: a meta&#x2019;omic analysis environment. Bioinformatics 34, 1235&#x2013;1237 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR59\" id=\"ref-link-section-d4054757e3491\" rel=\"nofollow noopener\" target=\"_blank\">59<\/a> to eliminate human reads from human rRNA-depleted total RNA sequencing data of tumor tissues, encompassing human and bacterial transcripts. Relative abundance of bacterial reads was calculated as the ratio of non-human reads to all reads in each sample. The total RNA was extracted from fresh\u2013frozen tumor tissues using the miRNAeasy kit (Qiagen). A total of 250\u2009ng of total RNA was used to generate libraries for total RNA sequencing. The Universal Plus Total RNA-seq kit with human rRNA depletion (Tecan) was used for library preparation, followed by sequencing on a NovaSeq 6000 in PE 150 base mode.<\/p>\n<p>Decomposition of tumor immune environments<\/p>\n<p>Gene expression profiles, annotated using GENCODE v43lift37, were deconvoluted to estimate tumor immune cell composition. The proportions of 22 immune cell types were inferred using CIBERSORTx based on the LM22 reference signature matrix, which includes naive and memory B\u2009cells, plasma cells, CD8+ T\u2009cells, naive and memory CD4+ T\u2009cells (resting and activated) and follicular helper T\u2009cells. Additional immune cell subsets analyzed comprised Treg\u2009cells, \u03b3\u03b4 T\u2009cells, resting and activated natural killer cells, monocytes, M0, M1 and M2 macrophages, resting and activated dendritic cells, resting and activated mast cells, eosinophils and neutrophils. CIBERSORTx was executed in \u2018B-mode\u2019 with the \u2018disable quantile normalization\u2019 option enabled, as recommended by the developers for RNA-seq data. Only samples with a CIBERSORTx output P\u2009value of &lt;0.05 were retained for downstream analyses, ensuring robust estimation of immune cell fractions.<\/p>\n<p>T\u2009cell-inflamed GEP analysis<\/p>\n<p>The T\u2009cell-inflamed GEP score<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 60\" title=\"Cristescu, R. et al. Pan-tumor genomic biomarkers for PD-1 checkpoint blockade-based immunotherapy. Science 362, eaar3593 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR60\" id=\"ref-link-section-d4054757e3520\" rel=\"nofollow noopener\" target=\"_blank\">60<\/a> was calculated as the weighted sum of the normalized expression values of 18 genes, each multiplied by its corresponding coefficient: CCL5 (0.008346), CD27 (0.072293), CD274 (0.042853), CD276 (\u22120.023900), CD8A (0.031021), CMKLR1 (0.151253), CXCL9 (0.074135), CXCR6 (0.004313), HLA-DQA1 (0.020091), HLA-DRB1 (0.058806), HLA-E (0.071750), IDO1 (0.060679), LAG3 (0.123895), NKG7 (0.075524), PDCD1LG2 (0.003734), PSMB10 (0.032999), STAT1 (0.250229) and TIGIT (0.084767). The resulting composite score reflects the degree of T\u2009cell-mediated inflammation within the tumor microenvironment.<\/p>\n<p>Quantification of the clbP gene by qPCR<\/p>\n<p>Quantification of the clbP gene (copies per ml) in genomic DNA extracted from tumor tissues or fecal samples was performed by Adenoprevent<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 32\" title=\"Hirayama, Y. et al. Activity-based probe for screening of high-colibactin producers from clinical samples. Org. Lett. 21, 4490&#x2013;4494 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR32\" id=\"ref-link-section-d4054757e3539\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a>. qPCR was conducted using the KAPA SYBR Fast qPCR Kit (Roche Molecular Systems) on the Thermal Cycler Dice Real-Time System (Takara Bio). The copy number of the clbP gene was determined from cycle threshold values using a standard calibration curve generated from clbP gene standards (Adenoprevent)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 30\" title=\"Kawanishi, M. et al. In vitro genotoxicity analyses of colibactin-producing E.&#x2009;coli isolated from a Japanese colorectal cancer patient. J. Toxicol. Sci. 44, 871&#x2013;876 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR30\" id=\"ref-link-section-d4054757e3549\" rel=\"nofollow noopener\" target=\"_blank\">30<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Kawanishi, M. et al. Genotyping of a gene cluster for production of colibactin and in vitro genotoxicity analysis of Escherichia coli strains obtained from the Japan Collection of Microorganisms. Genes Environ. 42, 12 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#ref-CR31\" id=\"ref-link-section-d4054757e3552\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>.<\/p>\n<p>Immunohistochemistry<\/p>\n<p>Formalin-fixed, paraffin-embedded tissue sections (4\u2009\u03bcm thick) were deparaffinized and treated with 3% hydrogen peroxide for 15\u2009min to quench endogenous peroxidase activity. Antigen retrieval was performed using an autoclave in 10\u2009mM citrate buffer (pH\u20096.0) for 10\u2009min. Slides were incubated with a mouse monoclonal anti-FOXP3 antibody (clone 236\u2009A\/E7; dilution 1:100; ab20034, Abcam) for 1\u2009h at room temperature (20\u201325\u2009\u00b0C), followed by detection using the EnVision system (Dako) with mouse linker (Dako), according to the manufacturer\u2019s protocol. Signal visualization was achieved using the Chem-Mate EnVision detection method (Dako). Immunohistochemical staining for FOXP3 was performed in 28 CRC cases classified as subtype 3. For each case, five representative tumor regions exhibiting lymphocytic infiltration were randomly selected and imaged at \u00d7200 magnification. FOXP3-positive lymphocytes were manually counted within each field.<\/p>\n<p>Statistical analyses<\/p>\n<p>When multiple comparison groups were present, overall differences were assessed using the Kruskal\u2013Wallis test. Pairwise comparisons between individual groups were subsequently performed using the Wilcoxon rank-sum test or Fisher\u2019s exact test, as appropriate. Correction for multiple hypothesis testing was performed using the Benjamini\u2013Hochberg false discovery rate procedure, and adjusted P\u2009values are reported as q\u2009values.<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02692-x#MOESM2\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"Patients and samples This study was approved by the Research Ethics Committees of the National Cancer Center, the&hellip;\n","protected":false},"author":2,"featured_media":71924,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[2451,46331,14060,14061,32330,46330,1907,46329,8,17],"class_list":["post-71923","post","type-post","status-publish","format-standard","has-post-thumbnail","category-japan","tag-agriculture","tag-animal-genetics-and-genomics","tag-biomedicine","tag-cancer-research","tag-colorectal-cancer","tag-gene-function","tag-general","tag-human-genetics","tag-japan","tag-japanese"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/posts\/71923","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/comments?post=71923"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/posts\/71923\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/media\/71924"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/media?parent=71923"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/categories?post=71923"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/japan\/wp-json\/wp\/v2\/tags?post=71923"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}