{"id":93965,"date":"2026-07-30T14:49:11","date_gmt":"2026-07-30T14:49:11","guid":{"rendered":"https:\/\/www.europesays.com\/britain\/93965\/"},"modified":"2026-07-30T14:49:11","modified_gmt":"2026-07-30T14:49:11","slug":"gsk-backs-relation-therapeutics-with-110m-to-build-biological-data-factory","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/britain\/93965\/","title":{"rendered":"GSK Backs Relation Therapeutics With $110M to Build Biological Data Factory"},"content":{"rendered":"<p>Relation Therapeutics <a href=\"https:\/\/www.biospace.com\/press-releases\/relation-announces-expanded-collaboration-with-gsk\" rel=\"nofollow noopener\" target=\"_blank\">announced a strategic research collaboration<\/a> with GSK focused on generating large-scale human cellular perturbation data \u2014 paired with the unveiling of MORGAN, its cellular biology foundation model. But the deal&#8217;s structure reveals something the headline obscures. GSK is not paying to license MORGAN. It is paying Relation to manufacture the biological training data that will make MORGAN work \u2014 a distinction that matters more than the dollar amount.<\/p>\n<p>In AI drug discovery, the dominant commercial model has been model licensing: a biotech builds an AI platform, a pharma company pays to use it. This deal inverts that. GSK is paying Relation to run automated cellular experiments at scale, generating the petascale multi-omic datasets MORGAN will be trained on. The data, not the model, is what GSK is buying into.<\/p>\n<p>Under the terms of the agreement, Relation may receive up to $110 million in upfront and success-based milestone payments. In return, Relation will <a href=\"https:\/\/www.biospace.com\/press-releases\/relation-announces-expanded-collaboration-with-gsk\" rel=\"nofollow noopener\" target=\"_blank\">generate large-scale human cellular perturbation datasets<\/a> using advanced human cellular disease models and experimental systems \u2014 covering disease areas including immunology, inflammation, fibrotic diseases, and osteoarthritis. Those datasets will be used to investigate biological pathways and to train foundation models including MORGAN.<\/p>\n<p>Why Biological Training Data Is Scarcer Than AI Models<\/p>\n<p>The race to build a &#8220;foundation model of the cell&#8221; \u2014 a general-purpose AI that predicts how human cells respond to drugs and genetic changes \u2014 has attracted well-funded competitors from Isomorphic Labs to Noetik to insitro. What every one of them has in common is a data problem.<\/p>\n<p>Protein structure prediction succeeded in part because the Protein Data Bank had accumulated <a href=\"https:\/\/www.bvp.com\/atlas\/building-biology-native-data-infrastructure-for-the-ai-era\" rel=\"nofollow noopener\" target=\"_blank\">more than 200,000 experimentally determined protein structures<\/a> over decades of publicly funded science. There is no equivalent repository for cellular perturbation data \u2014 the molecular readouts that show what happens across the genome, transcriptome, and proteome simultaneously when a cell is disrupted by a drug or a genetic edit.<\/p>\n<p>Experts at the <a href=\"https:\/\/www.amacad.org\/publication\/daedalus\/building-drug-discovery-engine-future-ai-empowered-nodal-biology\" rel=\"nofollow noopener\" target=\"_blank\">American Academy of Arts and Sciences<\/a> put the problem plainly in a May 2026 analysis: &#8220;The key bottleneck in solving this foundational problem is obtaining cell perturbation data of large enough magnitude and quality to train cell-prediction AI models. Deep-learning models require data of a scale that is unprecedented in cell biology.&#8221;<\/p>\n<p>The existing public datasets \u2014 libraries like the LINCS L1000 and Perturb-seq screens \u2014 were generated before AI training requirements were understood. Bessemer Venture Partners described the structural gap in an April 2026 market analysis: much of the biological data available today was generated before the explosion of AI biology models, meaning it often lacks the traits that make it useful for machine learning. Annotations are incomplete, important experimental context is rarely captured, data tends to be siloed by modality, and the batch effects introduced by day-to-day laboratory variation make cross-experiment comparison unreliable at the scale required to train a foundation model.<\/p>\n<p><a href=\"https:\/\/alitheagenomics.com\/how-high-throughput-transcriptomics-solves-the-ai-training-data-bottleneck-in-drug-discovery\/\" rel=\"nofollow noopener\" target=\"_blank\">Alithea Genomics<\/a> was more specific in a March 2026 industry analysis: &#8220;Public datasets aren&#8217;t built for AI training and are too heterogeneous, sparse in perturbations, and biased in their sample composition to robustly train models that reliably generalize.&#8221;<\/p>\n<p>Relation&#8217;s answer is not to find better public data. It is to manufacture proprietary data at a scale and quality that public repositories cannot provide.<\/p>\n<p>What MORGAN Is and Why the Data Behind It Is the Point<\/p>\n<p>MORGAN \u2014 <a href=\"https:\/\/www.biospace.com\/press-releases\/relation-announces-expanded-collaboration-with-gsk\" rel=\"nofollow noopener\" target=\"_blank\">Multi-Omic Regulatory Genomics using Artificial Neural Networks<\/a> \u2014 is designed as a general-purpose model of cellular perturbation response, applicable across cell types and disease contexts rather than trained on a single indication. The model&#8217;s goal: given a cell and a perturbation, predict the full molecular response before running the lab experiment.<\/p>\n<p>The technical foundation is multi-omic perturbation data. Where a standard RNA-sequencing experiment captures only the transcriptome (which genes are expressed), a multi-omic readout simultaneously captures the genome, transcriptome, and proteome in the same cell population \u2014 a far richer picture of how the cell has responded. &#8220;Time-resolved&#8221; means the experiment doesn&#8217;t produce a single snapshot but a temporal series: how does the cell&#8217;s molecular state change over minutes, hours, and days following the intervention? This temporal dimension matters because many therapeutically relevant responses \u2014 signaling cascades, feedback loops, compensatory adaptations \u2014 unfold over time and are invisible in single-timepoint assays.<\/p>\n<p>What Relation calls &#8220;superhuman consistency&#8221; refers to eliminating the batch effects \u2014 the systematic variation introduced by human operators, day-to-day instrument calibration drift, and plate-to-plate differences \u2014 that make biological datasets unreliable for model training at scale. Automated laboratory systems run the same protocols identically, at any hour, without the sources of noise that human-run experiments unavoidably introduce.<\/p>\n<p>&#8220;Deepening our understanding of the underlying biology of disease starts with richer data, to be used in models that can give us greater confidence in the discovery of therapeutic targets that can ultimately yield medicines,&#8221; said <a href=\"https:\/\/www.biospace.com\/press-releases\/relation-announces-expanded-collaboration-with-gsk\" rel=\"nofollow noopener\" target=\"_blank\">David Roblin, Relation&#8217;s chief executive<\/a>, in the announcement.<\/p>\n<p>The honest caveat: a 2025 paper in Nature Methods found that current cellular foundation models &#8220;perform no better than much simpler models for perturbation response prediction.&#8221; That finding applies to models trained on existing public data \u2014 which is precisely Relation&#8217;s argument for why proprietary, high-quality, petascale data changes the equation. Whether it does remains to be demonstrated.<\/p>\n<p>What GSK Is Actually Paying For<\/p>\n<p>This structure differs meaningfully from GSK&#8217;s other recent AI deals. In January 2026, the company paid <a href=\"https:\/\/www.businesswire.com\/news\/home\/20260108468293\/en\/GSK-Licenses-Noetiks-AI-Foundation-Models-in-Anchor-Partnership-to-Transform-Cancer-Therapeutic-Research-and-Development\" rel=\"nofollow noopener\" target=\"_blank\">Noetik&#8217;s OCTO-VC virtual cell models<\/a> $50 million upfront for a five-year license to access Noetik&#8217;s models in non-small cell lung cancer and colorectal cancer \u2014 models already built on Noetik&#8217;s existing spatial biology dataset. That is model licensing: Noetik made the data, trained the model, and GSK paid to use what existed.<\/p>\n<p>The Relation deal is upstream of that. GSK is paying before the data exists, to fund its creation. The arrangement treats biological perturbation data the way a pharmaceutical company might treat a chemical library \u2014 as a proprietary asset worth building, not just purchasing access to.<\/p>\n<p>GSK is, in effect, co-funding Relation&#8217;s data engine in exchange for access to the results.<\/p>\n<p>GSK&#8217;s Systematic AI Strategy and the Phase II Problem<\/p>\n<p>For GSK, this is the latest in a deliberate sequence of AI investments oriented around a specific clinical problem: Phase II trial attrition. Most drugs that clear preclinical testing still fail when they reach human studies \u2014 a pattern that makes drug development expensive, slow, and structurally difficult to improve without fundamentally better target selection.<\/p>\n<p>GSK Chief Scientific Officer Tony Wood has stated this objective explicitly, <a href=\"https:\/\/www.cio.inc\/blogs\/gsk-ai-driven-science-factory-p-4121\" rel=\"nofollow noopener\" target=\"_blank\">identifying Phase II attrition as the exact problem AI must solve<\/a>. The logic: if the disease targets chosen for drug development are genetically and mechanistically validated \u2014 if there is rich data showing exactly how diseased human cells respond to disrupting a given target \u2014 then the odds of a Phase II drug working are higher than if the target was selected on weaker biological evidence.<\/p>\n<p>The portfolio GSK is building around this thesis is now coherent: a multi-year partnership with Helix for access to GenoSphere cohort genomic data (January 2026); a $50 million deal with Noetik for spatial oncology virtual cell models (January 2026); an integration with Microsoft Discovery for <a href=\"https:\/\/www.cio.inc\/blogs\/gsk-ai-driven-science-factory-p-4121\" rel=\"nofollow noopener\" target=\"_blank\">AI-assisted drug development workflows<\/a>; and now the Relation expanded collaboration for multi-omic perturbation data generation. Each acquisition addresses a different data modality. Together they are an attempt to build a data-rich target selection infrastructure.<\/p>\n<p>GSK has also committed to spending <a href=\"https:\/\/www.gsk.com\/en-gb\/media\/press-releases\/gsk-to-invest-30-billion-in-rd-and-manufacturing-in-the-united-states-over-next-5-years\/\" rel=\"nofollow noopener\" target=\"_blank\">at least $30 billion in US research and development and manufacturing<\/a> over the next five years, with artificial intelligence named as a component of that investment.<\/p>\n<p>What This Tells You About the Cellular AI Race<\/p>\n<p>The cellular foundation model field has attracted significant capital and well-funded competitors. Isomorphic Labs, the Google DeepMind spinout, has pursued partnerships with Eli Lilly, Novartis, and Johnson &amp; Johnson valued above $3 billion collectively. insitro integrates human cell data generation with machine learning and recently saw Bristol-Myers Squibb nominate additional ALS drug targets from their collaboration. Chai Discovery is pursuing a generalist protein and small molecule design platform backed by major pharma. <a href=\"https:\/\/www.bvp.com\/atlas\/building-biology-native-data-infrastructure-for-the-ai-era\" rel=\"nofollow noopener\" target=\"_blank\">Anthropic acquired Coefficient Bio in April 2026<\/a> for approximately $400 million, signaling that frontier AI labs are now making direct bets on drug discovery.<\/p>\n<p>What the Relation deal makes explicit is the competitive axis that will separate the leaders from the followers in this space. In an environment where the AI model architectures for cellular biology are rapidly improving and increasingly open-source \u2014 scGPT, Geneformer, BioGPT \u2014 the durable advantage belongs to whoever controls the proprietary training data those models learn from. Bessemer Venture Partners, surveying the field in April 2026, named <a href=\"https:\/\/www.bvp.com\/atlas\/building-biology-native-data-infrastructure-for-the-ai-era\" rel=\"nofollow noopener\" target=\"_blank\">biology-native data at scale<\/a> as the first and most important of three principles that will define the next generation of leading biotechs.<\/p>\n<p>The Relation\/GSK deal is the first major pharma contract to make that thesis explicit in its payment structure: the money flows toward the data generation, not toward the algorithm.<\/p>\n<p>Can MORGAN Deliver What It Promises?<\/p>\n<p>Relation has made ambitious claims. The company told <a href=\"https:\/\/endpoints.news\/relation-promises-transformationally-different-success-rates-in-expanded-gsk-pact\/\" rel=\"nofollow noopener\" target=\"_blank\">Endpoints News<\/a> that its approach could deliver &#8220;transformationally different success rates&#8221; in drug discovery. That remains to be demonstrated. No cellular foundation model \u2014 not MORGAN, not scGPT, not OCTO-VC \u2014 has yet produced an approved medicine or even a clinical candidate that entered human trials from a model-first discovery process.<\/p>\n<p>The underlying scientific claim \u2014 that high-quality, petascale, time-resolved multi-omic perturbation data will allow MORGAN to outperform simpler models and produce drug targets with better clinical success rates \u2014 is plausible but unproven. The Nature Methods 2025 finding that current foundation models don&#8217;t outperform baselines for perturbation prediction is a genuine caveat; Relation&#8217;s answer is that the quality of the training data, not the architecture, is what was missing.<\/p>\n<p>What Thursday&#8217;s announcement confirms is that a major pharmaceutical company has evaluated Relation&#8217;s argument and found it credible enough to pay $110 million for it \u2014 before MORGAN has produced a single validated drug target.<\/p>\n<p>In AI drug discovery, the model increasingly matters less than what it learns from. Relation is betting its entire competitive strategy on that thesis, and GSK just paid to test it.<\/p>\n<p>Frequently Asked QuestionsWhat is a cellular foundation model, and why does training data matter so much?<\/p>\n<p>A cellular foundation model is an AI system trained on large amounts of biological data to learn general representations of how human cells work, then adapted to predict specific things \u2014 like how a cell will respond to a drug \u2014 for drug discovery. Like large language models trained on text, these models are only as good as what they learn from. The key problem is that the biological data needed to train them \u2014 specifically, large-scale multi-omic readouts of how cells respond to genetic and pharmacological interventions \u2014 doesn&#8217;t exist in public repositories at anywhere near the required scale or quality. The Protein Data Bank (which helped train AlphaFold) contains over 200,000 protein structures accumulated over decades. No equivalent repository exists for cellular perturbation responses. Companies like Relation are trying to build that data themselves.<\/p>\n<p>What does &#8220;multi-omics&#8221; mean, and why is it different from standard biological data?<\/p>\n<p>Standard biological experiments typically measure one thing at a time \u2014 for example, RNA sequencing captures which genes are active (the transcriptome). Multi-omic experiments simultaneously capture multiple layers of molecular information in the same cells: the genome, the transcriptome, the proteome, and potentially the epigenome. This gives a much richer picture of how a cell has responded to a perturbation. For AI training, multi-omic data is more valuable because the interactions between these molecular layers often explain how drugs work and why they fail \u2014 information that single-readout experiments miss entirely.<\/p>\n<p>Has any AI-designed drug been approved by a major regulatory agency?<\/p>\n<p>No AI-designed drug has received approval from the FDA, the European Medicines Agency, or any other major regulatory body as of July 30, 2026. Insilico Medicine&#8217;s rentosertib \u2014 a TNIK inhibitor for idiopathic pulmonary fibrosis in which both the disease target and the drug compound were identified and designed by generative AI \u2014 entered Phase III clinical trials in China in July 2026 following positive Phase IIa results. That is currently the most advanced AI-designed drug candidate in clinical testing. MORGAN and the Relation\/GSK collaboration are at the preclinical target-identification stage, several years earlier in the development process.<\/p>\n<p>Why is GSK paying for data generation rather than just licensing Relation&#8217;s model?<\/p>\n<p>Because in cellular AI, the training data is the primary source of competitive advantage. AI model architectures for biology are increasingly available as open-source research, meaning any company can access the same algorithms. What cannot be easily replicated is proprietary, large-scale, high-quality biological data generated under controlled experimental conditions. GSK&#8217;s payment structure \u2014 funding Relation to build a data factory \u2014 reflects the judgment that owning rights to petascale multi-omic perturbation data is more strategically valuable than licensing a model trained on data that any competitor could also access. This is the same logic that drives investment in proprietary compound libraries or patient biobanks: the data asset, not the analytical tool, is where the durable advantage lies.<\/p>\n","protected":false},"excerpt":{"rendered":"Relation Therapeutics announced a strategic research collaboration with GSK focused on generating large-scale human cellular perturbation data \u2014&hellip;\n","protected":false},"author":2,"featured_media":93966,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[20784],"tags":[40096,40095,27811,15935,20785,40094,22666,40093],"class_list":["post-93965","post","type-post","status-publish","format-standard","has-post-thumbnail","category-gsk","tag-ai-drug-discovery","tag-biological-training-data","tag-biotech","tag-drug-discovery","tag-gsk","tag-morgan-foundation-model","tag-multi-omics","tag-relation-therapeutics"],"share_on_mastodon":{"url":"https:\/\/pubeurope.com\/@UnitedKingdom\/117009485874262594","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/posts\/93965","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/comments?post=93965"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/posts\/93965\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/media\/93966"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/media?parent=93965"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/categories?post=93965"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/britain\/wp-json\/wp\/v2\/tags?post=93965"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}