GSK Invests $110 Million in Relation Therapeutics to Build AI Drug Discovery Data Factory

MORGAN stands for Multi-Omic Regulatory Genomics using Artificial Neural Networks, and Relation describes MORGAN as a general‑purpose foundation model applicable across multiple cell types and disease areas rather than being tailored to a single indication.
The collaboration with GSK builds on two prior 2024 deals between the companies focused on identifying and validating targets for fibrotic diseases and osteoarthritis, signaling an ongoing, data‑driven expansion of their partnership.
Industry observers note the absence of a comprehensive cellular perturbation data repository, underscoring why data generation is central to MORGAN and why such datasets are currently more scarce than large-scale protein structure data.
Relation emphasizes the scale and consistency of its data generation approach, describing petascale, time-resolved multi-omic readouts produced via integrated automated laboratories designed for high-throughput experiments with exceptional reproducibility to train MORGAN.
GSK and Relation Therapeutics have struck a deal worth up to $110 million to build what they call a "Biological Data Factory," according to TechTimes. The goal is to generate massive amounts of human cellular data and use it to train MORGAN, Relation's AI model for drug discovery.
Unlike most AI deals in pharma, GSK is not paying to license an existing model. Instead, it is paying Relation to produce the training data itself, according to Fierce Biotech. The partnership covers disease areas including immunology, inflammation, fibrotic diseases, and osteoarthritis.
MORGAN stands for Multi-Omic Regulatory Genomics using Artificial Neural Networks. Relation describes it as a general-purpose foundation model — meaning it is built to work across many cell types and disease areas, not just one. Think of it like a scientific version of ChatGPT, but trained on biology instead of text.
The problem is that good training data for this kind of model barely exists. Industry observers note there is no comprehensive cellular perturbation data repository today, according to Digital Health News. Large-scale protein structure data is far more available. That scarcity is exactly why this deal is centered on data generation rather than model licensing.
Relation will run what it calls petascale, time-resolved multi-omic experiments. In plain terms, the company will measure how human cells respond to thousands of different interventions — at enormous scale and with precise timing. Automated labs handle the work to keep results consistent, according to Bioxconomy.
That consistency matters. Biological experiments are notoriously hard to repeat. By using integrated automation, Relation says it can produce high-throughput data with "exceptional reproducibility." That is the kind of clean, reliable data needed to train a foundation model that actually works in drug discovery.
This is not the first time GSK and Relation have worked together. The two companies signed two separate deals in 2024, both focused on finding and validating drug targets for fibrotic diseases and osteoarthritis, according to Fierce Biotech. This new $110 million agreement builds directly on that foundation.
The deal signals a broader shift in AI-powered drug discovery. Companies are no longer just licensing AI tools — they are investing in generating the raw data needed to make those tools powerful. Law firm Mishcon de Reya advised Relation on the transaction, according to Mishcon de Reya.
GSK's decision to fund data production — not just buy access to a finished model — reflects a bigger idea: in AI-driven biology, the data is the moat. Relation, which is backed by Nvidia, is positioning itself as the company that owns that data, according to Digital Health News.
For GSK, the payoff is increased confidence in picking the right drug targets before spending billions on clinical trials. The cellular perturbation data produced will feed directly into MORGAN and support target discovery across multiple disease areas. The $110 million price tag reflects just how hard that data is to get anywhere else.
Publishers
11
Articles
13
Reach
24