Artificial intelligenceAugust 3, 2026· via AI News

GSK bets $110M on AI-driven biological data to speed drug discovery

GSK bets $110M on AI-driven biological data to speed drug discovery

Image : AI News

GSK is placing a $110 million bet on biological data as the new backbone of AI-powered drug discovery. In a multi-year collaboration with Relation Therapeutics, the pharmaceutical giant will fund the generation of large-scale datasets that capture how human cells respond to genetic changes and drug interventions. The goal is to feed these datasets into Relation’s MORGAN AI platform and other models, training them to pinpoint drug targets with greater precision.

From lab bench to algorithm

Relation’s approach blends laboratory experimentation with computational analysis through its Lab-in-the-Loop system. The workflow includes tissue profiling, single-cell and spatial transcriptomics, sequencing, and perturbation experiments that measure how genetic tweaks alter cellular traits linked to disease. Machine learning then steps in to prioritize targets, validate findings, and even suggest next experiments. This closed-loop method builds on earlier GSK-Relation projects in fibrotic diseases and osteoarthritis, where observational studies produced functional disease datasets for AI analysis.

Public biological databases remain vital inputs for training AI models, but their heterogeneity poses challenges. A 2025 review in Experimental & Molecular Medicine highlights repositories like CZ CELLxGENE and the Human Cell Atlas, which together offer over 100 million standardized single-cell profiles. Yet differences in sampling methods, sequencing protocols, and processing pipelines can introduce technical noise and redundancy. Overlapping datasets risk skewing model training or leaking data between test and validation sets—undermining reliability.

Quality trumps quantity in AI training

New research in Nature Methods underscores this point. In an analysis of 22.2 million cells and 400 models across 6,400 experiments, scientists found that single-cell foundation models often plateau after training on a fraction of available data. Bigger datasets alone don’t guarantee better performance; diversity, curation, and quality control are equally critical. The study suggests that assembling non-redundant, high-quality datasets may matter as much as model architecture itself.

For GSK and Relation, the collaboration signals a shift from data scarcity to data strategy. By anchoring AI development in experimentally generated biological insights, they aim to reduce failures in target identification—a persistent bottleneck in drug discovery. If successful, the approach could accelerate the path from lab to clinic, but only if the data feeding the models is as robust as the algorithms processing it.

Why it matters

This partnership elevates biological data from supporting role to strategic asset in AI drug discovery. For researchers and drug developers, it highlights that the next breakthroughs may not come from bigger models alone, but from richer, cleaner, and more experimentally grounded training data. The stakes are clear: better data could mean fewer dead ends in target selection, faster timelines, and ultimately, more therapies reaching patients. The real test will be whether these lab-trained models can translate into real-world clinical successes.


Source: AI News. AI-assisted editorial synthesis — TechnoExpress.

Read the original source on AI News →

← Back to home