How Artificial Intelligence Is Changing Drug Discovery
How Artificial Intelligence Is Changing Drug Discovery
Drug discovery is a search problem conducted in an almost unimaginably large chemical space. Artificial intelligence can help researchers navigate that space, but prediction is only the beginning of scientific evidence.
Developing a medicine involves identifying a biological target, finding molecules that affect it, optimising their properties, testing toxicity and efficacy, and conducting clinical trials. AI can contribute at several stages by learning relationships among molecular structure, biological activity and experimental observations.
Representing molecules for machine learning
A molecule can be represented as a sequence, a graph or a three-dimensional structure. Different models capture different information. Graph neural networks treat atoms as nodes and bonds as edges. Geometric models attempt to learn spatial relationships that influence binding and reactivity. Language-inspired models can process molecular strings or protein sequences as tokens.
The representation matters because molecules with similar components may behave differently when their geometry or environment changes.
Prediction, generation and optimisation
Predictive models estimate properties such as solubility, toxicity or binding affinity. Generative models propose new structures that satisfy specified objectives. Optimisation systems attempt to balance several properties simultaneously, because a molecule that binds strongly may still be unsuitable if it is unstable or toxic.
These methods can prioritise candidates and reduce the number of experiments required. They do not eliminate experiments. A generated molecule may be difficult to synthesise, interact differently in living systems or exploit weaknesses in the computational scoring method.
The limits of available data
Published datasets are not random samples of all possible chemistry. They reflect what researchers chose to study, what succeeded and what was reported. Negative results are often less visible. Measurement protocols differ, and duplicated or related compounds can leak between training and test sets, inflating apparent performance.
Prospective validation is especially valuable: the model selects candidates before the outcome is known, and laboratory experiments determine whether the prediction holds.
AI as part of a closed experimental loop
A promising direction combines models with automated laboratories. The model proposes an experiment, robotic systems perform it, and the result updates the next selection. This active-learning loop can concentrate experimental effort where new data are most informative.
Even then, scientific judgement remains essential. Researchers define objectives, inspect unexpected findings and decide whether a measured improvement is biologically meaningful.
Acceleration without certainty
AI may shorten parts of discovery and reveal patterns that are difficult to identify manually. But a computational candidate is not a treatment, and early laboratory success is not evidence of safety or clinical benefit.
The scientifically responsible claim is therefore not that AI “discovers drugs” on its own. It expands the search capabilities of multidisciplinary teams and helps them decide which hypotheses deserve scarce experimental attention.
Prediction is only the beginning
A computationally attractive molecule must still be synthesised, tested for activity and selectivity, and assessed for absorption, distribution, metabolism, excretion and toxicity. Each stage can invalidate an earlier prediction. AI is most useful when it helps researchers choose better experiments, not when it is presented as a substitute for experiments.
Prospective validation is crucial. Retrospective benchmarks can leak information when related chemical scaffolds occur in both training and test data. Time-split and scaffold-based tests, followed by laboratory confirmation, provide more realistic evidence. Negative results should be reported because a record containing only successful candidates exaggerates reliability.
How to assess an AI-discovery claim
- Was the evaluation prospective or based on information already known?
- Were candidates synthesised and tested in independent assays?
- Are novelty, selectivity, toxicity and manufacturability reported?
- Are data provenance and unsuccessful experiments visible?
- Did AI improve scientific yield against a credible conventional workflow?
Disease biology remains harder than molecular ranking. Structure prediction and generative chemistry can narrow a search, but they do not establish causal biology or clinical benefit. The final standard remains reproducible experimental evidence and, eventually, safe and effective treatment.
Further reading: Nature Chemical Biology: machine learning in preclinical drug discovery.
