Development of a random background to understand ligand optimization (paper 16th Sept 2026)

https://www.nature.com/articles/s41586-026-11013-5

chatGPT(6AstraMaxPaid):

Xu et al., “Development of a random background to understand ligand optimization”, published in Nature on 16 September 2026, examines how often small, undesigned chemical changes improve a compound’s biological activity.

My assessment: this is a valuable experimental benchmark for drug discovery. It shows that substantial potency improvements can arise surprisingly often from simple substitutions, while documenting the difficulty of combining potency with useful pharmacokinetic properties. Its strongest conclusions concern the compounds studied; generalisation across drug discovery requires further evidence.

The study asks a useful question: how much improvement should researchers expect without deliberately designing a better molecule?

The authors started with 18 active compounds spanning 11 unrelated chemical scaffolds and six targets: three receptors, the serotonin transporter, and two enzymes. They made 257 derivatives, changing one position at a time through methyl, hydroxyl, halogen or aromatic nitrogen substitutions. Selection depended on synthetic feasibility and cost, without using predicted activity to choose the substitutions.

They measured target activity alongside solubility, permeability, metabolic stability, plasma stability, plasma protein binding and hERG inhibition, a cardiac safety indicator.

The principal activity results were:

Improvement over starting compound Using the authors’ rounded values Using strict, unrounded values
At least threefold 69 of 257: 26.8% 62 of 257: 24.1%
At least tenfold 29 of 257: 11.3% 25 of 257: 9.7%

The headline 11.3% includes four compounds whose improvements were between 9.7-fold and 9.9-fold. The authors disclose this in Table 1. The distinction changes the precise percentage, while leaving the broad finding of approximately one substantial improvement per ten tested derivatives intact.

Other findings were:

  • Small changes sometimes produced large effects. Approximately tenfold improvements occurred for 10 of the 18 starting compounds and five of the six targets. Methyl and chlorine substitutions were particularly productive. Around 30% of derivatives were classified as at least tenfold less active.
  • Potency improvements frequently accompanied less favourable drug properties. None of the 29 compounds in the rounded tenfold-improvement group also improved all three highlighted properties: metabolic stability, permeability and plasma fraction unbound.
  • Structural explanations were easier to construct after measurement than to predict beforehand. Crystal structures and cryo-EM revealed altered pocket packing, protein rearrangements and alternative binding poses.
  • Computational predictions were useful but imperfect. Free-energy calculations anticipated some activity changes. AI tools performed better for certain properties, particularly protein binding, than for metabolic or plasma stability.
  • One derivative demonstrated an improvement in selected mouse pain assays. Adding a methyl group to an alpha2A receptor agonist produced compound 4905, with a reported 52-fold improvement in cellular potency. However, its measured in vivo half-life fell from 196 to 15.2 minutes.

The novelty lies principally in the experimental design and dataset. The authors build on established positional analogue scanning methods, explicitly citing earlier work from 2020 and 2022. Large effects from methyl substitutions and trade-offs between potency and pharmacokinetics were also already recognised.

The important advances are:

  1. An experimentally measured background success rate. Testing substitutions without selecting them for predicted activity reduces the design and publication biases that complicate retrospective databases.
  2. Joint measurement of activity and other drug properties. This allows the study to quantify how often gains in one dimension accompany losses elsewhere.
  3. A benchmark connecting chemistry, structure and prediction. The measurements and structural data provide useful challenges for computational methods, including cases where apparently simple chemical changes behave unexpectedly.

The study has substantial strengths. Parent compounds and derivatives were compared using consistent activity readouts within each series. Unsuccessful derivatives contribute to the results. Structural work adds mechanistic depth, computational predictions were blinded, and the mouse experiments included randomisation and blinded assessment.

My main reservations concern the scope and interpretation of the benchmark:

  1. The “random background” is strongly conditional on the selection process.
    The substitutions were undesigned with respect to activity, but the starting molecules, targets and permitted transformations were deliberately selected. Only about 7% of enumerated possibilities met the feasibility and price criteria, and approximately 80% of those were successfully synthesised. The resulting success rate therefore describes affordable, accessible modifications around this particular collection of active compounds. It should not be treated as a universal probability for an arbitrary chemical change.

  2. The 257 derivatives are clustered within relatively few starting compounds.
    Derivatives of the same parent share chemical and experimental characteristics. Around 62% of the compounds came from just two targets, SERT and alpha2A. Success also varied substantially between individual series. A benchmark intended to predict performance on new targets needs uncertainty estimates that account for clustering by parent and target, alongside broader independent replication.

  3. There is no direct comparison establishing superiority over expert or AI-guided design.
    Similarity to success rates in ChEMBL does not establish equivalent efficiency: the compounds, objectives and reporting processes differ. A stronger comparison would give each strategy the same starting compounds, synthesis budget and assays, then compare the quality of the resulting candidates. This study also examines single modifications rather than the repeated, adaptive decisions involved in an optimisation programme.

  4. Binding affinity and functional potency are different quantities.
    Some series were assessed through binding measurements; others used cellular EC50 values. EC50 depends on receptor expression, signalling amplification and the compound’s ability to activate the receptor, as well as binding. The authors acknowledge this and include expression controls. Nevertheless, the reported 52-fold cellular potency gain should not be interpreted as proof of 52-fold tighter binding. This distinction also complicates comparisons with calculated binding free energies.

  5. The pharmacokinetic interpretation is somewhat stronger than the measurements justify.
    Most results are laboratory proxies, with in vivo pharmacokinetics measured for only two compounds. Microsomal assays miss some metabolic pathways, and artificial membrane permeability does not reproduce transport across living tissues. Moreover, a lower fraction unbound does not automatically imply poorer therapeutic exposure, because clearance and distribution also matter. The decisive outcome is free drug concentration at the intended tissue over time, relative to the concentration needed for activity.

  6. Failure to improve every property simultaneously is a demanding success criterion.
    A useful candidate need not improve metabolic stability, permeability and protein binding together. Some properties may already be adequate, and a modest deterioration may be tolerable. The study convincingly demonstrates competing effects, but establishing their practical cost requires judging compounds against relevant exposure, efficacy and safety requirements.

  7. The analgesic example supports a narrower conclusion than some of the wording suggests.
    Compound 4905 improved performance in selected mouse assays, but its advantage was not statistically established in every comparison. In the hotplate experiment, the direct comparison with its parent was not significant: P = 0.2209. The roughly 13-fold shorter half-life also raises an unresolved question about duration of benefit. Small groups of young male mice, assessed at limited times, cannot establish an overall superior analgesic profile. Longer exposure-response studies and direct assessment of adverse effects would make that claim substantially stronger.