Skip to content
Research notesBiology11 min read

Lifespan targets are not rejuvenation targets, and the evidence separates cleanly

one shared genelifespan targetsrejuvenation targets11 + 10 targets, and they are not the same list

Most longevity drug targets in geroscience are chosen from one literature and then asked to answer a different question. The field talks about extending lifespan and about making cells younger as though they were one programme. On the evidence we can measure, they run on almost entirely different genes.

This matters practically. If a molecule is designed against a nutrient-sensing pathway, the honest claim attached to it is about lifespan in model organisms. It is not a claim about reversing a cell's biological age, and presenting one number as though it answered both questions is the sort of thing that survives right up until somebody checks.

Two lists, built from two literatures

Our first target list came from the geroprotector literature: mTOR, AMPK, the NAD axis, the insulin and IGF-1 axis, and the apoptotic and inflammatory nodes around them. Every one of them is a cytoprotective or nutrient-sensing pathway, and every one of them has a compound behind it that extended life in an animal.

The second list came from a different question entirely: which genes, when perturbed, move a cell's measured transcriptional age. That search returns epigenetic writers and readers, not metabolic sensors.

Our lifespan setRejuvenation-evidence set
Listed in the longevity-gene database11 of 110 of 6
Has aging-perturbation data behind itthe non-circular row1 of 116 of 6
The overlap between the two target sets, measured rather than assumed. The top row is circular by construction and is shown for completeness; the bottom row is the one that carries the argument.
1 of 11 vs 6 of 6
Targets with aging-perturbation evidence behind them

One gene sits in both programmes. That is why we describe our platform as covering 21 runnable targets rather than 22: eleven on the lifespan side, ten on the rejuvenation side, and one shared between them.

We tested whether our own model could rank real results, twice, and it could not

The obvious way to validate a longevity model is to ask whether it ranks compounds that worked above compounds that did not. Two large public programmes make that testable, so we pre-registered both analyses, wrote the specification into version control before running anything, and published what came back.

The mouse programme

Ranking the intervention-programme winners above its nulls gave an AUC of 0.645 with a 95% confidence interval from 0.447 to 0.826. The interval includes 0.50, so the result is null.

Underpowered by construction: at this sample size the design could only have reached significance at AUC 0.695 or better. It did not show the model fails. It failed to show the model works.

Ranking mouse-lifespan winners above nulls, n = 35. The point estimate sits above chance and the interval crosses it, which is why the number must never appear without the interval beside it.

The worm programme

A second, larger dataset in a different organism gave AUC 0.532, interval 0.391 to 0.671. Also null, and this time the more damaging number is the secondary one: the rank correlation against the continuous outcome was −0.0226 with a p-value of 0.796, across 133 compounds. The first test was underpowered by design. That one was not.

Our structure-based lifespan model has no demonstrated external ranking ability. We publish this on the product website because a result that only appears when someone asks is not really published.

The same question in a second organism, n = 133. Two pre-registered analyses, two datasets, two organisms, both null.

There is a detail from both runs we find more informative than either AUC. In each case the compound the model ranked most confidently was a compound that had been tested and had failed. Oxaloacetate topped the mouse list. Metformin topped the worm list. Same failure signature, twice, in two organisms.

What does hold up

The aging reference itself reproduces textbook direction on a pre-specified marker list: 22 of 24 canonical genes move the expected way with age in bulk data, p = 2 × 10⁻⁵, and 15 of 21 in single-cell data.

Bulk reference, canonical markers correct22 / 24
Single-cell reference, canonical markers correct15 / 21
Canonical aging markers moving in the expected direction. The aggregate signal is solid. What follows is the part that constrains how it can be used.

The most useful thing we found was a gap

When we scanned candidate targets against two criteria, having aging-perturbation evidence and having enough measured binders to anchor a design, most failed on the second. Ten genes carry real aging evidence and essentially no small-molecule chemistry at all. One transcription factor with strong human longevity genetics has zero binders in the public activity database above the usual potency threshold, and eight compounds ever tested against it in the public screening database.

We label that an opportunity rather than calling those targets undruggable, because a low binder count is a fact about the field's chemistry, not about the biology. It is also the clearest description of where de novo design is uniquely useful: it is the one method that needs no prior chemistry to start.

  • Three targets we investigated turned out to have no small-molecule chemistry at all when checked compound by compound across two databases. One had 33 actives on record and every single one parsed as a peptide.
  • We built family and pathway proxies for those three, showed they worked, and then rejected them. A ligand for a related protein is not chemistry for the target, and once mixed into a pool nothing downstream can tell which scaffold came from which protein.
  • They remain listed, marked as not runnable. The gap is the finding.

There is no regulatory indication called aging

Whichever programme a molecule belongs to, it eventually needs a named disease. No aging clock is an accepted surrogate endpoint, and the only agreed precedent is a composite of cardiovascular disease, cognitive impairment, cancer and death.

When we pulled live disease associations for our lifespan targets, ten of eleven had an age-related association, but eight of those resolved only to a generic neurodegenerative umbrella. Meanwhile the strongest association for most of them was oncology or a rare congenital syndrome. We report those in a separate column and never as the headline, because letting an oncology association stand in for a geroscience indication is exactly the inflation this layer exists to prevent.

What this changes about design

  • Carry both programmes and label which question each answers. We did not replace the lifespan list when the rejuvenation evidence pointed elsewhere. We kept both and made the interface say which one a candidate belongs to.
  • Do not show a lifespan number on a rejuvenation candidate. Those targets were admitted precisely because they are absent from the longevity-gene literature, so a lifespan-extension estimate is the other programme's endpoint wearing this one's label.
  • Say "designed toward" rather than "activator" or "inhibitor". Our activity signal is trained on binding data with no direction filter, so it scores affinity, not sign. The anchors do pull candidates toward the intended direction, which is a real bias worth stating, and it is not a measurement.
  • Expect the wet lab to be the arbiter. Two pre-registered nulls in two organisms is the field telling you that no public retrospective will settle this. One laboratory, one protocol, real animals.

The full evidence table, including everything above and the results that went against us, is on the benchmarks page. How candidates are actually assembled is covered in synthesizability by construction, and why we treat calibration as the deliverable is in the applicability domain problem.

About these numbers

Every figure here comes from a run on our own engine, and the measurement scripts sit in the repository beside the code they measure. Where a result is null, weak, or inconclusive we say so and publish the interval. Full detail is on the benchmarks page, and the design pipeline is written up in Methods. Source is available for audit on request.