Skip to content
Research notes

What we measured, including what did not work.

Longer write-ups on the parts of de novo longevity drug design that are hard to get right. Every number here comes from a run on our own engine, and the results that went against us are in the same posts as the ones that did not.
Method9 min

The applicability domain problem: why AI drug discovery models fail exactly where novelty lives

A model that scores 0.945 on a benchmark scored 0.62 on our own molecules. Same model, same target, same metric. Here is what changed, and why every generative chemistry team should measure it.

New here

Our own leave-one-out and generated-molecule measurements: known binders sit 0.41-0.53 from each other, our output sits 0.14-0.17 from them, and recovery AUC falls 0.945 to 0.62 on that slice. Also the negative result that 3D physics lost to 2D fingerprints in this regime.

Method10 min

Synthesizability by construction: design the route, not the molecule

If your generator emits SMILES, it cannot tell you how the molecule was made, because it never knew. Evolving the recipe instead of the structure changes what the model can and cannot say.

New here

A measured reaction-search result (+22.9% products per reaction evaluation, 95% CI [+16.2%, +33.1%], 24/24 paired runs) and a worked failure: an azole coupling template that could never fire, invisible because a dead template produces no wrong molecules, only missing ones.

Biology11 min

Lifespan targets are not rejuvenation targets, and the evidence separates cleanly

Every gene on our lifespan list appears in the longevity-gene database. None of the rejuvenation-evidence genes do. That separation is measurable, and it changes which question a molecule can answer.

New here

A measured, non-circular separation between the two target sets (1 of 11 versus 6 of 6 carry aging-perturbation data) plus two pre-registered null results published with their confidence intervals, which is the part almost nobody publishes.

Why these read the way they do

Publishing where we fail is the point. Two of our own validation studies returned null results and both are written up here with their confidence intervals, because a number quoted without its interval is a different claim from the one the data supports. If a figure on this site looks weaker than a competitor's, that is usually because it was measured on the hard case.