Skip to content
The moat, benchmarked

Not a faster generator. A categorically different one.

We match the strongest unconstrained generators on standard optimisation while carrying a constraint none of them do, and where accuracy breaks down, we publish the number rather than the regime that flatters it.

Synthesis route
on every candidate
Novelty
measured, not asserted
ICH M7 screen
across route reagents
Confidence
calibrated interval
Everything we measure, on one screen
strong, weak and null, same table
Synthesis route coverage
every candidate is decoded from a verified-reaction recipe; chimeric routes are dropped
100%by construction
Genotoxic alerts (ICH M7)
6/6 benign stay clean · screened over route reagents at design time
12/12complete
Learned activity on novel chemotypes
ROC-AUC where similarity scores 0.626 · maxTanimoto < 0.25. The regime our molecules actually occupy · Δ +0.175 [+0.080, +0.272]
0.801beats similarity on 5 of 6
Reproducibility
same job id regenerates the same molecules exactly; distinct chemistry across runs (56 distinct of 60 slots over 6 runs)
exactdeterministic
Binder recovery
ROC-AUC · scaffold-disjoint · chance 0.50
0.945strong
Early enrichment @ 1%
BEDROC 0.90
≈ 20×strong
Sample efficiency
PMO AUC-top10 · QED oracle · 10k calls
0.938ties SOTA 0.94–0.95
Interval calibration
measured coverage of the 90% band
90%calibrated
Osimertinib MPO
PMO · kinase-like, geroscience-adjacent
0.790in-domain
Lifespan prior
leave-one-out Spearman · 775 compounds
0.158weak, disclosed
Novel-chemotype recovery
ROC-AUC on unseen scaffolds · chance 0.50
0.62collapses
ITP mouse-lifespan ranking
ROC-AUC · n = 35 · 95% CI [0.447, 0.826]
0.645inconclusive
MMC worm-lifespan ranking
ROC-AUC · n = 133 · 95% CI [0.391, 0.671]
0.532null

Bars for ROC-AUC measure skill above the 0.50 chance floor, not the raw value, drawn from zero, a coin flip would look half full. Detail and methodology for every row is below.

How to read this page
01 · What was measured

The exact task, and the metric, with its chance floor where one exists.

02 · What was excluded, and why

Held-out splits, dropped compounds, and analogue-exclusion rules, stated before the run.

03 · How confident

An interval, or a plain statement that we do not have one.

04 · Where it breaks

The regime the number does not cover. Every benchmark here has one.

How we compare
category-level · not named vendors

Where we lead, and where we don’t.

Compared against the two categories that matter commercially: large pharma-AI platforms, and structure-first generators. Open any row for the reasoning.

We compare against categories rather than named companies on purpose: we will not publish a competitor’s figures we cannot independently verify.

Capability
GeroQubit
Large pharma-AI
SMILES-first
Route-space (recipe) novelty
GeroQubit

Novelty measured on the synthesis recipe itself. The signal that maps onto freedom-to-operate rather than onto structural resemblance.

The field

Without a ground-truth recipe there is nothing to measure against, so recipe-level novelty cannot be computed at all.

Design-time genotoxic (ICH M7) screen
GeroQubit

DNA-reactive reagents are flagged before synthesis, including a genotoxin that is consumed during the route and absent from the final product.

The field

Product-only mutagenicity screening cannot see reagent or intermediate risk, which then surfaces in process chemistry.

Calibrated interval + published failure modes
GeroQubit

Every efficacy call carries a measured confidence interval, and the regimes where the model breaks are published alongside the ones where it works.

The field

Accuracy is typically reported for the regime that flatters it. The novel-chemotype collapse has been reported independently on six ADMET tasks, and it is rarely printed.

Geroscience-native (aging is the objective)
GeroQubit

Designs are specified on target × tissue × hallmark, so aging biology is the objective function rather than a filter applied afterwards.

The field

General-purpose platforms retarget oncology and kinase chemistry, where the data and the chemical matter both already exist.

Candidates in seconds, not overnight
GeroQubit

A full run finishes in seconds and reaches its top molecules in roughly 1,000 oracle calls, so iterating on a target costs an afternoon, not a quarter.

The field

Overnight training runs put cost and turnaround per campaign in a different bracket.

Competitive on standard optimisation (PMO)
Clinical assets and scale today

Not listed as a lead: every candidate we deliver carries a synthesis route. That is real, and it is not a differentiator — reaction-space generation with guaranteed routes is publicly available, so it separates us from structure-first tools and from nobody else. We would rather say so than count it twice.

yes / clear lead partial or late-stage only not offered
And the standard benchmarks hold up
PMO · QED oracle · 10,000-call budget

Competitive with the field's best, while every molecule stays makeable.

0.938
GeroQubit. AUC-top10, PMO QED
measured · synthesis-constrained
0.94–0.95
Reported SOTA cluster on QED
unconstrained · Gao et al., NeurIPS 2022

QED is a near-saturated oracle where the strongest unconstrained methods bunch at ~0.94–0.95; we land in that band carrying a constraint none of them do. We don't publish per-method competitor decimals we can't independently verify.

0.945
Binder recovery. ROC-AUC
Retrospective recovery of confirmed actives vs decoys, on a scaffold-disjoint held-out split. ~20× enrichment at 1%.
≈ 1,000
Oracle calls to 90% of peak
Reaches its top molecules well inside PMO's 10k budget, about 25 s per run.
in-domain
Osimertinib MPO 0.790
Our strongest oracle after QED is a kinase-inhibitor task. The geroscience-adjacent chemistry the library is built for.
The part nobody else prints
calibrated honesty · the trust layer

Every de novo tool collapses on truly novel molecules. We're the one that tells you.

Accuracy on novel chemistry is a known hard problem, usually reported only where it looks good. We report both regimes.

Distance from a target’s known binders

Similarity to compounds confirmed to bind. Lower = further from anything known.

Our molecules sit roughly ≈ 3× further from binder space than known binders sit from each other.

The aging references disagree with each other

Aging models are built on published old-vs-young expression references. We hold four, so we can ask what most groups cannot: when two of them cover the same gene, do they agree on which direction it moves with age? Across 13,784 genes, 55% are carried by a single source, and of the 6,171 where agreement can be assessed at all, they agree 49.7% of the time.

That includes the two targets our own lifespan programme is built around: CD38 (3 sources) and MTOR (4 sources) are both contradicted. We publish it because it constrains our numbers as much as anyone else’s.

Not a claim we found this first — the field is already asking whether aging clocks are ready, and peer-reviewed reviews conclude they are not validated as clinical surrogate endpoints. We add a measurement. And agreement is not correctness: two mouse-derived sources can agree because they share a bias.

“Novel scaffold” is mostly recombination — including ours

Generators report up to 98% novel scaffolds, measured on the Bemis-Murcko framework. For any engine that assembles molecules from catalogued parts that number cannot fail: couple a known scaffold to a known building block and the combined framework is new by construction. Ours is 82% by the same definition.

One level down the question can fail. Against a complete seed catalogue of 29,124 structures and 710 distinct ring systems, 94% of our molecules contain no ring system the catalogue did not already have. Only 6.3% carry an uncatalogued one.

We deliver a new arrangement of known rings, which is what a reaction does. This corrects our own earlier figure: we previously reported 27.5%, flagged as an upper bound because the seed enumeration was incomplete. Completing it moved the number to 6.3% — the old one was 4.4× too high, and the error flattered us. “Uncatalogued” means absent from our own seed lists, never unknown to chemistry.

We tried more physics. It lost.

A property of the regime, not our implementation. So we detect and disclose rather than claim it away.

Ranking accuracy on novel chemotypes, our own head-to-head
3D shape + electrostatics + desolvation0.572
Plain 2D structural similarity0.610

A 2026 benchmark found robust training and much larger models move frontier error by ≈ zero. Bars show skill above the 0.50 floor.

Bare scaffold
0.115
Finished candidate
0.159

Decoration is not the cause, building the molecule out moves it toward binder space, not away.

0.945
Binder recovery (ROC-AUC)
89% of the range above chance
confirmed actives vs decoys, scaffold-disjoint held-out · chance = 0.50
≈ 20×
Early enrichment @ 1%
BEDROC 0.90, actives at the top of the list
0.62
Novel-chemotype recovery
24% of the range above chance
unseen scaffolds, barely above the 0.50 coin flip, far below usable · disclosed
Pre-registered · 2026-08-06
NIA ITP · mouse lifespan

We tested the lifespan model against confirmed mouse negatives.

0.645
95% CI [0.447, 0.826]
ROC-AUC · 16 winners vs 19 nulls
Inconclusive · underpowered

Above chance, but the interval spans it. Nothing is demonstrated either way.

Measured
Rank 16 winners above 19 nulls
chance = 0.50
Excluded
Self + analogues ≥ 0.70
fixed before the run
Power
n = 35 could only pass at ≥ 0.695
80% power needs n ≈ 129
Likely cause
Trained ~90% worm/fly, judged on mouse
worm↔mouse ρ −0.03
−0.046
Leakage effect
Removing analogues raised the score, not propped up by memorised neighbours.
7th
Rapamycin's rank
Top-ranked of 35 is oxaloacetate, a confirmed negative.
python backend/validation/itp_retrospective.py
Reproduce it
Spec committed before the analysis existed. History available for audit on request.
Pre-registered · 2026-08-07
Ora MMC · worm lifespan

Then we ran it again, on four times the compounds.

The ITP test was too small to settle anything, so we took the largest open lifespan screen that publishes its negatives 160 rows from Ora Biomedical · Million Molecule Challenge (open data), and asked the same question in C. elegans. The method was fixed in writing before the data was touched.

0.532
95% CI [0.391, 0.671]
ROC-AUC · 23 winners vs 110 nulls
Null · no ranking ability shown

Barely above the coin flip, and the interval swallows it whole.

−0.023
Spearman · predicted vs actual % lifespan change · p = 0.796

This is the number that matters, and it is worse than the AUC. The AUC test was underpowered. We calculated that in advance and it could only have passed above 0.632. The correlation was not underpowered: across all 133 compounds our published in-house figure of ρ 0.158 should have shown up. It did not. What the model predicts and what the worms did are unrelated here.

Measured
Rank 23 winners above 110 nulls
chance = 0.50
Excluded
Self + analogues ≥ 0.70
fixed before the run
Power
Bound by 23 positives, not 133 total
passable only at ≥ 0.632
Top-ranked
metformin, a tested negative
rapamycin ranked 20th

MMC is a single-lab, fixed-dose, 20 °C screen, and its own nulls include several established geroprotectors. That widens what a null here can mean. It does not turn it into a pass. Raw extracted table and the scoring code are committed alongside the result, so the extraction can be audited independently of us.

Where that leaves the efficacy signal

Two pre-registered external tests. Two nulls. The lifespan model has no demonstrated ranking ability outside its own training set, and we publish that rather than wait to be asked.

It does not touch the rest of the platform: every candidate still carries a verified synthesis route, a novelty measurement and a design-time genotoxicity screen, and none of those depend on the lifespan model. But if you are here to ask whether we can predict what will extend life, today, honestly, the answer is that we have tried twice to show it and failed twice, and the next test that could settle it is a wet-lab one.

Honest scope

A front-of-funnel engine that de-risks a program, not a full pipeline.

We lead on makeability, novelty, safety-at-design, and calibrated honesty. We're not a clinical organisation, and our efficacy signals are computational and retrospective today, prospective wet-lab validation is the roadmap. On chemistry outside geroscience a specialised library scores low by design PMO DRD2 is 0.000, a dopamine-receptor task our building-block library was never built to reach. All disclosed, none hidden.

We have no longevity-specific benchmark. Every benchmark below is a borrowed one.

PMO, MOSES and the TDC oracles measure molecular generation in general. None of them measures whether a molecule affects ageing.

One ageing-specific benchmark now exists. LongevityBench, 17 tasks across 15 models, and it is built for large language models, which we are not. So it is not a leaderboard we can honestly claim to top, and we will not pretend otherwise.

If we publish one, it follows the rule the two null results already followed: the specification is committed before the result exists, so the test cannot be shaped to its own outcome.