Not a faster generator. A categorically different one.
We match the strongest unconstrained generators on standard optimisation while carrying a constraint none of them do, and where accuracy breaks down, we publish the number rather than the regime that flatters it.
Bars for ROC-AUC measure skill above the 0.50 chance floor, not the raw value, drawn from zero, a coin flip would look half full. Detail and methodology for every row is below.
The exact task, and the metric, with its chance floor where one exists.
Held-out splits, dropped compounds, and analogue-exclusion rules, stated before the run.
An interval, or a plain statement that we do not have one.
The regime the number does not cover. Every benchmark here has one.
Where we lead, and where we don’t.
Compared against the two categories that matter commercially: large pharma-AI platforms, and structure-first generators. Open any row for the reasoning.
We compare against categories rather than named companies on purpose: we will not publish a competitor’s figures we cannot independently verify.
Route-space (recipe) novelty
Novelty measured on the synthesis recipe itself. The signal that maps onto freedom-to-operate rather than onto structural resemblance.
Without a ground-truth recipe there is nothing to measure against, so recipe-level novelty cannot be computed at all.
Design-time genotoxic (ICH M7) screen
DNA-reactive reagents are flagged before synthesis, including a genotoxin that is consumed during the route and absent from the final product.
Product-only mutagenicity screening cannot see reagent or intermediate risk, which then surfaces in process chemistry.
Calibrated interval + published failure modes
Every efficacy call carries a measured confidence interval, and the regimes where the model breaks are published alongside the ones where it works.
Accuracy is typically reported for the regime that flatters it. The novel-chemotype collapse has been reported independently on six ADMET tasks, and it is rarely printed.
Geroscience-native (aging is the objective)
Designs are specified on target × tissue × hallmark, so aging biology is the objective function rather than a filter applied afterwards.
General-purpose platforms retarget oncology and kinase chemistry, where the data and the chemical matter both already exist.
Candidates in seconds, not overnight
A full run finishes in seconds and reaches its top molecules in roughly 1,000 oracle calls, so iterating on a target costs an afternoon, not a quarter.
Overnight training runs put cost and turnaround per campaign in a different bracket.
Not listed as a lead: every candidate we deliver carries a synthesis route. That is real, and it is not a differentiator — reaction-space generation with guaranteed routes is publicly available, so it separates us from structure-first tools and from nobody else. We would rather say so than count it twice.
Competitive with the field's best, while every molecule stays makeable.
QED is a near-saturated oracle where the strongest unconstrained methods bunch at ~0.94–0.95; we land in that band carrying a constraint none of them do. We don't publish per-method competitor decimals we can't independently verify.
Every de novo tool collapses on truly novel molecules. We're the one that tells you.
Accuracy on novel chemistry is a known hard problem, usually reported only where it looks good. We report both regimes.
Similarity to compounds confirmed to bind. Lower = further from anything known.
Our molecules sit roughly ≈ 3× further from binder space than known binders sit from each other.
Aging models are built on published old-vs-young expression references. We hold four, so we can ask what most groups cannot: when two of them cover the same gene, do they agree on which direction it moves with age? Across 13,784 genes, 55% are carried by a single source, and of the 6,171 where agreement can be assessed at all, they agree 49.7% of the time.
That includes the two targets our own lifespan programme is built around: CD38 (3 sources) and MTOR (4 sources) are both contradicted. We publish it because it constrains our numbers as much as anyone else’s.
Not a claim we found this first — the field is already asking whether aging clocks are ready, and peer-reviewed reviews conclude they are not validated as clinical surrogate endpoints. We add a measurement. And agreement is not correctness: two mouse-derived sources can agree because they share a bias.
Generators report up to 98% novel scaffolds, measured on the Bemis-Murcko framework. For any engine that assembles molecules from catalogued parts that number cannot fail: couple a known scaffold to a known building block and the combined framework is new by construction. Ours is 82% by the same definition.
One level down the question can fail. Against a complete seed catalogue of 29,124 structures and 710 distinct ring systems, 94% of our molecules contain no ring system the catalogue did not already have. Only 6.3% carry an uncatalogued one.
We deliver a new arrangement of known rings, which is what a reaction does. This corrects our own earlier figure: we previously reported 27.5%, flagged as an upper bound because the seed enumeration was incomplete. Completing it moved the number to 6.3% — the old one was 4.4× too high, and the error flattered us. “Uncatalogued” means absent from our own seed lists, never unknown to chemistry.
A property of the regime, not our implementation. So we detect and disclose rather than claim it away.
A 2026 benchmark found robust training and much larger models move frontier error by ≈ zero. Bars show skill above the 0.50 floor.
Decoration is not the cause, building the molecule out moves it toward binder space, not away.
We tested the lifespan model against confirmed mouse negatives.
Above chance, but the interval spans it. Nothing is demonstrated either way.
Then we ran it again, on four times the compounds.
The ITP test was too small to settle anything, so we took the largest open lifespan screen that publishes its negatives 160 rows from Ora Biomedical · Million Molecule Challenge (open data), and asked the same question in C. elegans. The method was fixed in writing before the data was touched.
Barely above the coin flip, and the interval swallows it whole.
This is the number that matters, and it is worse than the AUC. The AUC test was underpowered. We calculated that in advance and it could only have passed above 0.632. The correlation was not underpowered: across all 133 compounds our published in-house figure of ρ 0.158 should have shown up. It did not. What the model predicts and what the worms did are unrelated here.
MMC is a single-lab, fixed-dose, 20 °C screen, and its own nulls include several established geroprotectors. That widens what a null here can mean. It does not turn it into a pass. Raw extracted table and the scoring code are committed alongside the result, so the extraction can be audited independently of us.
Two pre-registered external tests. Two nulls. The lifespan model has no demonstrated ranking ability outside its own training set, and we publish that rather than wait to be asked.
It does not touch the rest of the platform: every candidate still carries a verified synthesis route, a novelty measurement and a design-time genotoxicity screen, and none of those depend on the lifespan model. But if you are here to ask whether we can predict what will extend life, today, honestly, the answer is that we have tried twice to show it and failed twice, and the next test that could settle it is a wet-lab one.
A front-of-funnel engine that de-risks a program, not a full pipeline.
We lead on makeability, novelty, safety-at-design, and calibrated honesty. We're not a clinical organisation, and our efficacy signals are computational and retrospective today, prospective wet-lab validation is the roadmap. On chemistry outside geroscience a specialised library scores low by design PMO DRD2 is 0.000, a dopamine-receptor task our building-block library was never built to reach. All disclosed, none hidden.
We have no longevity-specific benchmark. Every benchmark below is a borrowed one.
PMO, MOSES and the TDC oracles measure molecular generation in general. None of them measures whether a molecule affects ageing.
One ageing-specific benchmark now exists. LongevityBench, 17 tasks across 15 models, and it is built for large language models, which we are not. So it is not a leaderboard we can honestly claim to top, and we will not pretend otherwise.
If we publish one, it follows the rule the two null results already followed: the specification is committed before the result exists, so the test cannot be shaped to its own outcome.
