Skip to content
Questions

The questions we get asked,
including the awkward ones.

Written for someone who does drug discovery but not aging, or aging but not chemistry. Every question opens with a plain answer; the science sits one click underneath. Where the honest answer is “we don’t know” or “we haven’t done that”, it says so.

The basics

What does GeroQubit actually do?

You pick one of eleven aging-related protein targets. It designs new small molecules aimed at that target and hands each one back together with the chemistry needed to make it — which building blocks to buy and which reactions to run.

11 targets · 10 tissues · 12 hallmarks

The longer answer

The unusual part is the order of operations. Most systems draw a molecule first and work out afterwards whether anyone can build it. Ours searches over recipes instead, so a molecule only exists in the output if a route to it already exists.

A design is specified on three axes rather than one: the target, the tissue you care about, and which kind of age-related damage you are trying to address. That combination is what makes it a geroscience tool rather than a general-purpose molecule generator.

Technical detail

The genome the search evolves is five integers — [scaffold, building-block 1, reaction 1, building-block 2, reaction 2] — decoded through verified reaction templates into a product. Selection is a four-objective Pareto front over composite quality, ADMET desirability, route quality and the lifespan prior, followed by a developability gate.

Who is it for?

Teams who need makeable starting matter for an aging programme and do not have a full computational chemistry group: longevity biotechs, geroscience labs with a validated target, and drug-hunting teams who want a ranked, route-bearing shortlist before spending on synthesis.

The longer answer

It is deliberately front-of-funnel. The job it does is turning "we believe this target matters in aging" into "here are candidate structures, each with a route, an ADMET profile and an explicit statement of how much to trust the numbers".

It is not a replacement for a medicinal chemistry team, a CRO, or an assay. It is the step before all three.

Why is designing for aging different from normal drug discovery?

In most programmes there is one disease and one target, and success means hitting that target hard. Aging is not one disease — it is a set of about twelve recognised kinds of damage, and the same protein can matter in one organ and be irrelevant in another.

11 targets · 10 tissues · 12 hallmarks of aging

The longer answer

So the useful question is not just "does this hit the target" but "does it hit the target, in the tissue that matters, for the damage type we care about". That is why every design here is aimed at a target, a tissue and a hallmark together.

The second difference is dosing. A geroprotector is a chronic, largely preventative medicine given to people who are not yet ill, so the safety bar is closer to a statin than to an oncology drug. That changes which trade-offs are acceptable — potency cannot buy back a liver signal — and it is why our scores multiply rather than average.

How it differs

How is this different from AlphaFold?

They solve different halves of the problem and do not compete. AlphaFold predicts what a protein looks like. GeroQubit designs a small molecule and the chemistry to make it. Neither one tells you what the molecule does in an animal.

The longer answer

Structure prediction has been genuinely transformative for knowing the shape of a target. It does not, on its own, tell you which molecule to make, whether that molecule can be made, or whether hitting the target helps.

We do not currently consume predicted structures at all, because our scoring is 2D and ligand-based. That is a deliberate limitation with a measured reason behind it — see the docking question below.

How is this different from traditional virtual screening?

Virtual screening ranks molecules that already exist in a catalogue. We generate molecules that do not exist yet, along with the route to make them. The trade-off is real: catalogue compounds can be ordered next week, ours have to be synthesised.

Binder recovery 0.945 on known chemistry · 0.62 on genuinely novel scaffolds (chance = 0.50)

The longer answer

The reason to accept that trade-off is coverage. Purchasable libraries are enormous but heavily biased toward chemistry that has already been made, which for most aging targets means the chemistry of a different, better-funded therapeutic area.

The reason to be sceptical of it is that generated molecules carry more uncertainty, and we quantify exactly how much: our own ability to tell a real binder from a decoy degrades sharply the further a molecule sits from known chemistry, and we publish that curve rather than hiding it.

How is this different from other generative chemistry models?

Most generative models emit a molecule as a text string or a graph and then ask a separate retrosynthesis tool whether it can be made. We never produce a molecule without a route, because the route is what the search is searching over.

Most generators
Design a moleculeSearch for a routeHope one exists

Synthesisability is an outcome to be checked after the fact — and sometimes there isn’t one.

GeroQubit
Pick real building blocksApply a verified reactionMolecule + route together

The route is an input to the search, so a molecule cannot exist in the output without one.

The longer answer

Designing molecules as synthesis routes is not unique to us — it is an active research area with several published methods, and we would be wrong to claim we invented it.

What we have not found anyone else doing is the combination: designing on a target-tissue-hallmark basis for aging specifically, screening for genotoxic risk across the reagents in the route rather than only the final compound, and publishing the regime where our own predictions stop working.

A reactive fragment on a starting material is a real handling and impurity risk even when it never appears in the product. Only a system that knows the route can see it.

Technical detail

Because the recipe is ground truth rather than a post-hoc guess, we can compute novelty in route space — a weighted edit distance over the five-slot recipe. A structure-first generator has no ground-truth recipe and so cannot compute this quantity at all.

Why don’t you use docking?

Partly because it needs a reliable 3D structure and a reliable pose, and for several aging targets that is not available. Mostly because we tested more physics on our own output and the simpler method won.

3D physics 0.572 vs plain 2D similarity 0.610, on novel chemotypes

The longer answer

We built a 3D shape-and-electrostatics comparison — including desolvation — and put it up against plain 2D structural similarity on exactly the kind of novel molecules we produce. The 2D method scored higher.

We are not claiming docking is useless. It clearly earns its place in many programmes, and running one docking pass over our final selections is still on our roadmap. We are reporting what happened in our own hands, on our own output, and being guided by it. Our honest expectation is that it fails again, and we intend to publish that too.

Technical detail

USRCAT + ElectroShape + a desolvation term scored 0.572 against plain 2D Tanimoto at 0.610 on novel chemotypes. This matches a broader pattern in the literature: on out-of-distribution chemistry, added model capacity and robust-training objectives have moved frontier error by approximately nothing.

The molecules

If you build from known reactions and catalogue building blocks, are the molecules actually new — or just recombinations?

This is the fair question. We are not inventing new chemistry — the reactions are established and the building blocks are purchasable — but the molecules that come out are genuinely far from known compounds, and we measure that rather than asserting it.

Similarity to a target’s known binders — ours 0.14–0.17, real binders to each other 0.41–0.53

The longer answer

When we compare our output to the compounds already known to bind a given target, our molecules sit about three times further away than those known compounds sit from each other. In chemistry terms they are not close analogues of anything in the reference set.

There is a real cost to that distance and we publish it: the further a molecule sits from known chemistry, the worse our ability to predict whether it will work. That trade-off is the honest centre of this whole approach, and it is why the platform reports a "frontier distance" next to every candidate.

One thing this rules out, incidentally, is the most common failure of aging-focused AI tools — quietly rediscovering rapamycin, metformin and quercetin and presenting them as discoveries.

How is novelty measured?

Two different ways, because they answer different questions. Structural novelty is how unlike known compounds the molecule is. Route novelty is how unlike the other recipes in the batch its recipe is.

The longer answer

Structural novelty is one minus the highest similarity to any compound in our reference sets of known actives and known lifespan-extending compounds. It is the number that matters for "is this a new chemotype".

Route novelty matters for a different reason: two molecules can look different and still be made the same way, which means they will fail the same way at the bench and are weak as a portfolio. Route novelty is the signal that catches that, and it maps directly onto freedom-to-operate thinking.

Technical detail

Structural novelty = 1 − max Tanimoto over 2048-bit Morgan fingerprints against DrugAge, CellAge and Open Targets actives. Route novelty = weighted edit distance over the five-slot recipe, with a scaffold swap weighted highest. Structure matching is run on standardized molecules — skipping standardization misses roughly 44% of true matches, which we found the hard way.

Can a chemist actually make these? Has anyone checked?

Every candidate carries a one- or two-step route built from verified reaction templates with named reagents, so a route exists by construction. No route has yet been run at a bench, and that is the check we most want.

The longer answer

"By construction" is a strong claim about the route existing and a weak claim about it working first time. Yield, selectivity and purification are not things our system predicts, and any real synthesis may need a chemist to change the conditions.

We have also found and fixed the failure mode where a template produces chemistry that cannot run: a reaction named for one bond that actually forms another, an aromatic substitution proposed on a ring that would not undergo it. Those are now gated and asserted rather than assumed, because they survived being read and only failed when run.

We are actively asking contract research organisations to judge sample routes as though they were incoming orders. A negative answer there would be decisive, and cheap.

Technical detail

Templates carry a reagent profile keyed to the template identifier, and name↔bond↔reagent agreement is asserted rather than commented. Aromatic substitution requires a verified activating group ortho or para to the leaving group; meta does not activate. Any template lacking a name is refused outright, so an unaudited reaction cannot re-enter the pool.

Can molecules generated this way be patented?

We are not lawyers and this is not legal advice. In general, a genuinely novel and non-obvious compound is patentable in most jurisdictions regardless of whether software helped design it — but machine involvement raises inventorship questions that are still being litigated, and you should get counsel.

The longer answer

The two practical points we can speak to are factual rather than legal. First, novelty: our output sits far from known chemistry by measurement, and we give you that number rather than an assurance. Second, prior art: we check each candidate against PubChem, and we report when a structure appears to be already known.

A caution on that second check, from our own history. It once reported "known compound (CID 0)" for genuinely novel molecules because of a truthiness bug in how the identifier was tested, which made novel matter look unoriginal. It is fixed, and our published credibility figure was re-baselined downward as a result rather than quietly left high.

The route-novelty signal is the one most directly relevant to freedom-to-operate conversations, because it describes how a compound is made rather than only what it is.

The science

What data is this built on?

Public, citable sources: GenAge and CellAge for aging genetics, Open Targets for human genetic association, DrugAge for measured lifespan effects, ChEMBL for bioactivity, GTEx for tissue expression, and the López-Otín hallmark framework for the aging axis.

The longer answer

Nothing proprietary and nothing scraped. That is partly principle and partly practicality: a claim you cannot point someone at is not a claim they can check.

The building-block and scaffold pools are assembled from approved-drug chemistry and fragmented public compounds, with structural-alert filtering applied before assembly rather than after.

Technical detail

Target prioritisation uses Open Targets genetic association scores combined with small-molecule druggability. That is genetic ASSOCIATION, not proven causation — Mendelian randomisation is a higher bar that most of these targets do not clear, and we say so in the product rather than only here.

Why DrugAge, and what is wrong with it?

Because it is the only curated, public database of compounds with measured lifespan effects — there is no better option. Its limitation for our purpose is that the data is dominated by worms and flies, and that turns out to matter a great deal.

The longer answer

To be clear about the framing: this is not a criticism of DrugAge, which does exactly what it says and is a genuine service to the field. It is a statement about what happens when we use it as a training label for a mammalian question.

Our own analysis found that the maximum-lifespan-extension label we train on is roughly 90% invertebrate. Invertebrate effect sizes run about three times larger, so when a compound has been tested in several species the maximum is almost always the worm or fly number and the mouse data is discarded by the aggregation.

We measured whether the invertebrate signal transfers. Across compounds tested in both, the rank correlation between worm and mouse effects was about zero, with an interval wide enough to include a strong effect in either direction. So we cannot show that it transfers, and we cannot show that it does not.

Technical detail

Worm↔mouse Spearman −0.03, 95% CI [−0.51, +0.42], n = 24. This is the leading candidate explanation for the null ITP result below, and the fix — a clade-split predictor — was attempted and did not deploy, because the mammal-only correlation crossed zero.

How are lifespan predictions estimated, and how good are they?

By similarity: we find the compounds in the reference set that most resemble the candidate and take a weighted average of their measured lifespan effects. It is a weak signal, and we report it as one — with an interval around every number.

Leave-one-out ρ 0.158 — weak, and stated as weak

The longer answer

The honest summary is that the fitted noise is roughly twice the signal. We say so in the interface rather than only in the methods, because a number with that much noise around it is dangerous if it is presented cleanly.

The interval is not a guess. It is calibrated so that the true value lands inside it about nine times out of ten, and we measured that coverage on held-out compounds rather than assuming it.

This prediction is deliberately kept out of the design loop. Letting it steer generation would pull every run toward things that already look like known geroprotectors, which is exactly the novelty trap the platform exists to avoid. It is computed afterwards, for display and for final ranking only.

Technical detail

Tanimoto-weighted k-NN regression over 775 DrugAge structures. Leave-one-out Spearman ρ = 0.158; conformal interval coverage 90%. This figure was re-baselined downward from 0.182 after we found that our own standardizer had collapsed several metal salts onto identical structures, handing those compounds a perfect-similarity neighbour. Removing that inflation lowered the headline and we kept the lower number.

What is the score next to each molecule, and what is it not?

It is a computed prior — a summary of how well a molecule matches what we are aiming for, combined from drug-likeness, ease of synthesis, resemblance to known actives, and predicted absorption and safety. It is not a measurement, a probability of working, or a prediction of lifespan.

A directional prior, not a measurement

The longer answer

No molecule we have designed has ever been made or tested, so no score on this platform has ever been checked against reality. That sentence is the most important one on this page.

The axes are combined by multiplying rather than averaging, on purpose. Averaging lets a molecule be excellent on one axis and poor on another and still look good. Multiplying means one weak axis drags the whole score down, so potency cannot buy back bad safety.

Because that multiplication includes three "does it look like a known binder" axes, a genuinely novel molecule scores low almost by definition. That is why the interface also shows a developability grade computed over only the similarity-independent axes — it answers "is this makeable, drug-like and safe?" separately from "does it resemble something already known?". A candidate can honestly be poor on the first number and strong on the second, and we show both rather than picking the flattering one.

Limits & honesty

Do you have any laboratory results?

No. Nothing designed by this system has been synthesised or tested in any organism or assay. We have no wet lab.

The longer answer

We put that on the homepage rather than in the footnotes, because it is the first thing a scientist should know and the last thing anyone should have to discover for themselves.

What we do have is a synthesis route for every candidate, which means the step from a design to a real compound is a purchase order rather than a research project.

Closing this gap is the roadmap, and the cost driver is synthesis rather than assay. Worm healthspan screening runs a few thousand dollars per compound; first-time synthesis of a novel molecule is the expensive part, and notably the price is close to flat between ten milligrams and five hundred, because labour dominates.

Has your prediction ever been tested against real animal data?

Yes. We wrote down how we would judge the test before running it, ran it against confirmed mouse-lifespan failures, and published the result — which was inconclusive. We cannot show the model picks winners, and we do not claim it does.

AUC 0.645, 95% CI [0.447, 0.826] — above chance in point estimate, but the interval spans it, so nothing is demonstrated either way

The longer answer

The US National Institute on Aging runs a programme that tests compounds for lifespan extension in mice at three independent sites and — unusually — publishes its failures as well as its successes. That gives a set of real, replicated negatives, which is the part most predictors never face.

We asked whether our model ranks the compounds that worked above the compounds that did not, scoring each compound with its own record and its close analogues removed. It scored above chance on the central estimate, but thirty-five compounds is a small set and the uncertainty around that estimate comfortably includes chance. So the test settles nothing in either direction: we cannot show the model has this ability, and neither did we show it lacks it. Being underpowered is not the same as failing, and we are careful not to describe it as either.

Two details matter more than the headline. Removing close analogues made the score go up rather than down, so the result is not propped up by the model recognising compounds it had already seen. And the ranking is unconvincing in its specifics: the top-ranked compound of all thirty-five is one that failed in mice, while rapamycin — the most firmly established geroprotector there is — lands seventh.

The likeliest explanation is our own label rather than our method, for the reason given in the DrugAge question above.

Technical detail

Primary analysis: ROC-AUC 0.645, 95% CI [0.447, 0.826], n = 35 (16 winners, 19 nulls), self and Tanimoto ≥ 0.70 analogues excluded. Secondary (self excluded only): 0.599 [0.405, 0.789]. Analogue-exclusion effect on AUC: −0.046. The specification — method, decision rule and exclusion criteria — was committed to version control before the analysis code existed, so it could not have been chosen to suit the outcome; that history is available for review on request. Reproduce with `python backend/validation/itp_retrospective.py`.

What are the limitations, stated plainly?

No wet-lab validation of anything we have designed. A weak lifespan signal that failed its hardest test. Predictions that degrade sharply on exactly the novel chemistry we produce. Direction of effect that is designed for, not measured.

The longer answer

That last one deserves unpacking, because it is the least obvious. When a programme is labelled as aiming to activate or inhibit a target, that describes the design intent: the reference compounds behind it all act in that direction and they drive the scoring. But the activity score itself measures resemblance to known binders without distinguishing activators from inhibitors. Nothing in the pipeline predicts functional direction, and only an assay can confirm it.

We take that seriously enough to have renamed a programme over it. One target was labelled as an activator while its reference set was led by two well-known inhibitors, which pulled candidates toward the opposite chemotype. Rather than relabel and move on, we found that the validated activator class for that target is essentially a single compound, so there is no honest five-compound activator set to build — and the programme is now named for the pathway instead of claiming a direction its references cannot support.

On chemistry outside geroscience, a specialised library scores low by design. On one standard benchmark task involving an unrelated receptor we score zero, and we publish that rather than omitting the row.

Does this replace experiments, or medicinal chemists?

Neither, and a tool that claimed otherwise would be worth distrusting. It replaces the blank page at the start of a programme, not the bench and not the chemist.

The longer answer

Everything this system produces is a hypothesis. The scores are priors computed from public data, not measurements, and the failure mode of this whole category of tool is that a clean-looking number gets treated as evidence.

What it genuinely saves is the expensive early loop: enumerating chemistry that turns out to be unmakeable, or unsafe in a way that only surfaces once someone tries to order the reagents. Catching a DNA-reactive starting material at design time is cheap; catching it in process chemistry is not.

The judgement calls — is this series worth pursuing, is this route sensible, is this target right — are chemistry and biology decisions, and we build the interface on the assumption that a person makes them.

How should these predictions be validated?

In this order: have a chemist review the routes, make a handful of compounds, run a target-engagement assay, then a cellular phenotype, and only then anything in an animal. Do not skip to lifespan.

Evidence ladder — where this platform actually stands today
  1. Route exists by constructionEvery candidate, by design
  2. Chemist review of routesIn progress with CROs
  3. Compound synthesisedNot yet
  4. Target engagement assayNot yet
  5. Cellular phenotypeNot yet
  6. Animal lifespanNot yet
Each rung should kill as many candidates as it can before the next, because the cost rises steeply. We are on rung two.
The longer answer

The reason for that order is cost asymmetry. A chemist rejecting a route costs an afternoon. A binding assay costs less than a synthesis. A worm healthspan run costs a few thousand dollars. A mouse lifespan study costs years. Each step should kill as many candidates as it can before the next one.

The most informative early experiment is not the most flattering one. Target engagement on a small set — including the candidates our own scores rank poorly — tells you far more about whether the platform works than testing only the top-ranked molecule.

If you want a decisive negative cheaply: send three routes to a contract research organisation as if they were incoming orders and ask whether they would accept them.

What would convince you that this does not work?

Compounds we designed being synthesised and showing no activity against their intended target. That is the test we most want to run.

The longer answer

Short of that: a chemist telling us the routes do not survive contact with a real bench. We are asking contract research organisations to judge sample routes as though they were incoming orders, precisely because a negative answer there would be decisive and cheap.

We have a track record of publishing the answers that did not go our way — a pre-registered mouse-data test that came back inconclusive, a physics-based scoring method that lost to a simpler one, and an accuracy figure we revised downward after finding a bug in our own data. That is the standard we intend to be held to.

Practical

Why does it run on ordinary processors instead of GPUs?

Because there is no neural network to train. The engine is a search over recipes, and the scoring is standard cheminformatics and closed-form calculation.

The longer answer

The practical consequence is cost. There is no cluster and no training bill, so the money in this project goes to chemistry rather than to compute. A full design run finishes in seconds on a laptop-class processor.

The honest counterpoint is that this is a constraint as well as a feature: it rules out approaches that genuinely need scale. We think that trade is correct for this problem, given that more model capacity has repeatedly failed to fix the frontier problem that actually limits us.

How do I get access, and what does it cost?

Access is by request while the platform is in research use. Pricing is per programme rather than per seat, because the unit of work is a target, not a user.

The longer answer

We would rather talk to you about the target than sell you a login. If your target is not among the eleven, tell us the mechanism — new programmes are scoped individually and the target list is not a hard limit.

Something not answered here? The awkward questions are the useful ones — ask us directly, or read the full methods and the benchmarks, including the one we failed.