Designing molecules against an aging state, and reporting what the controls killed.
A complete description of the engine as it runs today: how aging is represented, how candidates are assembled so a synthesis route exists by construction, which search operators survived their own adversarial controls, and which did not.
Molecule generation has outpaced its own credibility, and geroscience adds a second problem on top: there is no endpoint to optimise that anyone can measure in under three years. GeroQubit takes a different object as the thing being searched. Aging is carried as a 22-component state, 10 tissue components superposed with 12 hallmark components, built from 15,705 genes with a measured aging direction across 4 independent datasets. Nine quantum-derived operators read that state in a single pass, and the molecule is the actuator rather than the thing scored. Candidates are assembled as reaction recipes rather than emitted as structures, so a route exists before the molecule does. We report one operator with a measured steering effect, a +22.9% improvement in products per reaction evaluation with a 95% confidence interval of [+16.2%, +33.1%] over 24 paired runs, and we report in the same detail the three operators and one entire subsystem removed after their pre-declared controls failed to separate them from chance. Five pre-registered retrospective tests of our efficacy signal returned null or inconclusive, and all five are described here rather than omitted.
Most of the previous note describes a system we no longer run.
This is a rewrite rather than a revision. The parts that changed are the parts the old document was built on, so there was nothing structural left to amend.
| Component | v1.0 | v1.1 |
|---|---|---|
| Aging representation | An 8-gene signature in a 128-D embedding | A 22-component state over 15,705 genes |
| Search layer | None. A genetic algorithm on fitness alone | Nine quantum-derived operators reading the state in one pass |
| Reaction choice | Fixed template order | Ranked by measured template applicability. +22.9% more products per evaluation |
| Targets | 11 lifespan targets | 21 runnable across two programmes, with quality tiers published |
| Efficacy claim | A k-NN signal reported as the efficacy signal | Five pre-registered retrospectives, all null or inconclusive, published |
| Reproducibility | Not tested | Byte-identical across processes after two string-hash defects were fixed |
The aging state
Aging is not one number and we stopped producing one. The engine carries a state vector whose two axes are held in superposition rather than averaged, so a candidate that helps liver and does nothing for brain stays legible as exactly that, all the way to the end of the run. Collapsing it to a scalar is a measurement we defer until a human asks for one.
Every tissue is covered in the thousands, thymus included. That last one is worth a sentence: the curated aging database we started from returns zero thymus genes, so this platform ran with no thymus signal at all until a second, cell-type-resolved source was merged in.
That fix reached the aging direction reference and not the tissue clock. The clock that actually scores a molecule in thymus is synthesised from the immune clock by four hand-typed offsets and contains no thymus measurement. So thymus has 2,605 genes of aging direction behind it and zero behind its clock, and a thymus score is labelled derived-from-immune wherever we show one. The honest fix is thymus data, not better constants.
The twelve are not equally well observed, and we do not draw them as though they were. Chronic inflammation has 2,972 genes behind it and cellular senescence has 99.
Sources are merged by precedence rather than by vote, and we tested whether voting would be better. Of 6,175 genes with a direction from two or more sources, the signs agree on 3,070 and clash on 3,105, which is a coin flip. Keeping only the concordant genes improved the downstream metric, and so did keeping only the discordant ones, so the effect was coverage rather than agreement and the agreement claim was withdrawn.
Nine quantum operators, acting on one state
Each operator is a map on the aging state rather than a score appended to a molecule, which is the whole reason there are nine of them and not one weighted sum. They run as a single integrated pass, and none is a separate model bolted onto the side of the generator. Each acts at sampling or at selection; none touches a gate, a score, or the final ranking.
Advances the state under a candidate intervention.
Where separate interventions land on the same state.
Reads which gate a candidate died on, per molecule.
Novel and still inside the evidence we have.
Keeps the ten tissue components from collapsing together.
The covariance among the twelve hallmark channels.
Distance from the aged state toward the young one.
Which reaction template to try first on the fragment in hand. Applicability is whether the template APPLIES to the structure, not whether the reaction runs in a flask.
Combinations, rather than one molecule at a time.
One of the nine, QRC, is allowed to steer the engine, and it is opt-in. The other eight inform sampling and selection only. This is a deliberate constraint: every time a good signal has been folded into the fitness function in this project it has regressed the result, and the same signal applied at selection has shipped the gain.
Reaction-first assembly
The engine evolves recipes, not structures. A genome is five integers naming a scaffold, two building blocks and two reactions, and decoding it produces the molecule and the route together. A generator that emits a structure first has to guess a route afterwards, and sometimes there is not one.
Because the reactions are inside the genome, the route is not a post-hoc reconstruction. It is the thing that was optimised. Both assembly steps must take a real reaction template; a candidate where both fell back to a direct bond is dropped, because its route would be fictitious.
That refusal is the feature. A template whose name and reagent profile disagree with the bond it forms prescribes a route that cannot run, and we shipped one for months that had never once fired on anything.
QRC, projecting chemistry onto the reactions that can run
The template-applicability operator, and the only one of the nine granted a steering role. It projects the functional-group state of a fragment onto the reactions that state can support, and tries those first. Nothing is skipped and the reachable product set is identical; only the order changes. That constraint is deliberate, because an operator that changed what is reachable would be changing the chemistry rather than the search.
77 reaction templates ranked by measured success rate over 127,946 recorded decode attempts. 29,047 products, 22.7% overall. 16.1% of all attempts went to templates that have never once produced a product.
Held out on 9,173 reactant pairs never seen in training. The scale starts at chance, because a bar drawn from zero makes a coin flip look half competent.
We first scored these predictors on log loss and the hand-written rule table came out below the per-reaction averages, which cannot be right for a table that encodes real chemistry. The rule emits two probabilities and log loss punishes that calibration rather than the ranking, so the comparison was measuring the wrong thing. Repeated on a rank metric, which is what the engine actually consumes.
more products per reaction evaluation
95% CI [+16.2%, +33.1%]
- Pairs
- 24 (8 targets × 3 seeds)
- Record
- 24 won, 0 tied, 0 lost
- Reaction calls
- −4,675 [−5,479, −3,887]
- Molecules delivered
- unchanged (23 ties)
Separate process per arm, hash seed pinned, and the persisted caches snapshotted and restored around every arm. This is an efficiency result about a search. It says nothing about whether any molecule works in an animal.
What we do not claim from this run: a potency improvement. The peak interval excludes zero, but there are two regressions in the 24 pairs and the mean is carried by one of them. Scaffold diversity moved in a direction whose interval includes zero, so it is not a finding either way.
Tensor recipe search
A pathology we measured in our own engine: giving it more compute produced fewer and worse molecules. The cause was that the search optimises a softer feasibility surface than the verifier enforces, so a population can sit inside a region where everything is rejected and feel no gradient out of it.
Two low-rank tensor surrogates over the five-slot recipe: one for quality, one for the probability a candidate survives the hard gates.
The second is trained on the rejections the search used to discard, which was the largest unused signal in the system.
Forced to rank one, it degenerates into the per-slot bandits that failed here four times before. Interaction between slots is the whole mechanism, and it is testable rather than asserted.
One target is excluded from it by default, because it regressed there. A mechanism that adds exploitation helps a search that is stuck and hurts one that has already converged, and after five instances of that pattern we treat a lever that helps only a few targets as failing rather than as promising.
What the controls removed
Three operators and one entire subsystem were built, measured against controls declared before the measurement, and deleted. The slot each occupied is empty rather than filled with a third variant, because trying variants until one passes is how a null becomes a result.
Beat its permutation control only at grid resolutions that collapsed 25 molecules into 2–16 bins. Exactly 1.000 wherever they stayed distinct.
Bandwidth read off the data. Still below the control: effective n 35.62 against 38.06, −23.6 sd.
Held-out gain +0.0448 on pairs already seen and +0.0000 on unseen pairs. All of it was recall. Demoted to a cache.
Grover diffusion measured classical at shipped strength; the anti-collapse operator fired 4 times in a whole production run; the interference register gave byte-identical output on 8 of 9 paired runs. 0 wins.
The subsystem in the last row was 625 lines of production code with a scoreboard of zero wins, three measured negatives and seven mechanisms never tested at all. It is preserved at a version tag so every experiment stays reproducible by checkout, but negative results belong in a ledger and not in executable code.
What is quantum here, stated precisely
A molecule is amplitude-encoded, and the similarity kernel is built from the one-body reduced density matrix marginals of that state. The derivation is genuinely quantum. The object it produces is not, and we are the ones who measured that.
The projected quantum kernel is a radial basis function on L1-normalised count fingerprints. One scikit-learn call reproduces it to machine precision. Measured against the standard library implementation on our shipped reference set, the largest absolute difference is 1.44e-15.
Its +39% improvement over plain Tanimoto is real and stays. The cause is count weighting and a nonlinearity, both of them classical. The tell sat in our own documentation for months: the genuinely quantum forms gave no gain, and only the projected one worked.
Feature co-occurrence is invisible to any function of the one-body marginals alone. If anything here earns the name, it is this term. It moves a rank metric independently by +0.0112, 95% CI [+0.0083, +0.0145], and costs two matrix products through an identity that never forms the full density matrix.
Two-body structure has now been tested three times in this project. It failed twice: in one regime the sample size makes a null mean the coupling cannot be estimated, and in the other it does not. This is the single win, and we quote the tally rather than the win.
We describe this as quantum-derived and classically computed, and never as quantum computation or quantum advantage. A company named GeroQubit calling a classical kernel quantum would be exactly the inflation the rest of this document exists to prevent. The whole system runs on ordinary processors with no accelerator.
A density-matrix applicability domain
Our own retrospective binder recovery reaches 0.945 on held-out actives and collapses to 0.62 on genuinely novel chemotypes. That collapse is the regime we operate in by design, so the useful lever is detecting it rather than claiming to have solved it. Every applicability measure in common use is a distance to a nearest neighbour and is therefore blind to direction; treating the reference set as a density matrix is not.
Our molecules sit roughly three times further from binder space than real binders sit from each other. That distance is what a novel chemotype is, and it is also why any similarity-based score reads low on our output. We publish the low number and explain it rather than loosening the threshold that produces it.
Spectral leakage is the fraction of a query state lying outside the subspace the reference spans. About 60% of one of our molecules points where the reference has no mass, against about 20% for a real binder.
The same construction says a 200-compound reference set spans an effective rank of 8.45. Two hundred compounds are not two hundred independent directions, and nothing else we ship could say so.
A tie at n = 6 with no interval, so it is quoted as a tie. It earns its place as a description, not as a detector, and it never gates anything.
This collapse is not unique to us. It has been independently reported on six public ADMET tasks, and the robustness methods tested in that work did not resolve it. That report is a preprint on one domain, so it cannot carry a stronger word than this one, and neither can we.
Two programmes, and why they are not one list
Extending lifespan and pushing cells younger are different questions on different evidence. We measured the separation rather than assuming it, and it is close to total.
| Lifespan | Rejuvenation | Note | |
|---|---|---|---|
| In GenAge (known longevity genes) | 11 / 11 | 0 / 6 | circular. The 11 were selected from this literature |
| Has ageing-perturbation data | 1 / 11 | 6 / 6 | |
| Status | shipped | all ten rejuvenation-only targets run |
The top row is circular and we say so: the lifespan targets were selected from the longevity gene literature, so of course they are all in a longevity gene database. The rows that carry the finding are the other two. Nothing about how the rejuvenation targets were chosen referenced that database, and the perturbation data row was measured rather than selected on. Exactly one gene, SIRT6, appears in both programmes, which is why 11 and 11 give 21 runnable targets and not 22.
A flat count would hide that three sit under the anchor floor and two carry chemistry caveats, so the tier travels with the number.
Anchors buildable, and the available chemistry does not contradict the direction.
Real aging evidence, under the 50-binder floor a reliable anchor set needs.
Its best available anchors are KAT6A compounds, so they are not really its chemistry.
Thousands of binders and every one an inhibitor, while the biology asks for activation. Carried as a pathway, with no direction claimed.
No small-molecule chemistry exists at all. They stay visible because the gap is the finding, and no job can be launched against them.
You raise HIF-1α by inhibiting the prolyl hydroxylase, so EGLN1 is the target that actually runs.
A target with four human safety datasets, and an indication nobody pointed it at.
CXCR2 signalling reinforces cellular senescence from outside the cell: it is the receptor arm of the senescence-associated secretory phenotype (Acosta et al., Cell 133:1006, 2008). Four antagonists have been taken into humans against it. Every one was developed for lung disease or cancer, and every one was dosed systemically.
The recurring problem across the class is neutropenia. Blocking CXCR2 impairs neutrophil trafficking, so it is a consequence of hitting the target rather than of any particular molecule. A new structure does not fix it, and we do not claim it does.
Neutropenia limited the dose in every trial. The exposure that causes it is body-wide.
Osteoarthritis allows dosing directly into the knee. UBX0101 established that route in this exact indication, whatever else that trial did not show.
CXCR2 antagonism plus local joint delivery has not been tested. That is the open question, and it is a question rather than a result.
All three are 1-aryl-pyrazole-4-carboxamides built in two steps from commercial starting materials. The engine evolves the recipe, so the route exists before the molecule does rather than being proposed afterwards.
All three clear the hard gates: hERG and AMES both below 0.5, zero Rule-of-Five violations. DILI is read against a 0.6 floor rather than 0.5, because the model scores many safe approved drugs high on it.
GQ-SM-06 already exists in a make-on-demand catalogue. It is listed as 1-[1-(3-fluorophenyl)-1H-pyrazole-4-carbonyl]piperazine, supplied by Enamine through Chemspace as CSMS01090540378, at 90% purity with a 30-day lead time to India.
So the usual objection to a computational candidate does not apply to this one. Testing it against CXCR2 starts at $190 for 10 mg, with no synthesis step at all. The other two are not in any catalogue we can find and would need to be made.
This does not soften the novelty figure. The 0.185 is maximum Tanimoto to known CXCR2 binders: it says nobody has tested this chemotype against this target, not that nobody has ever made the molecule. Both are true at once.
| 1 mg | $163 |
| 2 mg | $166 |
| 5 mg | $175 |
| 10 mg | $190 |
| 20 mg | $220 |
Chemspace CSMS01090540378
List price only. Excludes shipping and duties, and a make-on-demand listing is a synthesis promise rather than a shelf item.
The obvious claim would be that these are structurally novel, so they escape what stopped the others. The measurement does not support it.
The four failed drugs are barely more similar to each other than ours are to them. CXCR2 antagonists are already a structurally diverse class, and all of them failed, so a new scaffold is not the argument. The argument is the indication and the route of administration, which is a different thing and is testable.
These are designs. There is no synthesis, no assay, no animal data, and no wet lab behind them.
That liability follows from blocking the target. Nothing here addresses it and nothing here should be read as addressing it.
The local-dosing argument is about the indication, not about these molecules. We do not model intra-articular pharmacokinetics, and two of the three carry basic amines that would clear a joint quickly.
Ten tissues are modelled and none of them is cartilage or synovium. These were designed in whole-body and immune context.
The activity score is similarity to known binders, and a target’s binder set contains both agonists and antagonists.
Similarity to known binders collapses on novel chemotypes by construction, which is why overall reads low. That is the measurement working, not a hidden problem.
Reproducibility
The engine did not reproduce itself across fresh processes, and we did not know until we checked. Three runs of an identical configuration returned three different answers.
The obvious suspect, floating point addition across threads, was not the cause. It was string hashing, which the language randomises per process by design.
- 1
BRICS.BRICSDecompose returns a set of SMILES strings, so building-block order, and therefore what a genome’s integer means — changed with every process.
- 2
A bond chosen by hash((smiles_a, smiles_b)), under a comment claiming two runs on the same pair produce the same bond. They did not.
Historical A/B deltas are unaffected: every rig in the repository already pinned PYTHONHASHSEED=0, so both arms of every comparison saw a deterministic engine. Absolute peak and yield values from before the fix do not reproduce and are not quoted as current.
Results, including the five that failed
Five retrospective tests of the efficacy signal were pre-registered, with the specification recorded before the result in every case. Four returned null and one is inconclusive. None of them is omitted here, and none of them is described as trending.
n = 35. A Hanley-McNeil calculation done afterwards shows the design could only ever have reached significance at AUC 0.695, so it was underpowered by construction. That does not rescue the model; it means this particular test can no longer settle it.
n = 133. The damaging figure is the secondary one, not the AUC: Spearman against the continuous outcome is -0.0226 at p = 0.796, and that arm was not underpowered. Ranking the same compounds by another laboratory’s measured lifespan scores no better, so the test had no headroom left in it.
The first run of this read p < 0.0001 and replicated beautifully. It was an artifact: the two contrasts being correlated share their Old group, and all 593 caloric-restriction records came from one study. Dropping that study collapsed the result. When two contrasts share a group, correlating them is not a test.
A real bug was found and fixed on the way: the harness took the median of six age brackets and correlated six points, discarding about 98% of the samples. Fixing it moved the number the predicted way and it is still null.
Not underpowered: every interval could have resolved a 0.60 effect and five of nine resolve a departure from chance. The diagnosis is exact, and it is in section 11.
Test five has an exact answer. Our aging reference returns the median change between an old group and a young group, which is a population-level contrast, and ordering individuals needs an individual-level weight. Fitting weights to individuals on the same samples, the same genes and the same protocol moves the score from 0.483 to 0.676, above chance in 45 of 45 donor-disjoint splits.
Every published aging clock fits weights to individuals. None reads them off a group difference. That is a category error rather than a tuning problem, and it explains all five tissues that sat below chance.
- Aggregate marker-level use of the reference, which reproduces textbook aging direction on a pre-specified gene list at p = 2e-5.
- Synthesizability by construction, which depends on no efficacy model.
- Route novelty, genotoxic screening at design time, and developability, for the same reason.
The efficacy signal itself reports a leave-one-out correlation of 0.158 over 775 compounds. That is weak in absolute terms and we say so wherever it appears.
Limitations
Stated in full, because a limitation a reader finds for themselves costs more than one they were handed.
No candidate this system has designed has been synthesised or assayed. We have no wet laboratory. Every number on this page is computational.
Two pre-registered retrospectives on two organisms found no demonstrated external ranking ability. The binding constraint is cross-dataset label disagreement rather than sample size, so a third public database would not settle it.
Every model here is static. The aging state is a difference between an old group and a young group, with no rate and no relaxation time, because we have no longitudinal data to attach one to. Claiming a dynamical result from static contrasts would be inflation.
A tissue coupling matrix and a hallmark conservation table are hand-assigned and uncited. Their directions reflect real consensus; their decimals are not measured and are never published.
One has a measured steering effect and one survives its statistical controls. The rest inform the search without a paired experiment demonstrating that they help.
One of the four aging sources is licensed for academic use only and we are commercial. It covers 5% of entries, no tissue depends on it for coverage, and it is the only one of the four that passes its own validation test. That is a business conversation, not a code change.
Data, code and availability
Validation draws on public datasets. The engine internals do not ship.
DrugAge for measured lifespan, ChEMBL for bioactivity, GTEx for tissue expression, Aging Atlas and Tabula Muris Senis for aging direction, Open Targets and GenAge for target evidence, and the López-Otín hallmark framework.
Scoring weights, the search operators, the generator encoding and the reaction template set are proprietary. Aggregate validation results are reported in full.
Representative candidate structures with their routes, under a research agreement. The code is available for audit on request.
No human or animal subjects were involved. This platform is a research tool for hypothesis generation and is not intended for clinical, diagnostic or therapeutic use. This document has not been peer reviewed.
- 1.López-Otín C, Blasco MA, Partridge L, Serrano M, Kroemer G. Hallmarks of aging: an expanding universe. Cell. 2023;186(2):243–278.
- 2.López-Otín C, Blasco MA, Partridge L, Serrano M, Kroemer G. The hallmarks of aging. Cell. 2013;153(6):1194–1217.
- 3.Gems D, de Magalhães JP. The hoverfly and the wasp: a critique of the hallmarks of aging as a paradigm. Ageing Res Rev. 2021;70:101407.
- 4.Flórez-Ablan M, Roth M, Schnabel J. On the similarity of bandwidth-tuned quantum kernels and classical kernels. arXiv:2503.05602. 2025.
- 5.Barardo D, et al. The DrugAge database of aging-related drugs. Aging Cell. 2017;16(3):594–597.
- 6.Harrison DE, et al. Rapamycin fed late in life extends lifespan in genetically heterogeneous mice. Nature. 2009;460:392–395.
- 7.Swanson K, et al. ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries. Bioinformatics. 2024;40:btae416.
- 8.Gao W, Fu T, Sun J, Coley CW. Sample efficiency matters: a benchmark for practical molecular optimization. NeurIPS. 2022.
- 9.Bickerton GR, Paolini GV, Besnard J, Muresan S, Hopkins AL. Quantifying the chemical beauty of drugs. Nat Chem. 2012;4:90–98.
- 10.Ertl P, Schuffenhauer A. Estimation of synthetic accessibility score of drug-like molecules. J Cheminform. 2009;1:8.
- 11.Rogers D, Hahn M. Extended-connectivity fingerprints. J Chem Inf Model. 2010;50:742–754.
- 12.Deb K, Pratap A, Agarwal S, Meyarivan T. A fast and elitist multiobjective genetic algorithm: NSGA-II. IEEE Trans Evol Comput. 2002;6:182–197.
- 13.Oh HS-H, et al. Organ aging signatures in the plasma proteome track health and disease. Nature. 2023;624:164–172.
- 14.Pridham G, Rutenberg AD. Network dynamical stability and biological aging. J Gerontol A Biol Sci Med Sci. 2024;79(10):glae021.
- 15.Kedlian VR, Dönertaş HM, Thornton JM. The widespread increase in inter-individual variability of gene expression in the human brain with age. Aging. 2019;11(8):2253–2280.
- 16.Pyrkov TV, et al. Longitudinal analysis of blood markers reveals progressive loss of resilience and predicts human lifespan limit. Nat Commun. 2021;12:2765.
- 17.Sen P, et al. Histone acetyltransferase p300 induces de novo super-enhancers to drive cellular senescence. Mol Cell. 2019;73(4):684–698. PMID 30773298.
- 18.Freund A, Patil CK, Campisi J. p38MAPK is a novel DNA damage response-independent regulator of the senescence-associated secretory phenotype. EMBO J. 2011;30(8):1536–1548. PMID 21399611.
- 19.Acosta JC, et al. Chemokine signaling via the CXCR2 receptor reinforces senescence. Cell. 2008;133(6):1006–1018. PMID 18555777.
- 20.Wang Q, et al. Aging Atlas: a multi-omics database for aging biology. Nucleic Acids Res. 2021;49(D1):D825–D830.
- 21.Tabula Muris Consortium. A single-cell transcriptomic atlas characterizes ageing tissues in the mouse. Nature. 2020;583:590–595.
The original methods note. It is kept available because withdrawing a document people may have read is worse than labelling it, but it describes an earlier engine: an eight-gene signature instead of a state, no search layer, and an efficacy figure that has since been re-baselined downward after we found a data bug that had inflated it. Read this page instead.
↓ v1.0 PDFThis page as a typeset preprint: abstract, numbered sections, tables and references, formatted for submission rather than for the browser. Six pages.
↓ v1.1 preprint PDFBenchmarks, the pre-registered nulls in full, and every interval are on the benchmark page.
