The applicability domain problem: why AI drug discovery models fail exactly where novelty lives
A model that scores 0.945 on a benchmark scored 0.62 on our own molecules. Same model, same target, same metric. Here is what changed, and why every generative chemistry team should measure it.
Our own leave-one-out and generated-molecule measurements: known binders sit 0.41-0.53 from each other, our output sits 0.14-0.17 from them, and recovery AUC falls 0.945 to 0.62 on that slice. Also the negative result that 3D physics lost to 2D fingerprints in this regime.
