A good docking score is not binding. Binding is not potency. Each layer must be independently confirmed — and this discipline is what separates a real candidate from a computational artifact.
Principle 01
Distrust single metrics
No one score decides. Binding prediction, dynamic stability, and free energy each say something different — and any of them can reverse the verdict.
Principle 02
Independent cross-validation
We confirm with methods built on different principles. Synthesizability is checked several ways; free energy is computed by two mathematical routes at once.
Principle 03
Prediction meets measurement
We compare predictions against real measurements and correct the next prediction with the gap. The platform grows more accurate the more it is used.
What the gates actually do
Three million molecules in. Nineteen out.
A recent multi-target campaign in neurodegeneration, stage by stage.
Structure-based generation
3,000,000
Passed novelty filter
424,387
Passed property & CNS suitability
34,847
Docked against all three targets
34,755
Strong binding at every target
1,645
Completed AI binding prediction
1,644
Top tier, carried forward to extended dynamics
19
Every stage is an independent gate, and each one uses a different principle from the last. The overall pass rate is 0.0006%. We treat that number as a measure of rigor, not a weakness — a pipeline that passes most of what enters it is not filtering anything.
Scored against published experiments
We do not run experiments. We are graded by them.
Everything below is a case where the answer already existed in the literature and we went at it anyway. Some we got. Some we did not, and those are listed with the rest. Our own pipeline targets stay undisclosed; the scoring is done on public benchmarks, so anyone can repeat it.
01Imatinib resistance mutations, predicted blindFor six known resistance mutations in Abl kinase, we computed how far each one shifts drug binding. The measured answers were taken from ChEMBL and hashed shut before any calculation began, then opened once the predictions were fixed. Classification was correct in 4 of 5; mean error against experiment was 0.32 kcal/mol across 4 mutations — the same spread two laboratories get measuring the same mutation.
02And where that run fell shortThe sixth mutation was a proline substitution, which does not converge under our current protocol; we excluded it rather than report it. Our discrimination limit on this method is 0.31 kcal/mol, and we do not claim to separate differences smaller than that. One of the five calls was wrong, at the boundary between neutral and mild — which is where our resolution runs out.
03A prediction sealed before the answer was openedPredictions on published structures were recorded and closed before we looked at the measured values. One landed within 0.3 percentage points. One missed, predicting 4.3 points low. Both are here, because a record that holds only hits is not a record.
04A known binder, found againA compound with experimentally confirmed binding went through the procedure as a control. Nine runs, nine calls, all correct. Nothing unknown enters the pipeline until the known answer comes back out of it.
05A crystal structure held with no restraintsFrom crystallographic coordinates, 100 ns of dynamics with nothing restrained. Every interaction the structure paper identifies survived. Under restraints it would have held by construction; without them, it is evidence the system was built correctly.
06Bilayer thickness against scattering dataWithin 0.2% of the measured value. This is the first quantity to drift when a system is assembled wrongly, which is why it is the first thing we score. We re-check it on every build and do not run one that misses.
07The point where our method stops seeingTwo compounds whose measured cooperativity differs 34-fold are not separated by our membrane metric; the gap sits inside ±6.6 percentage points of measurement uncertainty. We publish the limit, not only the reach.
08Our own metrics, audited for convergenceAcross twelve trajectories, one of four metrics met our convergence criterion in all of them. The other three failed in more than half. Those three are not used for ranking — by us, on our own work.
09A full log audit after a reported bugWhen a bug was reported in software we rely on, all 264 stored logs were checked against it rather than the recent ones. None were affected. The audit is on file.
We have found errors in our own pipeline, withdrawn the conclusions that rested on them, and recomputed. A validation system working does not mean errors are absent. It means errors are found.