methods & evaluation

The evidence behind DiscoveraVS

DiscoveraVS is built around computational methods that must be evaluated, bounded and interpreted in context. This page documents how each method is tested, what evidence supports its use, and where its outputs should guide prioritisation rather than be read as experimental truth.

panel built 2026-08-16 · 15 targets · 78,125 training records

what is evidenced here

Four evidence domains

criteria & alerts

Ligand-based screening

How compound sets are narrowed using physicochemical constraints, structural alerts and ligand-derived criteria when no target structure is required.

docking interpretation

Structure-based screening

How docking and pose analysis are used to rank compounds against a defined molecular target, and how docking scores should be interpreted.

model validation

Off-target liability prediction

How supervised models are trained and evaluated to flag structural patterns associated with selected off-target liabilities.

assay follow-up

Experimental validation

How computational prioritisation was followed by compound acquisition and experimental assay in an internal validation campaign.

The sections below document the off-target panel, its validation protocol and how its outputs are read. They also define how docking scores, liability flags and experimental outcomes should be interpreted. The screening campaign is on the product page.

protocol

How model performance was evaluated

Each off-target model was evaluated with five-fold cross-validation using scaffold and similarity-aware splitting. Compounds were grouped by Bemis–Murcko scaffold and by molecular similarity cluster, then whole groups were assigned to folds, so that closely related analogues were less likely to appear in both training and test sets.

This design is stricter than a random split. Random splits can inflate performance when analogue series are shared across training and test data; scaffold and similarity-aware folds give a more demanding estimate of how the model behaves on less familiar chemistry.

For each target, thresholds and out-of-fold predictions were derived only from the cross-validation folds. The deployed model is then retrained on all available data for that target. Operating thresholds are selected from out-of-fold predictions to maximise MCC, not from predictions made on the final training set.

the metric, and why

MCC is the primary metric because the off-target datasets are imbalanced across the 15 targets, with prevalence ranging from 0.613 to 0.868. MCC combines true positives, true negatives, false positives and false negatives into a single balanced statistic, which makes it more informative than accuracy when one class dominates.

Accuracy can mislead when the data are imbalanced, because a model can look good numerically by predicting the majority class. ROC-AUC measures ranking quality across thresholds but does not define an operating threshold. Balanced accuracy, precision and recall are therefore read alongside MCC to show class-specific and threshold-dependent behaviour.

Because scaffold and similarity-aware folds can differ in difficulty, the cross-validation spread is kept beside each MCC rather than hidden behind a single mean. Per-target results are not published on this page for now; the panel is being retrained.

applicability

Flags are reported, not confidence-weighted

DiscoveraVS exposes no confidence score for off-target predictions, and that is a result rather than an omission. On the previous panel, where it was measured, out-of-fold accuracy was flat across similarity to the training set: the within-band spread was far larger than anything between the bands. Familiarity did not predict whether a prediction was right.

A confidence derived from it would have done one thing reliably — discount the flag on an unfamiliar compound, which on a liability panel is the failure that matters. The platform therefore reports each above-threshold prediction as a flag to review, rather than weighting it by apparent confidence.

For each compound, DiscoveraVS reports flagged targets and near misses separately. Compounds can be ranked by the number of flagged liabilities, while the target-level decisions stay visible for inspection and follow-up.

what you get per compound

  • Target score. A model score for each of the 15 off-target panel members.
  • Flag. A target-level call when the score is above that target’s own operating threshold — 0.434 on one target and 0.897 on another.
  • Near miss. A below-threshold score falling within the predefined near-miss margin for that target, set in proportion rather than as a fixed distance, so each target gets a warning band on its own scale.
  • Compound ranking. Compounds are ranked by the number of flagged targets. Near misses are shown separately, so they inform review without being counted as full flags.
  • No confidence weighting. Flags are not weighted by score distance, chemical familiarity or an exposed confidence estimate.

interpretation

How to interpret the outputs

  • Docking scores rank poses within a target. Docking engines return empirical scores, often reported in kcal/mol-like units, but these are not measured binding free energies. Use them to compare poses and compounds within the same target and protocol, not as direct estimates of thermodynamic affinity or cross-target potency.
  • Off-target flags support liability triage. The panel estimates whether a compound matches structural patterns associated with activity against selected liability targets. A flag should prompt review, redesign, deprioritisation or experimental follow-up. It should not be read as a safety verdict.
  • The panel is intentionally scoped. It covers 15 targets selected for data density and clinical relevance — receptors, transporters and channels relevant to early risk assessment. It is not complete safety coverage: hERG and the CYPs are not in it and are handled elsewhere in your process. Further targets are in development.
  • The evidence has two layers. The model metrics report retrospective out-of-fold performance under scaffold and similarity-aware cross-validation. Separately, an internal screening campaign connected computational prioritisation to experimental testing — 11.3 billion compounds down to 8 bought and 2 confirmed active.

reproducing this

Trace results back to the model manifest

Each DiscoveraVS model release includes a manifest recording the evaluation split, target-specific thresholds, performance metrics, software versions and panel composition. Predictions can therefore be interpreted against the exact model version and validation protocol that produced them. Technical documentation describes how screening runs store inputs, parameters, software versions, intermediate files and ranked outputs.

Request technical documentation

reproducibility

The model manifest documents what was evaluated. The screening run record documents how a project was executed. Together they make predictions and project outputs traceable from model version to compound-level result.

models/offtarget_panel2/manifest.jsonships with access →