Parameter Reference
Complete reference for every parameter in Hedgehog’s YAML configuration files.
config.yml
The main configuration file. Controls input/output paths, parallelism, and references to all stage-specific configs.
| Parameter | Type | Default | Description |
|---|---|---|---|
generated_mols_path | string | src/hedgehog/configs/examples/moses_1000.csv | Path to the CSV file containing generated molecules |
target_mols_path | string | src/hedgehog/configs/examples/target_mols.csv | Path to the CSV file containing reference molecules |
alignment.enabled | bool | false | Calculate numeric thresholds from target_mols_path before evaluating generated molecules; structural policy remains fixed |
alignment.target_coverage_percent | float | 95 | Percentage of target molecules the generated thresholds retain at each stage; must be greater than 0 and at most 100 |
folder_to_save | string | results/run | Output directory where all pipeline results are saved |
n_jobs | int | -1 | Number of parallel workers for CPU-bound tasks (-1 = all available cores. Prefer an explicit smaller number such as 4 or 8.) |
sample_size | int | 10000 | Number of molecules to sample from the input file (null = use all) |
save_sampled_mols | bool | true | Whether to save the sampled molecule subset to disk |
evaluate_docked_coordinates | bool | false | For SDF inputs, additionally evaluate docked coordinates; requires docking.run: true and at least one selected docking tool |
large_dataset_mode | bool | false | Enable streaming chunked processing for very large pre-docking dataset statistics |
large_dataset_chunk_rows | int | 250000 | Rows per processing chunk in large dataset mode |
large_dataset_single_csv_limit | int | 1000000 | Maximum row count for also materializing compatibility CSV files from shard outputs |
large_dataset_output_format | string | csv.gz | Shard file format for large dataset row-level intermediate tables |
large_dataset_filter_data | bool | false | In large dataset mode, whether filter pass/fail results should remove molecules from downstream outputs |
large_dataset_enable_all_filters | bool | true | In large dataset mode, enable configured descriptor/structural filters as calculations even when they do not filter outputs |
ligand_preparation_tool | string | (proprietary path) | Absolute path to an external ligand preparation binary |
protein_preparation_tool | string | (proprietary path) | Absolute path to an external protein preparation binary |
config_mol_prep | string | src/hedgehog/configs/config_mol_prep.yml | Path to the preprocessing stage config |
config_descriptors | string | src/hedgehog/configs/config_descriptors.yml | Path to the descriptors stage config |
config_structFilters | string | src/hedgehog/configs/config_structFilters.yml | Path to the structural filters stage config |
config_synthesis | string | src/hedgehog/configs/config_synthesis.yml | Path to the synthesis stage config |
config_docking | string | src/hedgehog/configs/config_docking.yml | Path to the docking stage config |
config_docking_filters | string | src/hedgehog/configs/config_docking_filters.yml | Path to the docking filters stage config |
config_weighted_score | string | src/hedgehog/configs/config_weighted_score.yml | Path to the weighted model assessment config |
config_moleval | string | src/hedgehog/configs/config_moleval.yml | Path to the MolEval reporting config |
Target-based alignment is controlled directly by the master config:
target_mols_path: data/reference_molecules.csv
alignment:
enabled: true
target_coverage_percent: 95
# descriptor_bounds_mode: expand # keep source borders; widen only where targets need it
# descriptor_bounds_mode: target # replace each border with the target-derived value
descriptor_bounds_mode: expandWhen enabled, HEDGEHOG calculates numeric thresholds with respect to the target molecules before evaluating generated_mols_path. Structural hard filters are not percentile-calibrated. Descriptor alignment supports descriptor_bounds_mode: target, which replaces each configured descriptor border with the target-derived border, and descriptor_bounds_mode: expand, which applies min(source_min, target_min) and max(source_max, target_max). In expand, generic ZINC+ChEMBL descriptor envelopes are preserved and can only become wider. If descriptor_bounds_mode is omitted, expand is used. The target structural probe evaluates every calculate_* filter and every common-alert ruleset without removing molecules. Synthesis is skipped during target calibration and does not participate in target-cohort selection; the candidate uses the source synthesis config unchanged, including SA/RA/SYBA thresholds and retrosynthesis settings. The calibration probe may calculate every structural method for diagnostics, but the generated candidate config copies the source filter_* flags, Common Alerts enforcement lists, and structural parameters unchanged. Structural pass masks do not participate in target-cohort selection or the target_coverage_percent guarantee. structural_filter_failures.csv records every measured failure. Set enabled: false to use the stage configs without recalculation. --align-config PERCENT enables alignment and overrides target_coverage_percent for one run. Generated aligned master configs set alignment.enabled: false to prevent recalibration when reused, and add the target-run and threshold-audit paths as provenance. Each aligned stage config is created immediately after its alignable stage completes. Continuous thresholds use outward rounding, while integer-valued thresholds are rounded up.
config_mol_prep.yml
Preprocessing. Standardizes molecules before any descriptor computation. This stage aims to produce “clean” molecules by:
- removing salts and solvents and keeping the largest fragment
- disconnecting metals and normalizing/reionizing structures
- preserving formal charge by default (
steps.standardize_mol.uncharge: false) - preserving stereochemistry by default (
steps.remove_stereochemistry: false) - applying atom, radical, and single-fragment filters while retaining isotope labels
General Settings
| Parameter | Type | Default | Description |
|---|---|---|---|
run | bool | true | Enable or disable preprocessing |
n_jobs | int | -1 | Worker count for molecule preparation |
steps.standardize_mol.uncharge | bool | false | Neutralize formal charges during MolPrep |
steps.remove_stereochemistry | bool | false | Remove atom/bond stereochemical annotations |
filters.allowed_atoms | list[string] | [C, N, O, S, F, Cl, Br, I, P, H, Si] | Allowed atom symbols |
filters.require_single_fragment | bool | true | Reject multi-fragment molecules |
filters.reject_radicals | bool | true | Reject molecules with radical electrons |
filters.reject_isotopes | bool | false | Reject isotopically labeled molecules instead of preserving them for later stages |
output.write_duplicates_removed | bool | true | Write duplicates_removed.csv when duplicates are dropped |
config_descriptors.yml
Descriptors. Controls descriptor calculation and filtering borders. Plot presentation has code defaults and only needs configuration for custom reports.
General Settings
The descriptors stage always computes the complete descriptor table. When filter_data is true, configured filters are applied as follows:
bordersdefine generic descriptor ranges such asmolWt,logP,TPSA,hbd,hba,n_rings, andfsp3.structural_constraintsare converted into additional upper bound checks on derived descriptor columns.
Use borders for the production filtering contract. The optional legacy-compatible structural_constraints block can add motif caps when its own enabled field is true.
Optional Plot Overrides
Normally these fields should be omitted: plotted columns are derived from borders, while discrete-column metadata and labels come from code defaults (DESCRIPTOR_DISPLAY_NAMES and related constants in src/hedgehog/descriptors/constants.py). There is no YAML renamer.
| Parameter | Type | Description |
|---|---|---|
filtered_cols_to_plot | list[string] | Descriptor columns to include in distribution plots |
discrete_features_to_plot | list[string] | Columns treated as discrete |
Default borders are the ZINC250k + ChEMBL34 calibration envelope (q0.29–q99.71). See Descriptors.
config_structFilters.yml
Structural Filters. Calculates structural alerts and medicinal-chemistry diagnostics. calculate_<name> controls diagnostics; filter_<name> independently controls Stage 3 survival.
General Settings
| Parameter | Type | Default | Description |
|---|---|---|---|
run | bool | true | Enable or disable the structural filters stage |
n_jobs | int | 16 | Worker count for parsing and every structural filter (-1 = all available cores; falls back to master n_jobs when omitted) |
filter_data | bool | true | Whether to apply the configured hard-filter decision to downstream molecules |
filter_NIBR | bool | true | Include the published NIBR severity policy in survival |
filter_molgraph_stats | bool | true | Include MolGraph severity policy in survival |
filter_common_alerts | bool | true | Enforce the configured Common Alerts subset |
include_rulesets | all | list | null | all | Rulesets to calculate: all = full catalog, []/null = none, or an explicit list |
exclude_smarts | list[string] | [] | Exact SMARTS strings dropped from calculation |
common_alerts_filter_include_rulesets | list | [PAINS] | Rulesets enforced when filter_common_alerts is true; empty = every calculated ruleset |
common_alerts_filter_exclude_rulesets | list | [] | Calculated rulesets excluded from survival; exclusions win |
filter_lilly | bool | true | Include Lilly (cutoff from lilly_demerit_cutoff) in survival |
filter_protecting_groups | bool | true | Reject curated protecting-group matches |
filter_molcomplexity, filter_bredt, filter_ring_infraction, filter_halogenicity, filter_symmetry | bool | false | Optional hard-filter flags; calculations remain independent |
filter_stereo_center | bool | false | Optionally enforce the total-stereocenter diagnostic cutoff |
filter_undefined_stereo_center | bool | true | Reject underspecified structures above stereo_max_undefined; reuses the stereo-center calculation |
lilly_demerit_cutoff | int | 160 | Lilly demerit threshold passed as dthresh |
nibr_max_severity | int | 10 | Inclusive NIBR rejection threshold; pass requires accumulated severity below this value |
molgraph_max_severity | int | 5 | Inclusive MolGraph rejection threshold; pass requires maximum pattern severity below this value |
write_per_filter_outputs | bool | true | Write per-filter output folders and CSVs |
write_structural_liability_profile | bool | true | Write one molecule-level table containing the hard decision and every diagnostic field |
generate_plots | bool | true | Generate structural filter plots |
generate_failure_analysis | bool | true | Generate failure-analysis outputs |
ring_infraction_hetcycle_min_size | int | 4 | Largest small-ring cutoff checked by the ring-infraction rule |
stereo_max_centers | int | 4 | Diagnostic cutoff for total stereocenters; the total count is reported independently of hard enforcement |
stereo_max_undefined | int | 2 | Inclusive maximum for undefined stereocenters; values above it fail the undefined-stereo hard policy |
halogenicity_thresh_F | int | 6 | Inclusive fluorine-count limit |
halogenicity_thresh_Br | int | 3 | Inclusive bromine-count limit |
halogenicity_thresh_Cl | int | 3 | Inclusive chlorine-count limit |
symmetry_threshold | float | 0.8 | Inclusive maximum MedChem symmetry score |
Target-aware runs keep the complete generic structural gate. Alignment audits cannot change explicit filter_<name> flags. Any future target-specific relaxation must be an explicit, separately justified exception for a concrete hard rule. Scheduler settings such as lilly_scheduler remain execution controls.
Structural Filter Profiles
Three shipped configs share the same diagnostic calculations. Hard gates differ. The former balanced profile is removed.
| Profile | Config | Hard gate summary |
|---|---|---|
| default | config_structFilters.yml | PAINS + NIBR<10 + Lilly160 + MolGraph<5 + protecting groups + undefined stereo ≤2 |
| exploration | config_structFilters_exploration.yml | NIBR + MolGraph + protecting groups + undefined stereo ≤3 |
| strict | config_structFilters_strict.yml | PAINS + LD50-Oral + Toxicophore + Skin + MLSMR + NIBR + Lilly100 + Bredt + MolGraph + protecting groups + molcomplexity + undefined ≤2 + total stereo <5 |
Smoke on 20 MOSES molecules (results/profile_smoke_struct_filters/summary.csv): default 15/20, exploration 18/20, strict 0/20.
Prevalence evidence: results/zinc250_thresholds/iteration_24_structural_ruleset_prevalence/REPORT.md.
config_synthesis.yml
Synthesis Feasibility. Controls the retrosynthesis feasibility stage, including synthesizability score thresholds.
| Parameter | Type | Default | Description |
|---|---|---|---|
run | bool | true | Enable or disable the synthesis stage |
n_jobs | int | 64 | Number of AiZynthFinder worker processes; reduce this on smaller hosts |
enabled_scores | list or all | [sa, syba, rascore] | Synthesis score calculators to run. Use scalar all (or ['all']) for every scorer: sa, syba, rascore, sync, scscore, nonpher, fsscore, gasa. Optional scorers return NaN with warnings when dependencies are unavailable |
run_retrosynthesis | bool | true | Run AiZynthFinder retrosynthetic analysis |
filter_solved_only | bool | true | Keep only molecules for which a retrosynthetic route was found |
aizynthfinder_max_transforms | int | 15 | Maximum retrosynthetic route depth; higher values can sharply increase search cost |
aizynthfinder_time_limit | float | 2000 | Search time limit per target representation, in seconds |
aizynthfinder_iteration_limit | int | 300 | Maximum MCTS iterations per target representation |
aizynthfinder_return_first | bool | true | Stop each representation after its first solved route |
aizynthfinder_charge_mode | string | both | Search the preserved SMILES, a stereochemistry-preserving neutralized SMILES, or both |
aizynthfinder_retry_unsolved | bool | false | Enable a second AiZynthFinder pass only for molecules unsolved across all charge forms in pass 1 |
aizynthfinder_reuse_pass1 | bool | false | Reuse an existing retrosynthesis_variants_pass1.json checkpoint when restarting an interrupted retry |
aizynthfinder_retry_time_limit | float | unset | Time limit per representation in the unsolved-only second pass |
aizynthfinder_retry_iteration_limit | int | unset | MCTS iteration limit per representation in the unsolved-only second pass |
sa_score_min | float | 1 | Minimum synthetic accessibility score (Ertl) |
sa_score_max | float | 4.5 | Maximum synthetic accessibility score (lower = easier to synthesize) |
syba_score_min | float | 0 | Minimum SYBA score |
syba_score_max | float | inf | Maximum SYBA score |
ra_score_min | float | 0.5 | Practical retrosynthetic accessibility floor |
ra_score_max | float | 1 | Maximum retrosynthetic accessibility score |
sync_auto_install | bool | true | Download the SYNC checkpoint automatically when it is missing |
sync_device | string | cpu | Torch device for SYNC inference |
sync_conformer_seed | int | 61453 | RDKit ETKDG conformer seed for SYNC inputs |
fsscore_python | string | null | null | Python interpreter for isolated FSScore worker environment |
fsscore_model_path | string | null | null | Explicit FSScore checkpoint path (*.ckpt) |
fsscore_repo_path | string | null | null | Optional FSScore checkout path used to resolve models/pretrain_graph_GGLGGL_ep242_best_valloss.ckpt |
fsscore_batch_size | int | 128 | Batch size passed to fsscore.score |
fsscore_num_workers | int | null | null | Optional dataloader worker count passed to fsscore.score |
score_filters | object | {} | Optional min/max filters for additional score columns such as sync_score, sc_score, nonpher_complexity_score, fs_score, or gasa_score |
gasa.command | string | null | Optional local command template for batch gasa scoring using {input} and {output} placeholders |
gasa.executable | string | null | Optional local executable path/name used for gasa scoring (<exe> --smiles <SMILES>) |
gasa.api_url | string | null | Optional local loopback HTTP endpoint for gasa scoring (POST {"smiles": ...}) |
gasa.timeout_seconds | float | 30 | Timeout per gasa backend call |
config_docking.yml
Docking. Controls molecular docking using SMINA, GNINA, Matcha, or any explicit combination of them. Defines the receptor, search box, and engine-specific parameters.
General Settings
| Parameter | Type | Default | Description |
|---|---|---|---|
run | bool | true | Enable or disable the docking stage |
tools | string or list | [smina, gnina] | Explicit docking engines; Matcha is opt-in because it requires a trained checkpoint |
receptor_pdb | string | examples/7EW9_apo.pdb | Receptor PDB; relative paths are resolved from the docking config directory |
autobox_ligand | string | examples/05C_from_7EW9.sdf | Shared reference ligand for every selected engine |
autobox_add | float | 4 | Shared autobox padding in Angstroms |
auto_run | bool | true | Automatically start docking after ligand preparation |
run_in_background | bool | false | Run docking as a background process |
prepare_ligands | bool | false | Whether ligand_preparation_tool is actually invoked; false uses the direct SDF/RDKit path |
per_molecule_docking | bool | true | Generate one isolated engine config and result per molecule |
gnina_per_process_cpu | int | gnina_config.cpu | CPU threads per GNINA process in per molecule mode |
gnina_parallel_jobs_max | int or null | null | Optional override; by default parallelism is derived from CPU budget and capped at two jobs per visible GPU |
calculate_score_thresholds_from_targets | bool | true | Run configured docking tools on target molecules and generate tool-specific score cutoffs |
score_thresholds | mapping | max: -6.5 per tool | Lower-is-better minimizedAffinity upper bounds for SMINA, GNINA, and Matcha. Target alignment uses max(target_calibrated_max, configured_max), so it may relax the configured cutoff but never tighten it. Candidate molecules must pass every configured tool cutoff; missing scores fail |
SMINA Configuration (smina_config)
| Parameter | Type | Default | Description |
|---|---|---|---|
bin | string | smina | Path or name of the SMINA binary (resolved via PATH if not absolute) |
cpu | int | 1 | CPU threads per SMINA process |
seed | int | 42 | Random seed for reproducibility |
exhaustiveness | int | 8 | Search exhaustiveness (higher = more thorough, slower) |
num_modes | int | 1 | Maximum number of binding modes to generate per ligand |
GNINA Configuration (gnina_config)
| Parameter | Type | Default | Description |
|---|---|---|---|
bin | string | gnina | Path or name of the GNINA binary (resolved via PATH if not absolute) |
cpu | int | 8 | Number of CPU threads for docking |
seed | int | 42 | Random seed for reproducibility |
no_gpu | bool | false | Disable GPU acceleration (false keeps GPU enabled when available) |
num_modes | int | 9 | Binding modes generated per ligand; Hedgehog retains the pose with the lowest minimizedAffinity before target-threshold calibration |
Matcha Configuration (matcha_config)
Default path is the official Matcha CLI (LigandPro/Matcha), checked out under modules/matcha_remote. Hedgehog runs uv run --project <checkout> matcha ....
| Parameter | Type | Default | Description |
|---|---|---|---|
checkout_dir | string | modules/matcha_remote | Managed Matcha checkout (cloned/updated from GitHub on first use) |
autobox_ligand | string | shared value | Optional per-Matcha override of the shared reference ligand |
device | string | auto | Matcha device selection (auto, cpu, cuda, cuda:N, mps) |
n_samples | int | Matcha default (20) | Poses sampled per ligand (--n-samples) |
scorer | string | gnina | Pose scorer (gnina, custom, none) |
scorer_minimize | bool | true | Minimize poses during Matcha GNINA scoring |
keep_workdir | bool | false | Preserve Matcha internal work directory after the run |
checkpoints | string | Matcha package default | Optional override of the Matcha checkpoints folder |
backend | string | matcha_cli | Optional; set docking only for the LigandPro/docking screening adapter |
Optional Matcha CLI knobs also accepted when set: n_confs, docking_batch_limit, num_workers, prefetch_factor, persistent_workers, gnina_batch_mode, scorer_path, config, run_name, repo_url.
The optional backend: docking path needs checkpoint_root and checkpoint_run (training config is always <checkpoint_root>/<checkpoint_run>/config.yaml). Prefer $HEDGEHOG_MATCHA_CHECKPOINT_ROOT in YAML instead of host absolute paths. Sampling/GPU knobs for that backend come from the docking repo itself, not from Hedgehog YAML.
When prepare_ligands is true, one input molecule may produce several
prepared ligands. This can change row counts and downstream mapping. Keep it
false for the default 1:1-oriented docking path unless you explicitly need an
external preparation workflow.
config_docking_filters.yml
Three-Dimensional Filters. Stage 6 evaluates model-provided SDF coordinates by default and can additionally evaluate docked coordinates. Five independent filters can be combined with all (every filter must pass) or any (at least one must pass) aggregation.
General Settings
| Parameter | Type | Default | Description |
|---|---|---|---|
run | bool | true | Enable or disable the docking filters stage |
input_sdf | string | null | null | Explicit coordinate SDF override; model-provided SDF input otherwise takes priority, followed by docking output |
receptor_pdb | string | null | null | Path to receptor PDB; if null, uses docking config value |
Aggregation
| Parameter | Type | Default | Description |
|---|---|---|---|
mode | string | all | all = molecule must pass every enabled filter; any = pass at least one |
save_metrics | bool | true | Save detailed per molecule metrics to a CSV file |
save_failed | bool | true | Save molecules that failed filtering to a separate file |
Pose Quality (pose_quality)
posecheck-fast only. Legacy PoseCheck keys (strain_forcefield, clash_tolerance) are not used.
| Parameter | Type | Default | Description |
|---|---|---|---|
enabled | bool | true | Enable pose-quality checks |
clash_cutoff | float | 0.75 | Relative VDW clash cutoff |
volume_clash_cutoff | float | 0.075 | Volume overlap cutoff |
max_distance | float | 5.0 | Maximum minimum ligand–protein distance (Å) |
short_circuit | bool | true | Skip later filters on fail when mode is all |
Interactions (interactions)
| Parameter | Type | Default | Description |
|---|---|---|---|
enabled | bool | true | Enable ProLIF interaction checks |
reference_ligand | string | null | null | SDF for fingerprint similarity; required when similarity_threshold > 0 |
similarity_threshold | float | 0.0 | Tanimoto on ProLIF bits vs reference (0 disables). Without a reference SDF, a positive threshold only warns and does not filter |
min_hbonds | int | 0 | Minimum hydrogen bonds |
required_residues | list | [] | Residues that must interact |
forbidden_residues | list | [] | Residues that must not interact |
interaction_types | list | ProLIF defaults | Interaction types to detect |
reporting.enabled | bool | true | Write interaction reporting artifacts |
Shepherd Score (shepherd_score)
| Parameter | Type | Default | Description |
|---|---|---|---|
enabled | bool | false | Requires reference_ligand |
reference_ligand | string | null | null | Reference SDF |
min_shape_score | float | 0.5 | Minimum Gaussian-overlap Tanimoto |
alpha | float | 0.81 | Gaussian width |
align_before_scoring | bool | true | Align pose onto reference with RDKit AlignMol / GetBestRMS before scoring |
backend | string | auto | auto, worker, or inprocess |
auto_install_worker | bool | true | Auto-install worker environment when missing |
Conformer Deviation (conformer_deviation)
| Parameter | Type | Default | Description |
|---|---|---|---|
enabled | bool | true | Enable conformer-deviation check |
num_conformers | int | 50 | ETKDG conformers to generate |
conformer_method | string | ETKDGv3 | ETKDG, ETKDGv2, or ETKDGv3 |
max_rmsd_to_conformer | float | 3.0 | Maximum RMSD (Å) |
optimize_conformers | bool | false | UFF-relax generated ETKDG conformers before RMSD (not the docked pose); failed UFF steps are skipped |
backend | string | symmetry_rmsd | symmetry_rmsd or naive |
use_nvmolkit | bool | true | Prefer nvMolKit when available |
Deduplication
Docking can produce multiple poses per molecule. After filtering, the pipeline deduplicates to unique molecules:
- All passing poses are saved to
filtered_poses.csv - Poses are sorted by affinity (best first)
- For each unique
mol_idx, only the best-scoring pose is kept - Deduplicated molecules are saved to
filtered_molecules.csv
SMILES for the output are taken from the original ligand table rather than regenerated from 3D coordinates.
config_weighted_score.yml
Controls the post-run Generator Reality Assessment used by HTML reporting and RUN_INFO.md.
The scorecard is explainable and intended to rank generator behavior, not to estimate hit probability. It also reports a secondary Final Candidate Pool Quality score for the survivor set.
General Settings
| Parameter | Type | Default | Description |
|---|---|---|---|
run | bool | true | Enable or disable weighted model scoring output |
version | string | v1 | Internal scorecard schema version |
mode | string | generator_reality | Scoring mode label for the gate-aware generator score |
target_final_count | int | 100 | Target final count retained for secondary candidate-pool yield scoring |
target_final_retention | float | 0.10 | Target final retention rate for generator yield scoring |
confidence.min_final_molecules_high | int | 100 | Minimum final molecules for high confidence |
confidence.min_final_molecules_medium | int | 30 | Minimum final molecules for medium confidence |
Component Weights (weights)
| Parameter | Type | Default | Description |
|---|---|---|---|
weights.yield | float | 0.30 | Weight for final retention against target |
weights.physchem | float | 0.15 | Weight for descriptor all pass gate survival |
weights.structural | float | 0.25 | Weight for structural stage survival |
weights.synthesis | float | 0.10 | Weight for synthesis component |
weights.docking_pose | float | 0.15 | Weight for docking/pipeline pose component |
weights.diversity | float | 0.05 | Weight for diversity metrics component |
Weights are normalized over all configured components before scoring. When one component is unavailable, it is simply excluded, and the effective average is recomputed from the remaining available components.
physchem is measured from stages/02_descriptors_initial/filtered/pass_flags.csv as an all pass descriptor gate rate, so it reflects the early generated set rather than the final survivor pool. The mean flag pass rate is retained as evidence only. structural uses the stage survival rate from filtered and failed molecules, with the weakest structural filter as supporting evidence. Final descriptor files are used only as a fallback for older or partial runs. synthesis and docking_pose similarly prefer full stage evaluation artifacts before filtered or final survivor files.
Secondary Candidate Pool Weights (candidate_pool_weights)
candidate_pool_weights control the secondary Final Candidate Pool Quality score. It keeps the older survivor-pool interpretation: final-count yield saturation, mean descriptor flag pass rate, mean structural flag pass rate, and the same synthesis/docking/diversity formulas.
Yield and Structural Settings
| Parameter | Type | Default | Description |
|---|---|---|---|
yield.mode | string | retention | Use final retention for the generator score; absolute restores count-saturation yield |
yield.target_final_retention | float | 0.10 | Retention rate that maps to a full yield score |
yield.count_weight | float | 0.70 | Count-saturation weight for secondary candidate-pool yield |
yield.retention_weight | float | 0.30 | Log-retention weight for secondary candidate-pool yield |
structural.stage_pass_weight | float | 0.80 | Weight for structural stage survival |
structural.worst_filter_weight | float | 0.20 | Weight for the weakest structural filter pass rate |
Hard Caps (hard_caps)
Hard caps prevent a model from receiving a high generator score when an early AND-gate rejects most molecules.
| Parameter | Type | Default | Description |
|---|---|---|---|
hard_caps.structural_stage_pass_rate_below | float | 0.20 | Trigger threshold for structural stage survival |
hard_caps.structural_stage_pass_rate_cap | float | 60.0 | Maximum score after structural cap trigger |
hard_caps.descriptor_all_pass_rate_below | float | 0.50 | Trigger threshold for descriptor all pass survival |
hard_caps.descriptor_all_pass_rate_cap | float | 70.0 | Maximum score after descriptor cap trigger |
hard_caps.final_retention_rate_below | float | 0.05 | Trigger threshold for final retention |
hard_caps.final_retention_rate_cap | float | 70.0 | Maximum score after retention cap trigger |
Docking Thresholds (docking)
| Parameter | Type | Default | Description |
|---|---|---|---|
docking.bad_affinity | float | -6.0 | Affinity at which docking contribution starts to approach zero |
docking.good_affinity | float | -9.0 | Affinity at which docking affinity contribution reaches upper bound |
docking.bad_cnnscore | float | 0.35 | GNINA CNN score lower bound |
docking.good_cnnscore | float | 0.85 | GNINA CNN score upper bound |
docking.bad_cnnaffinity | float | 4.5 | CnnAffinity lower bound |
docking.good_cnnaffinity | float | 6.5 | CnnAffinity upper bound |
Increase strictness by moving bad_* upward and good_* downward, or relax by widening the interval.
Synthesis Thresholds (synthesis)
| Parameter | Type | Default | Description |
|---|---|---|---|
synthesis.sa_min | float | 1.0 | Easier-to-synthesize SA floor |
synthesis.sa_max | float | 4.5 | Harder-to-synthesize SA ceiling |
synthesis.ra_min | float | 0.5 | Minimum retrosynthetic accessibility minimum |
synthesis.ra_max | float | 1.0 | Retrosynthetic accessibility maximum |
synthesis.syba_midpoint | float | 0.0 | Sigmoid midpoint for SYBA |
synthesis.syba_scale | float | 50.0 | Sigmoid width for SYBA |
synthesis.target_search_time_sec | float | 30.0 | Reference retrosynthesis search time |
synthesis.search_time_scale_sec | float | 20.0 | Search-time penalty scale |
Raise or lower these to bias toward faster/easier synthetic routes.
config_moleval.yml
Controls generative evaluation metrics computed during report generation. These metrics assess diversity, scaffold coverage, and basic filter pass rates across pipeline stages.
General Settings
| Parameter | Type | Default | Description |
|---|---|---|---|
run | bool | true | Enable or disable MolEval metric computation |
n_jobs | int | -1 | Number of parallel workers for metric computation (-1 = all available cores) |
device | string | cpu | Compute device: cpu or cuda:0 (for neural metrics) |
max_molecules | int | 2000 | Subsample threshold for O(N^2) metrics; datasets larger than this are subsampled |
Metric Groups
Each flag enables or disables a group of related metrics.
| Parameter | Type | Default | Description |
|---|---|---|---|
validity | bool | false | Compute validity rate (disabled by default — always 1.0 after RDKit parsing) |
uniqueness | bool | false | Compute uniqueness rate (disabled by default — always 1.0 after deduplication) |
internal_diversity | bool | true | Compute IntDiv1 and IntDiv2 (intra-set Tanimoto diversity) |
se_diversity | bool | true | Compute sphere-exclusion diversity (SEDiv) |
scaffold_diversity | bool | true | Compute ScaffDiv and ScaffUniqueness (Murcko scaffold analysis) |
functional_groups | bool | true | Compute functional group diversity ratio (FG) |
ring_systems | bool | true | Compute ring system diversity ratio (RS) |
filters | bool | true | MCF + PAINS passage rate from vendored mcf.csv and wehi_pains.csv (not Stage 3 Common Alerts PAINS; no pains_file_path / mcf_file_path keys) |
mce18 | bool | true | Compute mean MCE-18 molecular complexity score |