Skip to Content
ConfigurationParameter Reference

Parameter Reference

Complete reference for every parameter in Hedgehog’s YAML configuration files.


config.yml

The main configuration file. Controls input/output paths, parallelism, and references to all stage-specific configs.

ParameterTypeDefaultDescription
generated_mols_pathstringsrc/hedgehog/configs/examples/moses_1000.csvPath to the CSV file containing generated molecules
target_mols_pathstringsrc/hedgehog/configs/examples/target_mols.csvPath to the CSV file containing reference molecules
alignment.enabledboolfalseCalculate numeric thresholds from target_mols_path before evaluating generated molecules; structural policy remains fixed
alignment.target_coverage_percentfloat95Percentage of target molecules the generated thresholds retain at each stage; must be greater than 0 and at most 100
folder_to_savestringresults/runOutput directory where all pipeline results are saved
n_jobsint-1Number of parallel workers for CPU-bound tasks (-1 = all available cores. Prefer an explicit smaller number such as 4 or 8.)
sample_sizeint10000Number of molecules to sample from the input file (null = use all)
save_sampled_molsbooltrueWhether to save the sampled molecule subset to disk
evaluate_docked_coordinatesboolfalseFor SDF inputs, additionally evaluate docked coordinates; requires docking.run: true and at least one selected docking tool
large_dataset_modeboolfalseEnable streaming chunked processing for very large pre-docking dataset statistics
large_dataset_chunk_rowsint250000Rows per processing chunk in large dataset mode
large_dataset_single_csv_limitint1000000Maximum row count for also materializing compatibility CSV files from shard outputs
large_dataset_output_formatstringcsv.gzShard file format for large dataset row-level intermediate tables
large_dataset_filter_databoolfalseIn large dataset mode, whether filter pass/fail results should remove molecules from downstream outputs
large_dataset_enable_all_filtersbooltrueIn large dataset mode, enable configured descriptor/structural filters as calculations even when they do not filter outputs
ligand_preparation_toolstring(proprietary path)Absolute path to an external ligand preparation binary
protein_preparation_toolstring(proprietary path)Absolute path to an external protein preparation binary
config_mol_prepstringsrc/hedgehog/configs/config_mol_prep.ymlPath to the preprocessing stage config
config_descriptorsstringsrc/hedgehog/configs/config_descriptors.ymlPath to the descriptors stage config
config_structFiltersstringsrc/hedgehog/configs/config_structFilters.ymlPath to the structural filters stage config
config_synthesisstringsrc/hedgehog/configs/config_synthesis.ymlPath to the synthesis stage config
config_dockingstringsrc/hedgehog/configs/config_docking.ymlPath to the docking stage config
config_docking_filtersstringsrc/hedgehog/configs/config_docking_filters.ymlPath to the docking filters stage config
config_weighted_scorestringsrc/hedgehog/configs/config_weighted_score.ymlPath to the weighted model assessment config
config_molevalstringsrc/hedgehog/configs/config_moleval.ymlPath to the MolEval reporting config

Target-based alignment is controlled directly by the master config:

target_mols_path: data/reference_molecules.csv alignment: enabled: true target_coverage_percent: 95 # descriptor_bounds_mode: expand # keep source borders; widen only where targets need it # descriptor_bounds_mode: target # replace each border with the target-derived value descriptor_bounds_mode: expand

When enabled, HEDGEHOG calculates numeric thresholds with respect to the target molecules before evaluating generated_mols_path. Structural hard filters are not percentile-calibrated. Descriptor alignment supports descriptor_bounds_mode: target, which replaces each configured descriptor border with the target-derived border, and descriptor_bounds_mode: expand, which applies min(source_min, target_min) and max(source_max, target_max). In expand, generic ZINC+ChEMBL descriptor envelopes are preserved and can only become wider. If descriptor_bounds_mode is omitted, expand is used. The target structural probe evaluates every calculate_* filter and every common-alert ruleset without removing molecules. Synthesis is skipped during target calibration and does not participate in target-cohort selection; the candidate uses the source synthesis config unchanged, including SA/RA/SYBA thresholds and retrosynthesis settings. The calibration probe may calculate every structural method for diagnostics, but the generated candidate config copies the source filter_* flags, Common Alerts enforcement lists, and structural parameters unchanged. Structural pass masks do not participate in target-cohort selection or the target_coverage_percent guarantee. structural_filter_failures.csv records every measured failure. Set enabled: false to use the stage configs without recalculation. --align-config PERCENT enables alignment and overrides target_coverage_percent for one run. Generated aligned master configs set alignment.enabled: false to prevent recalibration when reused, and add the target-run and threshold-audit paths as provenance. Each aligned stage config is created immediately after its alignable stage completes. Continuous thresholds use outward rounding, while integer-valued thresholds are rounded up.


config_mol_prep.yml

Preprocessing. Standardizes molecules before any descriptor computation. This stage aims to produce “clean” molecules by:

  • removing salts and solvents and keeping the largest fragment
  • disconnecting metals and normalizing/reionizing structures
  • preserving formal charge by default (steps.standardize_mol.uncharge: false)
  • preserving stereochemistry by default (steps.remove_stereochemistry: false)
  • applying atom, radical, and single-fragment filters while retaining isotope labels

General Settings

ParameterTypeDefaultDescription
runbooltrueEnable or disable preprocessing
n_jobsint-1Worker count for molecule preparation
steps.standardize_mol.unchargeboolfalseNeutralize formal charges during MolPrep
steps.remove_stereochemistryboolfalseRemove atom/bond stereochemical annotations
filters.allowed_atomslist[string][C, N, O, S, F, Cl, Br, I, P, H, Si]Allowed atom symbols
filters.require_single_fragmentbooltrueReject multi-fragment molecules
filters.reject_radicalsbooltrueReject molecules with radical electrons
filters.reject_isotopesboolfalseReject isotopically labeled molecules instead of preserving them for later stages
output.write_duplicates_removedbooltrueWrite duplicates_removed.csv when duplicates are dropped

config_descriptors.yml

Descriptors. Controls descriptor calculation and filtering borders. Plot presentation has code defaults and only needs configuration for custom reports.

General Settings

The descriptors stage always computes the complete descriptor table. When filter_data is true, configured filters are applied as follows:

  • borders define generic descriptor ranges such as molWt, logP, TPSA, hbd, hba, n_rings, and fsp3.
  • structural_constraints are converted into additional upper bound checks on derived descriptor columns.

Use borders for the production filtering contract. The optional legacy-compatible structural_constraints block can add motif caps when its own enabled field is true.

Optional Plot Overrides

Normally these fields should be omitted: plotted columns are derived from borders, while discrete-column metadata and labels come from code defaults (DESCRIPTOR_DISPLAY_NAMES and related constants in src/hedgehog/descriptors/constants.py). There is no YAML renamer.

ParameterTypeDescription
filtered_cols_to_plotlist[string]Descriptor columns to include in distribution plots
discrete_features_to_plotlist[string]Columns treated as discrete

Default borders are the ZINC250k + ChEMBL34 calibration envelope (q0.29–q99.71). See Descriptors.


config_structFilters.yml

Structural Filters. Calculates structural alerts and medicinal-chemistry diagnostics. calculate_<name> controls diagnostics; filter_<name> independently controls Stage 3 survival.

General Settings

ParameterTypeDefaultDescription
runbooltrueEnable or disable the structural filters stage
n_jobsint16Worker count for parsing and every structural filter (-1 = all available cores; falls back to master n_jobs when omitted)
filter_databooltrueWhether to apply the configured hard-filter decision to downstream molecules
filter_NIBRbooltrueInclude the published NIBR severity policy in survival
filter_molgraph_statsbooltrueInclude MolGraph severity policy in survival
filter_common_alertsbooltrueEnforce the configured Common Alerts subset
include_rulesetsall | list | nullallRulesets to calculate: all = full catalog, []/null = none, or an explicit list
exclude_smartslist[string][]Exact SMARTS strings dropped from calculation
common_alerts_filter_include_rulesetslist[PAINS]Rulesets enforced when filter_common_alerts is true; empty = every calculated ruleset
common_alerts_filter_exclude_rulesetslist[]Calculated rulesets excluded from survival; exclusions win
filter_lillybooltrueInclude Lilly (cutoff from lilly_demerit_cutoff) in survival
filter_protecting_groupsbooltrueReject curated protecting-group matches
filter_molcomplexity, filter_bredt, filter_ring_infraction, filter_halogenicity, filter_symmetryboolfalseOptional hard-filter flags; calculations remain independent
filter_stereo_centerboolfalseOptionally enforce the total-stereocenter diagnostic cutoff
filter_undefined_stereo_centerbooltrueReject underspecified structures above stereo_max_undefined; reuses the stereo-center calculation
lilly_demerit_cutoffint160Lilly demerit threshold passed as dthresh
nibr_max_severityint10Inclusive NIBR rejection threshold; pass requires accumulated severity below this value
molgraph_max_severityint5Inclusive MolGraph rejection threshold; pass requires maximum pattern severity below this value
write_per_filter_outputsbooltrueWrite per-filter output folders and CSVs
write_structural_liability_profilebooltrueWrite one molecule-level table containing the hard decision and every diagnostic field
generate_plotsbooltrueGenerate structural filter plots
generate_failure_analysisbooltrueGenerate failure-analysis outputs
ring_infraction_hetcycle_min_sizeint4Largest small-ring cutoff checked by the ring-infraction rule
stereo_max_centersint4Diagnostic cutoff for total stereocenters; the total count is reported independently of hard enforcement
stereo_max_undefinedint2Inclusive maximum for undefined stereocenters; values above it fail the undefined-stereo hard policy
halogenicity_thresh_Fint6Inclusive fluorine-count limit
halogenicity_thresh_Brint3Inclusive bromine-count limit
halogenicity_thresh_Clint3Inclusive chlorine-count limit
symmetry_thresholdfloat0.8Inclusive maximum MedChem symmetry score

Target-aware runs keep the complete generic structural gate. Alignment audits cannot change explicit filter_<name> flags. Any future target-specific relaxation must be an explicit, separately justified exception for a concrete hard rule. Scheduler settings such as lilly_scheduler remain execution controls.

Structural Filter Profiles

Three shipped configs share the same diagnostic calculations. Hard gates differ. The former balanced profile is removed.

ProfileConfigHard gate summary
defaultconfig_structFilters.ymlPAINS + NIBR<10 + Lilly160 + MolGraph<5 + protecting groups + undefined stereo ≤2
explorationconfig_structFilters_exploration.ymlNIBR + MolGraph + protecting groups + undefined stereo ≤3
strictconfig_structFilters_strict.ymlPAINS + LD50-Oral + Toxicophore + Skin + MLSMR + NIBR + Lilly100 + Bredt + MolGraph + protecting groups + molcomplexity + undefined ≤2 + total stereo <5

Smoke on 20 MOSES molecules (results/profile_smoke_struct_filters/summary.csv): default 15/20, exploration 18/20, strict 0/20.

Prevalence evidence: results/zinc250_thresholds/iteration_24_structural_ruleset_prevalence/REPORT.md.


config_synthesis.yml

Synthesis Feasibility. Controls the retrosynthesis feasibility stage, including synthesizability score thresholds.

ParameterTypeDefaultDescription
runbooltrueEnable or disable the synthesis stage
n_jobsint64Number of AiZynthFinder worker processes; reduce this on smaller hosts
enabled_scoreslist or all[sa, syba, rascore]Synthesis score calculators to run. Use scalar all (or ['all']) for every scorer: sa, syba, rascore, sync, scscore, nonpher, fsscore, gasa. Optional scorers return NaN with warnings when dependencies are unavailable
run_retrosynthesisbooltrueRun AiZynthFinder retrosynthetic analysis
filter_solved_onlybooltrueKeep only molecules for which a retrosynthetic route was found
aizynthfinder_max_transformsint15Maximum retrosynthetic route depth; higher values can sharply increase search cost
aizynthfinder_time_limitfloat2000Search time limit per target representation, in seconds
aizynthfinder_iteration_limitint300Maximum MCTS iterations per target representation
aizynthfinder_return_firstbooltrueStop each representation after its first solved route
aizynthfinder_charge_modestringbothSearch the preserved SMILES, a stereochemistry-preserving neutralized SMILES, or both
aizynthfinder_retry_unsolvedboolfalseEnable a second AiZynthFinder pass only for molecules unsolved across all charge forms in pass 1
aizynthfinder_reuse_pass1boolfalseReuse an existing retrosynthesis_variants_pass1.json checkpoint when restarting an interrupted retry
aizynthfinder_retry_time_limitfloatunsetTime limit per representation in the unsolved-only second pass
aizynthfinder_retry_iteration_limitintunsetMCTS iteration limit per representation in the unsolved-only second pass
sa_score_minfloat1Minimum synthetic accessibility score (Ertl)
sa_score_maxfloat4.5Maximum synthetic accessibility score (lower = easier to synthesize)
syba_score_minfloat0Minimum SYBA score
syba_score_maxfloatinfMaximum SYBA score
ra_score_minfloat0.5Practical retrosynthetic accessibility floor
ra_score_maxfloat1Maximum retrosynthetic accessibility score
sync_auto_installbooltrueDownload the SYNC checkpoint automatically when it is missing
sync_devicestringcpuTorch device for SYNC inference
sync_conformer_seedint61453RDKit ETKDG conformer seed for SYNC inputs
fsscore_pythonstring | nullnullPython interpreter for isolated FSScore worker environment
fsscore_model_pathstring | nullnullExplicit FSScore checkpoint path (*.ckpt)
fsscore_repo_pathstring | nullnullOptional FSScore checkout path used to resolve models/pretrain_graph_GGLGGL_ep242_best_valloss.ckpt
fsscore_batch_sizeint128Batch size passed to fsscore.score
fsscore_num_workersint | nullnullOptional dataloader worker count passed to fsscore.score
score_filtersobject{}Optional min/max filters for additional score columns such as sync_score, sc_score, nonpher_complexity_score, fs_score, or gasa_score
gasa.commandstringnullOptional local command template for batch gasa scoring using {input} and {output} placeholders
gasa.executablestringnullOptional local executable path/name used for gasa scoring (<exe> --smiles <SMILES>)
gasa.api_urlstringnullOptional local loopback HTTP endpoint for gasa scoring (POST {"smiles": ...})
gasa.timeout_secondsfloat30Timeout per gasa backend call

config_docking.yml

Docking. Controls molecular docking using SMINA, GNINA, Matcha, or any explicit combination of them. Defines the receptor, search box, and engine-specific parameters.

General Settings

ParameterTypeDefaultDescription
runbooltrueEnable or disable the docking stage
toolsstring or list[smina, gnina]Explicit docking engines; Matcha is opt-in because it requires a trained checkpoint
receptor_pdbstringexamples/7EW9_apo.pdbReceptor PDB; relative paths are resolved from the docking config directory
autobox_ligandstringexamples/05C_from_7EW9.sdfShared reference ligand for every selected engine
autobox_addfloat4Shared autobox padding in Angstroms
auto_runbooltrueAutomatically start docking after ligand preparation
run_in_backgroundboolfalseRun docking as a background process
prepare_ligandsboolfalseWhether ligand_preparation_tool is actually invoked; false uses the direct SDF/RDKit path
per_molecule_dockingbooltrueGenerate one isolated engine config and result per molecule
gnina_per_process_cpuintgnina_config.cpuCPU threads per GNINA process in per molecule mode
gnina_parallel_jobs_maxint or nullnullOptional override; by default parallelism is derived from CPU budget and capped at two jobs per visible GPU
calculate_score_thresholds_from_targetsbooltrueRun configured docking tools on target molecules and generate tool-specific score cutoffs
score_thresholdsmappingmax: -6.5 per toolLower-is-better minimizedAffinity upper bounds for SMINA, GNINA, and Matcha. Target alignment uses max(target_calibrated_max, configured_max), so it may relax the configured cutoff but never tighten it. Candidate molecules must pass every configured tool cutoff; missing scores fail

SMINA Configuration (smina_config)

ParameterTypeDefaultDescription
binstringsminaPath or name of the SMINA binary (resolved via PATH if not absolute)
cpuint1CPU threads per SMINA process
seedint42Random seed for reproducibility
exhaustivenessint8Search exhaustiveness (higher = more thorough, slower)
num_modesint1Maximum number of binding modes to generate per ligand

GNINA Configuration (gnina_config)

ParameterTypeDefaultDescription
binstringgninaPath or name of the GNINA binary (resolved via PATH if not absolute)
cpuint8Number of CPU threads for docking
seedint42Random seed for reproducibility
no_gpuboolfalseDisable GPU acceleration (false keeps GPU enabled when available)
num_modesint9Binding modes generated per ligand; Hedgehog retains the pose with the lowest minimizedAffinity before target-threshold calibration

Matcha Configuration (matcha_config)

Default path is the official Matcha CLI (LigandPro/Matcha), checked out under modules/matcha_remote. Hedgehog runs uv run --project <checkout> matcha ....

ParameterTypeDefaultDescription
checkout_dirstringmodules/matcha_remoteManaged Matcha checkout (cloned/updated from GitHub on first use)
autobox_ligandstringshared valueOptional per-Matcha override of the shared reference ligand
devicestringautoMatcha device selection (auto, cpu, cuda, cuda:N, mps)
n_samplesintMatcha default (20)Poses sampled per ligand (--n-samples)
scorerstringgninaPose scorer (gnina, custom, none)
scorer_minimizebooltrueMinimize poses during Matcha GNINA scoring
keep_workdirboolfalsePreserve Matcha internal work directory after the run
checkpointsstringMatcha package defaultOptional override of the Matcha checkpoints folder
backendstringmatcha_cliOptional; set docking only for the LigandPro/docking screening adapter

Optional Matcha CLI knobs also accepted when set: n_confs, docking_batch_limit, num_workers, prefetch_factor, persistent_workers, gnina_batch_mode, scorer_path, config, run_name, repo_url.

The optional backend: docking path needs checkpoint_root and checkpoint_run (training config is always <checkpoint_root>/<checkpoint_run>/config.yaml). Prefer $HEDGEHOG_MATCHA_CHECKPOINT_ROOT in YAML instead of host absolute paths. Sampling/GPU knobs for that backend come from the docking repo itself, not from Hedgehog YAML.

When prepare_ligands is true, one input molecule may produce several prepared ligands. This can change row counts and downstream mapping. Keep it false for the default 1:1-oriented docking path unless you explicitly need an external preparation workflow.


config_docking_filters.yml

Three-Dimensional Filters. Stage 6 evaluates model-provided SDF coordinates by default and can additionally evaluate docked coordinates. Five independent filters can be combined with all (every filter must pass) or any (at least one must pass) aggregation.

General Settings

ParameterTypeDefaultDescription
runbooltrueEnable or disable the docking filters stage
input_sdfstring | nullnullExplicit coordinate SDF override; model-provided SDF input otherwise takes priority, followed by docking output
receptor_pdbstring | nullnullPath to receptor PDB; if null, uses docking config value

Aggregation

ParameterTypeDefaultDescription
modestringallall = molecule must pass every enabled filter; any = pass at least one
save_metricsbooltrueSave detailed per molecule metrics to a CSV file
save_failedbooltrueSave molecules that failed filtering to a separate file

Pose Quality (pose_quality)

posecheck-fast only. Legacy PoseCheck keys (strain_forcefield, clash_tolerance) are not used.

ParameterTypeDefaultDescription
enabledbooltrueEnable pose-quality checks
clash_cutofffloat0.75Relative VDW clash cutoff
volume_clash_cutofffloat0.075Volume overlap cutoff
max_distancefloat5.0Maximum minimum ligand–protein distance (Å)
short_circuitbooltrueSkip later filters on fail when mode is all

Interactions (interactions)

ParameterTypeDefaultDescription
enabledbooltrueEnable ProLIF interaction checks
reference_ligandstring | nullnullSDF for fingerprint similarity; required when similarity_threshold > 0
similarity_thresholdfloat0.0Tanimoto on ProLIF bits vs reference (0 disables). Without a reference SDF, a positive threshold only warns and does not filter
min_hbondsint0Minimum hydrogen bonds
required_residueslist[]Residues that must interact
forbidden_residueslist[]Residues that must not interact
interaction_typeslistProLIF defaultsInteraction types to detect
reporting.enabledbooltrueWrite interaction reporting artifacts

Shepherd Score (shepherd_score)

ParameterTypeDefaultDescription
enabledboolfalseRequires reference_ligand
reference_ligandstring | nullnullReference SDF
min_shape_scorefloat0.5Minimum Gaussian-overlap Tanimoto
alphafloat0.81Gaussian width
align_before_scoringbooltrueAlign pose onto reference with RDKit AlignMol / GetBestRMS before scoring
backendstringautoauto, worker, or inprocess
auto_install_workerbooltrueAuto-install worker environment when missing

Conformer Deviation (conformer_deviation)

ParameterTypeDefaultDescription
enabledbooltrueEnable conformer-deviation check
num_conformersint50ETKDG conformers to generate
conformer_methodstringETKDGv3ETKDG, ETKDGv2, or ETKDGv3
max_rmsd_to_conformerfloat3.0Maximum RMSD (Å)
optimize_conformersboolfalseUFF-relax generated ETKDG conformers before RMSD (not the docked pose); failed UFF steps are skipped
backendstringsymmetry_rmsdsymmetry_rmsd or naive
use_nvmolkitbooltruePrefer nvMolKit when available

Deduplication

Docking can produce multiple poses per molecule. After filtering, the pipeline deduplicates to unique molecules:

  1. All passing poses are saved to filtered_poses.csv
  2. Poses are sorted by affinity (best first)
  3. For each unique mol_idx, only the best-scoring pose is kept
  4. Deduplicated molecules are saved to filtered_molecules.csv

SMILES for the output are taken from the original ligand table rather than regenerated from 3D coordinates.


config_weighted_score.yml

Controls the post-run Generator Reality Assessment used by HTML reporting and RUN_INFO.md.

The scorecard is explainable and intended to rank generator behavior, not to estimate hit probability. It also reports a secondary Final Candidate Pool Quality score for the survivor set.

General Settings

ParameterTypeDefaultDescription
runbooltrueEnable or disable weighted model scoring output
versionstringv1Internal scorecard schema version
modestringgenerator_realityScoring mode label for the gate-aware generator score
target_final_countint100Target final count retained for secondary candidate-pool yield scoring
target_final_retentionfloat0.10Target final retention rate for generator yield scoring
confidence.min_final_molecules_highint100Minimum final molecules for high confidence
confidence.min_final_molecules_mediumint30Minimum final molecules for medium confidence

Component Weights (weights)

ParameterTypeDefaultDescription
weights.yieldfloat0.30Weight for final retention against target
weights.physchemfloat0.15Weight for descriptor all pass gate survival
weights.structuralfloat0.25Weight for structural stage survival
weights.synthesisfloat0.10Weight for synthesis component
weights.docking_posefloat0.15Weight for docking/pipeline pose component
weights.diversityfloat0.05Weight for diversity metrics component

Weights are normalized over all configured components before scoring. When one component is unavailable, it is simply excluded, and the effective average is recomputed from the remaining available components.

physchem is measured from stages/02_descriptors_initial/filtered/pass_flags.csv as an all pass descriptor gate rate, so it reflects the early generated set rather than the final survivor pool. The mean flag pass rate is retained as evidence only. structural uses the stage survival rate from filtered and failed molecules, with the weakest structural filter as supporting evidence. Final descriptor files are used only as a fallback for older or partial runs. synthesis and docking_pose similarly prefer full stage evaluation artifacts before filtered or final survivor files.

Secondary Candidate Pool Weights (candidate_pool_weights)

candidate_pool_weights control the secondary Final Candidate Pool Quality score. It keeps the older survivor-pool interpretation: final-count yield saturation, mean descriptor flag pass rate, mean structural flag pass rate, and the same synthesis/docking/diversity formulas.

Yield and Structural Settings

ParameterTypeDefaultDescription
yield.modestringretentionUse final retention for the generator score; absolute restores count-saturation yield
yield.target_final_retentionfloat0.10Retention rate that maps to a full yield score
yield.count_weightfloat0.70Count-saturation weight for secondary candidate-pool yield
yield.retention_weightfloat0.30Log-retention weight for secondary candidate-pool yield
structural.stage_pass_weightfloat0.80Weight for structural stage survival
structural.worst_filter_weightfloat0.20Weight for the weakest structural filter pass rate

Hard Caps (hard_caps)

Hard caps prevent a model from receiving a high generator score when an early AND-gate rejects most molecules.

ParameterTypeDefaultDescription
hard_caps.structural_stage_pass_rate_belowfloat0.20Trigger threshold for structural stage survival
hard_caps.structural_stage_pass_rate_capfloat60.0Maximum score after structural cap trigger
hard_caps.descriptor_all_pass_rate_belowfloat0.50Trigger threshold for descriptor all pass survival
hard_caps.descriptor_all_pass_rate_capfloat70.0Maximum score after descriptor cap trigger
hard_caps.final_retention_rate_belowfloat0.05Trigger threshold for final retention
hard_caps.final_retention_rate_capfloat70.0Maximum score after retention cap trigger

Docking Thresholds (docking)

ParameterTypeDefaultDescription
docking.bad_affinityfloat-6.0Affinity at which docking contribution starts to approach zero
docking.good_affinityfloat-9.0Affinity at which docking affinity contribution reaches upper bound
docking.bad_cnnscorefloat0.35GNINA CNN score lower bound
docking.good_cnnscorefloat0.85GNINA CNN score upper bound
docking.bad_cnnaffinityfloat4.5CnnAffinity lower bound
docking.good_cnnaffinityfloat6.5CnnAffinity upper bound

Increase strictness by moving bad_* upward and good_* downward, or relax by widening the interval.

Synthesis Thresholds (synthesis)

ParameterTypeDefaultDescription
synthesis.sa_minfloat1.0Easier-to-synthesize SA floor
synthesis.sa_maxfloat4.5Harder-to-synthesize SA ceiling
synthesis.ra_minfloat0.5Minimum retrosynthetic accessibility minimum
synthesis.ra_maxfloat1.0Retrosynthetic accessibility maximum
synthesis.syba_midpointfloat0.0Sigmoid midpoint for SYBA
synthesis.syba_scalefloat50.0Sigmoid width for SYBA
synthesis.target_search_time_secfloat30.0Reference retrosynthesis search time
synthesis.search_time_scale_secfloat20.0Search-time penalty scale

Raise or lower these to bias toward faster/easier synthetic routes.


config_moleval.yml

Controls generative evaluation metrics computed during report generation. These metrics assess diversity, scaffold coverage, and basic filter pass rates across pipeline stages.

General Settings

ParameterTypeDefaultDescription
runbooltrueEnable or disable MolEval metric computation
n_jobsint-1Number of parallel workers for metric computation (-1 = all available cores)
devicestringcpuCompute device: cpu or cuda:0 (for neural metrics)
max_moleculesint2000Subsample threshold for O(N^2) metrics; datasets larger than this are subsampled

Metric Groups

Each flag enables or disables a group of related metrics.

ParameterTypeDefaultDescription
validityboolfalseCompute validity rate (disabled by default — always 1.0 after RDKit parsing)
uniquenessboolfalseCompute uniqueness rate (disabled by default — always 1.0 after deduplication)
internal_diversitybooltrueCompute IntDiv1 and IntDiv2 (intra-set Tanimoto diversity)
se_diversitybooltrueCompute sphere-exclusion diversity (SEDiv)
scaffold_diversitybooltrueCompute ScaffDiv and ScaffUniqueness (Murcko scaffold analysis)
functional_groupsbooltrueCompute functional group diversity ratio (FG)
ring_systemsbooltrueCompute ring system diversity ratio (RS)
filtersbooltrueMCF + PAINS passage rate from vendored mcf.csv and wehi_pains.csv (not Stage 3 Common Alerts PAINS; no pains_file_path / mcf_file_path keys)
mce18booltrueCompute mean MCE-18 molecular complexity score
Last updated on