Skip to Content
Pipeline StagesPreprocessing

Preprocessing

Stage 1 standardizes input molecules and rejects entries that are unsuitable for later stages. It defines the clean population that descriptors, structural filters, synthesis, and docking see.

Behavior

Controlled by config_mol_prep.yml.

Normalization (steps.*):

  • Parse SMILES to RDKit molecules (to_mol)
  • Optional fix / sanitize
  • Remove salts and solvents; keep the largest fragment
  • Standardize: disconnect metals, normalize, reionize
  • Preserve formal charge by default (standardize_mol.uncharge: false)
  • Preserve stereochemistry by default (remove_stereochemistry: false)
  • Write standardized SMILES

Hard filters (filters.*):

KeyDefaultEffect
allowed_atomsC,N,O,S,F,Cl,Br,I,P,H,SiReject unsupported elements
reject_radicalstrueReject radical electrons
reject_isotopesfalseKeep isotopes by default
require_single_fragmenttrueReject multi-fragment leftovers

set_allowed_atoms_from_targets: true replaces filters.allowed_atoms with the union of atoms seen in target_mols_path during target alignment only. Outside alignment it has no effect.

output.write_duplicates_removed: true writes duplicates_removed.csv in normal mode. In large-dataset mode the duplicates file is always written; this flag is ignored there.

Output

Under stages/01_mol_prep/:

  • filtered_molecules.csv — survivors
  • failed_molecules.csv — rejects
  • duplicates_removed.csv — when duplicate writing is enabled

If MolPrep finishes with zero molecules, the pipeline exits early.

Usage

uv run hedgehog uv run hedgehog --stage mol_prep uv run hedge --stage mol_prep
Last updated on