Preprocessing
Stage 1 standardizes input molecules and rejects entries that are unsuitable for later stages. It defines the clean population that descriptors, structural filters, synthesis, and docking see.
Behavior
Controlled by config_mol_prep.yml.
Normalization (steps.*):
- Parse SMILES to RDKit molecules (
to_mol) - Optional fix / sanitize
- Remove salts and solvents; keep the largest fragment
- Standardize: disconnect metals, normalize, reionize
- Preserve formal charge by default (
standardize_mol.uncharge: false) - Preserve stereochemistry by default (
remove_stereochemistry: false) - Write standardized SMILES
Hard filters (filters.*):
| Key | Default | Effect |
|---|---|---|
allowed_atoms | C,N,O,S,F,Cl,Br,I,P,H,Si | Reject unsupported elements |
reject_radicals | true | Reject radical electrons |
reject_isotopes | false | Keep isotopes by default |
require_single_fragment | true | Reject multi-fragment leftovers |
set_allowed_atoms_from_targets: true replaces filters.allowed_atoms with the union of atoms seen in target_mols_path during target alignment only. Outside alignment it has no effect.
output.write_duplicates_removed: true writes duplicates_removed.csv in normal mode. In large-dataset mode the duplicates file is always written; this flag is ignored there.
Output
Under stages/01_mol_prep/:
filtered_molecules.csv— survivorsfailed_molecules.csv— rejectsduplicates_removed.csv— when duplicate writing is enabled
If MolPrep finishes with zero molecules, the pipeline exits early.
Usage
uv run hedgehog
uv run hedgehog --stage mol_prep
uv run hedge --stage mol_prepLast updated on