Skip to Content
CLI Reference

CLI Reference

HEDGEHOG provides a Typer-based command-line interface. The primary command is hedgehog, with hedge available as a short alias. Both are interchangeable.

uv run hedgehog --help uv run hedge --help

Commands

CommandDescription
infoDisplay pipeline information and available stages
reportRegenerate the HTML report from an existing run
runRun the pipeline explicitly as a subcommand
setupInstall optional external tools and assets
tuiLaunch the interactive TUI (Terminal User Interface)
versionDisplay version information

hedge is a short alias for hedgehog. All commands work identically with either name:

uv run hedge --stage docking uv run hedgehog --stage docking

hedgehog (default pipeline command)

The root CLI command executes the evaluation pipeline by default. It runs all enabled stages using the configuration at src/hedgehog/configs/config.yml.

uv run hedgehog [OPTIONS]

The same pipeline command is also available as an explicit subcommand:

uv run hedgehog run [OPTIONS]

Options

OptionShortTypeDefaultDescription
--config-cTEXTcontext-dependentMaster YAML config path. Normally uses src/hedgehog/configs/config.yml; with --continue, uses the saved run config unless explicitly overridden.
--mols-mTEXTNoneInput molecule table path or glob. Overrides config input. Prefer CSV/TSV with a smiles header.
--out-oTEXTNoneOutput directory naming base; each run receives a fresh numbered folder.
--stage-sSTAGENoneRun one or more stages only. Repeat --stage to select multiple stages. Uses config molecules unless --mols is provided. See hedgehog info for stage descriptions.
--reuseFLAGFalseReuse existing results folder.
--continueRUN_FOLDERNoneContinue an unfinished run from its first incomplete enabled stage.
--force-newFLAGFalseExplicitly request the default fresh-folder behavior.
--auto-installFLAGFalseAuto-install missing optional tools without prompts.
--progressFLAGFalseShow live progress bar.
--large-datasetFLAGFalseStream large molecule libraries in chunks, write row-level .parts shards, and skip plots/report-heavy outputs. Filters are calculated by default but do not remove molecules unless large_dataset_filter_data: true.
--evaluate-docked-coordinatesFLAGFalseFor SDF inputs, also apply Stage 6 to docked coordinates. Input SDF coordinates are always the primary Stage 6 input.
--align-configPERCENTNoneEnable target alignment and override alignment.target_coverage_percent from the master config for this run.

The --reuse and --force-new flags are mutually exclusive. Using both will produce an error.

The --out option cannot be used together with --reuse or --force-new.

--continue cannot be combined with --mols, --out, --stage, --reuse, --force-new, --large-dataset, --evaluate-docked-coordinates, or --align-config. You may combine it with --config when you intentionally want to replace the saved master config.

Every run creates a fresh numbered results folder by default, including stage-only runs and runs using --out. For example, --out results/my_run creates results/my_run_1, then results/my_run_2. An existing run directory is used only when --reuse is passed explicitly.

The progress bar is disabled by default in CLI runs. Use --progress when you want live stage progress rendering.

For SDF inputs, Stage 6 evaluates the indexed source coordinates in input/ligands.sdf by default and intersects them with the latest enabled upstream stage by mol_idx. Set --evaluate-docked-coordinates or evaluate_docked_coordinates: true to additionally evaluate redocked poses under stages/06_docking_filters/docked/. Pose-level outputs contain pose_source=input or pose_source=docked. The root filtered_molecules.csv keeps molecules that pass either coordinate branch.

Use --large-dataset for PubChem/Enamine-scale statistics runs. It stores intermediate row-level tables as compressed shard directories such as filtered_molecules.parts/ and descriptors_all.parts/, uses large_dataset_chunk_rows from the master config, and skips plots plus HTML report generation. Without an explicit --stage selection, large-dataset mode runs the scalable statistics path: mol_prep, descriptors, struct_filters, and synthesis. Descriptor, structural, and synthesis filters are calculated by default, but downstream large-mode outputs keep all molecules unless large_dataset_filter_data: true. The synthesis stage skips AiZynthFinder retrosynthesis in large-dataset mode.

Set alignment.enabled: true and alignment.target_coverage_percent: PERCENT in the master config to align numeric score thresholds to target_mols_path. The --align-config PERCENT option provides a one-run enable-and-percentage override. HEDGEHOG first runs the reference molecules with numeric filtering relaxed and every structural filter and common-alert ruleset enabled in measurement-only mode. Structural results are diagnostic during calibration: the source structural hard gate, filter_* flags, Common Alerts enforcement rulesets, and structural parameters are copied unchanged and never participate in percentile-cohort selection. Synthesis is skipped during target calibration and does not participate in threshold derivation or cohort selection; the generated candidate uses the source synthesis config unchanged. Numeric stages select one shared central/best subset containing ceil(PERCENT / 100 * molecule_count) rows, then derive descriptor bounds from that subset. alignment.descriptor_bounds_mode accepts target or expand. expand preserves the source ZINC+ChEMBL descriptor envelope and only widens bounds required by the target. Docking pivots poses to one affinity per unique molecule and configured tool, selects one shared best subset using the worst normalized tool rank, and writes one minimizedAffinity.max cutoff per tool; candidate molecules must pass all generated tool cutoffs. Generated continuous thresholds use outward rounding; integer-valued thresholds are rounded up. MolPrep categorical rules are preserved. Alignment cannot be combined with large_dataset_mode or --large-dataset.

Aligned stage configs are created one by one under target_alignment/aligned_configs/: descriptors after the descriptor stage, an unchanged structural config plus diagnostics after structural filtering, docking after docking, and docking filters after the docking-filter stage. No complete config bundle is generated upfront or regenerated in bulk at the end. The aligned master and alignment_thresholds.yml are updated after each alignable stage, while unfinished stages remain pending_metrics. The structural diagnostic pass also writes structural_filter_failures.csv, with one row for every target molecule/rule failure. After the aligned candidate run is initialized, its reusable master config and self-contained support files are published under results/custom_configs/; the run-local continuation snapshot remains under <run>/configs/.

When you request stages with --stage, only the selected stages run. Mol Prep is not forced on automatically; include --stage mol_prep if you want standardization first.

Every run writes a transient .RUN_INCOMPLETE marker in the results folder. The marker is removed on successful completion. If it remains present, the run was interrupted or failed and the run log should be checked.

Continue an Interrupted Run

Stop the active process first with Ctrl-C, edit the copied YAML files under the unfinished run’s configs/ directory, and continue in place:

uv run hedge run --continue /absolute/path/to/results/run_1

Continuation loads input/sampled_molecules.csv; it does not preprocess or sample the input again. Enabled stages are checked in pipeline order. Completed stages are skipped, and execution starts at the first stage without a completion marker or complete output artifact. Every newly completed stage writes .STAGE_COMPLETE.yml, so later interruptions can be resumed reliably.

By default, continuation loads configs/master_config_resolved.yml and the run-local stage configs, so changes such as reducing AiZynthFinder workers should be made in those copied files. To use a different master config instead, pass it explicitly:

uv run hedge run --continue /absolute/path/to/results/run_1 \ --config /absolute/path/to/changed_config.yml

If the interruption happened inside target alignment, you may pass the outer run folder; HEDGEHOG resumes the nested target run, updates the remaining aligned stage configs, and then starts the candidate pipeline in the original folder. If alignment.target_coverage_percent was changed in the saved target-run master, completed aligned stages are recalibrated from their saved metrics without rerunning those stages.

Synthesis continuation reuses a compatible AiZynthFinder work directory. Finished shards are not launched again, and partial JSONL checkpoints resume from their last complete molecule. HEDGEHOG refuses continuation when .RUN_INCOMPLETE is absent or another continuation process holds the run lock.

Available Stages

Each --stage occurrence accepts one of the following values:

StageDescription
mol_prepStandardize and filter molecules
descriptorsCompute 28 physicochemical descriptors per molecule
struct_filtersApply structural filters
synthesisEvaluate synthetic accessibility using retrosynthesis (AiZynthFinder) and other metrics
dockingCalculate docking and binding affinity scores
docking_filtersFilter docking poses by quality and interactions
final_descriptorsRecompute descriptors on the final filtered set

Stage Failure and Skip Semantics

Stage completion is not binary “success or crash” across the whole pipeline. Some stages end the run early, while others can be skipped and let the pipeline continue.

  • Early exit with a completed upstream pipeline state: if mol_prep finishes but leaves zero molecules, the pipeline stops immediately.
  • Hard failure / early stop: structural filters stop the run if they do not complete successfully.

Examples

# Run the full pipeline with default config uv run hedgehog # Run with a custom master config uv run hedgehog --config src/hedgehog/configs/config.yml # Run with a custom molecule set uv run hedgehog --mols input/my_molecules.csv # Use a custom output naming base (creates results/my_run_1, then _2, ...) uv run hedgehog --out results/my_run # Run with a glob pattern uv run hedgehog --mols "input/generated/*.csv" # Run only the docking stage uv run hedge --stage docking # Run descriptors and structural filters in one invocation uv run hedge --stage descriptors --stage struct_filters # Run a specific stage with custom molecules uv run hedge --stage descriptors --mols input/candidates.csv # Rerun into the same results folder uv run hedge --reuse # Continue an interrupted run after editing its copied configs uv run hedge run --continue /absolute/path/to/results/run_1 # Explicit form of the default fresh-folder behavior uv run hedge --stage synthesis --force-new # Enable live progress bar uv run hedge --progress # Evaluate input coordinates and additionally evaluate docked coordinates uv run hedge --mols input/generated_3d.sdf --evaluate-docked-coordinates # Retain at least 95% of targets at each aligned stage; docking uses one # shared molecule subset and requires all generated tool cutoffs uv run hedge --align-config 95 # Stream a large library without monolithic intermediate CSVs or plots uv run hedge --large-dataset --mols input/pubchem.csv

hedgehog info

Displays a table of all available pipeline stages with their descriptions. Useful for checking stage names before using --stage.

uv run hedgehog info

hedgehog report

Regenerates the reporting artifacts from an existing pipeline run directory without re-running any pipeline stages. Useful when report templates, plotting logic, or the stage-audit notebook have been updated.

uv run hedgehog report <RESULTS_DIR>

RESULTS_DIRis a required argument. Set a path to an existing pipeline results directory (e.g.,results/run_10`).

The command loads the saved configuration from configs/master_config_resolved.yml inside the results directory and regenerates:

  • report.html
  • report_data.json
  • stage_filter_audit.ipynb

hedgehog setup aizynthfinder

Installs the upstream aizynthfinder package into the project environment and downloads its public data into modules/aizynthfinder/.

uv run hedgehog setup aizynthfinder

hedgehog setup sync

Downloads the SYNC 3D synthesizability classifier checkpoint to modules/sync/classifier_emb.ckpt.

uv run hedgehog setup sync

hedgehog setup fsscore

Clones upstream FSScore checkout into modules/fsscore so you can point synthesis configuration to the pretrained checkpoint without adding heavy torch dependencies to the base HEDGEHOG environment.

uv run hedgehog setup fsscore

After checkout, configure:

  • HEDGEHOG_FSSCORE_PYTHON to an isolated FSScore environment
  • HEDGEHOG_FSSCORE_MODEL_PATH (or HEDGEHOG_FSSCORE_REPO_PATH)

hedgehog setup gasa

Installs the optional GASA scorer checkout and an isolated worker environment used by the synthesis stage.

uv run hedgehog setup gasa

After setup, configure HEDGEHOG_GASA_COMMAND or synthesis.gasa.command with the printed worker command template, or use gasa.executable / gasa.api_url in config_synthesis.yml.

hedgehog setup nonpher-check

Validates optional Nonpher runtime for the synthesis nonpher scorer. This command does not install dependencies. For portable setup, first use uv-isolated environments under a writable per-host HEDGEHOG_OPTIONAL_ENV_ROOT (for example ~/work/hedgehog_optional_envs), then probe with --python.

uv run hedgehog setup nonpher-check

Options:

OptionTypeDefaultDescription
--pythonTEXTNoneExternal interpreter to probe (for example ~/work/hedgehog_optional_envs/nonpher/bin/python).
--probe-smilesTEXTCCOProbe molecule used for runtime validation.

Examples:

# Probe uv-only isolated Nonpher env uv run hedgehog setup nonpher-check --python ~/work/hedgehog_optional_envs/nonpher/bin/python # If uv-only bootstrap is blocked by native deps, probe a validated external runtime uv run hedgehog setup nonpher-check --python /mnt/ligandpro/shared_storage/data/nikolenko/hedgehog_optional_envs/nonpher-hybrid-py38-v2/bin/python

hedgehog setup nvmolkit-worker

Creates an isolated virtual environment at .venv-nvmolkit-worker and installs the optional nvMolKit worker there.

uv run hedgehog setup nvmolkit-worker

hedgehog setup shepherd-worker

Creates an isolated virtual environment at .venv-shepherd-worker and installs the optional Shepherd-Score backend there.

uv run hedgehog setup shepherd-worker

Environment Variables

The CLI supports a few environment variables that are useful in automation and CI:

VariableEffect
HEDGEHOG_AUTO_INSTALL=1Auto-accept optional dependency downloads and setup prompts.
HEDGEHOG_NON_INTERACTIVE=1Auto-decline download/setup prompts instead of waiting for interactive input.
HEDGEHOG_PLAIN_OUTPUT=1Disable Rich formatting and emit plain console output without the banner styling.

HEDGEHOG_AUTO_INSTALL is what the --auto-install flag sets internally for a pipeline run. The setup commands also use it when their --yes flag is enabled.

hedgehog tui

Launches the interactive Terminal User Interface (TUI) for visual pipeline configuration and management.

Requirements and behavior:

  • Requires Node.js >= 18 and npm.
  • Requires a source checkout that contains the repository tui/ directory.
  • If the TUI bundle does not exist yet, HEDGEHOG automatically runs npm install and npm run build inside tui/ before launching.
uv run hedgehog tui

Options:

OptionShortTypeDefaultDescription
--session-sTEXTNoneResume TUI directly into results for the specified job id.

Examples:

# Start TUI normally uv run hedgehog tui # Resume directly into an existing job uv run hedgehog tui --session 1a2b3c4d

For TUI-specific behavior such as editable config copies, job history storage, preflight checks, and keyboard shortcuts, see the TUI documentation.

The TUI can also be launched directly from the TUI package:

cd tui npm run tui

hedgehog version

Prints the current HEDGEHOG version and project tagline.

uv run hedgehog version
Last updated on