CLI Reference
HEDGEHOG provides a Typer-based command-line interface. The primary command is hedgehog, with hedge available as a short alias. Both are interchangeable.
uv run hedgehog --help
uv run hedge --helpCommands
| Command | Description |
|---|---|
info | Display pipeline information and available stages |
report | Regenerate the HTML report from an existing run |
run | Run the pipeline explicitly as a subcommand |
setup | Install optional external tools and assets |
tui | Launch the interactive TUI (Terminal User Interface) |
version | Display version information |
hedge is a short alias for hedgehog. All commands work identically with either name:
uv run hedge --stage docking
uv run hedgehog --stage dockinghedgehog (default pipeline command)
The root CLI command executes the evaluation pipeline by default. It runs all enabled stages using the configuration at src/hedgehog/configs/config.yml.
uv run hedgehog [OPTIONS]The same pipeline command is also available as an explicit subcommand:
uv run hedgehog run [OPTIONS]Options
| Option | Short | Type | Default | Description |
|---|---|---|---|---|
--config | -c | TEXT | context-dependent | Master YAML config path. Normally uses src/hedgehog/configs/config.yml; with --continue, uses the saved run config unless explicitly overridden. |
--mols | -m | TEXT | None | Input molecule table path or glob. Overrides config input. Prefer CSV/TSV with a smiles header. |
--out | -o | TEXT | None | Output directory naming base; each run receives a fresh numbered folder. |
--stage | -s | STAGE | None | Run one or more stages only. Repeat --stage to select multiple stages. Uses config molecules unless --mols is provided. See hedgehog info for stage descriptions. |
--reuse | FLAG | False | Reuse existing results folder. | |
--continue | RUN_FOLDER | None | Continue an unfinished run from its first incomplete enabled stage. | |
--force-new | FLAG | False | Explicitly request the default fresh-folder behavior. | |
--auto-install | FLAG | False | Auto-install missing optional tools without prompts. | |
--progress | FLAG | False | Show live progress bar. | |
--large-dataset | FLAG | False | Stream large molecule libraries in chunks, write row-level .parts shards, and skip plots/report-heavy outputs. Filters are calculated by default but do not remove molecules unless large_dataset_filter_data: true. | |
--evaluate-docked-coordinates | FLAG | False | For SDF inputs, also apply Stage 6 to docked coordinates. Input SDF coordinates are always the primary Stage 6 input. | |
--align-config | PERCENT | None | Enable target alignment and override alignment.target_coverage_percent from the master config for this run. |
The --reuse and --force-new flags are mutually exclusive. Using both will produce an error.
The --out option cannot be used together with --reuse or --force-new.
--continue cannot be combined with --mols, --out, --stage, --reuse, --force-new, --large-dataset, --evaluate-docked-coordinates, or --align-config. You may combine it with --config when you intentionally want to replace the saved master config.
Every run creates a fresh numbered results folder by default, including stage-only runs and runs using --out. For example, --out results/my_run creates results/my_run_1, then results/my_run_2. An existing run directory is used only when --reuse is passed explicitly.
The progress bar is disabled by default in CLI runs. Use --progress when you want live stage progress rendering.
For SDF inputs, Stage 6 evaluates the indexed source coordinates in input/ligands.sdf by default and intersects them with the latest enabled upstream stage by mol_idx. Set --evaluate-docked-coordinates or evaluate_docked_coordinates: true to additionally evaluate redocked poses under stages/06_docking_filters/docked/. Pose-level outputs contain pose_source=input or pose_source=docked. The root filtered_molecules.csv keeps molecules that pass either coordinate branch.
Use --large-dataset for PubChem/Enamine-scale statistics runs. It stores intermediate row-level tables as compressed shard directories such as filtered_molecules.parts/ and descriptors_all.parts/, uses large_dataset_chunk_rows from the master config, and skips plots plus HTML report generation. Without an explicit --stage selection, large-dataset mode runs the scalable statistics path: mol_prep, descriptors, struct_filters, and synthesis. Descriptor, structural, and synthesis filters are calculated by default, but downstream large-mode outputs keep all molecules unless large_dataset_filter_data: true. The synthesis stage skips AiZynthFinder retrosynthesis in large-dataset mode.
Set alignment.enabled: true and alignment.target_coverage_percent: PERCENT in the master config to align numeric score thresholds to target_mols_path. The --align-config PERCENT option provides a one-run enable-and-percentage override. HEDGEHOG first runs the reference molecules with numeric filtering relaxed and every structural filter and common-alert ruleset enabled in measurement-only mode. Structural results are diagnostic during calibration: the source structural hard gate, filter_* flags, Common Alerts enforcement rulesets, and structural parameters are copied unchanged and never participate in percentile-cohort selection. Synthesis is skipped during target calibration and does not participate in threshold derivation or cohort selection; the generated candidate uses the source synthesis config unchanged. Numeric stages select one shared central/best subset containing ceil(PERCENT / 100 * molecule_count) rows, then derive descriptor bounds from that subset. alignment.descriptor_bounds_mode accepts target or expand. expand preserves the source ZINC+ChEMBL descriptor envelope and only widens bounds required by the target. Docking pivots poses to one affinity per unique molecule and configured tool, selects one shared best subset using the worst normalized tool rank, and writes one minimizedAffinity.max cutoff per tool; candidate molecules must pass all generated tool cutoffs. Generated continuous thresholds use outward rounding; integer-valued thresholds are rounded up. MolPrep categorical rules are preserved. Alignment cannot be combined with large_dataset_mode or --large-dataset.
Aligned stage configs are created one by one under target_alignment/aligned_configs/: descriptors after the descriptor stage, an unchanged structural config plus diagnostics after structural filtering, docking after docking, and docking filters after the docking-filter stage. No complete config bundle is generated upfront or regenerated in bulk at the end. The aligned master and alignment_thresholds.yml are updated after each alignable stage, while unfinished stages remain pending_metrics. The structural diagnostic pass also writes structural_filter_failures.csv, with one row for every target molecule/rule failure. After the aligned candidate run is initialized, its reusable master config and self-contained support files are published under results/custom_configs/; the run-local continuation snapshot remains under <run>/configs/.
When you request stages with --stage, only the selected stages run. Mol Prep is not forced on automatically; include --stage mol_prep if you want standardization first.
Every run writes a transient .RUN_INCOMPLETE marker in the results folder. The marker is removed on successful completion. If it remains present, the run was interrupted or failed and the run log should be checked.
Continue an Interrupted Run
Stop the active process first with Ctrl-C, edit the copied YAML files under the unfinished run’s configs/ directory, and continue in place:
uv run hedge run --continue /absolute/path/to/results/run_1Continuation loads input/sampled_molecules.csv; it does not preprocess or sample the input again. Enabled stages are checked in pipeline order. Completed stages are skipped, and execution starts at the first stage without a completion marker or complete output artifact. Every newly completed stage writes .STAGE_COMPLETE.yml, so later interruptions can be resumed reliably.
By default, continuation loads configs/master_config_resolved.yml and the run-local stage configs, so changes such as reducing AiZynthFinder workers should be made in those copied files. To use a different master config instead, pass it explicitly:
uv run hedge run --continue /absolute/path/to/results/run_1 \
--config /absolute/path/to/changed_config.ymlIf the interruption happened inside target alignment, you may pass the outer run folder; HEDGEHOG resumes the nested target run, updates the remaining aligned stage configs, and then starts the candidate pipeline in the original folder. If alignment.target_coverage_percent was changed in the saved target-run master, completed aligned stages are recalibrated from their saved metrics without rerunning those stages.
Synthesis continuation reuses a compatible AiZynthFinder work directory. Finished shards are not launched again, and partial JSONL checkpoints resume from their last complete molecule. HEDGEHOG refuses continuation when .RUN_INCOMPLETE is absent or another continuation process holds the run lock.
Available Stages
Each --stage occurrence accepts one of the following values:
| Stage | Description |
|---|---|
mol_prep | Standardize and filter molecules |
descriptors | Compute 28 physicochemical descriptors per molecule |
struct_filters | Apply structural filters |
synthesis | Evaluate synthetic accessibility using retrosynthesis (AiZynthFinder) and other metrics |
docking | Calculate docking and binding affinity scores |
docking_filters | Filter docking poses by quality and interactions |
final_descriptors | Recompute descriptors on the final filtered set |
Stage Failure and Skip Semantics
Stage completion is not binary “success or crash” across the whole pipeline. Some stages end the run early, while others can be skipped and let the pipeline continue.
- Early exit with a completed upstream pipeline state: if
mol_prepfinishes but leaves zero molecules, the pipeline stops immediately. - Hard failure / early stop: structural filters stop the run if they do not complete successfully.
Examples
# Run the full pipeline with default config
uv run hedgehog
# Run with a custom master config
uv run hedgehog --config src/hedgehog/configs/config.yml
# Run with a custom molecule set
uv run hedgehog --mols input/my_molecules.csv
# Use a custom output naming base (creates results/my_run_1, then _2, ...)
uv run hedgehog --out results/my_run
# Run with a glob pattern
uv run hedgehog --mols "input/generated/*.csv"
# Run only the docking stage
uv run hedge --stage docking
# Run descriptors and structural filters in one invocation
uv run hedge --stage descriptors --stage struct_filters
# Run a specific stage with custom molecules
uv run hedge --stage descriptors --mols input/candidates.csv
# Rerun into the same results folder
uv run hedge --reuse
# Continue an interrupted run after editing its copied configs
uv run hedge run --continue /absolute/path/to/results/run_1
# Explicit form of the default fresh-folder behavior
uv run hedge --stage synthesis --force-new
# Enable live progress bar
uv run hedge --progress
# Evaluate input coordinates and additionally evaluate docked coordinates
uv run hedge --mols input/generated_3d.sdf --evaluate-docked-coordinates
# Retain at least 95% of targets at each aligned stage; docking uses one
# shared molecule subset and requires all generated tool cutoffs
uv run hedge --align-config 95
# Stream a large library without monolithic intermediate CSVs or plots
uv run hedge --large-dataset --mols input/pubchem.csvhedgehog info
Displays a table of all available pipeline stages with their descriptions. Useful for checking stage names before using --stage.
uv run hedgehog infohedgehog report
Regenerates the reporting artifacts from an existing pipeline run directory without re-running any pipeline stages. Useful when report templates, plotting logic, or the stage-audit notebook have been updated.
uv run hedgehog report <RESULTS_DIR>RESULTS_DIRis a required argument. Set a path to an existing pipeline results directory (e.g.,results/run_10`).
The command loads the saved configuration from configs/master_config_resolved.yml inside the results directory and regenerates:
report.htmlreport_data.jsonstage_filter_audit.ipynb
hedgehog setup aizynthfinder
Installs the upstream aizynthfinder package into the project environment and downloads its public data into modules/aizynthfinder/.
uv run hedgehog setup aizynthfinderhedgehog setup sync
Downloads the SYNC 3D synthesizability classifier checkpoint to modules/sync/classifier_emb.ckpt.
uv run hedgehog setup synchedgehog setup fsscore
Clones upstream FSScore checkout into modules/fsscore so you can point synthesis
configuration to the pretrained checkpoint without adding heavy torch dependencies
to the base HEDGEHOG environment.
uv run hedgehog setup fsscore After checkout, configure:
HEDGEHOG_FSSCORE_PYTHONto an isolated FSScore environmentHEDGEHOG_FSSCORE_MODEL_PATH(orHEDGEHOG_FSSCORE_REPO_PATH)
hedgehog setup gasa
Installs the optional GASA scorer checkout and an isolated worker environment used by the synthesis stage.
uv run hedgehog setup gasa After setup, configure HEDGEHOG_GASA_COMMAND or synthesis.gasa.command with the printed worker command template, or use gasa.executable / gasa.api_url in config_synthesis.yml.
hedgehog setup nonpher-check
Validates optional Nonpher runtime for the synthesis nonpher scorer.
This command does not install dependencies. For portable setup, first use
uv-isolated environments under a writable per-host HEDGEHOG_OPTIONAL_ENV_ROOT
(for example ~/work/hedgehog_optional_envs), then probe with --python.
uv run hedgehog setup nonpher-checkOptions:
| Option | Type | Default | Description |
|---|---|---|---|
--python | TEXT | None | External interpreter to probe (for example ~/work/hedgehog_optional_envs/nonpher/bin/python). |
--probe-smiles | TEXT | CCO | Probe molecule used for runtime validation. |
Examples:
# Probe uv-only isolated Nonpher env
uv run hedgehog setup nonpher-check --python ~/work/hedgehog_optional_envs/nonpher/bin/python
# If uv-only bootstrap is blocked by native deps, probe a validated external runtime
uv run hedgehog setup nonpher-check --python /mnt/ligandpro/shared_storage/data/nikolenko/hedgehog_optional_envs/nonpher-hybrid-py38-v2/bin/pythonhedgehog setup nvmolkit-worker
Creates an isolated virtual environment at .venv-nvmolkit-worker and installs the optional nvMolKit worker there.
uv run hedgehog setup nvmolkit-workerhedgehog setup shepherd-worker
Creates an isolated virtual environment at .venv-shepherd-worker and installs the optional Shepherd-Score backend there.
uv run hedgehog setup shepherd-workerEnvironment Variables
The CLI supports a few environment variables that are useful in automation and CI:
| Variable | Effect |
|---|---|
HEDGEHOG_AUTO_INSTALL=1 | Auto-accept optional dependency downloads and setup prompts. |
HEDGEHOG_NON_INTERACTIVE=1 | Auto-decline download/setup prompts instead of waiting for interactive input. |
HEDGEHOG_PLAIN_OUTPUT=1 | Disable Rich formatting and emit plain console output without the banner styling. |
HEDGEHOG_AUTO_INSTALL is what the --auto-install flag sets internally for a pipeline run. The setup commands also use it when their --yes flag is enabled.
hedgehog tui
Launches the interactive Terminal User Interface (TUI) for visual pipeline configuration and management.
Requirements and behavior:
- Requires Node.js >= 18 and npm.
- Requires a source checkout that contains the repository
tui/directory. - If the TUI bundle does not exist yet, HEDGEHOG automatically runs
npm installandnpm run buildinsidetui/before launching.
uv run hedgehog tuiOptions:
| Option | Short | Type | Default | Description |
|---|---|---|---|---|
--session | -s | TEXT | None | Resume TUI directly into results for the specified job id. |
Examples:
# Start TUI normally
uv run hedgehog tui
# Resume directly into an existing job
uv run hedgehog tui --session 1a2b3c4dFor TUI-specific behavior such as editable config copies, job history storage, preflight checks, and keyboard shortcuts, see the TUI documentation.
The TUI can also be launched directly from the TUI package:
cd tui
npm run tuihedgehog version
Prints the current HEDGEHOG version and project tagline.
uv run hedgehog version