HPC deployment
AmorphGen runs CPU and GPU workloads on HPC clusters through Slurm. Install the package and the backend used by your job first; see Installation.
Generate a portable workflow
amorphgen-slurm generates Slurm job scripts from one YAML workflow. Use it
for a single job, independent array tasks, or a dependency chain on any Slurm
cluster. It writes one .slurm file per job and a submit.sh script; generation
does not submit jobs.
Start with the example workflow, edit its paths and resource requests, then generate and submit it from the repository:
amorphgen-slurm examples/slurm_workflow.yaml --output-dir jobs
bash jobs/submit.sh --account=your-project
The example chains random generation, relaxation and analysis. The submission
script creates the log directory, submits prerequisites before their dependants,
and passes the returned job IDs to Slurm. It records successfully submitted
job IDs in jobs/submitted.*.tsv, including when a later submission fails.
Additional submit.sh arguments are passed to every sbatch call, so account
and partition choices can stay local.
Running it again submits a new workflow: wait for the previous one to finish or
cancel it before resubmitting into the same output directories. Use --force
to regenerate existing script files after editing the YAML.
A small CPU workflow illustrates the format:
version: 1
profile: generic
root: .. # relative to this YAML file's directory
resources:
ntasks: 1
cpus-per-task: 2
mem: 4G
time: "01:00:00"
environment:
modules: []
venv: /path/to/your/venv
checkpoint:
signal_seconds: 120
requeue: false
jobs:
- name: generate
array: "0-9%4"
work_dir: runs/generate/{task_id}
commands:
- [amorphgen, --random-gen, --composition, "SiO2*16", "-n", "10",
--seed, "{task_id}", --work-dir, "{work_dir}", --resume]
- name: analyse
array: "0-9%4"
work_dir: runs/analyse/{task_id}
needs: [generate]
dependency: aftercorr
commands:
- [amorphgen, --analyse,
--input-dir, "{root}/runs/generate/{task_id}/random_initial",
--save-report, "{work_dir}/report.txt"]
root defaults to the YAML file’s directory. Relative work_dir paths are
resolved under that root. Commands run from root, and the script creates
work_dir before starting them. Array jobs must include {task_id} in
work_dir, so each task can write to its own directory; use {work_dir} in
output arguments to keep those outputs isolated. Commands are lists of arguments,
with {root}, {work_dir} and {task_id} placeholders; shell expansions,
pipes and redirects are not interpreted. Include --resume explicitly on
AmorphGen commands that should resume. The generator does not add CLI flags.
resources uses Slurm option names, including ntasks, cpus-per-task,
mem, time, gres, partition, account and qos. Request GPUs with,
for example, gres: gpu:1. Each job can override the shared resources with
its own mapping. The generic defaults are one task, four CPUs, 8 GB of memory
and one hour. environment.modules lists modules to load and
environment.venv activates the chosen virtualenv (a relative path resolves
under root). Without venv, the generic profile uses AMORPHGEN_VENV if
set, otherwise it keeps the existing environment.
Use --profile bluebear to select the BlueBEAR preset, or
profile: bluebear in the YAML. This loads bear-apps/2024a/live and
Python/3.12.3-GCCcore-13.3.0, and chooses bbgpu QoS for GPU requests or
bbdefault otherwise. Explicit resource and module settings take precedence.
Set environment.venv or export AMORPHGEN_VENV for this profile.
Environment and checkpoint settings apply to the whole workflow.
Arrays and dependencies
An array expression such as 0-9%4 runs ten tasks with at most four active
at once. Each task receives the same resource request; {task_id} becomes its
Slurm array index. needs names prerequisite jobs in the workflow. Choose one
dependency condition for each dependent job:
|
When the dependent job becomes eligible |
|---|---|
|
All prerequisite jobs complete successfully; for arrays, all tasks must succeed. |
|
The matching task in each prerequisite array succeeds. Requires matching array task IDs. |
|
All prerequisite jobs finish, regardless of success. |
|
The prerequisite jobs finish unsuccessfully. |
Use aftercorr for a generation → relaxation array where task 3 consumes only
task 3’s outputs. Use afterok for a collector that needs an entire array.
Slurm’s job-array guide explains the
scheduler’s dependency rules. If an upstream job fails, downstream afterok
jobs cannot run; inspect their state before submitting a replacement chain.
Checkpointing on signals and preemption
Generated jobs request --signal=B:USR1@120 by default. The batch script
turns USR1 or TERM into checkpoint requests for AmorphGen/Python child
processes, enables AmorphGen’s checkpoint handler, and stops before starting
further commands. AmorphGen exits with status 75 after a checkpoint interruption, so an unfinished run does not
release an afterok dependency as successful.
The warning interval comes from checkpoint.signal_seconds. This Slurm option
requests a warning before the walltime limit; the signals and grace period
provided during actual preemption depend on the cluster’s configuration.
See the Slurm signal option.
Choose an interval long enough for the slowest checkpoint boundary and file
writes on your workload. An immediate SIGKILL cannot be handled.
On a handled signal, MD stops at a normal trajectory checkpoint boundary
(usually every 100 steps). Completed structures and stages remain available;
an interrupted optimisation restarts its current structure or batched chunk,
and random generation resumes from completed structures. Restarted MD does not
restore every thermostat, barostat or random-number state, so it is not a
bitwise continuation of the interrupted trajectory. These guarantees apply to
directly launched AmorphGen commands that support resume. Local Python API
programs must wrap their work in amorphgen.utils.preemption.checkpoint_signals()
to enable the same handlers. Native programs and remote srun job steps need
their own signal forwarding and restart integration; the bundled supervisor
forwards warnings only to local AmorphGen/Python processes.
The default checkpoint.requeue: false disables explicit requeue requests.
With requeue: true, a USR1 interruption requests scontrol requeue after the
child command stops; the cluster must permit user requeue. TERM does not
explicitly request requeue, leaving cancellation and preemption decisions to
Slurm. Keep --resume in resumable commands and preserve the same inputs,
settings and work directories when a job starts again.
Configuring the bundled examples
The 27 existing examples/*.slurm scripts remain available for the original
BlueBEAR experiments. They use BlueBEAR module and QoS names. Adapt those
settings to your cluster, choose your own allocation account at submission,
and export the path to a virtualenv containing AmorphGen and the required
backends:
export AMORPHGEN_VENV=/path/to/your/venv
export AMORPHGEN_ROOT=/path/to/AmorphGen
cd "$AMORPHGEN_ROOT"
mkdir -p logs
sbatch --account=your-project examples/run_ensemble_resume_bluebear.slurm
AMORPHGEN_VENV is required; the scripts stop with a setup message if it is
unset or empty. Use a virtualenv compatible with the Python module loaded by
the script (Python 3.12 or newer for torch-sim). Scripts that read files from
the repository use AMORPHGEN_ROOT, defaulting to SLURM_SUBMIT_DIR (or the
current directory when run directly). Follow each script’s input/config
instructions. Create logs/ before calling sbatch, so Slurm can open its
log files.
The scripts have no embedded allocation account. Pass --account=your-project
on each submission, or set export SBATCH_ACCOUNT=your-project when submitting
several jobs, including dependency chains. Keep these settings in your shell
or a local submission wrapper. #SBATCH directives do not expand shell variables.
The scripts share a signal supervisor and request a walltime warning. As with
generated jobs, a handled interruption exits with status 75 by default. Set
AMORPHGEN_SLURM_REQUEUE=1 to request requeue after USR1 when permitted by
your site. The scripts’ --requeue directive allows scheduler-managed requeue;
it does not, by itself, resubmit every interrupted job. Override module names
with AMORPHGEN_BEAR_MODULE and AMORPHGEN_PYTHON_MODULE, or set
AMORPHGEN_SKIP_MODULES=1 when the environment is already prepared.
For the existing SiO₂ experiment, submit its four generation tasks with a concurrency limit and release relaxation only once all four succeed:
generate_id=$(sbatch --parsable --account=your-project --array=0-3%2 \
examples/run_sio2_gen_array_bluebear.slurm)
sbatch --account=your-project --dependency="afterok:${generate_id%%;*}" \
examples/run_sio2_relax_bluebear.slurm
Each SiO₂ generation task keeps its checkpoints in
sio2_100/placements/task_ID/, publishing its completed structures through
disjoint links in sio2_100/random_initial/ for the relaxation job. Preserve
the task index ranges defined by the scripts: they often select a specific
material or structure slice. Singleton scripts can have shared output names;
use the YAML generator with {task_id} work directories for new arrays.
Updated generation scripts use deterministic seeds for reproducible restarts. If an older run used an unspecified seed or different settings, its saved metadata may reject resume with the new script. Use a fresh output directory for that protocol; preserve the old checkpoints instead of overwriting them.
Resuming timed-out jobs
--resume reuses completed stage outputs and saved MD frames in the work
directory. Resubmit with the same input, model, configuration and work directory;
use a new directory when changing the simulation protocol. Pipeline and random
generation runs verify saved settings before reusing outputs and reject
concurrent writers with an exclusive .amorphgen.lock in the output directory.
The lock is released automatically when a job exits or is killed. Leave the
lock file in place; its existence does not indicate an active job. Runs from
older versions without complete settings provenance need a fresh directory.
Pipeline mode
amorphgen POSCAR \
--stages 1 4 5 6 7 \
--config my_config.yaml \
--work-dir my_run/ \
--resume
If stages 1 and 4 are already complete, AmorphGen picks up from stage 5 using the stage4_eq.xyz checkpoint. If all stages are done, it exits immediately.
Batch quench mode
amorphgen --batch-quench \
--snapshot-dir snapshots/ \
--model mace-mpa-0 \
--device cuda \
--resume
This skips already-completed structures and continues from where the previous job left off.
Ensembles on the torch-sim engine
amorphgen --hybrid-ensemble \
--input-dir random_structures/random_opt/ \
--config hybrid.yaml \
--model mace-mpa-0 --device cuda \
--engine torchsim --batch-size auto \
--work-dir hybrid_run/ --resume
With --engine torchsim (see Calculator backends) the structures are processed
in batched chunks. Relaxed structures are written after every chunk and MD
trajectories every 100 steps. After interruption, --resume skips finished
runs, restarts unfinished relaxation chunks, continues a partly done MD stage
from the last valid frame common to the chunk, and reuses the
chunk size recorded in quench_runs/batch_size.json under the hybrid work
directory. Keep the input list and chunk size unchanged when resuming.
Resubmitting the same job script until the log reports the ensemble complete
is the intended way to run a large ensemble through a short queue. Put
export PYTHONUNBUFFERED=1 in the script, otherwise the progress messages
only appear in the log when the job ends. The engine compiles kernels when it
starts, so the compute node needs a C/C++ compiler on PATH: if g++ --version
fails there, load one in the script (module load GCC; the name varies by
site). See
the installation page.
Python API
from amorphgen import MeltQuenchPipeline
pipe = MeltQuenchPipeline(
input_file="POSCAR",
work_dir="my_run",
cfg_override={"model": "mace-mpa-0", "device": "cuda"},
)
atoms = pipe.run(stages=[1, 4, 5, 6, 7], resume=True)