Changelog
Unreleased
Added
Declared per-descriptor precision tolerances and ensemble convergence reports via
--convergence, repeatable--tolerance NAME=VALUE, YAML and Python. Order-independent Student-t half-width planning curves, per-descriptor and overall additional-structure estimates, and text/JSON/CSV/PNG/PDF exports retain missing-data counts and forecast assumptions. Seeded void sampling now assigns draws in a canonical geometry order so shuffled input structures also preserve its convergence observations. Automatic neighbour cutoffs likewise use a canonical order while per-structure output retains input order.--mq-ensemblenow defaults to adaptive burn-in and snapshot spacing based on energy/volume autocorrelation and the slowest species’ diffusion.snapshot_sampling.{json,txt}reports the selected frames and estimated effective independent count; short holds can produce fewer snapshots than requested. Resume checks preserve the snapshot-to-quench mapping. Explicit--select uniform/lastretain legacy selection; standalone extraction and batch trajectories can opt in with--select decorrelated.Opt-in void distributions, bridging/non-bridging oxygen speciation, stress-derived elastic tensors and Voigt/Reuss/Hill moduli, and harmonic vibrational DOS through Python, CLI and YAML. Each supports per-structure JSON, CSV and figure exports. Elasticity and VDOS use the selected live ASE/MLIP calculator; geometric descriptors remain calculator-free.
Fixed
Stage 3 preserves the input cell and bond lengths before heating. Cubic reshaping now requires an explicit
melt: make_cubic: truesetting.Ensemble collection handles flat single-input batch outputs, including the default
--mq-ensemble -n 1and one-input array jobs. Missing expected finals raise an error. The quench-array example uses a separate directory per task.Stage 7 inherits all
opt:settings and applies partialfinal_opt:or CLI overrides without dropping unrelated YAML optimisation settings.The pipeline and standalone stage/batch entry points forward
--dtype/default_dtypeto their calculators instead of using the backend’s default precision.Conversion refuses destinations that would overwrite an input, including symlink and hardlink aliases, before writing any converted files.
Torch-sim rejects an invalid explicit
--model-pathinstead of falling back to a foundation model.Pipeline and random generation resume refuse changed settings before modifying saved outputs, including seeds, density/placement parameters, stage protocols, and model identities. Pipeline resume also checks input contents, stage completion records and readable checkpoints. Old outputs without complete settings provenance require a fresh directory. Exclusive
.amorphgen.lockfiles prevent concurrent writers and release ownership automatically on exit or process termination.
v1.0.0rc4 (2026-09-24)
Added
Random-gen placement 4–10x faster. Profiling a 600-atom SiO2 placement showed 96 % of the 39 s in the min-CN repair pass, a Python loop making 1.4 million small NumPy calls. Its candidate search is now batched (128 trial positions per distance evaluation) and an atom whose environment yields no candidate for 10 consecutive batches is treated as saturated for that pass instead of exhausting the 2500-draw budget. 600-atom SiO2: 33 s -> 9 s; 350-atom IGZO: unchanged 3.3 s; repair outcome within the seed-to-seed scatter of the old pass. Behaviour change: the random draws of the repair pass differ, so a given
seedgives different (still reproducible) structures than rc3-era builds for compositions that need repair.Batched MD through torch-sim (
--hybrid-ensemble --engine torchsim): stages 4–7 run for all structures at once (NVT-Langevin with per-step temperature schedules for the quench, then batched relaxation). Same per-run files andfinal/collection as the ASE path; NVT only; seeded.amorphgen/utils/torchsim_md.py,batch_quench.run_torchsim. Resume is run-level and frame-level: a chunk killed inside an MD stage continues from the last frame common to all its runs (momenta included), so a walltime kill costs at most 100 steps.--batch-size auto(the new default for the torch-sim engine): the chunk size is taken from a GPU memory probe of the first structure (a short relaxation or MD block), using about 40% (relaxation) or 50% (MD) of the card. An integer still fixes it; on CPUautois 16.torchsim_engine.estimate_batch_size().--indices SPEC(80-90,0,5,7-9) for--random-genand--batch-opt, and--pattern GLOBfor--batch-opt: generate or relax only selected structure indices. Seeds are index-derived, so a split ensemble is identical to a full run.Optional torch-sim engine (
pip install "amorphgen[torchsim]",--engine torchsimorengine: torchsim):--batch-optand--random-gen --relaxrelax all structures in one batched torch-sim FIRE call (MACE, SevenNet, LJ; CUDA or CPU) instead of one after another through ASE; outputs are identical in name and format. CHGNet/Buckingham stay on the ASE engine. Python 3.12+ only; no Apple MPS. The optimiser (-O) carries over (LBFGS default) and, with a cell filter, convergence requires |P| <pressure_tol_gpa(default 0.02 GPa) as well as the force tolerance. Relaxation runs in chunks (--batch-size, defaultauto, see below) with outputs written per chunk, and--resumeskips already-relaxed structures. The GPU cache is cleared between chunks and a CUDA out-of-memory error splits the chunk in half and retries instead of aborting the job.--analysecounts each structure once when a directory holds the same stem in several formats (the optimiser writes.xyzand.cif); priority xyz > extxyz > vasp > cif.Global
seed/--seedfor end-to-end reproducibility. Previously only the random placement was seeded; the velocity initialisation and the Langevin thermostat of stages 2–6 used ASE’s unseeded generator, so two runs from the same seed gave different trajectories. A top-levelseed:now derives a separate NumPy stream per stage and per run directory (SeedSequence) for both. Limits (GPU MLIP forces, frame-level resume) are documented in the YAML guide.Polyhedral connectivity analysis (
--connectivity,StructureAnalyser.polyhedral_connectivity(), YAMLconnectivity: true): corner/edge/face sharing between cation-centred polyhedra and the fraction of cations in edge- or face-sharing pairs, the descriptor that separates a corner-sharing network glass from a random packing with the same short-range order. Validated on cristobalite (100 % corner) and rutile (2 edge + 8 corner links per Ti).--rings [PAIR]and--voronoi [ELEMENT]for--analyse(YAML:rings:,voronoi:in theanalysisblock). Ring statistics and Voronoi indices were previously reachable only from YAML and only printed; they are now printed, appended to--save-report, and written asanalysis_rings.{csv,png}/analysis_voronoi.csv.--versionprints the installed AmorphGen version.
Fixed
Review round 2 (~40 findings), the ones that changed results silently: temperature ramps ran one extra segment at
T_start, so the realised rate was n/(n+1) of the configured one (0.68× for T_step 1000); a negativeratecollapsed a quench to one step per segment; the block-average equilibration test compared block means with the whole-run SEM (√n_blocks too strict); Gaussian-smoothed g(r) fell to ~0.5 at rmax becausenp.convolve(mode='same')zero-pads; bond angles in mixed-cation oxides included cation–cation contacts (Mg–Zn–O); bond counts were doubled; the direct S(q) could miss q-vectors in skewed cells; the Voronoi/temperature running window ignored the frame stride. All fixed.Crashes / lost output: the optimiser
.trajwas never written (manualstep()loop bypassed ASE’s observers);traj_format: xyzbroke--resumeandlammps-dumpcould not be written (formats are nowextxyzandtraj,xyzan alias); a torn last trajectory frame discarded the whole checkpoint (now truncated to the complete frames);convertoverwrotes.xyz/s.cif→s.vasp; mixed-case model names (MACE-MPA-0) failed; Ewald raised an unhelpful error without a cell and had no neutralising background for non-neutral cells; LJ dropped self-image pairs beyond L/2; resumed MD logs restarted Step/Time at 0; a pair missing from a user cutoff dict fell back to the largest cutoff instead of “not bonded”.Config / CLI: values typed with their default value were ignored in several modes (
-C FrechetCellFilter,-n 1in--extract-snapshots,-f 0.01with--relax,--format,--cutoff,--sq-weighting); YAMLopt:was dropped by--hybrid-ensemble/--batch-optand--random-gen --relax;1e-3in YAML was read as a string;--mq-ensemble,--hybrid-ensemble,--batch-quenchand--extract-snapshotsall defaulted tomelt_quench_run/;_load_classicalmutated the caller’s config.Random-gen retry ladder: the cell-expansion rungs are skipped (with a log line) when the density or cell length is fixed by the user, instead of silently doing nothing; a density cut by expansion is now logged as a WARNING with the percentage; a reduced minsep no longer carries over after a skipped structure; covalent nonmetal–nonmetal bonds (P–O, S–O, C–N) are no longer softened by
reduce-minsep; atoms stay inside the box withpbc=False.Ring statistics used the wrong nodes. Auto-detection of the ring network sorted element symbols, so for SiO₂ it took O as the ring node and reported 3-rings for cristobalite; single-element networks (a-Si) were routed through a bridging atom of the same species and also gave 3-rings. Nodes are now the least electronegative element (cation / network former) with the most electronegative as bridge, and single-element networks use the direct bonds: cristobalite and diamond both report 6-rings. Re-run any ring statistics produced without an explicit
bond_pair.Heating/cooling ramps crashed with
npt_method: mtk(IsotropicMTKNPThas noset_temperature) and withparrinello-rahman(the trajectory hook wrapped the live atoms between segments, which ASE’s NPT integrator rejects). Aset_md_temperaturehelper updates every integrator’s target temperature, and the hook now writes a wrapped copy (with energy/forces) instead.--random-gendropped the automatic CN tolerance (target_cn, _ = ...), so CLI and batch runs placed with tolerance 0 while the Python API used the class default (e.g. 2 for alloys) and failed placement more often.Batch “reduce M-M minsep” rung used a CN-unaware table, replacing the CN-aware minseps generate_random builds (Si–O 1.33 → 1.44 Å for SiO₂), so the step meant to ease placement made it harder. The batch path now builds the same CN-aware table.
Pipeline defaults differed with and without
--config. CLI parser defaults (eq_high NVT 10 000 steps, eq_premelt 50 000, eq_low 10 000) disagreed withDEFAULT_CONFIG(NPT/MTK 20 000, 100 000, 20 000), so a YAML containing onlymodel:silently changed the protocol. Parser defaults now come fromDEFAULT_CONFIGand only typed flags are passed as overrides in both modes. Behaviour change: bare-CLI runs now use the documentedDEFAULT_CONFIGprotocol (stage 4 NPT/MTK, longer equilibrations); pass the flags explicitly to keep the previous values.--random-genresume is now seed-stable. Per-structure seeds are derived from the structure index (viaSeedSequence) rather than a running counter that advanced on every attempt and was not advanced for skipped indices, so a--resumerun now reproduces exactly the structures a fresh run would generate, while retries of a failed placement still draw fresh randomness. Behaviour change: the seed→structure mapping changed, so a givenseed(YAMLrandom_gen: seed:/ APIseed=) now produces different (but reproducible and resume-stable) structures than it did before this release. Regenerate rather than expecting old seeds to reproduce old structures.
Added after the rc4 upload (on GitHub main, 2026-09-24; not in the rc4 wheel on PyPI)
--sq-partials(YAMLsq_partials: true) for--analyse --sq: the Faber-Ziman partial structure factors S_ab(q) of every element pair, previously reachable only throughstructure_factor_direct(partials=True). First-peak positions are printed,s_<A-B>columns are appended toanalysis_sq.csvandanalysis_sq_partials.pngis written. Direct method only.--pair-panels(YAMLpair_panels: true): one small panel per element pair for the partial g(r) (analysis_rdf_panels.png) and, with--sq-partials, for the S_ab(q) (analysis_sq_partials_panels.png).plotting.plot_pair_panels()is the helper.Per-pair cutoff overrides on the CLI:
--cutoff "In-O=2.6,Zn-O=2.3"fixes the named pairs and keepsauto-rdffor the rest;"auto,In-O=2.6"or"2.4,In-O=2.6"change the base rule. A YAML/API cutoff dict that lists only some pairs is now completed fromauto-rdf(previously the unlisted pairs were silently “not bonded”); an optionaldefaultentry sets the base. The report header names the overrides.cutoff.parse_cutoff_spec()/resolve_cutoffs().Total coordination in the report for elements bonded to several partner types (O in IGZO: one line for O surrounded by Ga + In + Zn, next to the per-pair O-Ga, O-In, O-Zn entries).
StructureAnalyser.total_coordination(centre, partners);--total-cn O/--total-cn "O:In+Ga"(repeatable; YAMLtotal_cn:) request a chosen centre and partner set. Requested totals are also plotted (analysis_cn_total.png+ CSV).Total g(r) always in
analysis_rdf.csv(g(r)_Totalcolumn) for every system; the plot still draws it only with--total-rdf.Conda environment files in
build_tools/:environment.yml(envamorphgen) installs the checkout in editable mode with MACE + CHGNet, andenvironment_dev.yml(envamorphgen-dev) adds the torch-sim engine, pytest and the Sphinx toolchain. Both install through thepyproject.tomlextras, so the dependencies are declared in one place.conda env create -f build_tools/environment.yml.Wider CI. Every push and pull request to
mainanddevnow also runs the torch-free suite on macOS; the torch-sim engine, CHGNet and pymatgen tests on CPU-only PyTorch (previously skipped in CI), with a coverage report; the suite with every dependency at the lowest versionpyproject.tomlallows; the suite against the built wheel; and a ruff check for syntax errors and undefined names. The docs build treats Sphinx warnings as errors and runs fordevtoo, thebuild_tools/conda environments are built and tested when they or the extras change, and the weekly canary covers SevenNet (7net-0) as well as CHGNet. SeeCONTRIBUTING.md.Python 3.13 and 3.14. The torch-free suite now runs on Python 3.10 to 3.14 on Linux, and on 3.14 (was 3.12) on macOS, and the classifiers list 3.13 and 3.14. The
backendsjob and thebuild_tools/conda environments stay on 3.12: CHGNet 0.4.2 publishes wheels up to 3.12 only, so on 3.13 and 3.14 pip compiles it, which needs a C compiler.Supported platforms stated. AmorphGen supports Linux and macOS, the platforms CI tests; native Windows is not supported (use WSL). The installation docs and the classifiers now say so.
CITATION.cffandCODE_OF_CONDUCT.md. GitHub’s “Cite this repository” button now gives the reference the README asks for, in APA and BibTeX. The code of conduct adopts the Contributor Covenant 3.0 and says where to report a problem. A test checks thatCITATION.cffnames the current release, so a version bump has to update it.
Fixed after the rc4 upload (on GitHub main; not in the rc4 wheel on PyPI)
Which contacts count as first-shell bonds, and which run gets which MD seed (2026-09-29). Two rules that several parts of the package share were rebuilt after a sequence of code reviews.
Bonding pairs. The coordination report, the total coordination, the bond angles and the CN plot all ask one question: is an A-B contact a bond? The answer used to come from a fixed anion list, which failed whenever an element on it was acting as the cation. TeO2 and SO3 had no bonds at all (every pair was anion-anion, so the coordination table fell back to Te-Te and the angles came out empty, visible as a blank CN column for TeO2 in the class benchmark); sulfates, nitrates and hydroxides lost their S-O, N-O and O-H bonds the same way. Which elements are anions is now derived from the composition by charge balance: every element of the anion table starts as a candidate, and one is promoted to cation only when that brings the compound closer to neutrality. The most electronegative element is never promoted, so an off-stoichiometry or defective cell always keeps an anion. That single criterion reproduces the chemistry without a table of exceptions: tellurium is the anion in CdTe and the cation in TeO2 and in a tellurite, sulfur the anion in ZnS and in La2O2S but the cation in a sulfate, hydrogen the anion in LiH and the cation in a hydroxide. It also leaves real cells intact, because promoting their major anion would overshoot far past neutrality: the mixed-chalcogen glasses (Ge-S-Se, Ge-Se-Te) keep Ge-Se and Ge-Te, F-doped silica keeps Si-O, an oxygen impurity in NaCl keeps Na-Cl and LiPON keeps P-N. Verified on about 40 compositions.
Two further corrections to the same rule. Cation-cation contacts are no longer bonds even when the radii table calls them covalent, which used to inflate the total coordination in aluminosilicate and soda-lime glasses (Si-(Al+O) = 5) and add spurious angles. And a same-element pair is a bond only in a single-element system or a composition that is at least 70 % metal atoms, which keeps Ni-Ni in Ni80P20, Fe-Fe in Fe80B20 and Cu-Cu in CuZr while dropping Ga-Ga in GaAs, Ti-Ti in TiC and Be-Be in Be2C. The analysis call sites pass element counts rather than a bare set so that fraction can be computed.
MD seed streams. Runs that shared a
--seedcould share their velocities and thermostat noise: the run index came from therun_NNNN/directory name alone, and a single-snapshot run writes straight into the work directory, so every SLURM array task drew the same numbers. A run’s index is now itssnapshot_NNNNnumber when the filename has one and its position in the loop otherwise, which keeps a run’s seed stable when the input set changes or a resume selects differently. That local index is then banded by where the run’s scope came from. Each source has its own range: an explicit--run-indexon an ensemble, aSLURM_ARRAY_TASK_IDon an ensemble, the same two on the single-structure pipeline, and no scope at all. Two runs therefore share a seed stream only when they come from the same source with the same scope and the same local identity. That includes a pipeline run started inside arun_NNNN/directory, the usual SLURM patterncd run_$SLURM_ARRAY_TASK_ID && amorphgen ..., and the directory name now has to match exactly, so a folder calledmyrun_2is not mistaken for one. The banded value travels under its own config key, which a config file may not set, so a stage cannot mistake a local index for an explicit one and band it twice. Out-of-range values are refused rather than wrapped, since foldingsnapshot_100003ontosnapshot_0003would silently merge two streams, and every snapshot filename is checked before the first run starts rather than when its turn comes. A pipeline run started inrun_0003/and one given--run-index 3get the same index on purpose: both say “this is run 3” of the same workflow. Behaviour change: these indices are seed labels and they have moved, so a run resumed across this change draws different velocities and thermostat noise from that point on. Finish a running ensemble before updating, or regenerate it.Classification. The metal-rich metalloid-glass rule no longer reaches the s-block: Li3P, Na3Sb and Cs3Sb are Zintl phases and keep their pnictide treatment, while Ni80P20, Fe80B20 and Pd80Si20 remain alloys.
--tr: the total correlation function T(r) = 4 pi r rho g(r), the curve a diffraction paper plots beside S(q) (StructureAnalyser.total_correlation(),rdf.compute_total_correlation()). It follows the experiment’s own route: the weighted S(q), then the Fourier transform over the measured q range, then the 4 pi r rho factor, with--tr-qrangeand--tr-windowexposing the two choices that decide whether two curves can be compared at all. The reduced PDF G(r) comes with it, andrdf.coordination_from_Tr()integrates a peak of r*T(r) for the coordination number a diffraction paper would quote. This is NOT the unweighted--total-rdfcurve: for a multi-element system the scattering weights matter, and in IGZO the indium correlations dominate the X-ray weighted one.--tr-scan: sweeps the T(r) transform choices, qmax and the window, and reports how far the first peak and its integrated count move (rdf.scan_Tr_qmax(),rdf.format_Tr_scan()). Those two are properties of the measurement, not of the model, so a comparison should carry that spread rather than imply a precision the transform does not have. On a-IGZO the first peak sits at 2.07-2.12 A across qmax 12-25 with a Lorch window and 2.07-2.15 A without one, where the ripple narrows the integration window and the count drifts from 2.8 down to 1.7.The T(r) first-peak finder no longer locks onto the truncation ripple. A peak now has to stand out, with a prominence of at least a fifth of the largest in the curve, sit above 1 A, and have
g(r)above 0.5. Height alone could not do it:T(r)grows as4 pi rho r, so a fraction of the maximum is a threshold on r rather than on the structure, and the ripple that precedes the first shell cleared it. Over 96 transforms of eight amorphous systems at six values of qmax the old rule missed the first shell by more than 0.35 A in 25 cases, every one of them without a Lorch window; the new rule misses 2, both at qmax 12 and for different reasons: one skips a weak first shell in favour of the taller second, the other is residual ripple. The integration window is also clamped at 1 A, so a Lorch-broadened first shell with no clear minimum before it no longer prints a lower bound down at 0.6 A. The unwindowed transform of a-Al2O3 reported a first peak at 1.3 to 1.5 A with a coordination number of 0.01, against a real Al-O shell at 1.82 with 3.7. The bar forg(r)is deliberately well below 1, because with one heavy scatterer the real first shell can sit under the average: the Cu-O shell of Cu2O has a weightedgof 0.81. Everynonerow of--tr-scanand every--tr-window nonesummary was affected, and the scan’s advice line blamed the q range instead.Behaviour change, hybrid-ensemble seeds.
_run_seed_indexpromised that a run’s MD seed survives a change to the input set, but it read the index only from asnapshot_NNNNname, so therandom_NNNNfiles a hybrid ensemble is fed fell back to the loop position. It now readssnapshot_,random_,struct_andhybrid_names. A hybrid ensemble resumed across this change gets different seeds for its remaining runs than for those already finished; start such a set again if that matters.The T(r) summary now reaches the saved report.
--save-reportwrites the file before the S(q) and T(r) blocks run, so the peak line and the--tr-scantable were printed but never saved. Both are appended to the file now, as the ring statistics are.A relaxation no longer reports a false placement stall. The density line added with the warning above was printed after
--relaxhad moved the cell, so it compared the relaxed density with the placement target and claimed “placement stalled and the cell was expanded” on any relaxation that shifted the density by more than 2 %, which is most of them. The placement is now measured before the relaxation and reported asplaced at ...; the relaxed value keeps its ownFinal densityline.The seed reached the ASE engine only.
--hybrid-ensemble --engine torchsimseeded its batched MD from the stage, chunk and resume offset alone, so two SLURM array tasks running the same inputs with one--seeddrew identical momenta and thermostat noise: the problem the run-index work had just fixed for the ASE path. The job’s run index is now part of the torch-sim seed as well.SO3 and SeO2 densities were half their true value. The rule that gives a non-metal acting as an oxide cation its positive Shannon radius only fires when the table lists one, and sulfur and selenium had only their anionic state, so both fell back to an anion radius four times too large (SO3 1.10 against 1.92 g/cm3, SeO2 1.66 against 3.95). Their positive states are now in the table, and a non-metal that is promoted to cation but has no positive state falls back to its covalent radius. The rule is gated on the element actually being a cation in that compound, which charge balance decides, so the chlorine of an oxychloride, the nitrogen of an oxynitride and the hydrogen of a hydroxide keep their ionic radii: giving them covalent ones halved the cell and doubled the density (BiOCl 1.50 of its crystal value instead of 0.75, Si2N2O 1.98 instead of 0.78).
anion_elements()moved toutils.radii, where the radius selector can reach it, and is re-exported fromanalysis.structure.Placement no longer changes the density silently. The auto-expand and auto-retry:minsep messages were
logger.info, and the package configures no logging handler, so a 20-40 % density loss was invisible on the CLI. They are now warnings (visible with no logging setup), andbatch_randomprints and logs every structure’s achieved density, with an explicitWARNING: density X, -44% from the requested Yline whenever it misses the request by more than 2 %. Soft-packed structures are marked[soft-packed at the requested cell].Soft-pack instead of cell expansion when placement jams. Random sequential addition stalls near a 0.38 hard-sphere fraction, which the estimated density of dense oxides (MgO, BeO), alloys, borides and large-cation nitrides reaches, and the old response expanded the cell 5 % per retry, silently losing 20-30 % of the density (benchmark: MgO placed at 1.99 instead of 2.67 g/cm3, CuZr 5.13 instead of 5.94, ZrN 5.38 instead of 6.22). The first response is now a soft pack at the same cell: placement with floors x0.72, then iterative overlap removal (
_push_apart) until every pair reaches 0.985 of its floor, which succeeds up to a hard-sphere fraction of about 0.6. All five jamming benchmark classes now keep their estimated density;info["soft_pack"]marks the structures. Expansion remains the fallback.Density rules caught by the 100-system class benchmark (2026-09-28). Ni80P20, Fe80B20 and other metal-rich metalloid glasses were classed as pnictides / borides (Ni4P estimated at 0.48 of the glass density; now alloy, 0.90). P2O5 fell through to the default class and phosphorus took its P3- anion radius (0.36 of the glass density; now covalent_oxide with P5+, 0.85). Be2+ had no Shannon radius, so Be3N2 was routed to the large-cation nitride class (0.47 of crystal; now small_cation_nitride, 0.77). Sulfide and selenide network glasses (GeS2, GeSe2, As2S3, As2Se3) were 17-48 % too dense under the II-VI chalcogenide factor; a new
chalcogenide_glassclass (packing 0.23, Ge/Si/As/Sb/B/P cations, S/Se anions) puts them within 13 %. Tellurides, III-V pnictides, MB2 / MB6 borides and BeO are unchanged. Side effect: BeF2 now uses the Be2+ radius and estimates 1.13 of the glass density (was 0.97 through an accidental covalent fallback).--seeddid nothing on the torch-sim engine. torch-sim draws initial momenta and Langevin noise from its state generator, which started from torch-sim’s own fixed default seed, so every batch and every resume got identical velocities and noise whatever seed was set, and ensemble members were not independent. The state generator is now seeded from (seed, stage, chunk, resume offset); with no seed each batch gets fresh entropy. Test: seeds 1 and 2 now give different trajectories.SevenNet could not be built on the torch-sim engine: its wrapper is float32-only and multi-fidelity checkpoints need a modal; both are now set as on the ASE path.
--hybrid-ensemble --engine torchsimfailed with default settings (stage 4 defaults to NPT, which torch-sim does not run): unset MD stages now switch to NVT with a note; an explicit NPT is refused before anything starts.Resume could turn a
traj_format: trajtrajectory into extxyz. The torn-frame repair rewrote the file in the format guessed from the*_traj.xyzname; it now keeps the file’s real format (binary trajectory or extxyz), so the resumed stage can append.CN plot of multi-cation compounds showed a cation-cation pair. For IGZO the mirrored-bars layout picked Ga-In / In-Ga (second-shell contacts) because the reciprocal-pair search ran over every pair. The plot now uses bonded pairs only: binary oxides keep the mirrored layout, multi-cation compounds get one panel per cation-centred pair (Ga-O, In-O, Zn-O) plus the anion total (O-(Ga+In+Zn)), which also goes into
analysis_cn.csv.Lower bounds that could not work.
ase>=3.22andmatplotlib>=3.5allowed versions that break:import amorphgenneedsase.filters(ASE 3.23), themtkbarostat (the stage-4 default) needsIsotropicMTKNPT(ASE 3.25), and before 3.6.1 matplotlib’sviolinplot(the density panel of--analyse --save-plot) raisesIndexErrorwith numpy >= 1.24. The bounds are nowase>=3.25andmatplotlib>=3.6.1, and the newmin-depsCI job tests them.--random-genfailed on Windows.random_gen.logwas written in the locale encoding, and cp1252 cannot encode the→,ρandΔχof the auto-derive line, sobatch_randomraisedUnicodeEncodeErrorbefore placing a structure. The log is now written, and read back byrank_from_log, as UTF-8.Tutorial 1 stopped with a
NameErrorin its summary table:cn_all_unrelaxedwas no longer defined. The cell now computes the unrelaxed In-O coordination itself.ASE cross-references in the docs pointed at the retired wiki.fysik.dtu.dk inventory; they now resolve against docs.ase-lib.org.
License metadata in the PEP 639 form.
license = "MIT"(an SPDX expression) andlicense-files = ["LICENSE"]replace the{text = "MIT"}table and theLicense ::classifier, which setuptools deprecated and stops accepting on 2027-02-18. Building from source now needs setuptools >= 77 (pip’s isolated builds fetch it), and the distributions carryMetadata-Version: 2.4, so uploading them needs twine >= 6.1.No deprecated ASE MD calls. Initial momenta come from
thermalize_momenta, the ASE 3.29 name forMaxwellBoltzmannDistribution(the old name is used on ASE 3.25-3.28), and the NVT Langevin thermostat runs withfixcm=False, since ASE 3.28 deprecatesfixcm=Truefor not sampling NVT exactly. Behaviour change: NVT stages no longer pin the centre of mass. It diffuses (about 0.5 Å in 10 ps for 108 Cu atoms at 300 K), a rigid translation that leaves the structure, temperature and energy unchanged, but seeded NVT trajectories differ from earlier versions (they stay reproducible). ASE’s suggestedFixComconstraint is not used: it would stay on the atoms, andIsotropicMTKNPTrejects constrained atoms. The equilibration MSD (compute_msd) is now measured relative to the centre of mass, so a drift of the whole system, from this or from the net momentum that themtkandparrinello-rahmanintegrators conserve, no longer reads as diffusion.cell_filter: cubicwithout ASE’s deprecatedExpCellFilter. The isotropic cell relaxation, the default of--random-gen --relax,--batch-optand--hybrid-ensemble, now runs throughFrechetCellFilterwith hydrostatic strain, which gives the same forces asExpCellFilterfor this deformation. It usesexp_cell_factor=1,ExpCellFilter’s scale for the cell forces, so the convergence test still requires |P| V < fmax;FrechetCellFilter’s default divides the cell forces by the number of atoms, which would let a relaxation stop at about 0.1 GPa at fmax = 0.01 eV/Å. In a check on a strained 108-atom Cu cell, every ASE optimiser took the same number of steps with either filter. An explicit-C ExpCellFilterstill selects ASE’s deprecated filter.Bond-angle plots left out angles above 178°. The histogram bins of
--analyse --save-plotand of the ensemble comparison plots ended at 178°, so linear triplets (every Si-O-Si of ideal β-cristobalite, the trans O-M-O of an ideal octahedron) were not counted, and a triplet type with only such angles became an invisible NaN curve. The bins now end at 180°, which adds a 179° row to the comparison plots’angles.csv.File handles closed. The torch-sim batch quench wrote
batch_size.json, and on resume rewrote the stage logs, through file objects it never closed, leaving the flush to the garbage collector.A quiet test suite.
pytest test/reported about 9,900 warnings, nearly all NumPy 2.5’s deprecation of settingndarray.shape, which ASE 3.29 triggers insideAtoms.copy, the MD integrators and the.trajreader. That warning and the TorchScript andweights_onlynotices of loading a MACE model are now filtered inpyproject.toml, and the tests that raised warnings of their own (unclosed files, Berendsen MD started from rest, expected warnings not asserted, a class-scoped fixture written as a method, which pytest 10 rejects) are fixed.Next steps after
--random-genfound no structures.--random-genwrites to<work-dir>/random_initial/(andrandom_opt/with--relax), but the README, the--examplestext and several guides passed<work-dir>itself to--batch-optor--hybrid-ensemble(and<work-dir>/random_0000.xyzto the pipeline), and--batch-optthen exited 0 having done nothing. The examples now name the subdirectory,--batch-optexits 1 when it has nothing to optimise, and--batch-opt,--hybrid-ensembleand--batch-quenchgiven a random-gen work dir name the subdirectory that holds its structures.CHGNet with
default_dtype: float64failed only after hours of MD. The CHGNet configs in the MQ-ensemble and YAML guides set float64, which CHGNet does not support, and under--mq-ensemblethe error came in phase 3, after stages 1-4, because the melt-quench stages build their calculator withoutdefault_dtype. The configs now leave it atauto, and the CLI refuses the combination before any work starts. Behaviour change: the 7-stage pipeline, which ran such a config at float32 without saying so, now refuses it too.Flags and example files the docs referred to but that do not exist. The sweep example used a
--quench-rateflag (now--quench-steps-per-T), and the validation page, the MQ-ensemble guide and a docstring namedexamples/hybrid_stages_4567_cuda.yaml,examples/mq_stages_1234_cuda.yaml,mq_stages_567.yaml,examples/hpc/with its SLURM scripts, andexamples/test_structure_factor.py, none of which were ever committed. They now use configs and commands that exist. The validation page’s reproduction also generated 160 atoms instead of 400 and collected the results from the pre-rc2run_*/run_0000/layout.Plotting switched notebooks off inline figures.
StructureAnalyser.plot(),plot_sq,plot_rings,plot_pair_panelsand the equilibration plots (convergence_report,plot_msdand the rest) calledmatplotlib.use("Agg"). After one call a notebook showed no more figures,plt.show()only warning that FigureCanvasAgg is non-interactive, and a script lost its interactive backend the same way. The analysis plots,compare_ensemblesincluded, only write files: they now draw on amatplotlib.figure.Figureoutside pyplot, so they leave the backend and the open figures alone, and the files are unchanged. The equilibration plots return pyplot figures on the caller’s backend, so a notebook shows them inline, including those ofconvergence_report()withoutoutput_dir. Without a display, as on a compute node, matplotlib picks Agg by itself.The sdist’s tests could not run. setuptools puts only
test/test*.pyin the sdist, sotest/conftest.pywas missing: from an unpacked sdist 45 tests errored for want of their fixtures, and the MACE tests, which it skips unless--run-maceis given, ran and failed.MANIFEST.innow adds every module undertest/. ThepackageCI job, which ran the checkout’s copy of the suite, now runs the sdist’s.The torch-sim engine’s compiler requirement was undocumented. torch-sim’s neighbour list goes through
torch.compile, which builds kernels with the system C/C++ compiler, so without one the first relaxation stops withInvalidCxxCompiler. The installation page, the quickstart, the backends and HPC guides, the README andbuild_tools/README.mdnow say so, and the installation page gives the commands that install a compiler.README links that 404 on PyPI. The links to the licence,
build_tools/and the tutorials were relative, and PyPI resolves them against pypi.org. They now point at GitHub, and thepackagejob renders the README as PyPI does and fails on a relative link (twine checkdoes not render Markdown, so it passed them).draft-pdf.ymlremoved. It built the JOSS draft frompaper/, which was deleted in June, so it could no longer run. The README’s package layout no longer listspaper/.The stage functions ignored
work_dir=.opt_cell.run,equilibrate.run,melt_cell.run,quench.runandfinal_opt.runtook the keyword into**kwargsand wrote their log, trajectory and output structure to the current directory. They now write them intowork_dir, which is created if missing. Without it they write to the current directory as before, which is whereMeltQuenchPipelineandbatch_quenchrun each stage.Tutorials. Five notebooks left over from the old numbering (
T1_random_gen, the two inT2_MQ_via_7_steps,T3_mix_random_MQandT4_classical_potential) duplicated Tutorials 3-6 and are removed, with the logs that came with them. Tutorial 6 had never been run and now ships with its output; its coordination table used a 3.0 Å cutoff that counted second-shell oxygens (Si 4.6 instead of 4.1) and now closes the first shell. Tutorial 7 stopped with aZeroDivisionError(its colour scale divided by the number of temperatures minus one, and it had been run with one), so its Arrhenius cells never ran. Its trajectory analysis had found nothing anyway:equilibrate.runwrote the trajectory outside thework_dirit was given, and the time axis treated frames, which are saved every 100 MD steps, as one step apart. Its CLI commands used a--no-relaxflag that does not exist, a lower-case ensemble name, the default 0.5 fs timestep and file names from an older layout. The notebook now runs end to end.The
auto-rdfcutoff could stop inside the first shell. The first minimum of g(r) was the first point past the peak that was below half its height and no higher than its neighbours, and in a small cell that can be a flat step or a noise dip on the falling side of the peak. The shipped 64-atom a-Si was cut at 2.53 Å, inside its first shell, and gave a CN of 3.75 instead of 4.00; the 96-atom a-HfO₂ was cut at 2.31 Å, which left out its longer Hf–O bonds (CN 5.47 instead of 5.78). The minimum is now read from g(r) averaged over 0.25 Å, and it has to be the lowest point within 0.25 Å on either side. Where g(r) is zero over a range, the cutoff goes in the middle of it. Cutoffs move outwards, and bonded pairs change most in small cells: a-Si now gets 4.00 at 2.84 Å, and the Ir coordination of the 24-atom IrO₂ model goes from 4.0 to 5.4, the value its README lists. The O–O and cation–cation cutoffs of the non-bonded contacts move as well. Re-run any analysis done with the default cutoff.Ring statistics counted paths that cross the cell. The ring search followed bonds by atom index and ignored which periodic image a bond reached. So a path that came back to another image of its first atom, having crossed the cell, counted as a ring. The 8-atom diamond cell gave 100 % 4-rings, and the shipped 48-atom a-SiO₂ gave 84 % 4-rings; it now has 10 % 3-, 23 % 4-, 45 % 5- and 23 % 6-rings. Each bond now carries its cell offset, and a ring has to close on the image it started from, so a structure and its supercells give the same distribution. The shipped Sb₂O₃ and Sb₂O₅ ensembles (112-120 atoms) change as well; a-Si, a-HfO₂ and the 400-atom a-Ga₂O₃ do not. Re-run ring statistics of small cells.
--referencedropped the Si–O bond check. The analyser names a bond with its elements in alphabetical order (O-Si).examples/reference_a_SiO2.yamlwritesSi-O, so the check found no value and left its row out of the table without saying so. A bond, and the two end atoms of an angle, now match in either order. A coordination entry is still directional:Si-Ocounts O around Si. A reference metric the structures do not have is now listed asn/aand counted in the summary, instead of being left out. Of the shipped references only a-SiO₂ was affected.Random generation mangled oxoanion compounds. The radii rules took every nonmetal for an anion, including the centre of an oxoanion (P in a phosphate, S in a sulfate, C in a carbonate, N in a nitrate, Cl in a perchlorate, I in an iodate) and the H of a hydroxide. So P-O was kept 2.46 Å apart against a 1.53 Å bond (S-O 2.27, C-O 2.24, N-O 2.29, O-H 2.24 Å), and no generated P, S, C, N or H had an O within bonding distance. These centres had no target coordination. Counting S, N, Cl or H as anions in the charge balance gave Li+5 in Li2SO4, Ca+10 in CaSO4, Na+9 in NaNO3 and Mg+6 in Mg(OH)2, which put sulfates, nitrates and hydroxides in the high-valent-oxide class and perchlorates and iodates in the oxyhalide class. Charge balance now decides which nonmetals are cations (
radii.cation_nonmetals). They are the ones theanion_elementsrule promotes, plus C and P. C and P count only when an anion more electronegative than them is present and they balance the charge better as cations than as C4- or P3-. That test keeps the carbide C of SiOC and SiCN, carbides, phosphides and a-C:H as they were. A nonmetal cation bonds to its anions at its Shannon cation radius (P-O 1.26, S-O 1.22, C-O 1.06, N-O 1.04, O-H 0.82 Å, P-S 1.61 Å in Li3PS4). It targets the ligand count of its oxoanion: 4 in PO4, SO4 and ClO4, 3 in CO3, NO3, IO3 and a sulfite, 1 for H. It is solved for its oxidation state, and sized as that cation in the density estimate. Sulfate, carbonate, nitrate, perchlorate and hydroxide estimates go from 52-72 % to 79-88 % of the crystal density. Behaviour change:infer_oxidation_statenow returns the state of a nonmetal cation (P +5 in Li3PO4) instead ofNone.Borates and K/Ba silicates placed cations on top of each other. A metal-metalloid pair with Δχ ≥ 1 was classed as an ionic bond even when both are cations of an oxide. So Na-B was kept only 0.90 Å apart (Li-B 0.70, K-Si 1.31, Ba-Si 1.29 Å), and the coordination-aware placement counted the pair as a bond.
classify_bond(a, b, composition)now applies the compound’s roles, and two cations of a compound with anions are never an ionic bond. Na-B is 2.16 Å, as Na-Si already was. A nonmetal cation and another cation get the right-angle contact across the anion that M-M pairs use. Two of the same element (P-P, S-S, C-C) get their two bonds end to end, which is above their homonuclear bond, so--check-dimersstill flags a P-P or S-S bond. None of the 100 class-benchmark systems changes. The placement still has no seed for a centre (O carries no target CN), so a centre starts with about as many O as Si does in an alkali silicate (P 2.75 of 4, C 2.05 of 3 on average), and it relies on the relaxation to complete the polyhedron.Random generation could not build a-Si:H or a-C:H. H is on the anion table as the H- of LiH, so a-Si:H was classed as a hydride and sized with the Si4+ ionic radius: 15.1 g/cm³ for Si64H8 against a measured ~2.2 (a-Ge:H 32 g/cm³). a-C:H was a covalent carbide with H as its anion, at 3.5 g/cm³ whatever its H content against a measured 1.2-2.0, and H targeted 6 neighbours. As anions, C and H were kept 2.24 Å from everything, so no C-C or C-H bond could be placed, and the placement of both failed after four cell expansions. C, Si and Ge with H and nothing else, with at most one H per host atom, are now the
hydrogenated_networkclass: a-Si:H, a-Ge:H, a-C:H, a-SiC:H, a-SiGe:H. The host keeps what it has without H (Cordero radii, packing factor, minimum separations and bonds of a-Si, a-C, SiC or SiGe), so the estimate runs into those as the H goes to zero. The hosts target 4 bonds and H one, and a network’s C gets the three-bond floor of Si. H is kept at 0.8 of its bond from a host (C-H 0.86, Si-H 1.18 Å); two H can share a host (H-H 1.21 Å in a-C:H, 1.67 Å in a-Si:H) but cannot form H2. In the density estimate H is sized at 0.90 Å, about 10 ų per H, so the density falls with the H content: Si64H8 is 2.30 g/cm³, and a-C:H 2.19, 1.84, 1.52 and 1.24 g/cm³ at 20, 30, 40 and 50 % H. The cells now place at that density with every H bonded to a host; a MACE-MPA-0 cell relaxation takes two Si64H8 cells to 2.19 and 2.22 g/cm³, with every H on one Si.--check-dimersno longer flags the C-C bonds and CH2 pairs of a-C:H. The real hydrides (LiH, MgH2, NaAlH4, TiH2) and all 100 class-benchmark systems are unchanged. Behaviour change:infer_oxidation_statereturnsNonefor these networks (Si50H50 gave Si +1).
v1.0.0rc3 (2026-09-22)
Added
MLIP-optional install: torch is no longer a core dependency, the base
pip installis lightweight (random-gen + analysis + classical potentials), and MACE/CHGNet/SevenNet arrive only via extras ([mace],[chgnet],[sevennet]). With no torch present,--device autoresolves to CPU and calculator-requiring commands fail fast viarequire_backend()with the exact install line.--list-modelsshows every model with installed/missing markers.Numerical-divergence guard. MD and relaxation now raise a clear
DivergenceError(viaassert_finite) the moment an energy or force turns non-finite (before a NaN/Inf frame reaches disk) with an actionable message (the MLIP is out-of-distribution at high T, or the timestep is too large). Guards the most common high-T melt-quench failure mode of universal MLIPs.--retry-mode {expand, reduce-minsep, none}, placement-stall policy for--random-gen:expand(default; grow the cell 5% per retry, right when the density is an estimate),reduce-minsep(hold the cell/density exactly fixed and soften only non-bonded minseps, for fixed-density film / isochoric studies), ornone(no adjustment; a stall raises). Cation–anion bond minseps are never reduced.Min-CN floor (default). Every atom gets a hard coordination floor (auto anions = 2, cations = 3, each capped at the element’s target CN) and a post-placement
_repair_min_cn()pass relocates below-floor atoms, cutting dangling bonds (IrO₂ dangling-O ~22% → ~3% at placement, < 1% after relaxation).Frame-level MD resume:
--resumenow continues an interrupted MD stage (2–6) from the last frame of its trajectory (momenta carried), on top of the existing stage-level skip.Oxyhalide material class (e.g. BiOCl, NaTaOCl₄): packing factor interpolated by halogen fraction between metal-oxide and halide, with a 10% dopant gate so trace halogens (F-doped TiO₂ / FTO) keep their oxide routing.
Homonuclear dimer detection (
--check-dimers): flags peroxide-type O–O and other same-element close pairs, skipping metal self-pairs when anions are present. Plus--sq/--sq-weightingto expose the direct S(q) method on the CLI.--mq-ensemblemode: full melt-quench ensemble in one CLI command. Stages 1-4 from a crystalline input, then N independent quenches via auto-extracted snapshots from the stage-4 trajectory, collected tofinal/.--hybrid-ensemblemode: take a directory of disordered structures and run stages 4-5-6-7 on each.--rank-from-logmode: parse a random-gen log and rank structures by total energy (no calculator re-evaluation needed; works for VASP/CIF outputs that don’t carry per-atom energy).--extract-snapshotsmode: utility CLI to extract N uniformly-spaced frames from any trajectory file.--reference YAMLflag for--analyse: validate computed structural metrics against literature ranges defined in a reference YAML; produces a match/concern/fail verdict per metric.Polymorphic
--snapshot-dir: for--batch-quench: accepts either a directory of static structures or a single trajectory file (auto-extracts internally).SevenNet backend: integrated via the
sevennpackage. Supports the multi-fidelity foundation models (7net-mf-ompa,7net-l3i5,7net-omat,7net-0, …) with automaticmodalselection for multi-fidelity variants.Publication-quality plotting:
--save-pdf(vector PDF),--dpi N,--show-title, Okabe-Ito colour-blind-safe palette, clean spines, proper unit symbols (Å, °).--resumesupport for--random-gen: skips completed structures on disk and continues from the first missing index. Validates files are non-empty and ASE-readable. Writesrun_metadata.jsonand warns if composition changes between runs.Calculator pre-warm for
--random-gen --relax: model load + first-inference happen once before the loop, so per-structure timing reflects only relax cost, not setup.Per-structure wall-time logging in
--random-gen --relax: each structure’s log showsWall time: X.XX s (N steps, Y s/step)for diagnosing slowdowns.
Fixed
Buckingham+Coulomb: Ewald summation replaces the Wolf sum. The Wolf (damped-shifted) Coulomb sum used by
BuckinghamCalculatorwas checked against a reference Ewald sum on a 576-atom GeO2 melt and found to be ~10 % off in forces and ~100 meV/atom off in energy differences at its default alpha = 0.2, rc = 10 A, with no setting that fixes it inside a 20 A box. The Coulomb term is now a full Ewald sum (real-space over the pair cutoff with alpha = 3.5/rc, NumPy reciprocal-space sum, self term); it reproduces the NaCl Madelung energy to 4 decimals and a reference Ewald to 0.02 meV/atom. Atom-self-image pairs (cutoff > L/2) are now counted.coulomb_method: wolfkeeps the old scheme for comparison.Auto-RDF cutoff on unrelaxed structures.
auto_cutoff_rdftook the first bump of g(r) above a fixed threshold as the first peak, which on unrelaxed random placements latched onto noise on the rising edge and returned a cutoff below the bond length (Si–O CN ≈ 0.2 in--analyseof raw random-gen output). The first peak is now the first local maximum at ≥ 50 % of the strongest feature, and the first minimum must be a genuine depletion (g ≤ half the peak). Cutoffs on relaxed structures are unchanged; when no minimum exists the radii-table cutoff is used with a warning.Critical:
batch_quench.pystage-numbering bug: the dispatch loop used the old 6-stage numbering (if s==4: quench; s==5: eq_low; s==6: final_opt) instead of the canonical 7-stage numbering (s==4: eq_high; s==5: quench; s==6: eq_low; s==7: final_opt). With the CLI default of--batch-stages 5 6 7, this caused the controlled cooling step to be silently skipped, runs did NVT-eq-at-300K + final-opt instead of quench + eq_low + final_opt. Re-run any batch-quench output produced before this fix if methodology accuracy matters (e.g. publication). Unknown stage numbers now raiseValueErrorinstead of silent skip.Resume bug in equilibrate stages (2, 4, 6): the trajectory file and the stage’s final-output checkpoint shared the same default name
stage{N}_eq.xyz. This had two effects: (1) successful runs overwrote the trajectory data with a single-frame final state, losing trajectory history; (2) interrupted runs left a partial trajectory file at the checkpoint location, causing--resumeto wrongly skip the stage. Fixed by splitting the defaults: trajectory →stage{N}_eq_traj.xyz, final output →stage{N}_eq.xyz. Pre-fixstage{N}_eq.xyzfiles are ambiguous and should be deleted before resuming.--extract-snapshotsnow honours--format. Previously the mode was hardcoded to write extxyz.xyzfiles regardless of the--formatflag, which silently ignored--format vaspand--format cif. Now writes the correct format with the correct extension (POSCAR-style withsort=Trueforvasp).--extract-snapshotscount flag unified. Both-n/--n-structures(the standard count flag used everywhere else) and the legacy--n-runsnow control the snapshot count.--n-runsis preserved for backwards compatibility with existing scripts.Critical: FT structure factor
S(q)had two errors: (1)compute_structure_factorintegratedr·(g−1)·sinc(qr)instead of the 3D isotropicr²·(g−1)·sinc(qr)(one factor ofrshort. (2) The partial S(q) used the partner-species densityn_b/Vin the transform prefactor instead of the total number density ρ₀ that the Faber-Ziman definition requires, scaling every partial byc_b; since the x-ray/neutron weighted totals are built from these partials, they were off by a composition-dependent factor (exactly 0.5 for a 50/50 binary) with equal scattering lengths the weighted total must equal the unweighted one, and it didn’t). Both fixed; the equal-scattering-length identity is now a regression test.compute_structure_factor_directwas unaffected by either and remains the recommended method for the FSDP. Re-generate any S(q) produced via the FT method.Same-element partial
g(r)was a factor of 2 too low._compute_partial_rdf_frame(used byplot_rdf_time_windows) counted undirected same-element pairs (i<j) but normalised with the directed pair-density, so A–A partials asymptoted to ~0.5 instead of 1. Now counts both directions, matchinganalysis.rdf.compute_rdf.Trajectory-based time axes (and fitted diffusion coefficient) off by the trajectory stride:
compute_msd/plot_msd(and the other trajectory-fed diagnostics (plot_energy_convergence,plot_temperature,plot_block_averages,plot_rdf_time_windows,compute_cn_vs_time/plot_cn_vs_time,convergence_report)) assumed one MD step per frame, but AmorphGen writes one frame perTRAJ_LOG_INTERVAL(100) steps. The time axis was 100× too short (andD100× too large, which could misclassify a frozen system as liquid); running-average windows were likewise 100× too wide. All now take aframe_strideparameter (defaultTRAJ_LOG_INTERVAL). Behaviour change: on a trajectory input these functions now interpret the frame spacing astimestep_fs × frame_stride; a script analysing a non-AmorphGen trajectory that stores every step must passframe_stride=1to keep the old time axis. Log-file inputs (which carry a real time column) are unchanged.Melt/quench temperature ramps hardened. The melt ramp used
range()(crashed on a floatT_step, and overshot the endpoint on a non-divisible span); the quench ramp used awhileloop with no guard against a zero or mis-signedT_step(infinite loop) and dropped the endpoint on a non-divisible span. Both now useresolve_ramp, which infers direction from the endpoints, supports float steps, always lands exactly onT_end, and never overshoots.Classical calculators + variable-cell now fail clearly: Lennard-Jones and Buckingham implement only energy+forces (no stress), so an NPT stage (or a cell-filter relaxation, including the default
cell_filter='cubic'of--random-gen --relax) used to crash with an opaque ASEPropertyNotImplementedError. A capability guard now raises an actionable error up front (use a stress-capable MLIP, or a fixed cell + NVT).Melt/quench ramp is validated before any file is touched. The ramp schedule (including the zero-step check) is now resolved before
attach_outputsopens the log/trajectory, so a badT_stepraises without first truncating an existing trajectory.Diagnostic time axes default to the pipeline timestep. The trajectory diagnostics in
utils.equilibrationdefaultedtimestep_fs=1.0whileDEFAULT_CONFIGruns at 0.5 fs, giving a 2× time axis (and D/2) on default-config trajectories analysed with library defaults. The default is now sourced fromDEFAULT_CONFIGand the docstring corrected. Always pass your run’s actual timestep if it differs.Three neutron scattering lengths corrected.
_NEUTRON_B(theweighting="neutron"table) was checked entry-by-entry against the printed Sears (1992) Table 1. Fifty-two of fifty-five matched; three did not and are now the Sears values: Cd 5.1 → 4.87, W 4.755 → 4.86, Au 7.90 → 7.63 fm (2–5% errors). Neutron-weighted S(q) for systems containing these elements changes accordingly; all other elements, and all x-ray/unweighted results, are unaffected. The x-ray form-factor table was likewise spot-checked digit-for-digit against the printed Waasmaier–Kirfel table (N, O, F, Ni, Cu, Zn, Ga, Ge, As); two last-digit transcription differences (Ni b₂, Cu b₁, effect on f(q) ≈ 5×10⁻⁶) were aligned to the print. Both tables are now pinned to their printed sources by a regression test.--analyse --save-reportno longer aborts when the report’s folder doesn’t exist yet.save_reportnow creates the parent directory, matching--save-plot. Previously the run crashed withFileNotFoundErrorafter the structural summary but before S(q), the reference validation and the plots were written.--random-gen --relaxhonoursopt: cell_filterfrom YAML. Cell-filter precedence is now CLI >random_gen:> explicitopt:>cubic. The shippedexample_classical.yamlsetscell_filter: noneunderopt:(classical potentials have no stress); previously that was ignored for random-gen, so the new stress guard told users to set a value their YAML already contained. The guard’s message now also names where the setting goes (-C none, orcell_filter: noneunderopt:/random_gen:).
Added
--sq-method {direct,ft}for--analyse --sq(defaultdirect; alsosq_method:in the YAMLanalyse:block).directis the Debye sum at reciprocal-lattice q-vectors (no real-space truncation; resolves the FSDP; CSV includesn_per_bin).ftis the Fourier transform of g(r) truncated at L/2, smoother, but it damps and shifts the FSDP (a-Ga₂O₃: first peak 2.09 → 1.59, 2nd/1st ratio 0.84 → 1.17 vs 0.86 measured), so it is offered for comparison with FT-based codes rather than as the default; the CLI prints a truncation note when it is used. Both methods are Faber-Ziman weighted.--sq-smooth SIGMA_Q(andsigma_q=oncompute_structure_factor_direct/StructureAnalyser.structure_factor_direct;sq_smooth:in YAML). Gaussian re-binning of the direct-method S(q) in q, weighted by the number of q-vectors per shell, statistically a wider, softer bin, so it reduces the low-q speckle noise (each reciprocal-lattice vector is one noisy sample; shells hold few of them at low q) without moving peaks. The CLI and plots apply σ_q = 0.05 by default (DEFAULT_SQ_SMOOTH; pass--sq-smooth 0for the raw shell averages), while the library functions default to 0 so API callers get the exact values; the raw values are always kept in the CSV ass_q_raw. The S(q) plot’s y-axis now starts at 0 (the Faber–Ziman low-q dip below zero is clipped from the plot only). Keep it well below the FSDP width (~0.3 Å⁻¹): on the a-Ga₂O₃ ensemble σ_q = 0.05 halves the noise for a 5% first-peak cost (2nd/1st ratio 0.84 → 0.87 vs 0.86 measured), whereas ≥ 0.12 starts to damp the FSDP the way the FT method does.Partial structure factors from the direct method:
compute_structure_factor_direct(..., partials=True)(andStructureAnalyser.structure_factor_direct(partials=True)) returns the Faber–Ziman partials S_ab(q) for every element pair from the same reciprocal-lattice sum as the total, so they resolve the FSDP (the FT partials are L/2-truncated). Per q-vector, S_ab = 1 + (Re⟨F_aF_b*⟩/√(N_aN_b) − δ_ab)/√(c_ac_b) with F_a = Σ_{i∈a} e^{iq·rᵢ}; the weighted sum Σ(2−δ_ab)c_ac_bf_af_bS_ab/⟨f⟩² rebuilds the total to machine precision (tested at 10⁻¹⁴). Enables partial-by-partial comparison with isotope-substitution neutron data (e.g. GeO₂, Salmon et al., Nature 435, 75 (2005)).Reference YAMLs: for a-SiO₂, a-GeO₂ and a-HfO₂ added to
examples/(the a-HfO₂ ranges are deliberately broad, tighten against the model set you compare to).Custom-calculator injection into the full pipeline.
MeltQuenchPipeline(..., calc=<ase calculator>)now runs every stage with a user-supplied ASE calculator, bypassing theget_calculator()backend factory (subject to the stress guard above for NPT / cell-filter stages).
Changed
Default RDF smearing is now σ = 0.05 Å (was 0, the raw histogram) for
compute_rdf,StructureAnalyser.rdf(), the analysis plots/CSVs and the CLI--smearingflag, a single constant,amorphgen.analysis.rdf.DEFAULT_SMEARING. This is comparable to thermal/experimental broadening, so simulated g(r) compares naturally with diffraction data: peak positions are unchanged, heights drop and widths grow (a-Ga₂O₃ Ga–O peak 7.5 → 5.1, FWHM 0.125 → 0.203 Å). Pass--smearing 0/sigma=0.0for the raw histogram. S(q) is unaffected: both structure-factor methods always transform the raw g(r), since smearing would damp S(q) by exp(−q²σ²/2) (~16% at q = 12 Å⁻¹). Coordination cutoffs (auto-rdf) are also unaffected (they use their own histogram).X-ray S(q) now uses q-dependent atomic form factors: both
compute_structure_factor(FT) andcompute_structure_factor_direct(Debye) weighted x-ray totals with the constant atomic number Z. Real form factors f₀(q) fall off with q at element-specific rates (at q = 2.45 Å⁻¹, f/Z = 0.81 for Ga but 0.71 for O), so constant-Z systematically mis-weights the partials at finite q, for a-Ga₂O₃ it under-weights the Ga–Ga-dominated first peak by ~10%.weighting="xray"now uses the Waasmaier–Kirfel (1995) 5-Gaussian f₀(q) (98 elements, H–Cf; new public helperxray_form_factor(symbol, q)), with the Faber–Ziman weights and the Debye self-scattering offset evaluated per q. Validated: f₀(0) = Z for every element, and agreement with the International Tables for Crystallography Vol. C Cromer–Mann set to < 0.5% for oxide/semiconductor elements; the FT result was also cross-checked against an independent structure-factor code on the same structures (RMS 0.010, was 0.048 with constant Z). X-ray S(q) values change (neutron and unweighted are unaffected): for the PBE0 a-Ga₂O₃ ensemble the direct-method 2nd/1st peak-height ratio moves from 0.94 to 0.83, against 0.86 measured (Liu et al., Adv. Mater. 2023).Material classification expanded to 20 class-aware packing regimes, each drawing radii from the right source (Shannon ionic / Cordero covalent / Goldschmidt metallic) so a bare composition auto-derives a physical density and minseps. Adds a joint charge-balance oxidation-state solver for mixed cation/anion compounds and a high-valent d⁰ → CN=6 override; boride and oxyhalide packing factors calibrated against crystal densities.
amorphgen.analysispackage now exportsrank_from_log,format_log_ranking,validate_against_reference, andformat_validation_report(previously onlyStructureAnalyser).CLI help reorganised into argument groups (modes / calculator / optimisation / pipeline / random-gen / batch-quench / batch-opt / analyse) for readability. All flags continue to work; only the help-text layout changed.
Default MD timestep in
DEFAULT_CONFIGis now 0.5 fs (was 1.0 fs). Safer for heavy elements and unusual chemistries. Shipped example YAMLs explicitly set 1.0 fs which is fine for typical oxides with chgnet/MACE foundation models.make_cubicreshape moved from start of stage 3 (melt) to start of stage 4 (eq_high). Reshaping a fully molten liquid is benign (atoms diffuse and lose memory of the deformation in <1 ps); reshaping a still-crystalline structure at the start of stage 3 caused a small unphysical jolt at low T. The flag is now read from theeq_high:block, falling back tomelt.make_cubicfor one release as a backwards-compat bridge. Default behaviour is unchanged (cubic reshape on by default).
v1.0.0rc2 (2026-05-22)
Changed (breaking)
--random-genoutput layout. Initial structures are now written to<work_dir>/random_initial/and relaxed structures to<work_dir>/random_opt/(was: both flat in<work_dir>/). This makesamorphgen --analyse --input-dir <work_dir>/random_opt/work without any*_opt.vaspfiltering. The per-structurerandom_NNNN_opt.logfile moved intorandom_opt/alongside its structure.Default analysis cutoff changed from
"auto"(minsep-based) to"auto-rdf"(first-RDF-minimum). This is the standard convention in neutron-diffraction analysis of glasses, and avoids systematically under-counting coordination for materials with broad first-shell distributions.--default-dtypedefault changed from"float64"to"auto", which resolves per-backend:float32for CHGNet (its only supported dtype) and classical potentials;float64for MACE and SevenNet. Previously, users running CHGNet had to remember to pass--default-dtype float32explicitly, otherwise the run crashed withNotImplementedErrorfrom_load_chgnet. Explicitfloat32orfloat64flags continue to work as before.--dtypeis the new short form of--default-dtype.--default-dtypeis still accepted as a legacy alias, so existing scripts and YAML configs (which use the underlyingdefault_dtypekey) are unaffected.Single-snapshot
--batch-quench/--hybrid-ensembleruns no longer nest an extrarun_0000/subdirectory. When the work-dir contains exactly one snapshot (typical of SLURM array workflows where each task processes a single input), outputs land directly underwork_dir/instead ofwork_dir/run_0000/. Multi-snapshot runs still write towork_dir/run_NNNN/per run.Old layout:
hybrid_runs/task_0000/run_0000/final_amorphous.xyzNew layout:
hybrid_runs/task_0000/final_amorphous.xyz
Added
weightingparameter onstructure_factor():"unweighted"(default, current behaviour, a single FFT of the all-atom g(r)).amorphgen.analysis.compare_ensembles()andEnsembleSpec: multi-ensemble overlay plots (RDF, coordination, bond angles, density) with one call. Used by the new Validation docs page.Per-structure density violin: in
--analyse --save-plot: newanalysis_density.{png,pdf,csv}output alongside the existing RDF / CN / angles plots.Validation docs page (
/validation/) with four sub-tabs (a-Ga₂O₃, a-SiO₂, a-HfO₂, a-Si). Each tab includes results-vs-reference table, structure render, validation figure, and reproduce-it commands.
Fixed
Spurious M-M-M triplets in bond-angle analysis of binary oxides. In a multi-element compound with at least one ionic pair (e.g. HfO₂), same-element metallic pairs (Hf-Hf) are second-shell contacts mediated by the anion, not real first-shell bonds; they are now excluded from
compute_all_angles. Pure-metal alloy systems (NiTi, CuZr) keep their X-X first-shell bonds as before.Per-structure E/atom column: in
--analyse --per-structurewas alwaysN/Afor structure files that don’t carry energy in their header (VASP, CIF). The analyser now falls back to parsing the siblingrandom_gen.log.