Random structure generation
Generate ensembles of disordered structures by constrained random sequential placement with automated minimum separations from ionic, covalent and metallic radii.
How it works
Atoms are placed sequentially into a cubic cell. Pair-specific minimum separations are derived from radii and bonding-type classification. Composition-derived defaults provide starting structures; validate their density and local structure for the system being studied.
Automated minsep
For each element pair, the bond type is classified and the appropriate radii are used:
Bond type |
Radii source |
Scale factor |
Example |
|---|---|---|---|
Ionic (M-O, M-Cl) |
Shannon ionic (CN-aware) |
0.80 |
In-O: (0.80+1.40)*0.80 = 1.76 |
Metallic (M-M) |
Metallic radii |
0.85 |
Al-Al: (1.43+1.43)*0.85 = 2.43 |
M-M in oxide |
max(metallic, sqrt(2)*d(M-O)*0.85) |
cap 2.80 |
In-In: min(2.80, max(2.84, 2.64)) = 2.80 |
Metalloid (Si-Si) in oxide |
max(metallic, sqrt(2)*d(Si-O)*0.85) |
cap 2.80 |
Si-Si: max(1.99, 2.12) = 2.12 |
Small anion (O-O) |
Shannon ionic |
0.80 |
O-O: (1.40+1.40)*0.80 = 2.24 |
Large anion (Cl-Cl) |
Shannon ionic |
0.70 |
Cl-Cl: (1.81+1.81)*0.70 = 2.53 |
Nonmetal cation to anion (P-O, S-O, C-O) |
Shannon cation radius (top state, lowest CN) |
0.80 |
P-O: (0.17+1.40)*0.80 = 1.26 |
Nonmetal cation to cation |
sqrt(2)d(X-O), or 2d(X-O) for the same element |
0.85, cap 2.80 |
P-P: 2*(0.17+1.40)*0.85 = 2.67 |
When --target-cn is provided, CN-specific Shannon radii are used (e.g. Si CN=4: 0.26 A vs CN=6: 0.40 A), giving tighter minsep values.
Bond-type classifier
The classifier uses an element-type rule (nonmetal / metalloid / metal membership) with a Pauling electronegativity refinement: when the type rule would call a pair “ionic” but the Pauling Δχ is below 1.0, the pair is reclassified as covalent. This catches predominantly-covalent compounds whose Shannon ionic radii would otherwise give unrealistically small distances:
Pair |
Δχ |
Type rule |
Final class |
Why it matters |
|---|---|---|---|---|
Ga–As |
0.37 |
ionic (metal+metalloid) |
covalent |
Shannon radii would give minsep ≈ 0.68 Å, atomic overlap |
In–P |
0.41 |
ionic |
covalent |
III–V semiconductor |
Si–C |
0.65 |
ionic |
covalent |
Carbide |
B–N |
1.00 |
ionic |
ionic (at threshold) |
Sits right on the cutoff; Shannon radii give reasonable BN minsep |
Ga–N |
1.23 |
ionic |
ionic |
Stays ionic |
Li–F |
3.00 |
ionic |
ionic |
Stays ionic (clearly ionic) |
The Δχ refinement is applied to the metalloid–nonmetal, metalloid–metal, and metal–nonmetal branches. The metallic branch (metal–metal) and the pure-covalent branch (nonmetal–nonmetal, metalloid–metalloid) are unaffected because the type rule already gives the correct answer there. The 1.0 threshold is empirical, it cleanly separates III–V/II–VI semiconductors and carbides (covalent character ≥ 70% by Pauling’s formula) from the polar-ionic borderline cases (GaN, ZnO, BeO) where Shannon ionic radii give sensible minsep values.
The bond classifier is exposed as amorphgen.utils.radii.classify_bond(sym_a, sym_b) for inspection.
The electronegativity scale follows Pauling’s original definition from bond-dissociation energies (Pauling 1932) as standardised in his textbook (Pauling 1960), with the specific numerical values used in AmorphGen taken from Allred’s 1961 thermochemical revision, the values now in the CRC Handbook and most chemistry textbooks. References:
Pauling, L. J. Am. Chem. Soc. 54, 3570–3582 (1932), original EN derivation.
Pauling, L. The Nature of the Chemical Bond, 3rd ed., Cornell Univ. Press (1960), textbook scale.
Allred, A. L. J. Inorg. Nucl. Chem. 17, 215–221 (1961), revised values used in the code (
PAULING_ENtable inamorphgen/utils/radii.py).
Edge cases at the classifier boundary
The Δχ = 1.0 threshold is a heuristic, so pairs near it can receive different classifications from the per-pair bond rule and the composition-based density rule. Examples include:
System |
Pair |
Δχ |
Material class expects |
Pauling rule says |
Context |
|---|---|---|---|---|---|
MgH₂ |
Mg–H |
0.89 |
ionic (hydride) |
covalent |
Rutile structure, predominantly ionic |
MnS |
Mn–S |
1.03 |
covalent (chalcogenide) |
ionic |
α-MnS is rocksalt, ionic-leaning |
TiC |
Ti–C |
1.01 |
covalent (carbide) |
ionic |
Rocksalt interstitial carbide, mixed bonding |
BN |
B–N |
1.00 |
ionic (nitride) |
ionic (stays, strict |
Sits exactly on cutoff |
The bond classifier selects radii for the per-pair minsep; the material
classifier selects radii for the density estimate. Check both derived values
for these boundary cases. Subsequent relaxation may change the local
structure and density, but does not guarantee agreement with experiment.
Set --target-density explicitly when a validated density is available.
Nonmetal cations: oxoanions and hydroxides
In a phosphate, a sulfate or a carbonate, the P, S or C is a nonmetal but it is the cation of its oxoanion, and in a hydroxide the H is. Which nonmetals are cations is decided by charge balance (radii.cation_nonmetals): an element of the anion table is promoted to cation when that brings the compound closer to neutrality, and C and P are cations when an anion more electronegative than them is present and they balance the charge better as cations than as C⁴⁻ / P³⁻. So the P of Li₃PO₄ and Li₃PS₄, the S of Li₂SO₄, the C of CaCO₃ and the H of Mg(OH)₂ are cations, while the carbide C of SiOC, the P of InP and the S of La₂O₂S stay anions. A nonmetal cation:
bonds to its anions at its Shannon cation radius (P-O 1.26, S-O 1.22, C-O 1.06, O-H 0.82 Å, about 0.8 of the bond, as for Si-O);
targets the ligand count of its oxoanion: 4 in PO₄³⁻, SO₄²⁻ and ClO₄⁻, 3 in CO₃²⁻, NO₃⁻, IO₃⁻ and a sulfite, 1 for H;
takes its oxidation state from the same charge balance (S⁶⁺ in Li₂SO₄, which leaves Li⁺), and is sized as that cation in the density estimate.
With a composition, classify_bond(sym_a, sym_b, composition) applies these roles: a nonmetal cation and an anion are ionic, a nonmetal cation and another cation are cation-cation (a second-shell contact across the anion, never a bond), and two cations of a compound with anions are never ionic. That last rule is why Na-B in a borate and K-Si in a silicate are placed 2.16 and 2.75 Å apart, like Na-Si, rather than as the ionic bond their Δχ ≥ 1 would suggest.
Hydrogenated networks: a-Si:H, a-C:H
C, Si and Ge with H and nothing else, with at most one H per host atom, are the hydrogenated_network class: a-Si:H, a-Ge:H, a-C:H up to the polymer-like 50 % H, a-SiC:H and a-SiGe:H. H there caps a host atom through a covalent bond; it is not the H⁻ of LiH, MgH₂ or NaAlH₄, which stay hydride. The host keeps what it has without H: its Cordero radii and packing factor, its minimum separations (Si-Si 1.87 Å as in a-Si, C-C 1.22 Å as in a-C, the C-C anion packing of SiC) and its bonds. Each H targets one bond (the hosts target 4) and is kept at 0.8 of its bond from any host (C-H 0.86, Si-H 1.18 Å). Two H can share a host atom (H-H 1.21 Å in a-C:H, 1.67 Å in a-Si:H) but cannot form H₂.
In the density estimate H is sized at 0.90 Å, not its Cordero 0.31 Å, which would give it no volume: each H replaces a host-host bond and brings free volume with it, so the density falls as the H content rises. Si₆₄H₈ (11 % H) comes out at 2.30 g/cm³ (glow-discharge a-Si:H ≈ 2.2), and a-C:H at 2.19, 1.84, 1.52 and 1.24 g/cm³ for 20, 30, 40 and 50 % H (hard a-C:H 1.6–2.2 g/cm³ at 30–40 % H, polymer-like 1.2–1.6 g/cm³ at 40–50 %). The real density also depends on how the film was grown (sp³ fraction, voids), so use --target-density when the measured value is known.
Automated density
Cell volume is estimated by class-aware sphere packing: the composition is classified into a material class, each element is given a radius from the table appropriate to that class’s bonding (Shannon ionic, Cordero covalent, or Goldschmidt metallic), and the cell is sized so the spheres fill it to the class’s packing factor. Cation oxidation states, needed to select the right Shannon radius, are assigned automatically by charge balance, using a joint solver that resolves multivalent cations in mixed-cation / mixed-anion compounds (e.g. FeTiO3 -> Fe2+/Ti4+, SrTiO3 -> Sr2+/Ti4+) and sums the balance over every anion-former (oxynitrides, oxyfluorides).
Material class |
Radius source |
Packing factor |
Examples |
|---|---|---|---|
|
Shannon ionic CN6 |
0.66 |
TiO2, SnO2, RuO2, IrO2 |
|
Shannon ionic CN6 |
0.63 |
ZrO2, HfO2, CeO2, ThO2, UO2 |
|
Shannon ionic CN6 |
0.62 |
AlN, GaN, Si3N4, TiN |
|
Shannon ionic CN6 |
0.60 |
V2O5, Nb2O5, MoO3, WO3 (OS ≥ 5) |
|
Goldschmidt (cation) + Cordero (C) |
0.60 |
TiC, WC, ZrC |
|
Goldschmidt metallic |
0.60 |
NiTi, CuZr, brass, pure metals |
|
Shannon ionic CN6 |
0.58 |
LiF, NaCl, Li2ZrCl6 |
|
Shannon ionic CN6 |
0.52–0.58 (interpolated by halogen fraction of the halogen+O anion pool; dopant-level halogens, <10%, keep the oxide class) |
BiOCl, LaOCl, NaTaOCl4, ZrOCl2 |
|
Shannon ionic CN6 |
0.55 |
LiH, MgH2, NaAlH4 |
|
Shannon ionic CN6 |
0.52 |
In2O3, Al2O3, Ga2O3, MgO, ZnO |
|
Shannon ionic CN6 |
0.52 |
ZrN, HfN, ScN (large cation) |
|
Shannon ionic CN6 |
0.50 |
SiO2, GeO2, B2O3 |
|
Goldschmidt cation + Cordero B |
0.60 |
TiB2, MgB2, ZrB2 (LaB6-type cage borides run ~100% of crystal; use |
|
Cordero covalent |
0.35 |
BeO (small polarizing cation) |
|
Cordero covalent |
0.32 |
GaAs, InP, InAs |
|
Cordero covalent |
0.32 |
SiC, B4C |
|
Cordero covalent |
0.30 |
Si, Ge, C |
|
Cordero covalent (host), H 0.90 Å |
the H-free host’s (0.28–0.32) |
a-Si:H, a-Ge:H, a-C:H, a-SiC:H, a-SiGe:H |
|
Cordero covalent |
0.30 |
ZnS, CdTe, GeTe |
|
Cordero covalent |
0.23 |
GeS2, GeSe2, As2S3, As2Se3 (network glasses; tellurides stay |
|
Cordero covalent |
0.28 |
a-Se, a-Te, a-As, a-Sb, a-P |
The dioxide split (rutile vs fluorite) and the small- vs large-cation nitride
split are decided by a cation-radius rule; high_valent_oxide is gated on
oxidation state ≥ 5. Compositions that match no specific class fall back to a
generic packing factor of 0.52 (Shannon ionic).
The density estimate sets the starting cell. Subsequent cell relaxation can
change it according to the chosen potential; validate the relaxed density
separately. Use --target-density when an experimental value is known, and
--retry-mode reduce-minsep or none when placement must retain that density.
Coordination-aware placement (“SC”)
When coordination targets are available, AmorphGen uses SC (“Seed-Coordinate”) placement: each new atom is added as a bonded neighbour of an existing under-coordinated site (the seed) by placing it within that seed’s bonding shell (minsep ≤ d ≤ dmax), which coordinates it. Over-coordinating a neighbour is rejected. This builds short-range order directly into the placement instead of relying on relaxation alone, giving more physical structures.
SC is the placement half of the Seed-Coordinate-Anneal (SCA) algorithm of Youn et al., Comput. Mater. Sci. (2014). AmorphGen factors the Anneal step out into its own stages (the MLIP geometry optimisation (--relax) and the melt-quench / hybrid workflow) so the placement step is just Seed-Coordinate (hence “SC”, not “SCA”).
Targets are inferred from composition by default; --target-cn overrides them.
The targets guide placement and cap over-coordination but do not guarantee that
every atom reaches its target. Disable SC with --no-sc to fall back to plain
random rejection sampling (or pass target_cn={} to the Python API).
Transparency: the auto-derive log line
Every --random-gen run writes a single one-line summary at the top of random_gen.log capturing every chemistry-informed decision the auto chain made, so you can see why a particular minsep / density / target CN was used without reading the code. Example for Ga₂O₃:
[auto-derive] Ga16O24 → metal_oxide, OS{Ga:+3}, CN{Ga:5}, minsep{Ga-Ga:2.43 metallic | Ga-O:1.62 ionic Δχ=1.63 | O-O:2.24 anion-pack}, ρ=4.44 g/cm³ L=8.25 Å
Each field:
Field |
What it means |
|---|---|
|
Composition routed to the |
|
Cation oxidation state inferred from charge balance against the anion |
|
Target coordination auto-detected per cation |
|
Per-pair: bond class + Pauling Δχ (shown only when the ionic classification is at stake) + minsep value in Å |
|
Auto-estimated mass density and cubic cell length |
The line is grep-friendly: grep "auto-derive" random_gen.log retrieves it as a single line per generation run. Bond classes shown are ionic, covalent, metallic, cation-cation (a nonmetal cation and another cation, see “Nonmetal cations” above), and anion-pack (same-element anion pairs use a separate anion-packing scale factor, see “Bond-type classifier” above).
CLI examples
# Formula format (In2O3 * 16 formula units = 80 atoms)
amorphgen --random-gen --composition "In2O3*16" --n-structures 20
# Atom count format (equivalent)
amorphgen --random-gen --composition In=32,O=48 --n-structures 20
# With target CN and density
amorphgen --random-gen --composition "SiO2*16" --n-structures 20 \
--target-cn Si=4,O=2 --target-density 2.2
# With relaxation
amorphgen --random-gen --composition "In2O3*8" --n-structures 10 \
--relax --model mace-mpa-0 --cell-filter none
# Fixed-density study: hold the cell EXACTLY at the target density and
# soften non-bonded minseps on placement stalls instead of expanding
amorphgen --random-gen --composition "SiO2*16" --n-structures 10 \
--target-density 2.6 --retry-mode reduce-minsep
Placement-stall policy (--retry-mode)
Random sequential placement can stall when the requested density and minimum separations leave too little space for the remaining atoms.
The first response to a stall, in the default expand mode, is a soft pack:
the structure is placed again in the same cell with every floor scaled by
0.72, then every pair closer than its floor is pushed apart iteratively until
all pairs reach at least 98.5 percent of their floors. A successful soft pack
keeps the current cell and sets info["soft_pack"] = True. If soft packing
fails, expand grows the cell. The other modes skip soft packing and follow
their own retry policies:
Mode |
Cell |
Minseps |
Use when |
|---|---|---|---|
|
cell edge grows 5% per retry (up to 4 retries) |
requested floors; soft packing accepts 98.5% of them |
the starting density may change during placement |
|
held fixed |
same-element separations reduced 5% per retry (up to 4 retries, about 19% total); unlike-element cation–anion floors retained |
density must stay fixed during placement |
|
held fixed |
requested floors retained |
both cell and minimum separations must stay fixed; a stall raises an error |
Same-element separations may represent bonds in elemental networks or alloys,
so inspect the resulting contacts when using reduce-minsep. Run relaxation
at fixed cell (--cell-filter none) to preserve the density through that step.
The table describes retries inside generate_random. batch_random also
resamples seeds and, after repeated failures, can reduce the auto-estimated
density and then same-element metal/metalloid separations. Explicit density
or cell settings disable the batch density reductions, but expand can
still grow the cell inside generate_random. With reduce-minsep, the batch
never changes density; with none, it only resamples seeds before skipping
unplaceable indices. Inspect random_gen.log for adjustments and skipped
structures.
Python API
from amorphgen import generate_random, batch_random
# Single structure (auto minsep and material-class density estimate)
atoms = generate_random(
composition={"In": 16, "O": 24}, # atom counts (Python API)
seed=42,
)
# Batch generation (20 structures of SiO2)
paths = batch_random(
composition={"Si": 16, "O": 32},
n_structures=20,
target_density=2.2, # optional, auto if omitted
target_cn={"Si": 4, "O": 2}, # optional, auto if omitted
)
Note: The Python API takes atom counts as a dict. The CLI also accepts formula format:
--composition "SiO2*16"is equivalent to{"Si": 16, "O": 32}.
Accessing radii data
from amorphgen.utils.radii import get_ionic_radius, classify_bond, default_minsep
# Shannon ionic radius
get_ionic_radius("In", cn=6) # 0.80 A
get_ionic_radius("In", cn=4) # 0.62 A
# Bond classification
classify_bond("In", "O") # "ionic"
classify_bond("Si", "Si") # "covalent"
# Auto minsep for a composition
ms = default_minsep(["In"] * 16 + ["O"] * 24, target_cn={"In": 4})
# Pass the full symbol list so the rules see the actual composition.
print(ms)
See Random structure generation for the full API reference.
Selecting structure indices
When using --resume, keep the composition, seed, density and placement
settings, output format, and relaxation settings the same. Schema 2 of
run_metadata.json records those settings, including model identity and local
model file hashes. AmorphGen refuses changed settings before modifying existing
outputs or logs, and checks the elements and atom counts in saved structures.
Reusing a random-generation directory with incompatible settings is also refused
without --resume, so a partial rerun cannot relabel older structures.
Missing, unreadable or older metadata without complete settings causes an
error; restore compatible metadata or use a separate output directory.
You can increase the requested structure count or change the selected indices
while retaining the same settings. Use --batch-opt to relax structures from
an earlier generation run.
An exclusive .amorphgen.lock covers generation and optional CLI relaxation.
A concurrent writer fails immediately. The OS releases the lock after normal
exit, failure or process termination; leave the persistent lock file in place.
--indices SPEC restricts a run to given structure indices; inclusive ranges
and lists both work (80-90, 0,5,7-9). Every index has a seed derived from
--seed. With the same placement settings, split jobs reproduce the placements
from a full run unless batch-level retry escalation changes those settings.
Escalation state is not saved across resumes, so a resumed or split run that
hits that path is not guaranteed to match a continuous run.
Give concurrent jobs separate output directories to keep their logs and metadata separate, then collect their non-overlapping structure files:
amorphgen --random-gen --composition "InGaZnO4*50" -n 100 --seed 2026 --indices 0-49 -o igzo_a
amorphgen --random-gen --composition "InGaZnO4*50" -n 100 --seed 2026 --indices 50-99 -o igzo_b
mkdir -p igzo/random_initial
cp igzo_a/random_initial/*.xyz igzo_b/random_initial/*.xyz igzo/random_initial/
# relax only some of them later (files whose name ends in the index)
amorphgen --batch-opt --input-dir igzo/random_initial --indices 80-90 -m mace-mpa-0 -o igzo/random_opt
--batch-opt also takes --pattern GLOB to choose the input files. Both
combine with --resume and --engine torchsim.