An instrument · a magnet · the critical point

Fever

Here is a snapshot of a magnet — a grid of atoms, each one a tiny arrow pointing up or down — at some temperature you are asked to name. Hot, and it is salt-and-pepper noise. Cold, and it is one solid colour. In between it forms islands. Almost everywhere on the temperature axis this is an easy game — and then, within a percent or so of 2.269, it becomes impossible, and impossible for a reason: at that temperature the islands come in every size at once, so the picture looks the same at every zoom and carries no scale to measure against. The exact temperature where that happens, 2/ln(1+√2), was calculated by Onsager in 1944; this page locates it live, from the point where three curves cross.

And losing there is not a failure of your eyes. It is the fate of anyone who reads the picture's geometry — human or machine. Section V fields two machines against you: one that measures the same shapes you can see, and one that, at and above the critical point, does nothing but count how many neighbouring pairs agree. The second one sails through the critical point without noticing. Counting bonds sees through the fever, and counting bonds is not what looking is.

I · Look at it first

Six snapshots, four of them unlabelled

Each square below is a 256×256 lattice of spins, drawn after the simulation has been run to equilibrium at some fixed temperature. Nothing is drawn but the spins: one colour for up, one for down, one pixel each, no smoothing, no outlines, no shading. The two on the ends are controls whose temperatures you already know. The four in the middle are not labelled, and two of them straddle the critical temperature closely, one about three hundredths above it and one about seven hundredths below. Drag on any panel to move the magnifying lens; the lens is at ×4, and the point of it is that at the critical point magnifying the picture does not tell you anything new.

Equilibrated 256² lattices, sampled live liverefreshed one panel at a time · Wolff cluster updates · resampled 0 times
T = ∞
control: coin flips
A
hidden
B
hidden
C
hidden
D
hidden
T → 0
control: one colour
Lens: drag on a panel above, or focus one and use the arrow keys. ×4 magnification; the window is 24 to 128 spins across, depending on how wide the panels are.
spin upspin down
The two spin colours swap roles between the light and the dark theme — in the dark theme up is the pale bone colour on a dark field, in the light theme it is the dark slate on a pale one. The image is meant to read as a piece of material, so it keeps its contrast either way. Which colour is which never matters physically: the model is exactly symmetric under flipping every spin at once.

The four unlabelled panels are at fixed hidden temperatures and are re-equilibrated from a fresh random start every few seconds, so what stays the same between refreshes is the character of the picture, not the picture. Equilibration uses the kernel's energy-drift rule and then discards a further 50 sweep-equivalents before the snapshot is taken (Provenance). The controls are not simulations at all: T = ∞ is createLattice(init:'hot'), i.e. one independent fair coin per site with no dynamics; T → 0 is init:'cold', every spin up, with two spins flipped by hand so that there is something to see.

II · The statistic

Two exact curves and a living magnet

Turn the pictures into two numbers. The magnetisation |m| is how far the spins are from balanced: 1 if they all agree, 0 if up and down exactly cancel. The energy per spin u counts how many neighbouring pairs disagree: −2 if every pair agrees, 0 if they are uncorrelated. Both are exactly known for the infinite lattice — Yang solved |m| in 1952, Onsager the energy in 1944 — and both are being measured here, live, on a 64×64 lattice, by averaging over configurations as they accumulate.

The two curves are not the same kind of object. The energy is continuous through the critical point — its slope is not, and that divergence is what section V leans on — and the measurement follows it across. The magnetisation is exactly zero above Tc in the infinite lattice and the measurement is not zero — a 64×64 box always has a leftover ⟨|m|⟩ of a few percent, because a finite box can never quite average itself away. That tail is a property of the box, not of the material, and it is drawn as such.

⟨|m|⟩ and ⟨u⟩ at L = 64 against the exact curves liveconfigurations measured 0 · one sweep-equivalent apart live
measured, ±1 s.e.Yang 1952, exact exactYang inapplicable: ξ > L/7Tc
measured, ±1 s.e.Onsager 1944, exact exactTc
largest |⟨u⟩ − uexact| away from Tc largest |⟨|m|⟩ − mYang| where Yang applies finite-size tail ⟨|m|⟩ at T = 3.0 ⟨u⟩ at Tc vs the exact −√2the one place they part: the gap is the 64-box, not the model

β, fitted rather than assumed

ln⟨|m|⟩ against ln(1 − T/Tc), T ∈ [1.90, 2.24] livethe exponent is the slope
measured points in the fit windowfit to the measured pointsthe same fit run on Yang's exact curve
β fitted from this page's measurements  the same fit, same window, on the exact curve exact exponent0.125 points in the window

Read the two big numbers against each other before reading either against 0.125. β = 1/8 is the slope in the limit T → Tc; over a window that starts at 1.90 it is not the slope, and the honest control is to run the identical fit on Yang's exact formula over the identical window, which is the middle number. The gap between the middle number and 0.125 is what the window costs; the gap between the left number and the middle one is what the simulation costs. The window is the larger error by roughly an order of magnitude, which is the usual situation and the usual reason a critical exponent is hard to measure. Fitting closer to Tc would fix the window and break the box: at L = 64 the correlation length reaches L/7 at about T = 2.20 and the finite-size rounding takes over. By that measure the window’s last point, 2.24, already sits past the line — the exact ξ there is more than twice L/7 — so the gap between the left number and the middle one is not pure sampling noise: at that point it carries finite-size rounding too, which the control fit cannot separate out, Yang’s formula having no L in it at all.

III · The deeper theory

Where three curves cross

The problem with reading Tc off section II is that every curve there depends on the box size, and none of them has a feature sharp enough to point at. Binder's cumulant is the trick that gets around it: U₄ = 1 − ⟨m⁴⟩/3⟨m²⟩². Deep in the ordered phase the magnetisation barely fluctuates and U₄ → 2/3; deep in the disordered phase m is a Gaussian about zero — one number, not a vector — so ⟨m⁴⟩ = 3⟨m²⟩² and U₄ → 0. In between, at the critical point, U₄ becomes independent of the box size — so curves measured at different L all pass through the same point, and that point is Tc. Three boxes, three curves, one crossing, and no formula for Tc anywhere in the measurement.

Binder cumulant U₄ at L = 16, 32, 64 livemeasurements per point 0 · jackknife errors over up to 50 blocks
L = 16L = 32L = 64exact Tc = 2.269185 exactU* = 0.61069 numerical
crossing temperature, mean of the three pairs  crossing value Ux gap to the exact Tc pairwise crossings 16/32, 16/64, 32/64

Everything in the box above is this page's own data. The two reference lines are not: the vertical one is Onsager's exact Tc = 2/ln(1+√2), and the horizontal one is U* = 0.61069, which is a numerical literature value (Kamieniarz & Blöte 1993) with no closed form, so it is tagged as such and was not verifiable from inside this build. The kernel's own recorded run, with 15,000 measurements per point, put the three pairwise crossings at 2.26808, 2.26664 and 2.26587, mean 2.26686, and the crossing value at 0.61317 stored, kernel report §3.6. Each pairwise crossing now carries its own propagated statistical error: the jackknife error on U₄ at the bracketing temperatures, divided by the local slope of the difference between the two curves. At these statistics that comes to a thousandth or two for 16/64 and 32/64, and several times that for 16/32, whose two curves cross at a shallow angle. The pair-to-pair spread is of the same size as those errors, so the finite-size drift — which is real, since the crossing of two finite sizes only extrapolates to Tc as L → ∞ — is not resolved from the sampling noise here. Read the printed ± as a floor rather than a full budget: the propagation is linear, and the three pairs are built from the same three chains, so they are not independent of each other.

What the picture looks like at every distance

G(r) at Tc, L = 128, log–log liveconfigurations averaged 0
measured G(r), axis-averagedfit over r ∈ [1, 8]exact slope −η = −1/4 exactwhere the finite box flattens it
η fitted from this page's G(r)  exact η0.25 fit residual (rms in ln G) G(r) plateau at r = L/2

A straight line on log–log axes means a power law, and a power law is the same shape at every magnification — this is the scale invariance of the critical point written down as a number. The fit runs over r ∈ [1, 8] only. Beyond that the curve bends and then flattens onto a plateau at ⟨m²⟩, because in a 128-wide periodic box two spins 64 apart are also 64 apart the other way round and the correlation cannot keep falling. That plateau is finite-size, not physics. Expect the fitted η to come out a few percent below 1/4 for the same reason: even inside the fit window the plateau is already lifting the far end. Measured off-line with the same rule, 3,000 configurations per point and three independent seeds at each size, the fit gives 0.2464 ± 0.0014 at L = 256 against 0.2393 ± 0.0029 at L = 128 stored, this page's build check — those ± are the seed-to-seed spreads, and the bias, 1/4 minus the fitted value, shrinks from about 0.011 to about 0.0036 as it must. Note that the ± printed live beside η above is narrower than either: it is the linear fit's own slope error over eight points that share one set of configurations, and it does not include the finite-size bias, which is the larger number here.

IV · The dial

The ruler melts

The correlation length ξ is the size of the islands: the distance over which two spins still tend to agree. Away from the critical point it is a few lattice spacings and the picture has an obvious grain. Approach Tc from either side and ξ diverges — and once ξ passes about a quarter of the box, the box stops being able to measure it, which is the same sentence as “the picture stops having a scale”. Watch the number climb from 1.7 at T = 3.2 to forty-odd at 2.30, then past a hundred at 2.27 — already badly low there, the exact value being 1,580 — and then watch it stop being a measurement and become a floor, as the box runs out of room to hold what it is measuring. It measures at all only on the hot side: below Tc the whole lattice agrees with itself, and the quantity this estimator is built on stops describing the islands and starts describing the ordering. The exact curve is quoted on both sides so you can see the divergence whole.

The second instrument on the dial is what that costs the simulation. Metropolis flips one spin at a time, and near Tc the configuration takes longer and longer to forget where it started: the integrated autocorrelation time of |m| blows up. Wolff's algorithm flips a whole cluster at once, chosen so that the move is still exactly correct, and barely notices. The simulation gets hardest exactly where the physics is most interesting — unless you change what a move is.

1.5 frozen2.0 ordered2.269 Tc2.8 disordered3.5 hot
T = 2.27
128 × 128, equilibrated at this temperature live
correlation length ξ, measured above Tc accumulating…
ξ as a fraction of the box, L = 128
exact ξ at this temperatureOnsager ξ / L
Sampler for the autocorrelation instrument:
τint(|m|), measured on an L = 64 chainsweepsat Tc and L = 64 the recorded run gives 807.7 sweeps for Metropolis and 1.501 sweep-equivalents for Wolff stored §3.8 — the live chain needs to be long before it gets there series length n Sokal window W n / τa long chain is needed before this number can be believed
temperatures visited, once τ has stopped climbingTc

ξ is the second-moment estimator ξ² = [S(0)/S(qmin) − 1] / (2 sin π/L)², formed from configuration averages — a ratio of means, never a mean of ratios — using Wolff's improved cluster estimator for S(q). It is run only above Tc, which is the only regime where it is meaningful at all: S(0) = ⟨M²⟩/N is the magnetisation fluctuation only where ⟨M⟩ = 0, which above Tc it is, exactly, by symmetry. Below Tc the same expression is dominated by the long-range order and returns a number with nothing to do with ξ at all — tens to hundreds of times the width of the box: 12,850 at T = 1.5 and 2,750 at T = 2.0, where the true ξ is 0.63 and 2.19 stored, build check. Printing that would be worse than printing nothing, so the instrument declines. On a single snapshot the same formula is worthless: the kernel report quotes 10.7 from one configuration at L = 128, T = 2.6 where the true value is 4.27 stored, §3.7, which is why the number above carries a sample count and moves for a while after you let go of the dial. The exact comparison is Onsager's ξ−1 = ln coth β − 2β above Tc; below Tc the lowest excitation is a two-particle state, so the gap doubles and the exact ξ is half the naive continuation — the kernel's xiBelow. Measured against those, the estimator agreed to better than 1.2% at T = 2.5, 2.6, 2.8, 3.2 in the recorded run, always slightly low stored, §3.7. Closer in, once ξ passes L/4 — which happens at about T = 2.30 — the same estimator is bounded by the box and reads low by a large and fast-growing margin: about 4% low at 2.30, 36% at 2.28 and 93% at the dial’s own starting temperature of 2.27 stored, build check. That is why past that point the number above is shown as a lower bound rather than as a measurement. τint is measured on a separate L = 64 chain, because resolving an autocorrelation time needs a series many times longer than the time itself and L = 128 near Tc would take many minutes. A temperature only joins the chart once the estimate has stopped moving, tested by recomputing it separately on the first and second halves of the chain — two disjoint stretches, sharing no data — and requiring those two to agree to 12%; the usual n ≳ 100 τ rule is not sufficient on its own, because a chain that is far too short returns a truncated τ that then looks well sampled by its own standard. Near Tc with Metropolis that takes a minute or two of watching, and the number climbs the whole way. For scale, the recorded run at L = 16/32/64 gives Metropolis 33.9 / 140.0 / 807.7 sweeps and Wolff 1.175 / 1.358 / 1.501 sweep-equivalents stored, §3.8 — a factor 538 at L = 64. The growth of the Metropolis numbers fits an effective exponent 2.29 ± 0.14 over three sizes; that is not the dynamic exponent z, which is a different quantity defined in the thermodynamic limit and is about 2.17 in the literature.

V · The game

Name that temperature

Ten rounds. Each round is a fresh 128×128 lattice equilibrated at a temperature drawn from 1.6 to 3.4, with extra weight near 2.269. Set the slider, commit, and see how far off you were. Two machines answer the same picture at the same time.

The Eye reads the picture's geometry, which is what you are doing. It measures the correlation length from how fast agreement between two spins decays with the distance between them, and inverts Onsager's exact ξ(T); below Tc it reads the magnetisation instead and inverts Yang's exact curve. When the islands grow past about a quarter of the box, or the magnetisation lands in the range that a box this size produces at any near-critical temperature, it has nothing left to measure and says so — and then answers 2.269, a number it got from Onsager rather than from the picture.

The Bond-counter does something you cannot do by looking: it counts what fraction of the 32,768 neighbouring pairs agree, and inverts Onsager's exact energy. That is a local quantity, and although its own noise grows near Tc right along with the specific heat, the slope du/dT = c grows faster still: the error that noise implies in T goes as T/√(Nc), so the energy is most informative exactly where the picture is least readable. Below about T = 2.02 the energy alone reads more ordered than the truth, so there it also folds in the magnetisation — inverting Yang's exact curve, the way the Eye does — and the verdict line names the channel it used. From Tc upward, which is the whole of the argument here, the energy is all it has. It does not fail with you. That is the point of it.

Round 1 of 10 liveEquilibrating a fresh lattice…
1.62.02.2692.83.4
2.50
Set the slider and commit. The two machines answer the same picture.
Your mean error The Eye Bond-counter Rounds played 0 Rounds the Eye was blind 0

Error against the true temperature

youthe Eyehollow: the Eye was blind and fell back on 2.269the Bond-counterTc

Before this instrument was wired up, both machines were run over 100 independent equilibrated snapshots at each of 29 temperatures from 1.6 to 3.4, at the same L = 128 stored, build check. The Bond-counter's mean |error| is flat: 0.020 at 1.6 (the only one of those four that uses the magnetisation channel as well as the energy), 0.012 at 2.2, 0.0082 at Tc — its best — and 0.040 at 3.4. One asymmetry to hold on to when reading those two rows against each other: the Eye and you are both hard-bounded to the announced 1.6–3.4, the Eye by a clamp and you by the slider, while the Bond-counter is not bounded at all and is free to answer outside it. Clamping it to the same range would improve it at the two ends, to 0.0077 at 1.6 and 0.017 at 3.4, and leave every other temperature untouched; the numbers here are the unclamped ones, so the comparison is if anything harder on it than on the Eye. The Eye stays close from 1.6 to 2.20, never more than about two and a half thousandths behind it, and then falls apart: on the rounds where it can see at all its mean |error| is 0.014 at 2.20, 0.094 at Tc, 0.10 at 2.35 and 0.28 at 3.0 — seven to twelve times the Bond-counter's from Tc through 2.9, easing to three to six times over 3.2 to 3.4 as the Eye runs into the end of the slider. And in the band 2.24 to 2.29 it is blind on 55–90% of snapshots. The reason it is not also ten times worse in that band on the plot above is that a blind Eye answers 2.269, which is nearly the right answer for the wrong reason; it did not read that number off the picture. Ten rounds is far too few to reproduce any of this, so read the chart for its shape and not for its averages.

Provenance

What is measured, what is exact, and what is neither

Model and conventions

Square lattice L × L with periodic boundaries, spins si = ±1, coupling J = 1, kB = 1, no external field, so E = −Σ⟨ij⟩ sisj counting each bond once and β = 1/T. Temperatures are therefore in units of J/kB and Tc = 2/ln(1+√2) = 2.269185314213022 is a pure number. E and M are held incrementally as exact integers — every flip anywhere in the kernel goes through the same three lines — and the test suite recomputes both from scratch after thousands of updates and compares with ===. |m| means ⟨|M|/N⟩ per configuration: in a finite box at zero field ⟨M⟩ is exactly zero by symmetry, so the absolute value is the estimator, and it does not go to zero above Tc.

Sampling and equilibration

Wolff single-cluster updates with padd = 1 − e−2β do all the work; Metropolis single-spin sweeps at uniformly random sites are the second engine, used in section IV. Equilibration is a rule, not a guess: run in blocks of 10 updates, record the block mean of E/N, and after at least 24 blocks and 5 sweep-equivalents take the trailing window of 12 block means, split it in half, and declare the run drift-free when the difference between the two halves is smaller than 2s/√12. Three consecutive passes are required; one failure resets the count. Every snapshot used in section I and section V then discards a further 50 sweep-equivalents on top, because the recorded transient test at L = 128 near Tc passed with only 5% of its margin to spare kernel report §3.4, §4.8. The generator is sfc32 keyed by xmur3. The repeatable instruments are keyed by fixed strings — the hero panels by their index and the resample count shown above them, the section II–IV background sweeps by temperature and engine — while the dial adds an internal move counter on top of its temperature, and each game round and each self-test takes a fresh unrepeatable seed, so neither a dial snapshot nor a round can be replayed. The same generator is used in this suite's other pages.

Live, exact, stored

Run a small version of the kernel's own gold-standard test here, on fresh samples, in your browser: it enumerates all 216 states of a 4×4 lattice, computes ⟨e⟩ and ⟨|m|⟩ exactly by Boltzmann-weighted sum, then draws fresh Wolff samples and reports the discrepancy in units of its own standard error.

Exact enumeration of 65,536 states at T = 2.0 and T = 3.0 against 2,000 fresh Wolff samples each. A second or two.
Not run yet.
QuantityStatusEvidence
Tc = 2/ln(1+√2), u(T), c(T), ξ(T) above TcexactOnsager 1944; ξ from the transfer-matrix gap (Kaufman & Onsager 1949); the kernel's implementations agree with mpmath to max |δ| 8.9e−16 (u), 3.1e−15 relative (ξ) and 4.0e−15 relative (c), and with an independent transfer-matrix gap to 1.6e−14
m(T) = (1 − sinh−4(2/T))1/8 below TcexactYang 1952; agrees with mpmath to 1.8e−15, and the exponent recovered off the exact curve itself is 0.12499995
ξ below Tc: ξ−1 = 2(2β − ln coth β)exactthe two-particle gap; a factor 2 that is easy to get wrong, cross-checked against the transfer matrix in the kernel's tests
β = 1/8, ν = 1, γ = 7/4, η = 1/4, α = 0exact, assertedclassical results, stated not fitted; sections II and III fit their own values and report the gap
U* = 0.61069numericalKamieniarz & Blöte 1993 — a numerical literature value with no closed form, cited from knowledge with no network available and not verified against the source; the kernel flags it as numerical in its own API
Binder crossing 2.26686, Ux 0.61317measuredkernel report §3.6, 15,000 measurements per point; the live instrument in section III repeats it from scratch
ξ second-moment estimatormeasured, biasedagrees with Onsager to better than 1.2% at T = 2.5–3.2, always low by 0.2–1.2%; the bias is real and grows as ξ shrinks (Ornstein–Zernike is not exact on a lattice)
Metropolis τ growth, effective exponent 2.29 ± 0.14measured, over three sizes onlythis is an effective exponent for τint(|m|) at L = 16/32/64, not the dynamic exponent z; the literature z ≈ 2.17 (Nightingale & Blöte) refers to the exponential autocorrelation time in the thermodynamic limit
The Eye's error curvemeasured, this build100 snapshots at each of 29 temperatures at L = 128, run before the opponent was wired into the page; the Bond-counter was measured on the identical snapshots
Yang's curve near Tc at finite Lnot applicableat L = 64 and T = 2.269 the exact ξ below Tc is 3473, hundreds of times the box; the measured 0.60 against Yang's 0.377 is not a disagreement, it is a category error, and section II greys the curve out where ξ exceeds L/7

References

L. Onsager, Crystal statistics I, Phys. Rev. 65, 117 (1944) — Tc, the free energy, u and c; the correlation length follows from the transfer-matrix gap, made explicit in B. Kaufman & L. Onsager, Crystal statistics III, Phys. Rev. 76, 1244 (1949). C. N. Yang, The spontaneous magnetization of a two-dimensional Ising model, Phys. Rev. 85, 808 (1952) — the first published proof of a formula Onsager had announced, without one, in 1949. N. Metropolis, A. W. Rosenbluth, M. N. Rosenbluth, A. H. Teller & E. Teller, J. Chem. Phys. 21, 1087 (1953). U. Wolff, Collective Monte Carlo updating for spin systems, Phys. Rev. Lett. 62, 361 (1989). K. Binder, Z. Phys. B 43, 119 (1981) — the cumulant. G. Kamieniarz & H. W. J. Blöte, J. Phys. A 26, 201 (1993) — U* = 0.61069. M. P. Nightingale & H. W. J. Blöte (1996/2000) — z for Metropolis dynamics. A. D. Sokal, Cargèse lecture notes — the automatic windowing rule used for τint. R. H. Swendsen & J.-S. Wang, Phys. Rev. Lett. 58, 86 (1987) — the cluster idea Wolff's algorithm sharpens. All of these were cited from knowledge: this build had no network access, so titles, volumes and page numbers are unverified.

Contact the author: np56123@icloud.com