← Benchmarks

girih-bench · a model benchmark

Looks right. Is it right?

A thousand years ago, Muslim craftsmen were quietly building quasicrystals - aperiodic geometric order that Western physics would not describe until Penrose in 1974, and would not win a Nobel for until 2011. I asked two frontier AI models to rebuild that order, and scored them the way quasicrystals were actually discovered: by diffraction. Because a pretty picture can lie. A diffraction pattern cannot.

The secret in the tilework

Forbidden by nature

A repeating pattern can only carry 2, 3, 4, or 6-fold symmetry. 5 and 10-fold are mathematically forbidden for anything periodic.

Built anyway, in 1453

The Darb-i Imam shrine in Isfahan does the forbidden thing: a 10-fold quasicrystal, five centuries before the math existed to explain it.

The order under the surface

Islamic geometry turned away from copying the world and toward the rule beneath it. The girih you see is generated by a hidden network you do not.

Three rungs, climbing the physics of order

Each level is harder than the last: a periodic warm-up, then the forbidden 10-fold, then the aperiodic quasicrystal itself. Fable 5 and GPT-5.6 Sol got the exact same prompt at each rung. Look for yourself.

L2 · periodic order — the warm-up

8-fold star-and-cross

Fable 5 — 8-fold star-and-cross
Fable 5attempts the over-under weave
GPT-5.6 Sol — 8-fold star-and-cross
GPT-5.6 Solcrisper stars, no weave

Both build a clean 8-fold star on a repeating lattice. A tie, and the correct answer.

L3 · forbidden symmetry begins

10-fold girih strapwork

Fable 5 — 10-fold girih strapwork
Fable 5space-filling tiling
GPT-5.6 Sol — 10-fold girih strapwork
GPT-5.6 Sola single centered rosette

Both hit clean 10-fold. But Fable fills the whole field with strapwork, while Sol builds one perfect medallion in the middle.

L4 · aperiodic — the hardest rung

Darb-i Imam quasicrystal

Fable 5 — Darb-i Imam quasicrystal
Fable 5fills edge to edge
GPT-5.6 Sol — Darb-i Imam quasicrystal
GPT-5.6 Solrosette with gaps at the seams

Both produce clean 10-fold rosettes. Neither is provably the true aperiodic tiling the shrine actually uses — and telling the two apart is a genuinely open problem.

The result that fooled me

My first scorer rasterized each pattern to a small image and ran a Fourier transform. It handed me a dramatic answer: one model crushed the other on the hardest level. A clean headline.

Then I stress-tested my own benchmark - re-ran it at higher resolution. The result reversed. The “winner” flipped. The whole dramatic gap was an artifact of how I measured, not a fact about the models.

When I threw out the rasterizer and read the geometry directly - resolution-free - the truth was quieter: both models tie on the physics. Clean 10-fold symmetry, every level, both of them.

What is left is taste

Once the physics ties, the only difference is craft - and craft is your eye, not a number. Look back at Level 4. Fable fills the field and tries to weave the straps; Sol’s rosette is gorgeous but leaves gaps at the seams. To my eye Fable is the better craftsman. But that is a judgment, not a measurement - and one sample is not a benchmark. So I am not going to tell you a winner. I am going to show you both and let you decide.

the discipline

Everyone tells you to test AI on your own work. True. But your own test can lie to you just as easily as the model can.

The confident number is the thing to distrust. This tradition was built on exactly that instinct - do not take the surface at face value, verify what sits underneath it. The benchmark that only checks the pretty picture misses the point. Measurement is easy. Measurement integrity is the whole job.