Hand it to the next shire

22 September 2026 · three ET-SoC-1 cards, aifoundry2, aifoundry3 and aifoundry1 card 1 (every sweep re-run three times on each, 25 September) · a new probe, workloads/onchip (--test relay) · part of the ET-SoC-1 measurement reports

A shire's own scratchpad runs at 2.5 TB/s and another shire's at about 1 TB/s, while DRAM manages 76 GB/s for the whole chip. Those are bandwidths, not results. The question is whether a computation can be rearranged to collect them. It can. A chain of stages that hands each stage's output to another shire, instead of writing it to DRAM and reading it back, runs 12.2–12.4 times faster and uses a thirteenth of the energy per byte. The shire that receives it is the next one in ID order, which on the mesh is 1 to 10 hops away (3.5 on average). Over all 31 ring offsets, which shire it is moved the bandwidth by up to a quarter (), against the twelvefold gain over DRAM; the ring in ID order, used for the headline, is the slowest, together with its mirror image. Keeping the output in the shire's own scratchpad is 30.8–31.2 times faster. Both hold on all three cards () and in every repeat (). Both also need the data to outgrow the 32 MB L3: below that, the hand-off wins at most 1.4–1.5× and the shire's own scratchpad 1.3–3.2×.

Checked on three cards (26 September 2026): this page's claims were re-measured under a pre-registered plan, every sweep three times on aifoundry2, aifoundry3 and aifoundry1 card 1, and the sweep figures here are now the means of those passes (one sweep per card before). Section 5 now covers all 31 ring offsets (the line fitted to the first five predicted the other eleven ring geometries within 0.7% on every card). Corrected: the single-stage advantage (14.3–14.8×, not 15.3×) and the own scratchpad's read energy, no longer split by card. The energies per byte are now that check's too (six passes on each card). aifoundry3 lost four of its 351 relay launches to a host-process crash, so the registered test of section 5 is complete on the other two cards only. Of 9 claims tested here, this page counts 7 held and 2 corrected; the hub’s scoreboard, 1 “proven on the cards” and 8 “fewer than three repeats”. Record: docs/reports/data/2026-09-25-claims-v3.

Terms used on this page

The ET-SoC-1's cores are minions, 32 to a shire; each shire has 4 MB of SRAM, split on these cards into 0.5 MB of L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad that any shire can address. A tensor store writes a block from a minion's registers straight to memory, past its caches. The shires sit on a mesh network-on-chip, and a hop is one step between neighbouring mesh stops, 3.72 mm. More in the hub's glossary.

Passing each stage to the next shire
against writing it to DRAM and reading it back
Keeping it in the shire's own scratchpad
no data crosses a shire at all
Energy per byte, next shire vs DRAM
pooled over the three cards' passes
Where the advantage appears
of live intermediate data: it has to outgrow the 32 MB L3

1. One kernel, three places to put the answer

The computation is a relay. A slab of floats belongs to each shire; a stage reads every element, adds one to it, and writes it out; the next stage reads what the last one wrote. Thirty-two shires, 1,024 minions, eight stages, 1 MB per shire per stage, half a gigabyte of traffic. The same code, the same barriers and the same arithmetic run three ways, and the only thing that differs is where a stage's output goes:

The table, and how we know the data really crossed a shire

How we know the data really crossed a shire. Each shire starts its slab filled with its own number. After eight stages of adding one, a shire that kept its own data holds 8. In the hand-off run, shire 0 holds 32 — which is 24 plus 8, and 24 is the shire eight places back round the ring. Every element is checked against that expectation, on every run in this report; a run that did not really move its data would fail.

2. Power within a watt, a thirtieth of the work

Power was measured separately, in one session on aifoundry2, with each medium launched back to back for twelve seconds so the card reaches a steady state; the table's last column gives the three-card check's passes (six on each card, 26 September). The striking part is not that DRAM costs more energy per byte. It is that the three draw within about a watt of each other, and one of them gets thirty times as much done with it. Across those passes the own scratchpad draws 0.5–0.7 W more than DRAM, more at 99% on every card, and the next shire 0.3–0.5 W less, a difference resolved at 99% on aifoundry1 card 1 only (on aifoundry3 its interval just reaches zero):

Power, rate and energy per byte for the three places, side by side

The table, method note and cross-card re-measurement

3. Where the advantage comes from, and where there isn't one

The win is against DRAM, not against the memory hierarchy. Shrink the working set and the L3 quietly does most of the same job for free:

The same eight-stage relay at different working-set sizes

The numbers behind the L3 crossover

Two more sweeps: stages, and shires taking part

The same relay with one setting changed at a time: pipeline stages, and shires taking part

4. How much arithmetic it takes to stop mattering

All of this is a statement about moving data, so it fades as the arithmetic grows. Repeating the add w times per element walks the workload up the roofline:

Advantage over DRAM against arithmetic per element

5. How far the slab moves: the farthest hand-off sets the pace

Which shire reads from which, and what sets the bandwidth

The 31 offsets, card by card, and what the mesh distance costs

6. What this is good for

7. Method, and what is not established

Method, and what is not established, in full
Reproduce this
workloads/onchip/run_onchip.sh DATA                 # 88 configurations, milliseconds of card time each (22 September)
tools/claims-v3/lat/block.sh PASS                   # the three-card passes (unit rl: the 88 and ring offsets 1-31; its README)
tools/ettelem/run_onchip_power.sh DATA 12           # board power, one burst per medium
R=docs/reports/data/2026-09-25-claims-v3/raw
python3 workloads/onchip/analyze_onchip.py $R/aifoundry2/lat/p*/rl/sweep.jsonl $R/aifoundry3/lat/p*/rl/sweep.jsonl \
    $R/aifoundry1-c1/lat/p*/rl/sweep.jsonl --cards aifoundry2,aifoundry3,aifoundry1-c1 \
    --power docs/reports/data/2026-09-22-onchip-aifoundry2 --out onchip.json   # pass means; power from 22 September
python3 tools/ettelem/analyze_reruns.py docs/reports/data/2026-09-23-reruns-aifoundry2-warm \
    docs/reports/data/2026-09-23-reruns-aifoundry3 --v3-rl $R --out reruns.json   # the passes behind §2's ranges
Version history and provenance

Raw data and the analysis output: docs/reports/data/2026-09-22-onchip-aifoundry2 and …-aifoundry3 (22 September), and the three-card passes under docs/reports/data/2026-09-25-claims-v3/raw/<card>/lat/p*/rl/. Versions (one line per date; every wording is in the source's history): 22 September, first published; 24 September, the “next shire” is the next by ID, not a mesh neighbour, and section 4's flops-per-byte unit corrected; 25 September, the ring offset moves the bandwidth by up to a quarter (it had said distance made no measurable difference), and “the same power” replaced by what the passes show; 26 September, three passes on three cards, the one-stage advantage corrected (14.3–14.8×, not 15.3×) and the own-scratchpad read energy no longer split by card; 27 September, charts; 28 September, repeats cut and the note's counts given by both rules. Record: docs/reports/data/2026-09-25-claims-v3.