Hand it to the next shire
A shire's own scratchpad runs at 2.5 TB/s and another shire's at about 1 TB/s, while DRAM manages 76 GB/s for the whole chip. Those are bandwidths, not results. The question is whether a computation can be rearranged to collect them. It can. A chain of stages that hands each stage's output to another shire, instead of writing it to DRAM and reading it back, runs 12.2–12.4 times faster and uses a thirteenth of the energy per byte. The shire that receives it is the next one in ID order, which on the mesh is 1 to 10 hops away (3.5 on average). Over all 31 ring offsets, which shire it is moved the bandwidth by up to a quarter (), against the twelvefold gain over DRAM; the ring in ID order, used for the headline, is the slowest, together with its mirror image. Keeping the output in the shire's own scratchpad is 30.8–31.2 times faster. Both hold on all three cards () and in every repeat (). Both also need the data to outgrow the 32 MB L3: below that, the hand-off wins at most 1.4–1.5× and the shire's own scratchpad 1.3–3.2×.
Checked on three cards (26 September 2026): this page's claims were
re-measured under a pre-registered plan, every sweep three times on aifoundry2, aifoundry3 and aifoundry1 card 1,
and the sweep figures here are now the means of those passes (one sweep per card before). Section 5 now covers all
31 ring offsets (the line fitted to the first five predicted the other eleven ring geometries within 0.7% on every
card). Corrected: the single-stage advantage (14.3–14.8×, not 15.3×) and the own scratchpad's read energy, no longer
split by card. The energies per byte are now that check's too (six passes on each card). aifoundry3 lost four of its
351 relay launches to a host-process crash, so the registered test of section 5 is complete on the other two cards
only. Of 9 claims tested here, this page counts 7 held and 2 corrected; the hub’s scoreboard, 1 “proven on the cards” and 8 “fewer than three repeats”. Record:
docs/reports/data/2026-09-25-claims-v3.
Terms used on this page
The ET-SoC-1's cores are minions, 32 to a shire; each shire has 4 MB of SRAM, split on these cards into 0.5 MB of L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad that any shire can address. A tensor store writes a block from a minion's registers straight to memory, past its caches. The shires sit on a mesh network-on-chip, and a hop is one step between neighbouring mesh stops, 3.72 mm. More in the hub's glossary.
1. One kernel, three places to put the answer
The computation is a relay. A slab of floats belongs to each shire; a stage reads every element, adds one to it, and writes it out; the next stage reads what the last one wrote. Thirty-two shires, 1,024 minions, eight stages, 1 MB per shire per stage, half a gigabyte of traffic. The same code, the same barriers and the same arithmetic run three ways, and the only thing that differs is where a stage's output goes:
The table, and how we know the data really crossed a shire
How we know the data really crossed a shire. Each shire starts its slab filled with its own number. After eight stages of adding one, a shire that kept its own data holds 8. In the hand-off run, shire 0 holds 32 — which is 24 plus 8, and 24 is the shire eight places back round the ring. Every element is checked against that expectation, on every run in this report; a run that did not really move its data would fail.
2. Power within a watt, a thirtieth of the work
Power was measured separately, in one session on aifoundry2, with each medium launched back to back for twelve seconds so the card reaches a steady state; the table's last column gives the three-card check's passes (six on each card, 26 September). The striking part is not that DRAM costs more energy per byte. It is that the three draw within about a watt of each other, and one of them gets thirty times as much done with it. Across those passes the own scratchpad draws 0.5–0.7 W more than DRAM, more at 99% on every card, and the next shire 0.3–0.5 W less, a difference resolved at 99% on aifoundry1 card 1 only (on aifoundry3 its interval just reaches zero):
The table, method note and cross-card re-measurement
3. Where the advantage comes from, and where there isn't one
The win is against DRAM, not against the memory hierarchy. Shrink the working set and the L3 quietly does most of the same job for free:
The numbers behind the L3 crossover
Two more sweeps: stages, and shires taking part
4. How much arithmetic it takes to stop mattering
All of this is a statement about moving data, so it fades as the arithmetic grows. Repeating the add
w times per element walks the workload up the roofline:
5. How far the slab moves: the farthest hand-off sets the pace
The 31 offsets, card by card, and what the mesh distance costs
6. What this is good for
- Any chain of passes over data that outgrows the L3. Multi-stage filters, iterated stencils, a deep and narrow network, a stream that several kernels touch in turn. The standard shape writes the intermediate to memory between stages, which on this chip runs at 48 GB/s of reads and writes.
- The unit that matters is the shire, not the minion. Its scratchpad is 2.5 MB, addressable by every other shire, and written by a tensor store that bypasses both caches, so what one shire writes another reads with no coherence problem at all.
- Budget one barrier per stage. A chip-wide barrier costs about 5,000 cycles (8.3 µs at 600 MHz; see on-chip communication). At 1 MB per shire a DRAM stage is , so the barrier is under a fifth of even the fastest case. From four stages to thirty-two the advantage is flat (12.0–12.8× on the three cards); with one or two stages it is larger (14.3–14.8× and 13.3–13.8×), where the DRAM route reaches only 39–41 and 43–44 GB/s. Build the barrier with one atomic per shire and a credit release, never one counter every shire polls (the hot line): this relay's first barrier did that and hung in its one development run (card not recorded).
- Inside a shire, use the register path, one partner at a time.
TensorSend/TensorRecvis immune to shire-cache contention, but a minion that takes readies from two partners at once can hang until the chip is reset, as read from the RTL and seen once, on aifoundry2 (on-chip communication): give each link its own minion, or take the directions in barrier-separated phases.
7. Method, and what is not established
Method, and what is not established, in full
- The probe is
workloads/onchip, new for this. A stage's output is written with a tensor store, which bypasses the L1 and L2 caches, so the shire that reads it next cannot read a stale line. Inputs are read with plain vector loads from addresses the reading minion has never written; two buffers alternate, and a minion's 32 KB chunk is sixty-four times its 512 B of L1, so nothing it touched two stages ago survives. Every run verifies its own output element by element, which is what makes that argument checkable rather than asserted. Every run reported here passed. The check has failed, in one discarded attempt on aifoundry2 on 23 September: of the 16 DRAM-route runs in which the governor changed the minion clock mid-run, two returned 1,000 and 1,118 wrong elements; none of the 16 that stayed at 600 MHz did. One session does not establish that the clock change caused it, and aifoundry3, whose clock is fixed, cannot test it. - Rates come from the on-device cycle counter, converted at 600 MHz (the minion clock stayed there through every power burst), not from the wall clock: a kernel launch costs a few hundred microseconds to a millisecond (about 10 ms with the largest buffers) and an on-chip stage costs 27,000–68,000 cycles, so wall time would measure the host.
- Scratchpad addressing is the Programmer's Reference Manual's (PRM) format 0, shire in bits [29:23]. Offset 0 of a shire's scratchpad faulted during development (card not recorded); the buffers start 256 KB in, which also leaves room for two of them inside the 2.5 MB.
- The 88 configurations, and the hand-off at every ring offset from 1 to 31, ran three times on each of three cards on 25 September, in the pre-registered passes of the three-card check; the first sweep, on 22 September, ran each configuration once on aifoundry2 and aifoundry3, and the power session of §2 ran on aifoundry2 only. Every sweep figure on this page is the mean of those passes: . A result that names no card held on all three, in three independent runs or as the same deterministic value on each; where one rests on one card, one session or fewer runs, the text says so.
- Not established:
- Whether a working set larger than the 80 MB of scratchpad can be streamed through the same relay. That is the case where on-chip hand-off would be the only option rather than the faster one, and we did not build the flow control it needs.
- Whether a ring laid out along the mesh, with every hand-off one hop (such a ring exists on this map), is faster still and cheaper per byte. The trend over longest hand-offs of 6–10 hops (section 5) suggests it could be faster, and every offset measured here mixes near and far shires, so the energy per byte of a neighbour hand-off is not known either.
- A real pipeline. The arithmetic here is one vector add, chosen to make the measurement about data movement; a real pipeline would have real stages, and section 4 says what that costs.
- A cheaper barrier. The chip-wide barrier per stage is the simplest synchronisation, not the cheapest; a credit from each shire to the one that reads its output would do, and would help the on-chip media most.
- Other memory systems. The comparison is against this chip's DRAM, which streams 76 GB/s of plain reads and carries this relay's reads and writes at 48 GB/s, not against a memory system in general.
Reproduce this
workloads/onchip/run_onchip.sh DATA # 88 configurations, milliseconds of card time each (22 September)
tools/claims-v3/lat/block.sh PASS # the three-card passes (unit rl: the 88 and ring offsets 1-31; its README)
tools/ettelem/run_onchip_power.sh DATA 12 # board power, one burst per medium
R=docs/reports/data/2026-09-25-claims-v3/raw
python3 workloads/onchip/analyze_onchip.py $R/aifoundry2/lat/p*/rl/sweep.jsonl $R/aifoundry3/lat/p*/rl/sweep.jsonl \
$R/aifoundry1-c1/lat/p*/rl/sweep.jsonl --cards aifoundry2,aifoundry3,aifoundry1-c1 \
--power docs/reports/data/2026-09-22-onchip-aifoundry2 --out onchip.json # pass means; power from 22 September
python3 tools/ettelem/analyze_reruns.py docs/reports/data/2026-09-23-reruns-aifoundry2-warm \
docs/reports/data/2026-09-23-reruns-aifoundry3 --v3-rl $R --out reruns.json # the passes behind §2's ranges
Version history and provenance
Raw data and the analysis output: docs/reports/data/2026-09-22-onchip-aifoundry2 and …-aifoundry3 (22 September), and the three-card passes under docs/reports/data/2026-09-25-claims-v3/raw/<card>/lat/p*/rl/. Versions (one line per date; every wording is in the source's history): 22 September, first published; 24 September, the “next shire” is the next by ID, not a mesh neighbour, and section 4's flops-per-byte unit corrected; 25 September, the ring offset moves the bandwidth by up to a quarter (it had said distance made no measurable difference), and “the same power” replaced by what the passes show; 26 September, three passes on three cards, the one-stage advantage corrected (14.3–14.8×, not 15.3×) and the own-scratchpad read energy no longer split by card; 27 September, charts; 28 September, repeats cut and the note's counts given by both rules. Record: docs/reports/data/2026-09-25-claims-v3.
8. Related reports
- The energy manual, §5 — this relay's energy re-measured, with bars from three cards.
- Heat per millimetre — what each mesh hop costs in energy, measured with chosen bit patterns.
- Memory hierarchy — the 2.5 TB/s, 1 TB/s and 76 GB/s of the lede.
- On-chip communication — the shire map, 12 cycles per hop, the chip-wide barrier and the one-ready-flag trap.
- Ridge points — the roofline that section 4 climbs, for every level of memory.
- One hot line stops a shire — the companion: what happens when the structure several shires share is a single line rather than a slab.
- docs/findings/ — the findings index: where every claim here comes from, and the experiment register.