ET-SoC-1 matmul on silicon: 91–97% of the tensor unit's peak, and its energy against the A100

18 September 2026 · first measured in one session on aifoundry2; throughput, power and efficiency re-measured on three cards (aifoundry2, aifoundry3 and aifoundry1 card 1, four passes each) on 25–26 September · ET-SoC-1 PCIe cards with 32 GB LPDDR4X, minions at 600 MHz throughout · code: kernels/mmbench in the project repository · part of the ET-SoC-1 measurement reports

A tensor-unit matmul kernel on all 1,024 compute minions sustains 9.5 TFLOP/s in fp32, 19.0 TFLOP/s with fp16 inputs and 71.8 TOP/s in int8, the same on each of the three lab cards. That is 91–97% of the tensor unit's peak at this clock, and every result is checked bit-exact against a host reference. With tiles in L2 the cards draw 50–75 W doing it, their dies at 62–80 °C (four passes on each of the three cards, 25–26 September): each card's idle, 26–42 W, plus 24–28 W for the workload on aifoundry2 and aifoundry3 and 27–33 W on aifoundry1 card 1. Compared with the A100's spec sheet (peak ÷ TDP), the ET-SoC-1 is 3.3× more energy-efficient at full fp32 precision on these ±1/⁠±2 operands on aifoundry2, 3.9× on aifoundry3 and 2.8× on aifoundry1 card 1, each at its own die temperature. On random data, measured later with a variant of this loop that keeps both operands in the L1 scratchpad, it is about 3.0× on aifoundry2 and 2.4× on aifoundry1 card 1 at 80 °C, and 3.7× on aifoundry3, whose die ran at 57–58 °C. At fp16 and int8 the A100's tensor cores win on efficiency on every card (2.1–2.9× and 1.13–1.63×), and by a much wider margin on raw speed (16× and 8.7×).

Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry2 and aifoundry3 (firmware 1.3.1; aifoundry3 held at 600 MHz) and aifoundry1 card 1 (firmware 1.2.0), four passes per card. Of 19 claims tested here, this page counts 17 held and 2 differ by card; the hub’s scoreboard, 10 “proven on the cards”, 2 “differs by card”, 1 “one card only”, 4 “fewer than three repeats” and 2 “not a measurement”. The two that differ by card are how often the power reading refreshes and how much power rises within a run. The int8 variant in Later measurements held on all three cards, including "the difference is the kernel", and its energy figures there are now the energy manual's three-card values. With tiles in L2, throughput and cycle counts repeat to within 0.05% and every result is exact again; power and efficiency depend on each card's idle and die temperature, so every energy on this page is now the check's, per card (the three-card table), with the first run's as one note under Results. Record: docs/reports/data/2026-09-25-claims-v3.

Later measurements, on other operands

Later measurements. These ±1/⁠±2 operands have all-zero mantissas, which draw little power. In the Horace experiment's ablation, re-run on three cards on 26 September (four runs of each pattern per card), the same TensorFMA draws 38 W on zeros, 47 W on ones and 64 W on random-normal values on aifoundry2 at 80 °C, 51, 60 and 78 W on aifoundry1 card 1 at 80 °C (its idle is 50 W), and 27, 35 and 50 W on aifoundry3, whose die ran at 57–58 °C; and 52–53 W on aifoundry2 with only the signs or only the exponents random (these operands have both). These are the ablation's registered values, which carry a launch-temperature offset whose size depends on how the die sensor's whole-degree reading at launch maps to the die temperature (amendment C2): aifoundry2's read from 0.6 W high to 0.2 W low, aifoundry3's 0.9–1.4 W low and aifoundry1 card 1's from 0.7 W high to 0.05 W low; at the same die temperature aifoundry3 switches about 0.96–0.97 of aifoundry2's power.

On aifoundry2, a variant of this loop with both operands in the L1 scratchpad (546 cycles per op rather than 529), on random-normal data at 80 °C, gives 144 GFLOP/s per W in fp32 and 302 in fp16 (four runs each, 26 September; energy manual §3.2); its idle is 36.4 W, against 33.3–33.4 W before this loop's fp32 and fp16 workloads on the same card in the three-card check (launched at 76 °C). So the A100's bf16 tensor cores (Horace He's measured 779 GFLOP/s per W) are 5.4× more efficient per FLOP than aifoundry2's fp32 and 2.6× more than its fp16 (Why is the ET-SoC-1 low power?). On aifoundry1 card 1, also at 80 °C but idling at 50 W, the same loop gives 118 and 245 GFLOP/s per W (the A100 leads by 6.6× and 3.2×); on aifoundry3, whose die ran at 57–58 °C with a 26 W idle, 182 and 384 (4.3× and 2.0×). With amendment C2's offset these efficiencies are 2–3% lower on aifoundry3 and within 1.2% on the other two cards.

In int8 the energy manual's variant is more efficient than this loop even on random data: 1,365 GOP/s per W (63.1 TOP/s at 46.2 W, 80 °C), which narrows the A100 int8 spec's lead from 1.33× to 1.14× (aifoundry2, four passes of this loop against four runs of the variant). On aifoundry1 card 1 the variant gives 1,053 (the A100 leads by 1.48×), and on aifoundry3, at 57–58 °C with its 26 W idle, 1,852 (1,780–1,800 with amendment C2's offset), ahead of the A100's int8 spec. The difference is the kernel, not the operands: in the three-card check this loop, on its low-power operands, drew 18.0 W more above idle than the variant on random data on aifoundry2 and on aifoundry3, and 23.4 W more on aifoundry1 card 1 (99% ranges 17.0–19.1, 17.4–18.7 and 23.0–23.8 W; the variant's values carry amendment C2's offset, from 0.5 W high to 0.3 W low on aifoundry2, 0.8–1.3 W low on aifoundry3 and from 0.7 W high to 0.05 W low on aifoundry1 card 1, which moves each difference by 1.3 W at most). This loop streams both tiles from L2 at 280 cycles per op (7.3 B per minion-cycle, roughly 14 W at 3.1 pJ/B, an estimate), while that one keeps them in the L1 scratchpad (318 cycles per op); holding A there and streaming only B through the TenB buffer (270 cycles per op) adds 6.6, 6.0 and 7.5 W on random data (aifoundry2, aifoundry3, aifoundry1 card 1).

fp32 on the tensor unit
138–189 GFLOP/s per W
9.51 TFLOP/s at 50.4–69.1 W of board power; 161 on aifoundry2, 189 on aifoundry3, 138 on aifoundry1 card 1
fp16 in, fp32 accumulate
270–370 GFLOP/s per W
19.02 TFLOP/s at 51.4–70.4 W of board power; 316 on aifoundry2, 370 on aifoundry3, 270 on aifoundry1 card 1
int8 in, int32 accumulate
958–1,377 GOP/s per W
71.77 TOP/s at 52.1–74.9 W of board power; 1,172 on aifoundry2, 1,377 on aifoundry3, 958 on aifoundry1 card 1
Idle board power
26–42 W
just before each workload, 50–61% of the board power under load, much of it die leakage: 33 W on aifoundry2, 26 W on aifoundry3, 42 W on aifoundry1 card 1
Terms used on this page

The ET-SoC-1's compute cores are minions: small in-order RISC-V cores, 32 to a shire; the 32 shires that run kernels hold 1,024 of them. Each minion has two hardware threads (harts) and a tensor unit whose matrix multiply-accumulate instruction, TensorFMA, is what this benchmark times. Each shire's 4 MB of SRAM includes a 512 KB L2 cache, where the compute-bound runs keep their tiles, and each minion sets aside 3 KB of its 4 KB L1 data cache as an L1 scratchpad for tensor operands. More in the hub's glossary.

Energy efficiency against the A100

Is this chip more energy-efficient than an A100? It depends on which A100 unit it is set against, on the data and on the card. All three rows share one log axis, so equal ratios are equal distances in every row. The blue dot is the chosen card, measured with its own board-power telemetry (the three-card check, or the energy manual's random-normal variant); grey diamonds are the A100, open for the datasheet peak ÷ TDP (dense, not measured) and filled for the one measured A100 run (Horace He, bf16 on random data). Click an A100 marker, or focus it and press Enter, to compare against it.

fp32 ET: IEEE fp32 on the tensor unit (GFLOP/s per W)

fp16 inputs, fp32 accumulate ET: TensorFMA16A32 (GFLOP/s per W)

int8 inputs, int32 accumulate ET: TensorIMA8A32 (GOP/s per W)

The operand-and-card band, on other data

The operand-and-card band

Does the verdict above hold on other data? Same axis, same units, at each of the three operand patterns the energy manual measured (zeros, the ±1/⁠±2 family and random-normal), for the tiles-in-L1 variant behind the "randn" series above.

fp32

fp16

int8

A100 reference numbers, as a table

A100 reference (NVIDIA A100 datasheet, dense)

PrecisionA100 peakA100 per W (SXM4)ET per WET ÷ A100 efficiencyA100 ÷ ET speed
fp32 (A100 CUDA cores)19.5 TFLOP/s48.8138–1892.82–3.87×2.05×
tf32 tensor (10-bit mantissa)156 TFLOP/s390138–189 (fp32)0.35–0.48×16.4×
fp16 / bf16 tensor312 TFLOP/s780270–3700.35–0.47×16.4×
int8 tensor624 TOP/s1,560958–1,3770.61–0.88×8.7×

Per W is the A100 SXM4's datasheet peak divided by its 400 W TDP. The 250 W PCIe 40 GB card's per-watt figures are 1.6× higher (fp32 CUDA cores 78 GFLOP/s per W, so ET's lead there is 1.8–2.4×). ET per W spans the three cards of the check (four passes each, on ±1/⁠±2 operands); each card's values are in its table (the A100 keeps its int8 lead on all three). Both chips are 7 nm-class designs. The A100 sustains 1.4–1.8 TB/s from HBM (40 and 80 GB cards; 1.6–2.0 TB/s peak), about 20× the 76 GB/s this card streams from LPDDR4X (memory hierarchy). That gap is what the DRAM-streaming row in Results below measures.

Results

WorkloadThroughput% of peakBoard powerPer W (board)Per W above idleCheck
fp32 tensor, tiles in L29.51 TFLOP/s96.7–96.8%50.4–69.1 W138–189349–390exact
fp16→fp32 tensor, tiles in L219.02 TFLOP/s96.7%51.4–70.4 W270–370666–746exact
int8→int32 tensor, tiles in L271.77 TOP/s91.3%52.1–74.9 W958–1,3772,171–2,734exact
fp32 tensor, every tile from DRAM0.31 TFLOP/s3.2%34.4–50.6 W6.1–9.035–37exact
Idle, just before each workload––25.9–41.9 W–––

Each cell is the lowest and highest of the three cards, each the mean of four passes (25–26 September); every card is in the three-card table. First measured on 18 September in one session on aifoundry2, with each workload run for 12 s while the die warmed from 71 to 77 °C and its idle rose from 30.6 to 35.2 W (the power chart below): 56.6–61.7 W with tiles in L2, and 168, 321 and 1,162 GFLOP/s (GOP/s) per W (3.4× the A100's fp32 figure).

How these numbers were computed

"Per W" is GFLOP/s (GOP/s for int8) per watt. Above idle subtracts the idle measured just before each workload, from 3.5 to 1.8 s before its first timed launch (the windows are shaded on the power chart): 33.3–33.6 W on aifoundry2, 25.9–26.1 on aifoundry3 and 41.7–41.9 on aifoundry1 card 1. Peak at 600 MHz on 1,024 minions: 9.83 TFLOP/s fp32 (16 flop/cycle/minion), 19.66 fp16 (32), 78.6 TOP/s int8 (128). Measured op costs: 529 cycles per 16×16×16 fp32 or 16×16×32 fp16 op (512 ideal), and 280 cycles per 16×16×64 int8 op (256 ideal), the same in every launch on all three cards (529.00 and 280.35–280.37). Throughput is total work over the host-measured launch time. The minions' cycle counters agree with it to within 0.1% (implied clock 0.600 GHz; 0.5995–0.5999 GHz on all three cards). Launch-to-launch spread was below 0.03% in the L2 runs and 0.34% in the DRAM run (in the three-card check's shorter runs, up to 0.034% and 0.8%).

The three-card check

The kernel runs at the same speed on every card, so why does its work per watt differ from card to card? Each bar splits a card's board power into the idle it drew just before the workload and what the workload added on top of it.

WorkloadThroughputBoard powerAbove idle [99%]Per W (board) [99%]Per W above idle
aifoundry2 · firmware 1.3.1 · idle before each workload 33.3–33.6 W · die 77–80 °C
fp32 tensor, tiles in L29.51 TFLOP/s59.0 W25.7 W [24.2–27.2]161 [148–174]370
fp16→fp32 tensor, tiles in L219.02 TFLOP/s60.2 W26.7 W [25.1–28.4]316 [290–342]711
int8→int32 tensor, tiles in L271.77 TOP/s61.2 W27.9 W [26.6–29.3]1,172 [1,110–1,235]2,572
fp32 tensor, every tile from DRAM0.31 TFLOP/s42.4 W8.8 W [8.1–9.5]7.335
aifoundry3 · firmware 1.3.1, held at 600 MHz · idle 25.9–26.1 W · die 59–62 °C
fp32 tensor, tiles in L29.51 TFLOP/s50.4 W24.4 W [23.8–25.0]189 [179–199]390
fp16→fp32 tensor, tiles in L219.02 TFLOP/s51.4 W25.5 W [24.8–26.2]370 [346–395]746
int8→int32 tensor, tiles in L271.77 TOP/s52.1 W26.3 W [26.0–26.6]1,377 [1,307–1,446]2,734
fp32 tensor, every tile from DRAM0.31 TFLOP/s34.4 W8.4 W [8.2–8.5]9.037
aifoundry1 card 1 · firmware 1.2.0 · idle 41.7–41.9 W · die 71–76 °C
fp32 tensor, tiles in L29.51 TFLOP/s69.1 W27.2 W [26.6–27.9]138 [136–139]349
fp16→fp32 tensor, tiles in L219.02 TFLOP/s70.4 W28.6 W [28.3–28.9]270 [269–272]666
int8→int32 tensor, tiles in L271.77 TOP/s74.9 W33.1 W [32.6–33.5]958 [949–966]2,171
fp32 tensor, every tile from DRAM0.31 TFLOP/s50.6 W8.9 W [8.5–9.3]6.135

25–26 September, the same benchmark (tools/claims-v3/mmb/): four passes per card, each workload 6 s after a heat to 76 °C on the two cards whose governor is free; every launch exact, 600 MHz throughout. Means over the four passes; in brackets the 99% interval over them. "Above idle" subtracts the idle just before each workload, as the table above does. The cards differ in idle power and die temperature (aifoundry3's die runs cool, aifoundry1 card 1 idles high), and per-watt figures follow: aifoundry3's are 17% above aifoundry2's in every mode (Welch, significant at 99%), a difference of each card at its own operating temperature, not a property of the card. aifoundry3's above-idle watts are 5–6% below aifoundry2's; aifoundry1 card 1's are 1–7% above them, and 18% above in int8. Reduced by tools/claims-v3/mmb/reduce.py (items MMB-a to MMB-f).

Board power during the first run

What does "per W above idle" subtract, and why did idle rise from one workload to the next in the first run (18 September, aifoundry2)? The three-card check changed the method for that reason: it heated the die first on the two cards whose governor is free and ran each workload for 6 s, and the idle before each workload held within 0.3 W on each card. Board power was polled from the service processor about 8 times a second. The reading itself changes less often: with a poll like this one about every 126–139 ms on aifoundry2 and aifoundry1 card 1 and 223–224 ms on aifoundry3 (three passes per card, 26 September; faster pollers stretch it: Power and temperature). Shaded spans are the benchmark launches, about 12 s per workload, with idle gaps between them; the thin spikes just before each span are short calibration launches. Power creeps up within each span as the die warms: minion-shire temperature (the sensors' average) went from 71 °C to 77 °C, and the hottest sensor's high-water mark read 85 °C afterwards. For the same reason the idle in each gap sits higher than the one before (the grey steps).

How to read this chart

Pick a workload to shade its two windows: the idle it is measured against (3.5 to 1.8 s before its first timed launch, orange) and the span its power is averaged over (from 1 s after its first launch to its end, blue). The dashed lines are those two means.

How the benchmark works

Version history

Versions: published 18 September 2026; revised 24 September 2026 (per-workload idle, operand and A100 context, corrections) and 25 September 2026 (efficiency explorer, idle windows, corrections); 25 September (version 3): which card and how many runs each figure rests on, aifoundry3's values beside aifoundry2's in the later measurements, and the idle's leakage split as a range (record: docs/reports/data/2026-09-25-claims-v3); 26 September: the three-card check (aifoundry2, aifoundry3, aifoundry1 card 1) added to the lede, the KPIs, Results (its table), the power-meter refresh per card and the within-run rise per card; 27 September: the operand-and-card band (each card on board power), the A100 legend drawn as diamonds, and repeated text cut; later that day, each card's board power charted as its idle and what the workload added (the three-card check); 28 September: the meter's refresh note shortened; the note's counts given by both rules; then every energy on the page (the lede, the KPIs, the explorer, the A100 table and Results: board power, per W, the A100 ratios) the three-card check's, the first run one note under Results, and the power chart marked as the first run's.

Caveats

Caveats in full

Next steps

Reproduce

Reproduce this
# regenerate this page's numbers and charts from the committed data (no card)
python3 scripts/mmbench-report-data.py docs/reports/data/2026-09-18-aifoundry2 \
    --manual docs/reports/data/2026-09-23-energy-manual/manual.json \
    --embed docs/reports/2026-09-18-et-soc1-matmul-efficiency.html

# exact checks of every mode on the simulator, one shire, no card (about 40 s each)
scripts/vm make mmbench-check          # laptop VM; on a lab machine: cd ~/nekko && make mmbench-check

# on the card: copy the repository to the lab machine, build gp-sdk (pinned, patched) and mmbench there
scripts/deploy-lab-gpsdk.sh aifoundry2
# then on aifoundry2 (quit et-powertop first: the power logger needs the management node)
cd ~/nekko && nice make mmbench-check bench-power DEVICE=silicon JOBS=4
timeout 10 /opt/et/bin/it_test_code_loading --mode=pcie   # the card hello world
# simulator hello worlds: see Test drive
# copy ~/nekko/build/mmbench-power/ into a new dated docs/reports/data/ directory

The lab cards are shared, and the lab rule is to hold a device for at most 10 s. The first run kept the card busy for about 12 s per workload, before the project adopted that rule. scripts/mmbench-power.py now runs every launcher under timeout 10 and defaults to about 6 s per workload, which still leaves about 40 samples per workload after the settle time (about 35 distinct readings on aifoundry2, about 20 on aifoundry3). It also stores the idle windows it averages.

Raw data: docs/reports/data/2026-09-18-aifoundry2 (board-power samples, one record per launch and the per-workload summary). scripts/mmbench-report-data.py turns it into this page's throughput, efficiency, idle before each workload (the same window as idle_before() in scripts/mmbench-power.py) and the power chart; with --manual it adds the energy manual's random-normal variant to the efficiency explorer, and by default (--v3) it reads the three-card check's docs/reports/data/2026-09-25-claims-v3/results/mmb.json, which the explorer, the three-card chart and every energy on the page come from, and prints the values the tables quote. The tables are written by hand from its output.

Sources: aifoundry-org/et-platform, et-man (programmer's reference and errata), marty1885/et-testdrive, FOSDEM 2026, "Zero to matmul with the ET-SoC-1", the NVIDIA A100 Tensor Core GPU datasheet, and Horace He's "Strangely, Matrix Multiplications on GPUs Run Faster When Given 'Predictable' Data!" for the measured A100 run.

Running it on a lab machine

These notes are for running the benchmark on a lab machine, whose /opt/et is older than the gp-sdk that upstream documents. The laptop setup and the first scalar SGEMM, from the same day, are in Test drive.

  1. Toolchain was already there. /opt/et is et-platform built from commit 353f20e (30 Dec 2025): RISC-V GCC 15.1, sys_emu, runtime, firmware and the upstream test binaries. It is not on PATH by default. aifoundry2 is x86-64 Ubuntu 24.04 with one card (PCIe Gen4, 32 GB, firmware BL 0.20 / minion 0.23); in this session its minion shires ran at 600 MHz and its NoC at 400 MHz.
  2. Hello world: et-platform's it_test_code_loading --mode=pcie passes 3/3 on the card in 0.6 s (the simulator runs are in Test drive). It also passed on all three cards at the start of every pass of the three-card check.
  3. et-testdrive (marty1885's minimal host+kernel project) builds unmodified against /opt/et. On the card all 64 harts of shire 0 print "Hello World from hart N", decoded in-process from the trace buffer. One catch: it opens the management node, which only one process may hold, so it fails with "Device or resource busy" while et-powertop is running. The upstream tests open only the ops node.
  4. Writing my own kernels with gp-sdk needed four fixes against this older install. They are in patches/lab-gp-sdk-06605ab.patch, which scripts/deploy-lab-gpsdk.sh applies; patches/README.md says what each one fixes.

← All ET-SoC-1 measurement reports