ET-SoC-1 matmul on silicon: 91–97% of the tensor unit's peak, and its energy against the A100
A tensor-unit matmul kernel on all 1,024 compute minions sustains 9.5 TFLOP/s in fp32, 19.0 TFLOP/s with fp16 inputs and 71.8 TOP/s in int8, the same on each of the three lab cards. That is 91–97% of the tensor unit's peak at this clock, and every result is checked bit-exact against a host reference. With tiles in L2 the cards draw 50–75 W doing it, their dies at 62–80 °C (four passes on each of the three cards, 25–26 September): each card's idle, 26–42 W, plus 24–28 W for the workload on aifoundry2 and aifoundry3 and 27–33 W on aifoundry1 card 1. Compared with the A100's spec sheet (peak ÷ TDP), the ET-SoC-1 is 3.3× more energy-efficient at full fp32 precision on these ±1/±2 operands on aifoundry2, 3.9× on aifoundry3 and 2.8× on aifoundry1 card 1, each at its own die temperature. On random data, measured later with a variant of this loop that keeps both operands in the L1 scratchpad, it is about 3.0× on aifoundry2 and 2.4× on aifoundry1 card 1 at 80 °C, and 3.7× on aifoundry3, whose die ran at 57–58 °C. At fp16 and int8 the A100's tensor cores win on efficiency on every card (2.1–2.9× and 1.13–1.63×), and by a much wider margin on raw speed (16× and 8.7×).
Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry2 and aifoundry3 (firmware 1.3.1; aifoundry3 held at 600 MHz) and aifoundry1 card 1 (firmware 1.2.0), four passes per card. Of 19 claims tested here, this page counts 17 held and 2 differ by card; the hub’s scoreboard, 10 “proven on the cards”, 2 “differs by card”, 1 “one card only”, 4 “fewer than three repeats” and 2 “not a measurement”. The two that differ by card are how often the power reading refreshes and how much power rises within a run. The int8 variant in Later measurements held on all three cards, including "the difference is the kernel", and its energy figures there are now the energy manual's three-card values. With tiles in L2, throughput and cycle counts repeat to within 0.05% and every result is exact again; power and efficiency depend on each card's idle and die temperature, so every energy on this page is now the check's, per card (the three-card table), with the first run's as one note under Results. Record: docs/reports/data/2026-09-25-claims-v3.
Later measurements, on other operands
Later measurements. These ±1/±2 operands have all-zero mantissas, which draw little power. In the Horace experiment's ablation, re-run on three cards on 26 September (four runs of each pattern per card), the same TensorFMA draws 38 W on zeros, 47 W on ones and 64 W on random-normal values on aifoundry2 at 80 °C, 51, 60 and 78 W on aifoundry1 card 1 at 80 °C (its idle is 50 W), and 27, 35 and 50 W on aifoundry3, whose die ran at 57–58 °C; and 52–53 W on aifoundry2 with only the signs or only the exponents random (these operands have both). These are the ablation's registered values, which carry a launch-temperature offset whose size depends on how the die sensor's whole-degree reading at launch maps to the die temperature (amendment C2): aifoundry2's read from 0.6 W high to 0.2 W low, aifoundry3's 0.9–1.4 W low and aifoundry1 card 1's from 0.7 W high to 0.05 W low; at the same die temperature aifoundry3 switches about 0.96–0.97 of aifoundry2's power.
On aifoundry2, a variant of this loop with both operands in the L1 scratchpad (546 cycles per op rather than 529), on random-normal data at 80 °C, gives 144 GFLOP/s per W in fp32 and 302 in fp16 (four runs each, 26 September; energy manual §3.2); its idle is 36.4 W, against 33.3–33.4 W before this loop's fp32 and fp16 workloads on the same card in the three-card check (launched at 76 °C). So the A100's bf16 tensor cores (Horace He's measured 779 GFLOP/s per W) are 5.4× more efficient per FLOP than aifoundry2's fp32 and 2.6× more than its fp16 (Why is the ET-SoC-1 low power?). On aifoundry1 card 1, also at 80 °C but idling at 50 W, the same loop gives 118 and 245 GFLOP/s per W (the A100 leads by 6.6× and 3.2×); on aifoundry3, whose die ran at 57–58 °C with a 26 W idle, 182 and 384 (4.3× and 2.0×). With amendment C2's offset these efficiencies are 2–3% lower on aifoundry3 and within 1.2% on the other two cards.
In int8 the energy manual's variant is more efficient than this loop even on random data: 1,365 GOP/s per W (63.1 TOP/s at 46.2 W, 80 °C), which narrows the A100 int8 spec's lead from 1.33× to 1.14× (aifoundry2, four passes of this loop against four runs of the variant). On aifoundry1 card 1 the variant gives 1,053 (the A100 leads by 1.48×), and on aifoundry3, at 57–58 °C with its 26 W idle, 1,852 (1,780–1,800 with amendment C2's offset), ahead of the A100's int8 spec. The difference is the kernel, not the operands: in the three-card check this loop, on its low-power operands, drew 18.0 W more above idle than the variant on random data on aifoundry2 and on aifoundry3, and 23.4 W more on aifoundry1 card 1 (99% ranges 17.0–19.1, 17.4–18.7 and 23.0–23.8 W; the variant's values carry amendment C2's offset, from 0.5 W high to 0.3 W low on aifoundry2, 0.8–1.3 W low on aifoundry3 and from 0.7 W high to 0.05 W low on aifoundry1 card 1, which moves each difference by 1.3 W at most). This loop streams both tiles from L2 at 280 cycles per op (7.3 B per minion-cycle, roughly 14 W at 3.1 pJ/B, an estimate), while that one keeps them in the L1 scratchpad (318 cycles per op); holding A there and streaming only B through the TenB buffer (270 cycles per op) adds 6.6, 6.0 and 7.5 W on random data (aifoundry2, aifoundry3, aifoundry1 card 1).
Terms used on this page
The ET-SoC-1's compute cores are minions: small in-order RISC-V cores, 32 to a shire; the 32 shires that run kernels hold 1,024 of them. Each minion has two hardware threads (harts) and a tensor unit whose matrix multiply-accumulate instruction, TensorFMA, is what this benchmark times. Each shire's 4 MB of SRAM includes a 512 KB L2 cache, where the compute-bound runs keep their tiles, and each minion sets aside 3 KB of its 4 KB L1 data cache as an L1 scratchpad for tensor operands. More in the hub's glossary.
Energy efficiency against the A100
Is this chip more energy-efficient than an A100? It depends on which A100 unit it is set against, on the data and on the card. All three rows share one log axis, so equal ratios are equal distances in every row. The blue dot is the chosen card, measured with its own board-power telemetry (the three-card check, or the energy manual's random-normal variant); grey diamonds are the A100, open for the datasheet peak ÷ TDP (dense, not measured) and filled for the one measured A100 run (Horace He, bf16 on random data). Click an A100 marker, or focus it and press Enter, to compare against it.
The operand-and-card band, on other data
The operand-and-card band
Does the verdict above hold on other data? Same axis, same units, at each of the three operand patterns the energy manual measured (zeros, the ±1/±2 family and random-normal), for the tiles-in-L1 variant behind the "randn" series above.
A100 reference numbers, as a table
A100 reference (NVIDIA A100 datasheet, dense)
| Precision | A100 peak | A100 per W (SXM4) | ET per W | ET ÷ A100 efficiency | A100 ÷ ET speed |
|---|---|---|---|---|---|
| fp32 (A100 CUDA cores) | 19.5 TFLOP/s | 48.8 | 138–189 | 2.82–3.87× | 2.05× |
| tf32 tensor (10-bit mantissa) | 156 TFLOP/s | 390 | 138–189 (fp32) | 0.35–0.48× | 16.4× |
| fp16 / bf16 tensor | 312 TFLOP/s | 780 | 270–370 | 0.35–0.47× | 16.4× |
| int8 tensor | 624 TOP/s | 1,560 | 958–1,377 | 0.61–0.88× | 8.7× |
Per W is the A100 SXM4's datasheet peak divided by its 400 W TDP. The 250 W PCIe 40 GB card's per-watt figures are 1.6× higher (fp32 CUDA cores 78 GFLOP/s per W, so ET's lead there is 1.8–2.4×). ET per W spans the three cards of the check (four passes each, on ±1/±2 operands); each card's values are in its table (the A100 keeps its int8 lead on all three). Both chips are 7 nm-class designs. The A100 sustains 1.4–1.8 TB/s from HBM (40 and 80 GB cards; 1.6–2.0 TB/s peak), about 20× the 76 GB/s this card streams from LPDDR4X (memory hierarchy). That gap is what the DRAM-streaming row in Results below measures.
Results
| Workload | Throughput | % of peak | Board power | Per W (board) | Per W above idle | Check |
|---|---|---|---|---|---|---|
| fp32 tensor, tiles in L2 | 9.51 TFLOP/s | 96.7–96.8% | 50.4–69.1 W | 138–189 | 349–390 | exact |
| fp16→fp32 tensor, tiles in L2 | 19.02 TFLOP/s | 96.7% | 51.4–70.4 W | 270–370 | 666–746 | exact |
| int8→int32 tensor, tiles in L2 | 71.77 TOP/s | 91.3% | 52.1–74.9 W | 958–1,377 | 2,171–2,734 | exact |
| fp32 tensor, every tile from DRAM | 0.31 TFLOP/s | 3.2% | 34.4–50.6 W | 6.1–9.0 | 35–37 | exact |
| Idle, just before each workload | – | – | 25.9–41.9 W | – | – | – |
Each cell is the lowest and highest of the three cards, each the mean of four passes (25–26 September); every card is in the three-card table. First measured on 18 September in one session on aifoundry2, with each workload run for 12 s while the die warmed from 71 to 77 °C and its idle rose from 30.6 to 35.2 W (the power chart below): 56.6–61.7 W with tiles in L2, and 168, 321 and 1,162 GFLOP/s (GOP/s) per W (3.4× the A100's fp32 figure).
How these numbers were computed
"Per W" is GFLOP/s (GOP/s for int8) per watt. Above idle subtracts the idle measured just before each workload, from 3.5 to 1.8 s before its first timed launch (the windows are shaded on the power chart): 33.3–33.6 W on aifoundry2, 25.9–26.1 on aifoundry3 and 41.7–41.9 on aifoundry1 card 1. Peak at 600 MHz on 1,024 minions: 9.83 TFLOP/s fp32 (16 flop/cycle/minion), 19.66 fp16 (32), 78.6 TOP/s int8 (128). Measured op costs: 529 cycles per 16×16×16 fp32 or 16×16×32 fp16 op (512 ideal), and 280 cycles per 16×16×64 int8 op (256 ideal), the same in every launch on all three cards (529.00 and 280.35–280.37). Throughput is total work over the host-measured launch time. The minions' cycle counters agree with it to within 0.1% (implied clock 0.600 GHz; 0.5995–0.5999 GHz on all three cards). Launch-to-launch spread was below 0.03% in the L2 runs and 0.34% in the DRAM run (in the three-card check's shorter runs, up to 0.034% and 0.8%).
The three-card check
The kernel runs at the same speed on every card, so why does its work per watt differ from card to card? Each bar splits a card's board power into the idle it drew just before the workload and what the workload added on top of it.
| Workload | Throughput | Board power | Above idle [99%] | Per W (board) [99%] | Per W above idle |
|---|---|---|---|---|---|
| aifoundry2 · firmware 1.3.1 · idle before each workload 33.3–33.6 W · die 77–80 °C | |||||
| fp32 tensor, tiles in L2 | 9.51 TFLOP/s | 59.0 W | 25.7 W [24.2–27.2] | 161 [148–174] | 370 |
| fp16→fp32 tensor, tiles in L2 | 19.02 TFLOP/s | 60.2 W | 26.7 W [25.1–28.4] | 316 [290–342] | 711 |
| int8→int32 tensor, tiles in L2 | 71.77 TOP/s | 61.2 W | 27.9 W [26.6–29.3] | 1,172 [1,110–1,235] | 2,572 |
| fp32 tensor, every tile from DRAM | 0.31 TFLOP/s | 42.4 W | 8.8 W [8.1–9.5] | 7.3 | 35 |
| aifoundry3 · firmware 1.3.1, held at 600 MHz · idle 25.9–26.1 W · die 59–62 °C | |||||
| fp32 tensor, tiles in L2 | 9.51 TFLOP/s | 50.4 W | 24.4 W [23.8–25.0] | 189 [179–199] | 390 |
| fp16→fp32 tensor, tiles in L2 | 19.02 TFLOP/s | 51.4 W | 25.5 W [24.8–26.2] | 370 [346–395] | 746 |
| int8→int32 tensor, tiles in L2 | 71.77 TOP/s | 52.1 W | 26.3 W [26.0–26.6] | 1,377 [1,307–1,446] | 2,734 |
| fp32 tensor, every tile from DRAM | 0.31 TFLOP/s | 34.4 W | 8.4 W [8.2–8.5] | 9.0 | 37 |
| aifoundry1 card 1 · firmware 1.2.0 · idle 41.7–41.9 W · die 71–76 °C | |||||
| fp32 tensor, tiles in L2 | 9.51 TFLOP/s | 69.1 W | 27.2 W [26.6–27.9] | 138 [136–139] | 349 |
| fp16→fp32 tensor, tiles in L2 | 19.02 TFLOP/s | 70.4 W | 28.6 W [28.3–28.9] | 270 [269–272] | 666 |
| int8→int32 tensor, tiles in L2 | 71.77 TOP/s | 74.9 W | 33.1 W [32.6–33.5] | 958 [949–966] | 2,171 |
| fp32 tensor, every tile from DRAM | 0.31 TFLOP/s | 50.6 W | 8.9 W [8.5–9.3] | 6.1 | 35 |
25–26 September, the same benchmark (tools/claims-v3/mmb/): four passes per card, each workload 6 s after a heat to 76 °C on the two cards whose governor is free; every launch exact, 600 MHz throughout. Means over the four passes; in brackets the 99% interval over them. "Above idle" subtracts the idle just before each workload, as the table above does. The cards differ in idle power and die temperature (aifoundry3's die runs cool, aifoundry1 card 1 idles high), and per-watt figures follow: aifoundry3's are 17% above aifoundry2's in every mode (Welch, significant at 99%), a difference of each card at its own operating temperature, not a property of the card. aifoundry3's above-idle watts are 5–6% below aifoundry2's; aifoundry1 card 1's are 1–7% above them, and 18% above in int8. Reduced by tools/claims-v3/mmb/reduce.py (items MMB-a to MMB-f).
Board power during the first run
What does "per W above idle" subtract, and why did idle rise from one workload to the next in the first run (18 September, aifoundry2)? The three-card check changed the method for that reason: it heated the die first on the two cards whose governor is free and ran each workload for 6 s, and the idle before each workload held within 0.3 W on each card. Board power was polled from the service processor about 8 times a second. The reading itself changes less often: with a poll like this one about every 126–139 ms on aifoundry2 and aifoundry1 card 1 and 223–224 ms on aifoundry3 (three passes per card, 26 September; faster pollers stretch it: Power and temperature). Shaded spans are the benchmark launches, about 12 s per workload, with idle gaps between them; the thin spikes just before each span are short calibration launches. Power creeps up within each span as the die warms: minion-shire temperature (the sensors' average) went from 71 °C to 77 °C, and the hottest sensor's high-water mark read 85 °C afterwards. For the same reason the idle in each gap sits higher than the one before (the grey steps).
How to read this chart
Pick a workload to shade its two windows: the idle it is measured against (3.5 to 1.8 s before its first timed launch, orange) and the span its power is averaged over (from 1 s after its first launch to its end, blue). The dashed lines are those two means.
How the benchmark works
- Code.
kernels/mmbench(device),launchers/mmbench(host, checks every result),scripts/et-power-log.sh(board-power sampler) andscripts/mmbench-power.py(runs everything and averages power over the launch windows). It passed exact checks on the simulator on one shire in all modes, then on the card on all 32 shires, before the timed runs. The card checks passed again on all three cards before every pass of the three-card check. - The loop. On each of the 1,024 compute minions, hart 0 (the first of the minion's two hardware threads) runs the loop. Only hart 0 may issue TensorFMA and TensorLoads into the L1 scratchpad, so hart 1 idles. Each op issues a
tensor_loadof a 16×K A tile into one of two L1-scratchpad buffers, atensor_loadof the B tile through TenB (the tensor unit's path that streams the B operand from memory), atensor_wait, and atensor_fmathat accumulates a 16×16 C tile in the vector registers. The next op's loads overlap the running FMA. This is the inner loop of gp-sdk's auto-generated matmul and of the pipelined kernel in the FOSDEM 2026 "Zero to matmul" talk (10.25 TFLOP/s at 650 MHz). - Op size. One op is 16×16×K multiply-adds, with K = 16 (fp32), 32 (fp16) or 64 (int8). Operand layouts come from
sw-sysemu's tensor emulation, which serves as the spec. - Working set. In the compute-bound runs all minions share 16 A/B tile pairs (32 KB), so the data stays in each shire's L2 cache. In the DRAM run each minion has its own 64 tile pairs (128 MB in total), so every load goes to memory. Effective read bandwidth there was about 78 GB/s (2 KB per op), in line with the 76 GB/s the memory-hierarchy probe streams.
- Inputs and checking. Inputs are random values from {−2, −1, 1, 2} (low-power operands; see Later measurements). They are never zero, so no multiply can be skipped, and they are exact in fp16 and fp32. The host checks each minion's C tile for exact equality with iters × Σ A·B, capping iteration counts so fp32 accumulation stays exact. Every launch passed with zero bad minions, here and in all 92–96 timed launches per card of the three-card check.
- Power. Board power comes from the card's own telemetry (
DM_CMD_GET_MODULE_POWERthroughdev_mngt_service, polled about 8 times a second). It is averaged over the launch windows after skipping the first second. The idle is the mean just before each workload (from 3.5 to 1.8 s before its first timed launch), which the above-idle column subtracts; the first run's idle row was the median of the last 7 s of the 8 s idle before the first workload.
Version history
Versions: published 18 September 2026; revised 24 September 2026 (per-workload idle, operand and A100 context, corrections) and 25 September 2026 (efficiency explorer, idle windows, corrections); 25 September (version 3): which card and how many runs each figure rests on, aifoundry3's values beside aifoundry2's in the later measurements, and the idle's leakage split as a range (record: docs/reports/data/2026-09-25-claims-v3); 26 September: the three-card check (aifoundry2, aifoundry3, aifoundry1 card 1) added to the lede, the KPIs, Results (its table), the power-meter refresh per card and the within-run rise per card; 27 September: the operand-and-card band (each card on board power), the A100 legend drawn as diamonds, and repeated text cut; later that day, each card's board power charted as its idle and what the workload added (the three-card check); 28 September: the meter's refresh note shortened; the note's counts given by both rules; then every energy on the page (the lede, the KPIs, the explorer, the A100 table and Results: board power, per W, the A100 ratios) the three-card check's, the first run one note under Results, and the power chart marked as the first run's.
Caveats
Caveats in full
- The A100 figures are spec-sheet numbers, not measurements. The one measured A100 matmul (Horace He: bf16 on random data, 257 TFLOPS under a 330 W limit) matches peak ÷ TDP, so the tensor-core ratios are a fair estimate; no A100 fp32 CUDA-core GEMM was measured.
- Board power includes the whole card. Part of each card's idle is die leakage, which grows with temperature: aifoundry2's idle rises about 0.65 W per °C at 80 °C. The idle law in The DVFS loop and its leakage, fitted on aifoundry2, predicts 30.8 W at the 71 °C the first run started at and 34.0 W at 77 °C. The above-idle column still carries the leakage added as the die warms within each run, so it is not the compute alone.
- One clock point. Every run here was at 600 MHz (0.52 V; the first run because its die started at 71 °C): aifoundry2's governor steps the clock down whenever the die reads above 65 °C or board power exceeds 65 W, and raises it to 700 or 800 MHz (0.57 or 0.62 V) only below both (The DVFS loop and its leakage). Up-steps were seen at readings up to 66 °C, never at 67 °C or above. Other operating points, which the governor can choose on aifoundry2 and aifoundry1 card 1 (on aifoundry3 a boot service sets the TDP to 0 W at every boot, which holds it at 600 MHz), change both speed and efficiency. Power rises within a run as the chip warms: in the three-card check's 6 s runs with tiles in L2 it rose from the first to the last second by 1.5–1.9 W on aifoundry2, 1.0–1.2 W on aifoundry3 and 1.5–2.8 W on aifoundry1 card 1 (one of that card's four fp16 runs rose only 0.2 W); with every tile from DRAM it rose 0.2–0.3 W on aifoundry2 and aifoundry3 and fell 0.4 W on aifoundry1 card 1. In the first run it rose about 3 W (lowest to highest reading) within each 12 s run, so longer runs would read slightly worse.
- Precision differs. ET's fp32 tensor op is a full IEEE fp32 multiply-add, whose A100 equivalent is fp32 on CUDA cores; tf32 keeps only a 10-bit mantissa. ET's fp16 op accumulates in fp32 with round-toward-zero.
- Peak-rate microbenchmark. This measures the matmul engine on cached tiles. A full GEMM also pays for moving data, and the DRAM row (0.31 TFLOP/s) is the no-reuse floor.
Next steps
- A real GEMM (for example 4096³) with L2-scratchpad tiling and cooperative tensor loads, measured the same way.
- Put hart 1 to work prefetching into the L2 scratchpad (
TensorLoadL2Scpis allowed on hart 1). - A vector-unit fp32 baseline. Its energy per lane is now in the energy manual §3.1 (7.0 pJ per multiply-add on random data); its speed is FOSDEM's 2.94 TFLOP/s at 650 MHz.
- For AI Foundry: the lab's
/opt/etpredates the gp-sdk that current documentation points to. Upgrading it, or shipping a matching gp-sdk, would spare new users the four gp-sdk fixes under "Running it on a lab machine" below.
Reproduce
Reproduce this
# regenerate this page's numbers and charts from the committed data (no card)
python3 scripts/mmbench-report-data.py docs/reports/data/2026-09-18-aifoundry2 \
--manual docs/reports/data/2026-09-23-energy-manual/manual.json \
--embed docs/reports/2026-09-18-et-soc1-matmul-efficiency.html
# exact checks of every mode on the simulator, one shire, no card (about 40 s each)
scripts/vm make mmbench-check # laptop VM; on a lab machine: cd ~/nekko && make mmbench-check
# on the card: copy the repository to the lab machine, build gp-sdk (pinned, patched) and mmbench there
scripts/deploy-lab-gpsdk.sh aifoundry2
# then on aifoundry2 (quit et-powertop first: the power logger needs the management node)
cd ~/nekko && nice make mmbench-check bench-power DEVICE=silicon JOBS=4
timeout 10 /opt/et/bin/it_test_code_loading --mode=pcie # the card hello world
# simulator hello worlds: see Test drive
# copy ~/nekko/build/mmbench-power/ into a new dated docs/reports/data/ directory
The lab cards are shared, and the lab rule is to hold a device for at most 10 s. The first run kept the card busy for about 12 s per workload, before the project adopted that rule. scripts/mmbench-power.py now runs every launcher under timeout 10 and defaults to about 6 s per workload, which still leaves about 40 samples per workload after the settle time (about 35 distinct readings on aifoundry2, about 20 on aifoundry3). It also stores the idle windows it averages.
Raw data: docs/reports/data/2026-09-18-aifoundry2 (board-power samples, one record per launch and the per-workload summary). scripts/mmbench-report-data.py turns it into this page's throughput, efficiency, idle before each workload (the same window as idle_before() in scripts/mmbench-power.py) and the power chart; with --manual it adds the energy manual's random-normal variant to the efficiency explorer, and by default (--v3) it reads the three-card check's docs/reports/data/2026-09-25-claims-v3/results/mmb.json, which the explorer, the three-card chart and every energy on the page come from, and prints the values the tables quote. The tables are written by hand from its output.
Sources: aifoundry-org/et-platform, et-man (programmer's reference and errata), marty1885/et-testdrive, FOSDEM 2026, "Zero to matmul with the ET-SoC-1", the NVIDIA A100 Tensor Core GPU datasheet, and Horace He's "Strangely, Matrix Multiplications on GPUs Run Faster When Given 'Predictable' Data!" for the measured A100 run.
Running it on a lab machine
These notes are for running the benchmark on a lab machine, whose /opt/et is older than the gp-sdk that upstream documents. The laptop setup and the first scalar SGEMM, from the same day, are in Test drive.
- Toolchain was already there.
/opt/etis et-platform built from commit353f20e(30 Dec 2025): RISC-V GCC 15.1,sys_emu, runtime, firmware and the upstream test binaries. It is not onPATHby default. aifoundry2 is x86-64 Ubuntu 24.04 with one card (PCIe Gen4, 32 GB, firmware BL 0.20 / minion 0.23); in this session its minion shires ran at 600 MHz and its NoC at 400 MHz. - Hello world: et-platform's
it_test_code_loading --mode=pciepasses 3/3 on the card in 0.6 s (the simulator runs are in Test drive). It also passed on all three cards at the start of every pass of the three-card check. - et-testdrive (marty1885's minimal host+kernel project) builds unmodified against
/opt/et. On the card all 64 harts of shire 0 print "Hello World from hart N", decoded in-process from the trace buffer. One catch: it opens the management node, which only one process may hold, so it fails with "Device or resource busy" whileet-powertopis running. The upstream tests open only the ops node. - Writing my own kernels with gp-sdk needed four fixes against this older install. They are in
patches/lab-gp-sdk-06605ab.patch, whichscripts/deploy-lab-gpsdk.shapplies;patches/README.mdsays what each one fixes.
Related reports
- The Horace experiment — the same TensorFMA's power on 14 operand patterns at 80 °C on aifoundry2 (38 W on zeros, 64 W on random values in the three-card re-run), and on two more cards.
- Why is the ET-SoC-1 low power? — energy per FLOP against a measured A100 on its bf16 tensor cores, and where this card's watts go.
- The energy manual, §3.2 — the tensor unit's energy per multiply-add in fp32, fp16 and int8, on zeros, ones and random data, with ranges over repeated runs.
- The DVFS loop and its leakage — why idle power depends on temperature, and when the governor leaves 600 MHz.
- Ridge points — the reuse each memory level demands to keep this loop fed.
- Sparse compute — what the tensor unit does with zeros: it saves power but never a cycle.
- Test drive — the first scalar SGEMM, the laptop setup and the upstream fixes, the same day.