Power and temperature on the ET-SoC-1: a load step through every sensor, and the per-shire voltage map

20 September 2026, checked on three cards on 25–26 September · the load step and the voltage map are one session each on aifoundry2; the three-card check repeated both on aifoundry2, aifoundry3 and aifoundry1-c1 (card 1 of aifoundry1), and the text gives its results beside them · firmware source read at et-platform 353f20e (the cards' own log strings point to an older build; see Reproduce) · tools/ettelem, a telemetry client on the management library · part of the ET-SoC-1 measurement reports

The card has one power sensor path: a PMIC that reports 12 V board power and three of its rails (minion cores, SRAM, mesh), polled by the service processor once per pass: while ettelem samples at 10 Hz, a new value about every 156 ms on aifoundry2, 158 ms on aifoundry1-c1 and 263 ms on aifoundry3. Everything else is voltage or temperature. This report lists every way to read them, what each resolves, and what a small client on the management library adds over the stock tool: the per-rail snapshot the CLI refuses, every refresh instead of et-powertop's one poll a second, and a 34-shire on-die voltage map. Then what the numbers show. The main surprise is temperature: under load, the same work costs about 1 W more for every °C the die warms on aifoundry2 and aifoundry1-c1 (0.9–1.2 W/°C at 73–97 °C, four load steps on each) and 0.6 W on the cooler aifoundry3. The die idled at 72 °C here (62–74 °C on other days; aifoundry3 50–58 °C on 22–24 September and 55–57 °C since 25 September), and idle power after a load differs from idle before it by what the die's temperature did: 5 W higher here, with the die 9 °C warmer. So every energy comparison on this card has to control for heat.

Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry2, aifoundry3 and aifoundry1-c1 (four repeats of §2's load step and three of the meter and voltage-map tests on each card). Of 30 claims tested here, this page counts 6 held, 9 corrected, 13 differ by card and 2 not confirmed; the hub’s scoreboard, 1 “a test behind it failed”, 21 “differs by card” and 8 “fewer than three repeats”. The main correction is the busy slope: this session's 0.8 W/°C is the low end on aifoundry2, where the repeats gave 0.9–1.1; and the idle voltage map is not a per-monitor offset to subtract. The 20 September numbers stay, labelled as that session's (the record).
Later measurements. Idle power on aifoundry2 follows a fitted law, Pidle = W + W·e(T−80)/ (T the die temperature in °C), fitted to idle readings at °C and not re-measured by the three-card check (its slope and leakage share: the DVFS report, §5; the energy manual, §1). §2 compares this page's readings with it.
Idle
31.2 W
aifoundry2 at 72 °C die; 16 W on the three measured rails, 15 W elsewhere (12–13 W on aifoundry3, Limits §4.2)
Power slope under load
0.6–1.2 W/°C
board, the same matmul, four load steps per card: 0.90–1.12 on aifoundry2 (81–97 °C), 0.56–0.66 on aifoundry3 (61–70 °C), 1.10–1.16 on aifoundry1-c1 (73–83 °C); about half on the minion rail on all three. An idle card: 0.65 W/°C at 80 °C
Finest power record
10 mW · 156 ms
board power on aifoundry2, fresh each service-processor pass (every card: §1); the rails read to 1 mW but are the PMIC's running average (§2)
Finest voltage map
34 shires
3 rails each, 1 mV, with hardware low/high, from a user account; captured on all three cards
Terms used on this page

The ET-SoC-1's cores are minions, small RISC-V cores, 32 to a shire; of the 34 minion shires, 32 run kernels (1,024 minions), shire 32 is the master, which runs firmware, and shire 33 is a spare. The service processor (SP) is the on-die management core that reads the sensors, runs the clock governor and answers the host; its SPST trace is its per-pass stats log, and the MMST trace is the master shire's log of bandwidth. The PMIC is the board's power controller, which meters the 12 V input and three regulators (the minion, SRAM and mesh rails), and Moortec PVT monitors measure voltage and temperature on the die. More in the hub's glossary.

1. Every way to measure, by granularity

Space, time and resolution for each instrument. works now means from a user account with the tools in this repo; firmware means a service-processor change and a signed image; hardware means physical access; lab admin only means a change to the card that the lab's rules leave to its admin.

Asking has a price. Every board-power value is refreshed once per service-processor pass, and the pass stretches with what the host asks of it: a single power command costs it almost nothing, but ettelem's full snapshot lengthens it, and the faster ettelem samples, the longer each pass takes, so sampling harder returns fewer fresh readings, not more.

The service processor's pass under each way of polling the card, three passes per card (the three-card check, 26 September)

More instruments, and the improvement ladder

Beyond these meters. The Minion Debug Interface adds no power or temperature sensor: it reads the registers and memory of a halted hart (Limits of observability, §3). Its one use for energy work, writing the counter event selectors before a launch, is rung 17 of the hub's improvement ladder, which also ranks what firmware changes or physical access would add and marks what has since been done. A direct temperature reading needs the heatsink left on: Feasibility of running the ET-SoC-1 without its heatsink finds that in its model a 600 MHz card almost never holds a stable idle temperature (1 of 3,600 draws) without it, so a thermocouple taped to the heatsink base, or, with the lab's consent, one soldered into a groove in the lid, is the option, not a sensor pointed at the bare chip.

2. A load step, seen through every sensor

Twenty seconds idle, 58 s of fp32 matmul on all 1,024 minions, 40 s idle, 34 s of DRAM-bound loads, idle. Sampled at 10 Hz with ettelem; every load process held the card for under 10 s. The three-card check ran this load step four more times on each of aifoundry2, aifoundry3 and aifoundry1-c1 on 26 September, with the same phases, loads and flags (a card whose governor can move was first heated to 76 °C). The charts show the 20 September session; the text gives the repeats' results beside it.

Power (W): board total, the three measured rails, and what no rail meters
Die temperature (°C, mean of the 34 minion-shire sensors)

Time in seconds from the start of the recording, in 1 s means of the 10 Hz samples. The short ticks under the power chart mark the gaps between successive load processes, found in the 10 Hz stream. Hover or tap a chart, or move the time slider, for the readings at that moment; the slider also marks that second in the chart below.

How the load step proves the slope: power against temperature, and the per-card comparison
Board power against die temperature

Does idle power depend on temperature alone, and how much of the rise under load is the idle law's leakage? Each dot is one second.

What's in the unmetered remainder, and how it moves under DRAM and matmul loads
Why the remainder briefly looks wrong at each load edge, with the edge charts

3. The per-shire voltage map

With its log level raised to DEBUG, the service processor writes one line per shire per pass with the on-die voltage of each rail and the hardware's low/high capture. ettelem loglevel debug then ettelem sptrace reads them; the level is restored afterwards. The 4 KB dump holds all 34 shires only when it catches a whole pass: 22 of the three-card check's 27 idle, load and after dumps did, on all three cards. This is the finest spatial view of the power system the card gives: below is one pass over all 34 minion shires, for each rail its current reading and the lowest and highest the monitor captured: at idle on aifoundry2 on 20 September, or, from the three-card check, at idle and under a 7 s random-data load on each card. The 32 compute shires sit where marty1885's logical shire map (etTopoScan) puts them on the mesh, the layout the on-chip communication report found on aifoundry2 by latency. The four dashed cells hold the master, spare, I/O and PCIe shires in an order not measured here (labelled as the spatial brief infers them), so shires 32 (the master) and 33 (the spare) are drawn below the grid.

In this capture (aifoundry2, 20 September) current values span 517–521 mV across the die; the lowest captured minimum is 513 mV (shires 0–2 and 18). The three-card check's idle dumps read 517–520 mV on aifoundry2 (lowest minimum 514), 522–525 mV on aifoundry3 and 496–503 mV on aifoundry1-c1 (lowest minimum 492). The current readings differ by at most 4 mV, with no gradient across the grid (a plane fitted to them explains of the spread, no more than chance would). That holds in the three-card check's idle dumps on aifoundry2 (R² = 0.01) and aifoundry1-c1 (R² = 0.20, permutation p = 0.04: a hint only, above the bar for twelve views), but not on aifoundry3, whose minion rail leans across the grid (R² = 0.38, p = 0.001). aifoundry2's SRAM rail leans again in them (R² = 0.35 and 0.34, p = 0.002 and 0.003); aifoundry3's and aifoundry1-c1's do not (p = 0.37–0.51 and 0.94). Idle currents of about W per shire look too small to cause the spread, so each monitor might carry its own offset, a baseline to subtract from a map taken under load (the idea behind the hub's rung 5). The three-card check tested that, and it does not hold: under a 7 s load each shire's deviation from the common level changed with a standard deviation of 1.6–1.9 mV on aifoundry2 and aifoundry3 and 2.7 mV on aifoundry1-c1 (up to 4.5–7.1 mV in one shire), where a fixed offset would allow about 0.5 mV. So an idle map is not a baseline to subtract. The readings repeat: The three-card check's idle dumps, taken in passes hours apart, agree almost value for value, so exactly (their minion rail's readings correlate at 1.0 from pass to pass) that they may hold one repeated reading rather than new ones; they are not evidence that the map repeats. The lows are not idle values: they hold the lowest reading since the last stats reset, and when that was is not recorded for this capture. Whether they catch droops that the polling misses is not established: the three-card check's test needed a whole 34-shire dump at the end of a load window, which aifoundry2 and aifoundry3 never gave (on aifoundry1-c1, in the two passes that did, 18 of 34 shires' lows sat at least 2 mV under idle while the 10 Hz minion reading fell 1 mV). Per-shire temperature is sampled by the same controllers but never logged; that needs firmware (the spatial temperature brief sets out the options).

The offset test, drawn. If each monitor carried a fixed offset, a shire that reads high at idle would read high by the same amount under load, and every shire would sit on the diagonal below; shires do not.

Each shire's minion-rail reading, idle against a 7 s load, as a deviation from its card's 34-shire mean (the three-card check, 26 September)

4. Reproduce

Reproduce this
cmake -S tools/ettelem -B build/ettelem -DCMAKE_PREFIX_PATH=/opt/et -Wno-dev \
  && cmake --build build/ettelem
export LD_LIBRARY_PATH=/opt/et/lib
# JSON lines at 10 Hz; quit et-powertop first (the management node allows one opener)
build/ettelem/ettelem sample --seconds 10 --every-ms 100
# the loads run_thermal.sh launches
make all   # build/launchers/mmbench_launcher, build/kernels/nekko/mmbench.elf
cmake -S workloads/memprobe -B build/memprobe -DCMAKE_PREFIX_PATH=/opt/et -Wno-dev \
  && cmake --build build/memprobe
mkdir -p build/memprobe-data
python3 workloads/memprobe/gen_ops.py table --pattern dram_seq \
  --out-file build/memprobe-data/dram_seq.tbl
# the load step, about 3 minutes;
# writes thermal-telemetry.jsonl, thermal-phases.jsonl, thermal-loads.log
tools/ettelem/run_thermal.sh build/thermal1
# the voltage map: raise the SP's log level, dump its trace, restore
build/ettelem/ettelem loglevel debug
build/ettelem/ettelem sptrace sp.bin
build/ettelem/ettelem loglevel info   # restore
python3 tools/ettelem/parse_sptrace_voltage.py sp.bin > per-shire-voltage-idle.json
# from the repository root: reduce the session and build this page
python3 tools/ettelem/summarize_power_session.py docs/reports/data/2026-09-20-power-aifoundry2 \
  --out docs/reports/data/2026-09-20-power-aifoundry2/summary.json
python3 scripts/build-report.py power-temperature \
  docs/reports/data/2026-09-20-power-aifoundry2/summary.json \
  docs/reports/2026-09-20-et-soc1-power-temperature.html

Raw data: docs/reports/data/2026-09-20-power-aifoundry2/. Its raw/ folder holds the three service-processor trace dumps of 20 September (sp1.bin gave the map; the parser reproduces the committed per-shire-voltage-idle.json from it byte for byte) and the load step's thermal-loads.log, recovered from the lab machine's build tree on 25 September. The load step is session E5 of the experiment register, and the voltage map session E6. The three-card check's load steps (V3-MMB, item MMB-T) and meter and voltage-map tests (V3-TEL, items TEL-P1–P5, S, Q and R), raw and reduced, are in docs/reports/data/2026-09-25-claims-v3/; summarize_power_session.py copies the repeats' busy slopes (MMB-T/P1) into this page's data for the busy-drift chart. Every sensor and management command, with firmware line references: docs/research/power-telemetry.md.

Firmware. The service-processor code this page describes (the pass loop, what each management command returns, the governor) was read at et-platform 353f20e. The cards' own trace strings match an older build, from before et-platform commit 60b40c10f (24 September 2024), which rewrote the governor, so the governor details may differ on the card (the DVFS report). The cards' own build (BL2 0.20.0 of release 1.3.1, et-platform ffca4cbb4) was read on 27 September: its thermal response is a blocking loop of steps about 0.4 s apart that also acts on an idle card, and a climb goes to the top point in one call (the DVFS report). The measurements do not depend on it.

Version history

Versions. 20 September 2026: first edition. 21 September: the governor does act (§2; it first said it never does). 24 September: the rails are the PMIC's running average, not a 2 s moving average. 25 September: rebuilt by committed tools from the recovered raw records; the refresh period is per card (156 and 263 ms, not 133 ms), and the clause that busy power climbs faster per degree than idle is withdrawn. 26 September: the three-card check (the busy slope is 0.9–1.1 W/°C on aifoundry2, this session's 0.8 its low end; the voltage map's offset idea and the lows' "droops the polling misses" withdrawn). 27 September: the review's fixes and charts. 28 September: aifoundry1-c1's governor never raises the clock, aifoundry3's can only ask to step down, and nothing on the card limits the die's temperature (§2); later, the review's cuts. 29 September: the rails' filter measured rail by rail and undone (E58; §1's table and §2). What each version changed in full: this page's history in the repository.

Related reports