Power and temperature on the ET-SoC-1: a load step through every sensor, and the per-shire voltage map
The card has one power sensor path: a PMIC that reports 12 V board power and three of its rails (minion
cores, SRAM, mesh), polled by the service processor once per pass: while ettelem samples at 10 Hz, a new
value about every 156 ms on aifoundry2, 158 ms on aifoundry1-c1 and 263 ms on aifoundry3. Everything else is voltage or temperature. This report
lists every way to read them, what each resolves, and what a small client on the management library adds over the stock
tool: the per-rail snapshot the CLI refuses, every refresh instead of et-powertop's one poll a second, and a
34-shire on-die voltage map. Then what the numbers show. The main surprise is temperature: under load, the same work costs about 1 W more for every °C the die warms
on aifoundry2 and aifoundry1-c1 (0.9–1.2 W/°C at 73–97 °C, four load steps on each) and 0.6 W on the cooler
aifoundry3. The die idled at 72 °C here (62–74 °C on other days; aifoundry3 50–58 °C on 22–24 September and 55–57 °C since
25 September), and idle power after a
load differs from idle before it by what the die's temperature did: 5 W higher here, with the die 9 °C warmer. So every
energy comparison on this card has to control for heat.
Terms used on this page
The ET-SoC-1's cores are minions, small RISC-V cores, 32 to a shire; of the 34 minion shires, 32 run kernels (1,024 minions), shire 32 is the master, which runs firmware, and shire 33 is a spare. The service processor (SP) is the on-die management core that reads the sensors, runs the clock governor and answers the host; its SPST trace is its per-pass stats log, and the MMST trace is the master shire's log of bandwidth. The PMIC is the board's power controller, which meters the 12 V input and three regulators (the minion, SRAM and mesh rails), and Moortec PVT monitors measure voltage and temperature on the die. More in the hub's glossary.
1. Every way to measure, by granularity
Space, time and resolution for each instrument. works now means from a user account with the tools in this repo; firmware means a service-processor change and a signed image; hardware means physical access; lab admin only means a change to the card that the lab's rules leave to its admin.
Asking has a price. Every board-power value is refreshed once per service-processor pass, and
the pass stretches with what the host asks of it: a single power command costs it almost nothing, but ettelem's
full snapshot lengthens it, and the faster ettelem samples, the longer each pass takes, so sampling harder
returns fewer fresh readings, not more.
More instruments, and the improvement ladder
Beyond these meters. The Minion Debug Interface adds no power or temperature sensor: it reads the registers and memory of a halted hart (Limits of observability, §3). Its one use for energy work, writing the counter event selectors before a launch, is rung 17 of the hub's improvement ladder, which also ranks what firmware changes or physical access would add and marks what has since been done. A direct temperature reading needs the heatsink left on: Feasibility of running the ET-SoC-1 without its heatsink finds that in its model a 600 MHz card almost never holds a stable idle temperature (1 of 3,600 draws) without it, so a thermocouple taped to the heatsink base, or, with the lab's consent, one soldered into a groove in the lid, is the option, not a sensor pointed at the bare chip.
2. A load step, seen through every sensor
Twenty seconds idle, 58 s of fp32 matmul on all 1,024 minions, 40 s idle, 34 s of DRAM-bound loads, idle. Sampled at
10 Hz with ettelem; every load process held the card for under 10 s. The three-card check ran this load step four
more times on each of aifoundry2, aifoundry3 and aifoundry1-c1 on 26 September, with the same phases, loads and flags (a card
whose governor can move was first heated to 76 °C). The charts show the 20 September session; the text gives the repeats'
results beside it.
Time in seconds from the start of the recording, in 1 s means of the 10 Hz samples. The short ticks under the power chart mark the gaps between successive load processes, found in the 10 Hz stream. Hover or tap a chart, or move the time slider, for the readings at that moment; the slider also marks that second in the chart below.
How the load step proves the slope: power against temperature, and the per-card comparison
Does idle power depend on temperature alone, and how much of the rise under load is the idle law's leakage? Each dot is one second.
- Power climbs while the work stays constant. In this session the matmul, back-to-back runs of the
launcher (the short dips in the chart are the gaps between them), draws 56 W two seconds after it starts and 65 W at its
end, 56 s later, while the die warms from about 75 to 87 °C. It is the same kernel on the same 600 MHz clock at a steady
rate: every run did the same work, (
raw/thermal-loads.log), and the runs took from dip to dip, the late ones no longer than the early ones. The idle law accounts for W of the extra W as leakage. A straight line through board power against die temperature, from to the end, has a slope of W/°C in this session, W/°C of it on the minion rail. The idle law's slope over the same temperatures is W/°C. The strict 7 s runs of the Horace experiment found the same drift on this card: . The three-card check's repeats give steeper slopes on aifoundry2. In its four load steps per card the clock read 600 MHz in every sample and the power climbed with the die on every card. Fitted the same way (from 5 s after launch), aifoundry2 gave 0.90 and 0.93 W/°C at 81–89 °C and 1.11 and 1.12 W/°C at 86–97 °C (those two repeats started hotter than this session), aifoundry3 0.56–0.66 W/°C at 61–70 °C and aifoundry1-c1 1.10–1.16 W/°C at 73–83 °C. About half of it is on the minion rail on every card (0.44–0.53 of the board slope). So this session's 0.8 W/°C is the low end on aifoundry2, not a constant: the slope grows with the die's temperature, and so does the idle law's own slope over the same readings (0.74–0.91 W/°C in aifoundry2's repeats). Whether busy power climbs faster per degree than an idle card is not decided by four repeats.Busy drift per card: board power's rise per °C of die under the same load, in the strict runs and in the load step's repeatsThe upper rows are each card's strict 7 s Horace runs (21–22 September), the lower rows the three-card check's four repeats of this page's load step (26 September; small rings, one per repeat). The mark is the mean slope over a card's runs, the whisker ±2 standard errors of that mean. The two vertical lines are this session's matmul fit, which follows the "Fit starts" slider above, and the idle law's slope at 80 °C.
- Idle is not one number. In this session, before the load: 31.2 W at 72 °C. Forty seconds after it: 36.5 W at 81 °C. Both sit on the idle law measured later (31.3 and 36.5 W at those temperatures). The three-card check's repeats agree on aifoundry2: idle rose by 3.0–7.2 W across the load step (from 33.5–37.2 W at 76–82 °C to 36.5–44.4 W at 81–91 °C), and the law predicted each rise to within 0.54 W (on average 0.21 W low). On aifoundry3 it rose 1.5–1.7 W, the die 4–5 °C warmer. On aifoundry1-c1 it fell 2.0–2.8 W: that card was heated to 76 °C before the load step and cools fast, so 40 s after the matmul its die read 2–4 °C cooler than before it. The die sheds heat over tens of seconds (6 °C in the 40 s after this session's load; 6.0–7.5 °C on aifoundry2 and aifoundry3 in the repeats, 15–16 °C on aifoundry1-c1), so a baseline taken at another temperature misstates the idle power during and after a run: taken before it, it understates it. Energy measurements here should take their baseline at the same temperature, or fit a temperature term.
What's in the unmetered remainder, and how it moves under DRAM and matmul loads
- The three rails are not the whole card. Board minus rails is W at idle at °C ( W in the idle window at °C; two windows are too few to show a trend with temperature), W under matmul and W under DRAM load in this session (under the matmul in the three-card check's repeats: 20–25 W on aifoundry2, 17–20 W on aifoundry3 and 22–27 W on aifoundry1-c1). That remainder is the DDR and PCIe rails, the Maxions, the IO shire, and regulator losses, none of which has a sensor. In this session a DRAM-bound load moves it by W and the minion rail by only W (with the die °C warmer); a matmul moves the minion rail by W. In the repeats the DRAM load moved the remainder on every card, by 4.6 W on aifoundry3, 4.9–7.6 W on aifoundry2 and 5.6–7.8 W on aifoundry1-c1, and the minion rail by what the die's temperature did: 0.03–0.08 W on aifoundry2 and aifoundry3 where the die stayed put, 0.7–0.8 W on aifoundry2 where it warmed 1.4–1.8 °C, and −0.4 to −0.8 W on aifoundry1-c1, whose die cooled 0.9–1.8 °C. The matmul moved the minion rail by 16–22 W on aifoundry2, 15–18 W on aifoundry3 and 15–21 W on aifoundry1-c1. Every card shows the same for DRAM traffic: in the energy catalogue (three passes on each of the three cards), its 76 GB/s DRAM streams (tensor loads and stores, 1 KB row walks) move the remainder by 4.4–7.0 W on aifoundry2 and aifoundry3 and 5.9–8.1 W on aifoundry1-c1, and the minion rail by at most a quarter of a watt either way (Limits of observability, §4.2).
Why the remainder briefly looks wrong at each load edge, with the edge charts
- Why the remainder sometimes looks too small, or the rails seem to exceed the step. Board power is a fresh
reading at every service-processor pass. The rail figures are the PMIC's own running averages,
which the service processor copies each pass, roughly first-order with τ ≈ : a step reaches
(medians over ;
Limits of observability, §4.1). So board minus rails swings at every load edge until the rails catch up.
The filter, measured rail by rail (29 September). Experiment E58 (the hub's
rung 4) timed square bursts of 2
to 4 s: development on aifoundry1-c1 (00:01–00:13 PDT), then, with its predictions frozen
(
tools/claims-v3/tau/PREREG.md, sha2567d4fabad…), validation on aifoundry3 (00:34–00:45 PDT, three passes). Fitted as first-order averages, with the service processor's one-pass delay kept apart, the readings give τ = 1.06 s on aifoundry3's minion rail, 1.01 s on its SRAM rail, 1.04 s on its mesh rail and 1.05 s for the PMIC's board average; on aifoundry1-c1, 1.08 s on the minion and mesh rails, 1.10 s for the board average and 0.54 s on the SRAM rail, whose meter differs from the other cards'. The τ above is a single-parameter fit over the catalogue's bursts, which also absorbs that one-pass delay (0.16 s on aifoundry1-c1 and 0.26 s on aifoundry3 whileettelemsamples at 10 Hz), so it reads longer. Undoing the filter with the measured values (tools/ettelem/deconv.py) recovers each burst's level to within 3% of the step and its energy to within 1%, where the published readings understate a 2 s burst by 19–22% on the three rails (all but aifoundry1-c1's faster SRAM rail) and 16–17% for the board average. On aifoundry3 the first-order form, the per-rail values and the one-pass delay survived as predicted (T1–T3). The recovered top of the 2 s minion bursts scattered by 8.4% of the step against the predicted 6% at most, so the claim that the inverse comes out clean (T4) was falsified. On aifoundry1-c1 in development the minion rail rose with τ 1.32 s and fell with 1.02 s, so the first-order form (T1) failed there. Data:docs/reports/data/2026-09-29-tau-aifoundry3/(its README) and2026-09-29-tau-aifoundry1-c1/.The unmetered remainder at the load edgesThe three-card check's repeats show the same edges on every card: at τ = 0 the remainder overshoots its steady matmul level by 15–17 W as the matmul starts on aifoundry2 (16–18 W on aifoundry3, 15 W on aifoundry1-c1) and falls 17–19 W below idle as it stops (20 W on aifoundry3; 17–18 W on aifoundry1-c1, 29 W once). Filtered with the card's rail τ it steps cleanly on aifoundry2, and on aifoundry1-c1 (with aifoundry2's τ) in three of four repeats, but not on aifoundry3: filtered with its own τ, its remainder still falls 1.9–2.7 W below its idle level at the edges. At a steady idle the board reading and the PMIC's own board average agree within 0.03 W on aifoundry2 and aifoundry3; over every steady sample of the three-card check's load steps, idle or busy, the median gap was 0.04 W on aifoundry2, 0.01 W on aifoundry3 and 0.18 W on aifoundry1-c1. After a load step the average trails the reading just as the rails do. The same averaging is behind the gap between the service-processor trace and the host log in the memory-anatomy report, where the host log's totals came out 14–25% higher in each of its twelve runs on aifoundry2.The matmul starts (15–35 s)The matmul stops (73–95 s)
- Voltage barely droops at the die. The minion rail's on-die monitor reads 518 mV idle and 517 mV with 18 W more load (about 35 A more at 0.52 V), consistent with a regulator that senses at the die. The three-card check's load steps put the sag at 0.08 mV per watt of minion-rail rise on aifoundry2 and on aifoundry3 and 0.055 on aifoundry1-c1 (four repeats each). In the three-card check's repeats the idle reading fell 0.11–0.21 mV per °C on aifoundry2, 0.21–0.24 on aifoundry3 and 0.10–0.15 on aifoundry1-c1; under the DRAM load it sat 2.9–5.7 mV below that trend on aifoundry2, 2.9 mV on aifoundry3 and 4.1–7.0 mV on aifoundry1-c1; and under the matmul within 1.5 mV of it on aifoundry2 and aifoundry3 but 2.4–3.2 mV below it on aifoundry1-c1. That droop under DRAM traffic, which the energy catalogue finds on aifoundry2 and aifoundry3 (about 4 mV at 76 GB/s of DRAM streaming, none under compute), is now used as a meter for DRAM power (Limits of observability, §4.3).
- The clock stayed at 600 MHz because the die was hot. On aifoundry2 the governor steps down whenever the whole-degree die reading is above 65 °C or board power is above 65 W. Every run in this session had the die above 70 °C, so it held the lowest operating point (600 MHz at 0.52 V) even at 65 W. From a cool die it lifts a matmul to 800 MHz within about a second (seven cool-start runs on one morning, a demonstration: the Horace experiment, §6; the DVFS report reads the governor and classifies every clock change). aifoundry3 never leaves 600 MHz: a boot service sets its power limit to 0 W at every boot, so its governor can only ask to step down, which at the bottom point changes nothing. aifoundry1-c1 read 600 MHz throughout the three-card check (318,667 samples), although it ran busy at 45–65 W and at a die of 64 °C or less in 9,461 of them: its governor does not raise the clock.
- Nothing on the card limits the die's temperature. In the three-card check aifoundry2's mean reading passed
90 °C in eight of its telemetry files (12,176 samples), every one at 600 MHz. The hottest was the energy catalogue's hot
pass 11 on 26 September (02:26–02:31 PDT), which heats the die to at least 88 °C before each configuration and has no
upper stop: its mean read 90–103 °C (above 90 °C in 2,565 of 2,567 samples), the hottest sensor up to 106 °C and the
board up to 86.9 W (83.6 W at the first 103 °C reading), and the clock never moved and no safe state came. That is above
the 90 °C at which this work's long runs stop. The firmware's one hardware trip, a PMIC alarm at 75 °C or 75 W meant to
force a 300 MHz safe state, takes its temperature from the PMIC, not from the die's sensors. The service processor reads
that temperature only when it starts or resets its statistics, and the firmware notes that the PMIC reports it as 0
(the statistic read 0 on all three cards throughout the check; between resets the firmware feeds it a hard-coded 0, so
the samples show nothing more). The evidence that nothing acts is the telemetry above, 103 °C at 600 MHz with no safe
state. The host's
pmictemperature field is the minion-shire mean again (the spatial brief). At the bottom operating point only the runners' own caps stand between a workload and the die's temperature.
3. The per-shire voltage map
With its log level raised to DEBUG, the service processor writes one line per shire per pass with the on-die voltage of
each rail and the hardware's low/high capture. ettelem loglevel debug then ettelem sptrace reads
them; the level is restored afterwards. The 4 KB dump holds all 34 shires only when it catches a whole pass: 22 of the
three-card check's 27 idle, load and after dumps did, on all three cards. This is the finest spatial view of the power system the card gives: below is one
pass over all 34 minion shires, for each rail its current reading and the lowest and highest the monitor captured: at idle
on aifoundry2 on 20 September, or, from the three-card check, at idle and under a 7 s random-data load on each card.
The 32 compute shires sit where marty1885's logical shire map
(etTopoScan) puts them on the mesh, the layout
the on-chip communication report found
on aifoundry2 by latency. The four
dashed cells hold the master, spare, I/O and PCIe shires in an order not measured here (labelled as the spatial brief
infers them), so shires 32 (the master) and 33 (the spare) are drawn below the grid.
In this capture (aifoundry2, 20 September) current values span 517–521 mV across the die; the lowest captured minimum is 513 mV (shires 0–2 and 18). The three-card check's idle dumps read 517–520 mV on aifoundry2 (lowest minimum 514), 522–525 mV on aifoundry3 and 496–503 mV on aifoundry1-c1 (lowest minimum 492). The current readings differ by at most 4 mV, with no gradient across the grid (a plane fitted to them explains of the spread, no more than chance would). That holds in the three-card check's idle dumps on aifoundry2 (R² = 0.01) and aifoundry1-c1 (R² = 0.20, permutation p = 0.04: a hint only, above the bar for twelve views), but not on aifoundry3, whose minion rail leans across the grid (R² = 0.38, p = 0.001). aifoundry2's SRAM rail leans again in them (R² = 0.35 and 0.34, p = 0.002 and 0.003); aifoundry3's and aifoundry1-c1's do not (p = 0.37–0.51 and 0.94). Idle currents of about W per shire look too small to cause the spread, so each monitor might carry its own offset, a baseline to subtract from a map taken under load (the idea behind the hub's rung 5). The three-card check tested that, and it does not hold: under a 7 s load each shire's deviation from the common level changed with a standard deviation of 1.6–1.9 mV on aifoundry2 and aifoundry3 and 2.7 mV on aifoundry1-c1 (up to 4.5–7.1 mV in one shire), where a fixed offset would allow about 0.5 mV. So an idle map is not a baseline to subtract. The readings repeat: The three-card check's idle dumps, taken in passes hours apart, agree almost value for value, so exactly (their minion rail's readings correlate at 1.0 from pass to pass) that they may hold one repeated reading rather than new ones; they are not evidence that the map repeats. The lows are not idle values: they hold the lowest reading since the last stats reset, and when that was is not recorded for this capture. Whether they catch droops that the polling misses is not established: the three-card check's test needed a whole 34-shire dump at the end of a load window, which aifoundry2 and aifoundry3 never gave (on aifoundry1-c1, in the two passes that did, 18 of 34 shires' lows sat at least 2 mV under idle while the 10 Hz minion reading fell 1 mV). Per-shire temperature is sampled by the same controllers but never logged; that needs firmware (the spatial temperature brief sets out the options).
The offset test, drawn. If each monitor carried a fixed offset, a shire that reads high at idle would read high by the same amount under load, and every shire would sit on the diagonal below; shires do not.
4. Reproduce
Reproduce this
cmake -S tools/ettelem -B build/ettelem -DCMAKE_PREFIX_PATH=/opt/et -Wno-dev \
&& cmake --build build/ettelem
export LD_LIBRARY_PATH=/opt/et/lib
# JSON lines at 10 Hz; quit et-powertop first (the management node allows one opener)
build/ettelem/ettelem sample --seconds 10 --every-ms 100
# the loads run_thermal.sh launches
make all # build/launchers/mmbench_launcher, build/kernels/nekko/mmbench.elf
cmake -S workloads/memprobe -B build/memprobe -DCMAKE_PREFIX_PATH=/opt/et -Wno-dev \
&& cmake --build build/memprobe
mkdir -p build/memprobe-data
python3 workloads/memprobe/gen_ops.py table --pattern dram_seq \
--out-file build/memprobe-data/dram_seq.tbl
# the load step, about 3 minutes;
# writes thermal-telemetry.jsonl, thermal-phases.jsonl, thermal-loads.log
tools/ettelem/run_thermal.sh build/thermal1
# the voltage map: raise the SP's log level, dump its trace, restore
build/ettelem/ettelem loglevel debug
build/ettelem/ettelem sptrace sp.bin
build/ettelem/ettelem loglevel info # restore
python3 tools/ettelem/parse_sptrace_voltage.py sp.bin > per-shire-voltage-idle.json
# from the repository root: reduce the session and build this page
python3 tools/ettelem/summarize_power_session.py docs/reports/data/2026-09-20-power-aifoundry2 \
--out docs/reports/data/2026-09-20-power-aifoundry2/summary.json
python3 scripts/build-report.py power-temperature \
docs/reports/data/2026-09-20-power-aifoundry2/summary.json \
docs/reports/2026-09-20-et-soc1-power-temperature.html
Raw data: docs/reports/data/2026-09-20-power-aifoundry2/. Its raw/ folder
holds the three service-processor trace dumps of 20 September (sp1.bin gave the map; the parser reproduces
the committed per-shire-voltage-idle.json from it byte for byte) and the load step's
thermal-loads.log, recovered from the lab machine's build tree on 25 September. The load step is session E5
of the experiment
register, and the voltage map session E6. The three-card check's load steps (V3-MMB, item MMB-T) and meter and
voltage-map tests (V3-TEL, items TEL-P1–P5, S, Q and R), raw and reduced, are in
docs/reports/data/2026-09-25-claims-v3/; summarize_power_session.py copies the repeats' busy
slopes (MMB-T/P1) into this page's data for the busy-drift chart. Every sensor and management command, with firmware line references:
docs/research/power-telemetry.md.
Firmware. The service-processor code this page describes (the pass loop, what each management
command returns, the governor) was read at et-platform 353f20e. The cards' own trace strings match an older
build, from before et-platform commit 60b40c10f (24 September 2024), which rewrote the governor, so the
governor details may differ on the card (the
DVFS report). The cards' own build (BL2 0.20.0 of release 1.3.1, et-platform ffca4cbb4) was read on 27 September: its thermal response is a blocking loop of steps about 0.4 s apart that also acts on an idle card, and a climb goes to the top point in one call (the DVFS report). The measurements do not depend on it.
Version history
Versions. 20 September 2026: first edition. 21 September: the governor does act (§2; it first said it never does). 24 September: the rails are the PMIC's running average, not a 2 s moving average. 25 September: rebuilt by committed tools from the recovered raw records; the refresh period is per card (156 and 263 ms, not 133 ms), and the clause that busy power climbs faster per degree than idle is withdrawn. 26 September: the three-card check (the busy slope is 0.9–1.1 W/°C on aifoundry2, this session's 0.8 its low end; the voltage map's offset idea and the lows' "droops the polling misses" withdrawn). 27 September: the review's fixes and charts. 28 September: aifoundry1-c1's governor never raises the clock, aifoundry3's can only ask to step down, and nothing on the card limits the die's temperature (§2); later, the review's cuts. 29 September: the rails' filter measured rail by rail and undone (E58; §1's table and §2). What each version changed in full: this page's history in the repository.
Related reports
- The Horace experiment — what the values in a matrix do to power and temperature, measured with the same client. It is named after Horace He's post showing that GPU matrix multiplies run faster on predictable data.
- The DVFS loop and its leakage — the governor that picks the clock, and the idle law that refines this page's leakage slope.
- The energy manual — energy per instruction and per byte, measured with these meters; its §1 is the card at rest.
- Limits of observability, §4 — the meter chain read from the firmware, and what the unmetered remainder is made of.
- Anatomy of a memory access — each level's energy split by these meters on three cards, and the first runs' host-log and trace readings that §2 compares.
- Spatial temperature: a brief — what the 35 on-die temperature sensors hold, and what firmware would have to export for a per-shire temperature map like §3's voltage map.
- Feasibility of running the ET-SoC-1 without its heatsink — whether the die's temperature can be read directly: not with the heatsink off; a thermocouple soldered into a groove in the lid reads the lid, about 1–2 °C below the die mean at idle and 2–5 °C under load (an estimate), and one taped to the heatsink base reads the heatsink.
- Power telemetry research notes — every sensor, register and management command behind this page, with firmware line references.