Power records for AI & ML workloads

A record for each power change during a run

Training runs last days. When a power change hits mid-run, you want to know what happened and when.

Spark-XC records each GPU power-cap change it sees, whether your scheduler, your GPU tools, or anything else made it, and seals it into a Power Event Record: the limit before, the limit the GPU read back, and the time. Spark-XC keeps an attested, tamper-evident record of GPU power-cap changes, whoever makes them.

GPU & time Limit before / after Ed25519 signature SHA-256 chain
POWER EVENT · EXAMPLE · ILLUSTRATIVE
09:14:01 [RUN] training epoch 142/500
09:14:02 [CHG] scheduler sets cap 320W → 220W on example-node-07:gpu:3
09:14:03 [READ] readback 220W · 87°C · 214W draw recorded
09:14:03 [SIGN] Ed25519 signature added
09:14:03 [HASH] 4a91...7f30 chained
09:14:03 [OK] Power Event Record sealed
09:14:08 [RUN] training epoch 143/500
1
Record per recorded change
Before/after
Limit, as the GPU reports it
Ed25519
Signed records
SHA-256
Hash-chained

Example scenarios (illustrative)

Example · Training
Thermal spike mid-epoch
GPU temperature climbs during a long run and a protective throttle lowers the power cap.
What the records would show: the limit before, the readback, and the temperature.
Example · Inference
A scheduler changes power caps during serving
An automated scheduler lowers power caps across an inference pool during a busy hour.
What the records would show: each change, the readback on each GPU, and when it happened.
Example · Multi-GPU
Driver crash during distributed training
The GPU driver crashes on one node. Spark-XC can't read that node's GPUs until the driver is back, so it writes no records for that window and flags the gap.
What the records would show: the last limit recorded and its readback, then what the GPU reported after the driver came back.
Example · Post-incident
Investigating an odd training run
A run produces unexpected results and the team wants to rule power in or out.
What the records would show: replay the chain to see what happened to GPU power caps during the run, and when.

What a record holds, and why AI teams care

Part 1
Limit before and read back
For each recorded change, Spark-XC notes the limit before and reads the limit back from the GPU, with its power draw and temperature.
Why it matters for AI
A mismatch between requested and applied power can cause instability that shows up epochs later. The record shows what the GPU reported at the time.
Part 2
Changes by other tools
Records power-limit changes made by other tools, not just its own. A change made by hand on a node shows up next to the scheduler's changes.
Why it matters for AI
When several tools can touch power limits, you need one place that shows all the changes it saw.
Part 3
Ed25519 signature
Every record is signed with Ed25519. Auditors only need the public key.
Why it matters for AI
Anyone reviewing a run can check who sealed the record without access to your systems.
Part 4
Tamper-evident chain
Each record joins an append-only, SHA-256 hash-chained ledger. Edit a stored record and the chain no longer checks out when it's recomputed. Anyone with a copy can recompute the chain. Anchors go to a local write-once directory by default, and can be sent to off-site storage you control.
Why it matters for AI
When a run produces unexpected results, a natural question is whether power behaved normally. Replaying the chain shows what happened to GPU power caps, and when.

See the record for a power change in your run

We're looking for design partners among AI infrastructure teams. Send us one power event or curtailment scenario and we'll show you the record Spark-XC would keep for it.

Show me the record View the Pipeline