The record your GPU power-cap changes have been missing.

Spark-XC writes a signed, hash-chained Power Event Record for each GPU power-cap change it records: which GPU, when, the limit before, and the limit after as the GPU reports it. Records power-limit changes made by other tools, not just its own. Spark-XC keeps an attested, tamper-evident record of GPU power-cap changes, whoever makes them.

Explore the Architecture → Show me the record

What each record captures

Limit before and after
The power limit the GPU reported before the change and the limit it reports back after, read from the GPU through its management interface.
GPU telemetry
Power draw and temperature from the GPU at the time of the change, kept with the record.
Changes from other tools
Records power-limit changes made by other tools, not just its own. A change made outside Spark-XC still gets a Power Event Record.
Ed25519 signatures
Every record is signed with Ed25519. Auditors only need the public key.
Tamper-evident chain
Each record is SHA-256 hash-chained to the one before, so an edit to a stored record breaks the chain on recompute. Anyone with a copy can recompute the chain. Anchors go to a local write-once directory by default, and can be sent to off-site storage you control.
Fits your stack
Works alongside the GPU management, scheduling, and facility tools you already run. Nothing to rip out.

From a power-cap change to a record

Each recorded change becomes one Power Event Record. Here is what happens, step by step.

1
A power-cap change happens
A scheduler, an operator, a grid event, or Spark-XC itself changes a GPU's power limit.
2
Limit before is recorded
Spark-XC records the limit the GPU reported before the change.
3
Readback is recorded
Spark-XC reads the new limit back from the GPU, with its power draw and temperature. Changes made by other tools are picked up too.
4
The record is signed
Every record is signed with Ed25519. Auditors only need the public key.
5
The record is chained
The record is SHA-256 hash-chained to the one before, so an edit to a stored record breaks the chain on recompute.
6
Anyone with a copy can check it
Anyone with a copy can recompute the chain. Anchors go to a local write-once directory by default, and can be sent to off-site storage you control.
Record Flow
Power-cap change
CHANGE
↓
Your tools or Spark-XC set the limit
GPU
↓
SPARK-XC RECORD
GPU & time
DEVICE
Limit before / after
READBACK
Ed25519 signature
SIGN
SHA-256 chain link
CHAIN
↓
Power Event Record
SEALED

Spark-XC keeps the record. Your stack keeps running.

Spark-XC isn't in your workloads' data path. If it goes offline, your GPUs and schedulers carry on, and no Power Event Records are written for that window. If a record can't be written, Spark-XC flags it.

Spark-XC writes a Power Event Record for each power-cap change it records. It isn't in your workloads' data path.

Spark-XC offline: your tools and GPUs carry on, but no Power Event Records are written for that window.

Watch the records — read-only by design

The Spark-XC dashboard is a read-only view of your GPUs: power, temperature, and applied power limits alongside the Power Event Record stream. It pairs with the Prometheus and Grafana setup you already run.

Read-Only by Design
The dashboard shows telemetry and Power Event Records. It has no controls on it: looking at records and changing power stay separate.
Grafana & Prometheus
Per-GPU power draw, temperature, applied power limit, control actions, and safety counters are exposed on a Prometheus /metrics endpoint, with a sample Grafana dashboard.
Same records an auditor checks
The record stream on screen is the same signed, hash-chained set you can hand to an auditor. Every record is signed with Ed25519. Auditors only need the public key.

Without a record vs. with one

Without a record
  • Logs anyone with access can edit
  • Power-limit changes scattered across scheduler, driver, and facility tools
  • Changes made by other tools go unrecorded
  • No way to show a log hasn't changed since
SPARK-XC
  • One Power Event Record per recorded change
  • Limit before and after, as the GPU reports it, in one place
  • Records power-limit changes made by other tools, not just its own.
  • Every record is signed with Ed25519. Auditors only need the public key.
  • SHA-256 hash-chained; an edit to a stored record breaks the chain on recompute

What you need

Spark-XC runs on your own hardware, alongside the tools you already use. No kernel modifications. No driver replacements. No application changes. Each record can be checked with the public key on any machine, without access to your systems.

Prerequisites
✓ NVIDIA data-center GPUs, including H100
✓ Linux host (Ubuntu 20.04+, RHEL 8+)
✓ GPU driver 525+
✓ Root access on the GPU host
What You Get
→ A Power Event Record for each recorded change
→ Configuration in JSON/YAML
→ Records anchored to a local write-once directory by default
→ Zero application code changes required

See the record for one of your power events

We're looking for design partners among AI sites and GPU clouds with grid or curtailment obligations. Send us one power event or curtailment scenario and we'll show you the record Spark-XC would keep for it.

Show me the record → View Architecture