Who could have changed this record?
When an auditor or counterparty asks, the answer starts with the signing key: who holds it, how it's rotated, and what the chain shows when the agent isn't running.
Your key, held by you
Records are signed with Ed25519. Signing is required by default. You generate the key pair with the keygen tool that comes with Spark-XC, on a machine you choose. Spark-XC holds no copy of your key, and there's no key escrow.
- One key per agent. The agent on each node signs with a single key. On Kubernetes, the nodes in one install share that key unless you give them separate ones.
- Loaded from your secret store. The agent reads the key from Vault, a mounted secret file, AWS Secrets Manager, or an environment variable.
A new key starts a new chain
Rotate by starting the agent with a new key and a new chain file. A new signed chain begins, so each chain checks against a single key. Records can carry a key ID.
Whoever checks your records uses the public key you give them separately. They pin it in advance instead of trusting a copy carried in the chain.
When the agent stops or restarts
The chain records each agent start and each clean stop. Nothing is recorded while the agent isn't running, so that time shows as a gap between the timestamps of consecutive records. A start with no clean stop before it shows that the previous run didn't close cleanly.
If a chain was cut off by a crash, the verifier reports its unanchored end separately instead of passing it as complete.
One signed chain per node
Each node keeps its own signed chain, and each record names the GPU it applies to. Chains from different nodes aren't linked to each other.
Where you configure it, chain heads are anchored off the machine in write-once object storage (S3 with Object Lock) in your own bucket. Otherwise, anchors stay on the local machine. An off-box anchor is what lets someone checking the chain see that its end wasn't cut off.
Runs on your nodes
Install Spark-XC as a systemd service, or as a Kubernetes DaemonSet from the Helm chart or a plain manifest.
- Privileges. The observe-only systemd service runs as a non-root user with no added capabilities. The Helm chart runs as root with SYS_ADMIN and SYS_RAWIO, and installs in observe-only mode by default.
- Requirements. Python 3.10 or later, GPU driver 525 or later with the persistence daemon (nvidia-persistenced), and NVML access. No kernel modules and no driver changes: it uses the GPU driver's standard NVML management API. Enforcing caps needs the admin capabilities that NVML power-limit writes require; the observe-only systemd service runs without root.
- Network. No telemetry is sent to Spark-XC. Outbound connections go only to infrastructure the operator configures. Metrics listen on localhost only (127.0.0.1:9090).
- Files. Records and logs go to /var/log/spark-xc, and agent state to /var/lib/sparkxc. The chain rotates locally by size, 500 MB per file with seven older files kept, and the oldest file is deleted. Keep your own copy if you need longer retention.
What an export contains
Export the signed chain as JSON Lines, one record per line, and render Power Event Records from it as readable HTML or text.
Records in the chain carry UTC timestamps; telemetry records also carry the GPU's power draw and power limit in watts.
event_summary.timestamp_utcISO 8601 timestamp in UTC
ts, telemetry_raw.timestampEpoch timestamps, in seconds
gpu_idThe GPU the record applies to
power_w, power_limit_w, requested_limit_wPower draw, power limit, and requested limit, in watts
Power figures are instantaneous readings as the GPU reports them. Records don't include meter data or measured energy (kWh).
Check a record yourself
Recompute an example Power Event Record's hash chain in your browser, or ask us how keys and deployment would work on your nodes.