Measurement / Calibrated confidence

Confidence you can check for yourself.

A stated confidence of 0.7 carries an empirical coverage guarantee, not a model probability. It holds whether or not the underlying survival model is correctly specified.

Every feature derives from public state. No licensed or privileged data.

DefiLlama
Merkl
Blockscout
Alchemy
Foundry
EIP-712
DefiLlama
Merkl
Blockscout
Alchemy
Foundry
EIP-712

Why not the model's own number.

A survival model emits a probability, and it is tempting to publish that as confidence. For a bonded claim it is the wrong number.

01

A model probability inherits every misspecification in the model.

02

If the feature distribution shifts, a stated 0.9 can be an empirical 0.6.

03

Publishing that as confidence is a promise the system cannot keep.

How this works

Split conformal prediction gives distribution-free, finite-sample coverage under exchangeability. Market regimes are not exchangeable, so the mitigations matter as much as the method.

USDestETH / ETHReliability diagramCalibration stratum

Coverage by confidence bucket

Stated 0.9 – realised 0.88On diagonal22 Jun
Stated 0.8 – realised 0.79On diagonal22 Jun
Stated 0.7 – realised 0.72On diagonal22 Jun
Stated 0.6 – realised 0.58On diagonal22 Jun
Regime shift detected, re-stratifying

01

Scores nonconformity

Signed, so over-prediction of the horizon is the error that counts. That is the error that gets a user liquidated.

02

Stratifies by regime

Separate calibration sets per market regime, matched on supply direction, cross-chain flow, and realised volatility.

03

Degrades explicitly

Outside the support of every calibration stratum the system does not extrapolate. It flags OUT_OF_SUPPORT and lowers confidence. Refusing to make a confident claim is a supported output.

What it reads, derives,
and does with it.

What it reads

Resolved attestations, held back from training.

Calibration set
Realised breach times
Ensemble dispersion

What it derives

What calibration produces.

Nonconformity scores
Empirical quantile
Coverage guarantee
Shift weights

What it does with it

What ends up in the payload.

confidenceBps
Conformalised horizon
OUT_OF_SUPPORT flag

The scoring model is publishable and copyable, but the accuracy record is not. A competitor who reimplements the ensemble from this paper starts with an identical model and a calibration history of zero.

Cleaton whitepaper

Section 12 — Conclusion

Thirty-six surfaces across contract, API, agent, and automation

Durability attestation for autonomous capital. Read the horizon before the capital moves, not after.