Method

How the Sun is turned into a forecast, a sound, and a score — in plain language, every number traceable to a deposited record.

The ladder

The flare record is read as a cascade: a small stack of coupled layers, slow to fast, each either quiet or active. A slow layer holds a mood for weeks; a fast layer flickers day to day. The stack this record's likelihood prefers has six layers — a slow pair, a middle pair, a fast pair. It is a way of describing the record's multi-state structure, not a claim about the Sun's physical machinery.

This instrument's cascade controls are dial settings derived from a hierarchical cascade fitted to 50 years of solar flares (record v25). A fitted span is the ladder this record's likelihood prefers — a model-selection quantity, not a recovered physical hierarchy (Paper F §7.1–7.2).

The filter

A deterministic causal filter carries a belief over every combination of the six layers — sixty-four states in all — and updates it once per closed day from that day's count of M-class flares, corrected for how the detectors' sensitivity has changed across eras. The belief is carried forward from the start of the record: today's forecast is a function of a belief accumulated over the whole fifty years, not of the last few days alone.

Nothing is re-fitted live: the ladder is a frozen, sealed constant from the upstream record (record v26), and this site only runs the filter against it. The separate full-record fit named above (record v25) is what the instrument's dial settings derive from; the forecast never uses it.

The forecast

From the belief after the last closed day, the filter issues a one-step-ahead now-cast: the probability of at least one M-class flare in the next 24 hours. It is issued once per UTC day, the first run after the day's data has settled; every other hourly run only refreshes the flux reading and the clock — it never moves the forecast. A missed day is recorded as a visible gap, never back-dated.

When it is published, and how you can tell

A UTC day's forecast is issued at about 03:07Z — three hours into the very window it forecasts, and roughly twenty-one hours before that window closes. Those three hours are the settle margin: the official flare list is still receiving late reports for the day that just ended, and issuing before it settles would mean forecasting from an incomplete count. The forecast is computed only from days that closed before 00:00Z — it never sees a minute of the day it will be scored on.

Publishing a few hours into the window is a real compromise, stated plainly rather than smoothed over, and the timing does not have to be taken on trust.

How the issue time is made checkable — the ledger line, the signed tail, the external timestamp, and the commands
  • The issue timestamp is inside the ledger line and is hashed with it, so it cannot be moved afterwards without breaking the chain.
  • The chain tail is signed — an Ed25519 signature over the tail hash, the line count and the signing time, under the published production key ff-pack-2026-09, whose public half ships in this repository's trust set.
  • The tail is timestamped externally. An OpenTimestamps proof commits it to public calendars and thence to the Bitcoin blockchain — the one anchor here that is not our own word. It is taken at issue, hours before the window closes and longer still before it is resolved, which is exactly the claim a forecast written after the fact could not survive.
node tools/verify_ledger.mjs --anchor-state ledger/state.json
node tools/sign_tail.mjs --verify     # the signed tail (once the production key has signed one)
node tools/verify_ledger.mjs --ots    # optional, networked: upgrade and verify the timestamp proof
What “certified” means here — the anchor gate, the append-only ledger, the mirror rule, and the stated ceiling

The honesty box

The measured skill, stated once, with its caveats attached — verbatim, the same wording as the sealed records' own citation-gated sentence. Every number links to its deposit or its sealed record.

What is published is the sealed filter's forecast passed through a sealed recalibration map fitted on the training block alone (record v45); the map moves the probability only, never the belief the instrument plays.

On the 20-year held-out tail — a fixed-parameter causal filter, walked forward in state only — the M-class 24-h forecast published here is the sealed filter's forecast passed through a sealed two-parameter recalibration fitted on the training block alone (record v45), and it scores Brier 0.0784. Against the training-period climatology the sealed records use (a constant p = 0.419) that is BSS +0.605; against the verification period's own climatology (p = 0.144 — the reference a space-weather reader will assume) the same forecasts give +0.365. Before recalibration the identical filter scored Brier 0.0816 (+0.589 [0.521, 0.667] / +0.338) and over-issued in every reliability bin; the map moves probabilities only, never the belief the instrument plays. The ordering is identical under both references: the calibrated forecast leads a cycle-timescale rate baseline (+0.507 / +0.205) and recency persistence (+0.576 / +0.317), and the edge over persistence is a paired, interval-separated win (ΔBSS +0.029 [+0.020, +0.038] against the training reference; +0.047 [+0.035, +0.061] against the verification reference), while persistence keeps the higher peak TSS (0.633 vs 0.621). Two results cut the other way and are stated with the same prominence: a logistic regression on the same seven days of counts scores slightly better than the calibrated cascade (+0.373 against the verification reference; paired ΔBSS -0.009 [-0.017, -0.002], record v46), and once calibrated a four-layer ladder or a free five-state hidden Markov model forecasts as well as the sixty-four-state ladder (paired intervals cover zero, record v47) — the depth is not what buys the skill. Read at longer leads the same honesty holds: against an h-matched field the count-only logistic still scores better at every lead measured, one to fourteen days, and past about three days of lead the cycle-timescale rate overtakes the uncalibrated filter (record v49). The live board scores each epoch against its pinned reference and prints the training constant, the verification constant and the realised in-window base rate beside every score, so the live strips and the backtest strip stay comparable. The X-class tile beside the forecast carries a caveat of the same kind: the sixty-four-state belief carries X-relevant information only through the rate — one constant flare-level X fraction applied to the M-class rate reproduces the tile's skill, and the calibrated paired difference between the two covers zero (record v52). Most M/X predictive information lives in magnetograms this filter never sees. Research demonstration and baseline layer — not operational space-weather guidance. Only the M-class forecast is scored; the X-class tile is a display of the sealed real-label re-run, not a scored product.

Why two climatologies — which baseline a skill score is measured against, and why each score epoch pins its own

A skill score compares the forecast to a naive baseline, and the answer depends on which baseline. The sealed record uses the training-period climatology it was built on; a space-weather reader will instead assume the verification period's own base rate, which is lower because solar activity declined across the cycles in the tail. Both are shown, because quoting only the larger number beside a live board scored on the smaller one would make the two strips numerically incomparable — on a site whose whole thesis is comparability.

The live board scores each epoch against its pinned reference. Epoch zero is scored against the verification-period convention; epoch one, which begins with the first day issued under the recalibrated model (record v45), against a cycle-conditioned climatology, because a fixed reference from a quieter period flatters near cycle maximum. Every epoch keeps its reference forever, and the board prints the training constant, the verification constant and the realised in-window base rate beside every score, so every strip can be read on the same scale. Changing a reference starts a new score epoch — a stated, owner-confirmable choice, never a silent one.

Sunspot numbers: source WDC-SILSO, Royal Observatory of Belgium, Brussels (CC BY-NC 4.0), vendored and checksum-pinned; the cycle-conditioned constant is derived by tools/cycle_climatology.mjs and pinned in contracts/scoreboard.pins.json.

References

Retrospective numbers mirror a deposited record; they never first-publish here (SPECS §7). The three founding deposits are published and every DOI below resolves (launch gate G2, cleared); the later sealed records (v45 and after) are cited by sealed record id and commit, their deposits pending.