Method
How the Sun is turned into a forecast, a sound, and a score — in plain language, every number traceable to a deposited record.
The ladder
The flare record is read as a cascade: a small stack of coupled layers, slow to fast, each either quiet or active. A slow layer holds a mood for weeks; a fast layer flickers day to day. The stack this record's likelihood prefers has six layers — a slow pair, a middle pair, a fast pair. It is a way of describing the record's multi-state structure, not a claim about the Sun's physical machinery.
This instrument's cascade controls are dial settings derived from a hierarchical cascade fitted to 50 years of solar flares (record v25). A fitted span is the ladder this record's likelihood prefers — a model-selection quantity, not a recovered physical hierarchy (Paper F §7.1–7.2).
The filter
A deterministic causal filter carries a belief over every combination of the six layers — sixty-four states in all — and updates it once per closed day from that day's count of M-class flares, corrected for how the detectors' sensitivity has changed across eras. The belief is carried forward from the start of the record: today's forecast is a function of a belief accumulated over the whole fifty years, not of the last few days alone.
Nothing is re-fitted live: the ladder is a frozen, sealed constant from the upstream record (record v26), and this site only runs the filter against it. The separate full-record fit named above (record v25) is what the instrument's dial settings derive from; the forecast never uses it.
The forecast
From the belief after the last closed day, the filter issues a one-step-ahead now-cast: the probability of at least one M-class flare in the next 24 hours. It is issued once per UTC day, the first run after the day's data has settled; every other hourly run only refreshes the flux reading and the clock — it never moves the forecast. A missed day is recorded as a visible gap, never back-dated.
When it is published, and how you can tell
A UTC day's forecast is issued at about 03:07Z — three hours into the very window it forecasts, and roughly twenty-one hours before that window closes. Those three hours are the settle margin: the official flare list is still receiving late reports for the day that just ended, and issuing before it settles would mean forecasting from an incomplete count. The forecast is computed only from days that closed before 00:00Z — it never sees a minute of the day it will be scored on.
Publishing a few hours into the window is a real compromise, stated plainly rather than smoothed over, and the timing does not have to be taken on trust.
How the issue time is made checkable — the ledger line, the signed tail, the external timestamp, and the commands
- The issue timestamp is inside the ledger line and is hashed with it, so it cannot be moved afterwards without breaking the chain.
- The chain tail is signed — an Ed25519 signature over the tail hash, the line
count and the signing time, under the published production key
ff-pack-2026-09, whose public half ships in this repository's trust set. - The tail is timestamped externally. An OpenTimestamps proof commits it to public calendars and thence to the Bitcoin blockchain — the one anchor here that is not our own word. It is taken at issue, hours before the window closes and longer still before it is resolved, which is exactly the claim a forecast written after the fact could not survive.
node tools/verify_ledger.mjs --anchor-state ledger/state.json node tools/sign_tail.mjs --verify # the signed tail (once the production key has signed one) node tools/verify_ledger.mjs --ots # optional, networked: upgrade and verify the timestamp proof
What “certified” means here — the anchor gate, the append-only ledger, the mirror rule, and the stated ceiling
- Reproduces the seal. The forecast is the sealed record's construction, run live. The running code is a port, admissible only while it reproduces the sealed record's own day-by-day forecast to that record's re-derivation tolerance — an anchor gate checked on every change, with a deliberately-broken fixture proving it can fail.
- Scored in public, append-only. Every daily forecast is written to a hash-chained ledger and resolved against the official flare list. A wrong entry is corrected by a later dated entry, never edited; anyone can re-verify the whole chain from a clean clone.
- Mirrors, never first-publishes. Retrospective skill numbers appear here only after their deposit exists (the References below). The live forecast and its running score are this site's own first publication, and the two are never shown as one.
- States its ceiling. Most of the predictive information for M- and X-class flares lives in photospheric magnetograms this filter never sees. The claim is skill from the event record alone — a cheap, sensor-free baseline, not a replacement for operational space-weather guidance.
The honesty box
The measured skill, stated once, with its caveats attached — verbatim, the same wording as the sealed records' own citation-gated sentence. Every number links to its deposit or its sealed record.
What is published is the sealed filter's forecast passed through a sealed recalibration map fitted on the training block alone (record v45); the map moves the probability only, never the belief the instrument plays.
On the 20-year held-out tail — a fixed-parameter causal filter, walked forward in state only — the M-class 24-h forecast published here is the sealed filter's forecast passed through a sealed two-parameter recalibration fitted on the training block alone (record v45), and it scores Brier 0.0784. Against the training-period climatology the sealed records use (a constant p = 0.419) that is BSS +0.605; against the verification period's own climatology (p = 0.144 — the reference a space-weather reader will assume) the same forecasts give +0.365. Before recalibration the identical filter scored Brier 0.0816 (+0.589 [0.521, 0.667] / +0.338) and over-issued in every reliability bin; the map moves probabilities only, never the belief the instrument plays. The ordering is identical under both references: the calibrated forecast leads a cycle-timescale rate baseline (+0.507 / +0.205) and recency persistence (+0.576 / +0.317), and the edge over persistence is a paired, interval-separated win (ΔBSS +0.029 [+0.020, +0.038] against the training reference; +0.047 [+0.035, +0.061] against the verification reference), while persistence keeps the higher peak TSS (0.633 vs 0.621). Two results cut the other way and are stated with the same prominence: a logistic regression on the same seven days of counts scores slightly better than the calibrated cascade (+0.373 against the verification reference; paired ΔBSS -0.009 [-0.017, -0.002], record v46), and once calibrated a four-layer ladder or a free five-state hidden Markov model forecasts as well as the sixty-four-state ladder (paired intervals cover zero, record v47) — the depth is not what buys the skill. Read at longer leads the same honesty holds: against an h-matched field the count-only logistic still scores better at every lead measured, one to fourteen days, and past about three days of lead the cycle-timescale rate overtakes the uncalibrated filter (record v49). The live board scores each epoch against its pinned reference and prints the training constant, the verification constant and the realised in-window base rate beside every score, so the live strips and the backtest strip stay comparable. The X-class tile beside the forecast carries a caveat of the same kind: the sixty-four-state belief carries X-relevant information only through the rate — one constant flare-level X fraction applied to the M-class rate reproduces the tile's skill, and the calibrated paired difference between the two covers zero (record v52). Most M/X predictive information lives in magnetograms this filter never sees. Research demonstration and baseline layer — not operational space-weather guidance. Only the M-class forecast is scored; the X-class tile is a display of the sealed real-label re-run, not a scored product.
Why two climatologies — which baseline a skill score is measured against, and why each score epoch pins its own
A skill score compares the forecast to a naive baseline, and the answer depends on which baseline. The sealed record uses the training-period climatology it was built on; a space-weather reader will instead assume the verification period's own base rate, which is lower because solar activity declined across the cycles in the tail. Both are shown, because quoting only the larger number beside a live board scored on the smaller one would make the two strips numerically incomparable — on a site whose whole thesis is comparability.
The live board scores each epoch against its pinned reference. Epoch zero is scored against the verification-period convention; epoch one, which begins with the first day issued under the recalibrated model (record v45), against a cycle-conditioned climatology, because a fixed reference from a quieter period flatters near cycle maximum. Every epoch keeps its reference forever, and the board prints the training constant, the verification constant and the realised in-window base rate beside every score, so every strip can be read on the same scale. Changing a reference starts a new score epoch — a stated, owner-confirmable choice, never a silent one.
Sunspot numbers: source WDC-SILSO, Royal Observatory of Belgium, Brussels
(CC BY-NC 4.0), vendored and checksum-pinned; the cycle-conditioned constant is derived by
tools/cycle_climatology.mjs and pinned in contracts/scoreboard.pins.json.
References
Retrospective numbers mirror a deposited record; they never first-publish here (SPECS §7). The three founding deposits are published and every DOI below resolves (launch gate G2, cleared); the later sealed records (v45 and after) are cited by sealed record id and commit, their deposits pending.
-
Paper F — Hidden multi-state structure versus the piecewise-Poisson flare rate on 50 years of GOES/XRS solar flares (Paper F). the skill numbers, the two-reference climatology table, and the span-selection reading (§7.1–7.2).
https://doi.org/10.5281/zenodo.22180473 -
record v26 — Fenocosm record v26 — the now-cast construction (deterministic causal filter, frozen fitted ladder). the sealed backtest aggregate: Brier, both-reference BSS, the reliability diagram, the held-out-tail window.
https://doi.org/10.5281/zenodo.22180495 -
record v25 — Fenocosm record v25 — the full-record fit. the fitted ladder the instrument's cascade dial settings are derived from (§6).
https://doi.org/10.5281/zenodo.22180493 -
record v45 — Fenocosm record v45 — the calibrated re-seal (Platt map fitted on the training block; paired block-bootstrap differences). the calibrated Brier and both-reference BSS, the paired cascade-minus-persistence intervals, the sealed map the live forecast applies.
sealed 2026-09-09 · Fenocosm out/v45/p103_nowcast_calibration.json @ ae5ab048179e · Zenodo deposit pending -
record v46 — Fenocosm record v46 — the cheap-challenger tournament (a logistic regression on the same counts, the two-EWMA GLM, the stack). the paired calibrated-cascade-minus-logistic interval, and the logistic baseline the epoch-one board carries.
sealed 2026-09-09 · Fenocosm out/v46/p104_nowcast_challengers.json @ ae5ab048179e · Zenodo deposit pending -
record v47 — Fenocosm record v47 — the hierarchical-vs-flat ablation in forecast metrics (ladders of one to six layers; free Poisson-HMMs of two to six states). the paired intervals showing four layers, or a free five-state HMM, forecast as well as the sixty-four-state ladder once calibrated.
sealed 2026-09-09 · Fenocosm out/v47/p105_nowcast_ablation.json @ ae5ab048179e · Zenodo deposit pending -
record v49 — Fenocosm record v49 — the per-horizon flare skill curve (the frozen filter read at one to fourteen days of lead, against an h-matched field, with paired intervals). every number on the lead-time backtest strip: the per-horizon skill of the cascade raw and calibrated, h-matched persistence, cycle-rate and logistic baselines, and the paired differences.
sealed 2026-09-10 · Fenocosm out/v49/p107_nowcast_horizon.json @ 117a258a5966 · Zenodo deposit pending -
record v50 — Fenocosm record v50 — the real-label X-class re-run (five hundred sixty-eight real X-class flares; the per-state fraction read from the frozen filter's belief, calibrated). the X-class tile's read-out: its per-state fraction rule, its calibration map, and the paired evidence that it beats X-class persistence.
sealed 2026-09-10 · Fenocosm out/v50/p108_nowcast_xclass.json @ 117a258a5966 · Zenodo deposit pending -
record v52 — Fenocosm record v52 — the marked flare channel (does the latent activity regime set flare size, or only flare rate?). the X tile's mechanism: the sixty-four-state belief carries X-relevant information only through the rate — one constant flare-level X fraction on the M-class rate reproduces v50's X-day skill, and the calibrated paired difference covers zero.
sealed 2026-09-11 · Fenocosm out/v52/p110_flare_marks.json @ 8047648d8363 · Zenodo deposit pending