Scoreboard

A live, tamper-evident record of daily M-class forecasts, resolved in public against the official flare list — with climatology and persistence on the same board.

Live

Live board — this repo's own measurement

Every day the filter issues one forecast; at the day's close it is resolved against the official flare list and the running scores update. The board is append-only and hash-chained: a wrong entry is corrected by a later dated entry, never edited.

Running record on 5 resolved day(s): 0 verified M-class day(s), an observed base rate of 0.000, and a forecaster Brier of 0.037.

Additional baselines this score epoch pre-registered — recomputed from the committed ledger at render time.
baselinedays it could issue forits Brierforecaster Brier, same daysits BSS vs this epoch's reference
logistic on daily counts (record v46) 5 0.007 0.037

A cheap competitor that sees nothing but the daily M-class count history, with coefficients pinned from the sealed upstream record and never refitted here — a baseline that learned from the days it scores would not be a baseline. On that record's held-out tail it beat the calibrated forecaster this site ships, which is precisely why it is on the board rather than left off it. Its Brier and the forecaster's are computed over the same days, because two scores over two different day sets are not a comparison. Where the count history has a hole the feature window carries forward past it rather than inventing a quiet day, so across a gap its trailing windows reach further back than their names suggest.

Too few days for a skill score. A skill score divides by a reference Brier estimated from the same handful of days, so on a short board it measures which days happened to land in the sample rather than the forecast. This board prints no BSS — against climatology, against persistence, or against anything else — until it has thirty resolved days. The count, the hits, the base rate and the Brier above are the whole of what a board this short can honestly say. The verification board, folded below, still shows the shape of the forecasts — reliability, sharpness, and where the Brier score comes from — because those describe these days rather than score them against a reference.

Live daily forecasts (scored against the verification-period climatology, p = 0.144).
forecast day (UTC) score epoch P(M-class, next 24 h) resolved M-class events
2026-09-03 0 0.431 no flare 0
2026-09-04 0 0.353 no flare 0
2026-09-05 0 0.464 no flare 0
2026-09-06 0 0.520 no flare 0
2026-09-07 0 0.413 no flare 0
2026-09-08 0 0.353 no flare 0
2026-09-09 0 0.310 no flare 0
2026-09-10 1 0.223 no flare 0
2026-09-11 1 0.204 no flare 0
2026-09-12 1 0.189 no flare 0
2026-09-13 1 0.176 no flare 0
2026-09-14 1 0.166 no flare 0
2026-09-15 1 0.156 pending
How this board is referenced — what it is scored against, and what it will be
  • This epoch's scored reference is the verification-period constant, p = 0.144 — the M-class-day frequency of the sealed record's held-out tail. It is fixed for the life of the epoch. Days already scored under it are never re-scored under anything else.
  • Why the fixed reference flatters the model near cycle maximum. That p = 0.144 is the average of a held-out tail that spans a deep minimum as well as an active stretch. The Sun is not at that average now. A constant that under-forecasts the current rate is itself a badly calibrated forecast, so its reference Brier is large — and skill measured against a large reference Brier looks impressive without the forecaster doing anything. Conditioning the climatology on where the cycle actually is puts the reference at p = 0.298: the M-class-day frequency over the sealed fifty-year flare series, restricted to the months whose smoothed sunspot number sat in the same decile as today's. That is roughly twice the fixed constant, and it is a far harder thing to beat. It is not swapped in mid-epoch — changing the scored reference is exactly what starts a new epoch, which is why it is pre-registered above instead.
  • Sunspot numbers: source WDC-SILSO, Royal Observatory of Belgium, Brussels (CC BY-NC 4.0), vendored and checksum-pinned; the cycle-conditioned constant is derived by tools/cycle_climatology.mjs and pinned in contracts/scoreboard.pins.json.
For the verification reader — reliability, sharpness, the Murphy decomposition, discrimination, and the month-by-month strips

Every figure below is computed from the same committed ledger days as the scores, with the same functions the sealed backtest strip uses. It is a live board: it is expected to look thin, and it is shown thin rather than waited out.

Live reliability — mean forecast probability (horizontal) against observed frequency (vertical) across ten bins; larger dots hold more days. Near the diagonal is well-calibrated. On a board this short a single day moves a bin from one end to the other.
forecast probability → observed frequency →
Sharpness — how many forecasts land in each of the ten probability bins. A forecaster that never leaves the middle of the range is unsharp however well calibrated it is; the empty bins are the point.
forecast probability → forecasts →

Where the Brier score comes from

Murphy decomposition of the live Brier score — recomputed from the ledger at render time.
termvaluewhat it is
reliability (REL)0.037the calibration penalty — how far each bin's outcome frequency sits from the probability the bin promised. Lower is better; zero is perfect calibration.
resolution (RES)0.000how far the bins' outcome frequencies move away from the overall base rate — the useful part, the discrimination the forecaster adds. Higher is better.
uncertainty (UNC)0.000the base rate's own variance over these days. Nothing a forecaster does changes it; it is the difficulty of the sample.
REL − RES + UNC0.037equals the Brier score above — the identity is the check that this decomposition is of this board and not of something else.

Discrimination

Sweeping a yes/no decision threshold across the issued probabilities (the ROC sweep) gives a peak true skill statistic — but that is a skill score, and this board is below its minimum: too few days for a skill score. It appears once the board has thirty resolved days.

Month by month

Per-month strips — the resolved days grouped by calendar month, from the ledger.
month (UTC)resolved daysM-class daysbase rateBrier
2026-09 5 0 0.000 0.037

No skill score is shown per month: every month is far below the minimum this board requires before it prints one.

Backtest

BACKTEST (sealed v26 — a 20-year held-out tail, walked forward in state only)

The sealed record's own out-of-sample test: a fixed-parameter causal filter, walked forward in state only, over a 20-year held-out tail (7,305 days, ≈2005–2025) of the fifty-year GOES/XRS record, trained on the leading 60%. It is a backtest, kept on its own strip and never merged with the live board above; on that tail the cascade filter scores Brier 0.0816.

Sealed backtest skill — record v26, against two reference climatologies: the training-period constant the sealed records use (p = 0.419) and the verification period's own base rate (p = 0.144), which is the reference a space-weather reader assumes and the one the live board uses. Every number links to its deposit.
forecaster / baseline BSS vs training climatology 95% CI on that BSS BSS vs verification climatology peak TSS
cascade filter — record v26 +0.589 [0.521, 0.667] +0.338 0.621
cascade filter, calibrated — record v45 +0.605 +0.365 0.621
persistence — record v26 +0.576 [0.504, 0.662] +0.317 0.633
cycle-timescale rate — record v26 +0.507 [0.427, 0.601] +0.205

The cascade forecaster leads the cycle-timescale rate baseline and edges recency persistence on point estimates, within overlapping confidence intervals; persistence has the higher peak TSS. The cycle-timescale rate is named as what it is — a slow rate baseline — never as a win over the published forecaster.

On that same tail the calibrated re-seal of these forecasts scores Brier 0.0784 — the row above, labelled by its own record.

Reliability of the sealed cascade filter — mean forecast probability (horizontal) versus observed frequency (vertical), across ten bins; near the diagonal is well-calibrated (record v26).
forecast probability → observed frequency →
What the calibrated re-seal is, and what it did not fix

The second forecaster row is the same sealed forecasts, calibrated — a two-parameter map fitted on the record's training block alone and applied to the probabilities — sealed as its own record beside the raw one, never merged with it and never with the live board. Calibration is a strictly increasing map, so it moves no ranking at all: the peak true skill statistic is unchanged, and recency persistence keeps the higher one. It moves the probabilities, which is the defect it repairs. The record's own pre-registered reliability target was not met even so: the worst calibrated bin still misses the diagonal by more than the target allowed, and the residual error sits in the sparse high-probability bins. On that record, calibration bought skill — not a flat reliability diagram.

Record v45 — the calibrated re-seal of these same forecasts (no refit, no new observations). Sealed upstream and vendored here with a pinned SHA256 checksum (vendor/fenocosm/v45_calibration_anchor.json); no deposit has been minted for it yet, so it is cited by record id and pin rather than by a DOI it does not have.

Backtest

BACKTEST (sealed v49 — the same frozen filter read days ahead, h-matched field)

The sealed record reads the same frozen filter further ahead: for each lead time from one to fourteen days it issues the day's flare probability from counts up to the day minus the lead, against a field matched to the same information cut — recency persistence, the cycle-timescale rate, and a count-only logistic, each frozen at the matching lag. Same 7,305-day held-out tail as the strip above, scored against the verification-period climatology (p = 0.144 — the live board's convention), with paired block-bootstrap intervals throughout. The no-peeking property is measured, not assumed: at every lead the record probes that bumping a future day's count moves no earlier forecast, and at one day of lead the curve reproduces the sealed daily forecast trace exactly. A backtest, on its own strip, never merged with the live board.

Sealed skill by lead time — record v49, against the verification-period climatology. The last column is the paired difference, cascade minus h-matched persistence, with its bootstrap interval; at no lead time in the grid does it cover zero. Every number links to the sealed record.
lead time cascade BSS cascade BSS, calibrated persistence BSS cascade − persistence, 95% CI
one day +0.338 +0.365 +0.317 +0.021 [+0.012, +0.029]
two days +0.253 +0.297 +0.223 +0.030 [+0.018, +0.042]
three days +0.200 +0.259 +0.159 +0.041 [+0.026, +0.057]
five days +0.133 +0.212 +0.068 +0.066 [+0.045, +0.090]
seven days +0.112 +0.201 +0.040 +0.072 [+0.045, +0.103]
fourteen days +0.082 +0.184 -0.004 +0.086 [+0.052, +0.123]

The edge over persistence is at its smallest where the headline strip measures it (one day: +0.021 [+0.012, +0.029]) and widens with lead (fourteen days: +0.086 [+0.052, +0.123]): recency decays far faster than the filter — by fourteen days persistence scores below the climatology itself while the cascade still holds skill. Recalibration is worth more the further out the forecast is issued: the training-block map adds +0.026 at one day and +0.102 at fourteen.

Skill against the verification-period climatology (vertical) by lead time, one to fourteen days (horizontal) — the cascade filter against h-matched recency persistence and the h-matched cycle-timescale rate; the dashed horizontal is zero skill. The exact values are in the table above (record v49).
cascade persistence cycle rate lead time → skill →

Two results cut the other way, and are stated with the same prominence. The h-matched count-only logistic beats the cascade at every lead time — calibrated against calibrated, the paired difference is -0.009 [-0.017, -0.002] at one day and -0.029 [-0.053, -0.007] at fourteen — the same cheap competitor the live board carries as a baseline. And past roughly three days the raw filter's edge over the cycle-timescale rate is gone: their paired difference covers zero at three days and is negative from five (-0.100 [-0.147, -0.063] at fourteen) — at long lead, most of what the raw filter knows is cycle tracking, which is also why the calibrated form is the one worth reading there.

The record scores the recent bridge window separately and never pools it with this tail; on that short, unusually active window only the one-day lead stays positive against the window's own base rate, and the record quotes it as such. No lead-time forecast is issued live: the live board above is a one-day board, and this strip is the sealed record's measurement, not this repository's.

Verify this ledger yourself — the commands, from a clean clone, with no network and no build

The ledger is the audit trail — not the job's logs, which expire and are not content-addressed. Every line carries the hash of the line before it, so any edit to history breaks the chain.

git clone <repo> && cd <repo> && npm ci
node tools/verify_ledger.mjs --anchor-state ledger/state.json   # chain, structure, published anchor
python3 tools/verify_ledger.py ledger/scoreboard.jsonl          # an independent re-implementation
node tools/sign_tail.mjs --verify                               # the signed tail, once signed
node --test tests/ledger/*.test.ts  # re-derives the scores from the ledger and cross-checks them

The verifier ships as a standalone public artifact — the ledger plus verify_ledger.mjs, carrying no engine or coupling code — so the chain is third-party-checkable even while the main repository is private. A silently edited line, or a broken link, is rejected: the committed tamper fixtures are the demonstrated failure mode. A full re-seal that recomputes every following hash is caught against an external anchor — the append-only history, each published state file's recorded chain tail, the signed tail, and an OpenTimestamps proof of that tail taken at issue, which commits it to the Bitcoin blockchain and so is the one anchor not written by us. The second command above is a standard-library-only Python re-implementation of the same canonical form: two independent verifiers agreeing on the committed ledger is what makes the specification, not our code, the thing you have to trust. Resolution runs at each window's close plus six hours, against the pinned official flare list.

References

Retrospective numbers mirror a deposited record; they never first-publish here (SPECS §7). The three founding deposits are published and every DOI below resolves (launch gate G2, cleared); the later sealed records (v45 and after) are cited by sealed record id and commit, their deposits pending.