How the AI is used, and where every number comes from
Data provenance is itself a trust feature. Every stream on this platform is a public, documented feed with a fetch time; every forecast is produced by a model that was trained on that feed and evaluated on data it never saw; and every page says which of those feeds answered on this build.
Seven concrete uses, all running on this site
Learned forecasting on real archives
Each sphere has its own ensemble: a gated recurrent network (GRU) and a feed-forward network (MLP) trained by backpropagation on the ingested archive, alongside a ridge autoregression, Holt smoothing, a seasonal baseline and persistence. Members are weighted by validation error and vote on the next value. Training is deterministic, time-ordered, and reproducible from the committed data.
Distribution-free calibrated uncertainty
Every forecast ships a split-conformal interval: the 90 % quantile of absolute residuals on a calibration window the models never trained on, with the finite-sample correction. Because the world drifts, adaptive conformal inference (Gibbs & Candès, 2021) nudges the miscoverage level after every observed step, and the platform reports both the static and adaptive coverage it actually achieved.
Explainability by Input×Gradient
The neural members are differentiable end to end, so the platform computes the gradient of the forecast with respect to every input in the window, multiplies by the input, and sums per feature. Each driver on a sphere page — rainfall, dust, the 27-day solar rotation — is a signed contribution to the forecast, not a hand-written label.
Out-of-distribution detection and drift
Today's input row is standardised against the training window; the Trust Core takes the larger of its RMS distance and its worst single feature. Separately, the Population Stability Index compares recent day-to-day changes with the same season of the training years, so seasonal series are not mistaken for drift.
Abstention as a first-class output
If the members disagree beyond threshold, the input sits outside the training distribution, the calibrated band is too wide to act on, or the feed is stale, the sphere is withheld and excluded from the fused index. The rule is one function, classifyTrust, shared by the live pipeline and the interactive simulator.
Physical models where learning would be dishonest
The Canadian Fire Weather Index, the MOD17-style light-use-efficiency carbon model, and the concentration–discharge power law for dissolved solids are published equations, not fitted curves. They turn raw feeds into the quantities the spheres watch, and they are labelled as models, not measurements.
Self-audit on every build
The last 90 steps are re-forecast from the data available at the time and compared with what the feeds then reported. Coverage, error, drift and the reliability diagram on the Monitoring page are recomputed from those real forecasts, not stored from training.
| Sphere | Governed quantity | Model | Held-out | Freshness |
|---|---|---|---|---|
Sun | Next-day observed 10.7 cm solar radio flux (the official 20:00 UT reading) Penticton/DRAO archive via LISIRD (1947→) · NOAA SWPC live readings · SWPC Kp | GRU + MLP + AR ridge + Holt + seasonal ensemble, trained on 21 years of Penticton daily flux; split conformal + adaptive intervals | skill +4.4% cov 79% · n 1172 lag 1 d | Live |
Atmosphere | Next-day corridor-mean PM2.5 across the six sites CAMS via Open-Meteo (2022→) · ERA5 daily · NASA FIRMS VIIRS active fire | GRU + MLP ensemble on CAMS air quality with ERA5 weather and Canadian FWI covariates; NASA FIRMS detections as the live fire layer | skill +12.0% cov 93% · n 222 lag 0 d | Live |
Hydrosphere | Next-day GloFAS river discharge on the Bukhan River reach at Gapyeong (5 km grid cell) GloFAS via Open-Meteo (1984→) · ERA5 rain | GRU + MLP ensemble on GloFAS discharge with basin rainfall covariates; dissolved solids estimated from discharge by a concentration–discharge power law | skill +19.7% cov 89% · n 637 lag 0 d | Live |
Biosphere | Next 16-day composite corridor-mean NDVI MODIS MOD13Q1 via ORNL DAAC (2000→) · ERA5 | GRU + MLP ensemble pooled across six sites on QA-masked MODIS NDVI with ERA5 covariates | skill +50.4% cov 96% · n 28 lag 50 d | Scheduled pull |
Geosphere | Next-day corridor-mean net ecosystem exchange (positive = source to the atmosphere) ERA5 radiation/temperature/VPD · MODIS NDVI · NOAA Mauna Loa CO₂ context | MOD17-style light-use-efficiency model (ERA5 + MODIS) producing daily GPP, respiration and NEE; GRU + MLP ensemble forecasts the flux | skill +3.1% cov 86% · n 473 lag 0 d | Live |
Every feed the platform reads. All are publicly available; none requires a credential. NASA FIRMS accepts a free key for a bounded query and otherwise falls back to the public regional file.
Ingest → assemble → train → serve → govern → refresh
scripts/ingest.ts pulls each archive from its public endpoint into data/raw with the fetch time and URL; committed, so every build is reproducible.
src/lib/data/frames.ts aligns feeds by date, computes derived quantities (FWI, GPP/respiration, seasonal terms) and forward-fills short gaps — the same code for training and for the live site.
scripts/train.ts fits every member on a strict time-ordered split (60 % train, 10 % validation, 15 % calibration, 15 % test), writes weights, residuals and held-out metrics to data/models.
At request time the live tail of every feed is fetched, merged over the archive, and the trained ensembles are executed in TypeScript — no Python service, no GPU. Pages revalidate hourly.
The Trust Core wraps each output: interval, attribution, disagreement, OOD, staleness, verdict. Survivors fuse into the Ecosystem Stress Index.
A scheduled job re-ingests and retrains, commits the new archives and weights, and redeploys — so the numbers move with the world, not with the build date.
Solar flux (Penticton, NOAA), MODIS NDVI, VIIRS fire detections, Mauna Loa CO₂. Reanalysis and analysis products — ERA5 weather, CAMS air quality, GloFAS discharge — are model-assimilated observations, and are labelled as such.
Net carbon flux comes from a light-use-efficiency model, not a flux tower. Fire danger comes from the FWI equations. Dissolved solids come from a concentration–discharge law with literature parameters, and are labelled as an estimate.
Every stream on the platform is built from publicly available, documented data with a recorded fetch time. Where a quantity is estimated rather than measured, the page says estimate, never pretends.