How accurate is this WBGT — and how do we know?
Every live WBGT you see in an app is a calculation, not a thermometer reading. There's no black-globe sensor on your street corner — the number is modeled from weather data (temperature, humidity, wind, sunshine). So the fair question isn't "is it the real WBGT?" but "how close is the model?" Here's how we keep ours honest — in public.
We grade ourselves against an independent answer key
The U.S. National Weather Service publishes its own WBGT forecast for locations across the country. That gives us an independent answer key to check ourselves against. We compare our number to theirs continuously, and we show the raw comparison live on our Live Accuracy page — updated through the day — instead of making a one-time claim.
(An honest caveat: the NWS figure is itself a model, so a gap measures method-and-input disagreement, not absolute truth. But it's a rigorous, independent yardstick — and the best public one there is.)
The accuracy flywheel
Every comparison is saved — not just the two numbers, but the full set of conditions behind them (temperature, humidity, wind, sunshine, time of day…). Over time that builds a growing, retained record of "here's what we predicted, here's the reference, and here were the exact conditions."
That dataset is what powers the next two steps — and it compounds: the longer it runs, the more it can teach us.
Our error wasn't random — it depended on conditions
When we looked across 17,958 of these comparisons — 30 days of them, read on 4 October 2026 — a clear pattern emerged. The one-line average, −0.79 °C, hides most of what is actually going on:
- In strong sun, we run a touch warm — about +0.5 °C.
- After dark, we run cool by about −1.6 °C, and that is where our raw error is biggest. Reading low is the direction worth saying out loud, so we say it.
- The warm bias is worst in calm air — up to a few degrees — which is exactly the situation that occasionally produced a scary "Extreme" reading on a day that felt fine.
A single across-the-board nudge cannot fix that: the nudge that lifted the night would push the middle of the day further off. The fix had to depend on the conditions — chiefly how strong the sun is and how much wind there is.
So we correct it — but only when it's proven to help
We learned a small correction from the data: how much to add or subtract for a given amount of sun and wind. In strong, calm sun it gently pulls the number down; on a still night it nudges it up.
The careful part is how it's trusted. Before we apply the correction, we hold out whole days it has never seen — train on every other day, then check the day we held back. On the 4 October 2026 reading that cross-check cut error by 34.5% across 31 held-out days.
We hold out days rather than cities on purpose. Our reference locations share the same weather, so leaving one city out still leaves that day's pattern in the training data, and the result looks better than it is.
What that buys, in the place it matters most: on that reading our raw error after dark was 1.73 °C and the corrected error 0.99 °C. In daylight it went from 1.02 °C to 0.85 °C. And if the correction ever stopped passing the test, it would switch itself off.
Our ±1.5 °C target is for the number we show you, the corrected one, judged by that same held-out check twice: across all hours, and after dark on its own. We count the target as met only when both are inside it, so good daytime numbers can't hide a miss at night. On the 4 October 2026 reading the error averaged 0.93 °C across all hours, inside the target. After dark it averaged 0.99 °C, also inside the target. The raw model's 1.42 °C is the figure before correction, and we don't claim it meets the target (after dark it doesn't). Two honest limits: the target is an average, so a single reading can be further off, and while the correction isn't being applied (for about a minute if it fails to load, say), the number you see is the raw one.
You can see the exact correction — and that cross-check — on the Live Accuracy page.
Why a U.S. answer key helps the whole world
Here's the part we're genuinely proud of: the correction is keyed on sun and wind, and those exist everywhere. A lesson learned at U.S. reference cities — "in this much sun, with this little wind, the model tends to run a bit warm" — applies just as well in Lagos or Lahore. So an answer key that only exists for the United States quietly makes the reading better worldwide, including the many places where no one publishes a WBGT to check against.
What about the "in the shade" reading?
Everything above is about the number we lead with: the in-sun WBGT. That's the one with an answer key — we check it against the NWS reference and correct it from the data.
The shade toggle is different, and we want to be plain about it. There is no standard formula for "WBGT in the shade" — the official definition is where you put the instrument — and there is no independent shade reference to grade ourselves against. So the shade figure is a modeled best case: we remove the direct sun and assume the surroundings sit at air temperature. Real shade — dappled light, or shade next to hot pavement or a sunlit wall — reads higher. Treat it as "roughly how much cooler shade would be," not a validated reading, and let the in-sun number govern anything that has to be safe.
That's the whole idea: treat accuracy as a living system, not a fixed claim. Measure it in public, learn where it's off, correct it carefully, and prove the fix travels.