The CountSenate control · D 18% / R 82%Median seats 49D36 calls locked0 gradedBrier —
sleyor

Issue One · Vol. I — opinion that owes you a score

← Front page
The CountThe Count7 min read

Unbiased is not a claim. It is a score.

Every outlet in the country says it is fair, and almost none of them can be checked. A forecast can be. That is the whole reason this desk exists.

By Matt Cranford

One board can be checked. The other asks to be believed. Only one of them can ever be wrong out loud.The house drawing

There is a sentence that appears, in some form, on the about page of nearly every publication in the English-speaking world: we are fair, we are impartial, we go where the facts lead. It is almost always sincere and it is almost always useless, because it is unfalsifiable. Nobody has ever devised a way for a reader to check a masthead's interior life. You are asked to take the most important claim a news organisation makes entirely on trust, which is exactly the arrangement that trust in news has spent thirty years failing to survive.

A forecast is different, and the difference is the only interesting thing about forecasting. When a desk says a thing has a seventy percent chance of happening, it has written a cheque against a future that has not been negotiated with. The world then cashes it in public. Do that a few hundred times and something remarkable becomes available: not an assurance of honesty, but a measurement of it.

This is what calibration means, and it is simpler than it sounds. Take every event you gave a seventy percent chance. If roughly seventy percent of them happened, you were calibrated. If ninety percent happened, you were needlessly timid. If forty percent happened, you were fooling yourself or your audience, and the arithmetic says so regardless of how anyone felt about the coverage. There is no rhetorical move available. A forecaster cannot argue with their own scorecard, which is why so few of them publish one.

A number is not automatically honest, and this is where most data journalism goes wrong. Precision is a costume. A model can be scrupulously computed and still be a laundering operation for its author's assumptions: the choice of which polls to weight, which fundamentals to trust, when to intervene by hand. The remedy is not to avoid modelling. It is to publish the method before the result, log every override as an override, and keep the record up in the years the record is embarrassing. If the weights are only revealed after the outcome is known, what you were shown was not a method. It was a story about a number.

The map is where the lying is easiest. Colour a country by land area and you hand empty acreage the loudest voice in the picture, which produces the two most familiar and most misleading images in American politics — a sea of red that under-counts cities, and a coastal blue that under-counts everyone else. Both sides have used it and both sides know better. This desk uses equal tiles, one square per state, weight written in the square rather than drawn in the shape. It is less pretty and it is not arguable.

Every number carries its interval, or it does not run. A margin quoted without its uncertainty is a guess wearing a suit. Most of the public bitterness about forecasting after a surprising night is really an argument about a range that was published, technically, in grey type at the bottom. If a race sits inside the error, the honest rendering is grey — a real toss-up, not one quietly leaned in the direction the desk expects, and not a coin flip dressed as insight either.

Commentary never touches the model. The writing describes what the numbers do. It does not nudge the numbers toward what would be more satisfying to write about. This sounds obvious and is violated constantly, usually not by fraud but by drift: an input gets adjusted the week it starts producing an uncomfortable answer. So the adjustments are logged, dated, on the page, with the reasoning attached.

None of that makes this desk unbiased in the sense the word is usually meant. Every choice in a model is a judgement, and judgements come from somewhere; Sleyor's disposition is stated openly on the Standard page rather than hidden behind a neutral voice. What the scoring does is narrower and much more valuable: it makes the bias measurable. A lens you can see is a lens you can correct for. A Brier score does not care what anyone at this desk hopes will happen in November.

The board is live now, and it runs on certified returns rather than on anything this desk invented: every state margin, every vote total and every electoral vote on the page comes from the state canvasses, cited with the date they were retrieved. What is still a harness is the forecast layer — the simulated distribution, the tipping-point ranking, the calibration record — and it is labelled as a harness at the top of the page in the same size type as everything else, because a probability is a number and the Standard forbids publishing a number a reader cannot trace. It would have been trivially easy to seed those panels with plausible-looking values and let the design imply authority. That is the exact failure this desk was built to argue against.

When the model lands, the harness banner comes down and the grading begins, permanently. Ask us in a year how the calls went. The answer will already be on the page, and we will not have been the ones deciding what it says.


What works
  • Judge any forecaster by their calibration chart, not their reputation. If they do not publish one, that is the finding.
  • Ask when the method was released. Before the event is a method; after is a narrative.
  • Distrust any election map coloured by area. Tile cartograms and per-capita shading are widely available and everyone in the business knows it.
  • Treat a margin without an interval as unreported. The range is the reporting; the point estimate is the headline.

Every Sleyor piece ends here, per the standard. A critique without a working alternative doesn't run.

Also in this issue