When a probability scoring system says an event has a 72% chance of occurring, that number means something only if the system has established a track record. Specifically: have events it scored near 72% historically occurred about 72% of the time? If yes, the system is calibrated. If not, the number is precise but not reliable, which is arguably worse than a rough estimate that accurately represents its own uncertainty.
Calibration is what separates a probability instrument from a probability display. Building systems that produce numbers is straightforward. Building systems that produce accurate numbers, where the numerical scale corresponds to observed frequencies, is the harder problem, and it is the one that most AI-driven probability tools have yet to seriously confront.
What Calibration Means, Precisely
Calibration, in the formal sense used in forecasting research, describes the relationship between predicted probability and observed outcome frequency. A perfectly calibrated forecaster, one that assigned probabilities of 0.60 to many events, would see those events occur 60% of the time. At 0.80, they would occur 80% of the time. The diagonal line on a calibration plot represents this ideal relationship.
Most AI scoring systems deviate from this line in systematic ways. Overconfidence is the most common pattern: a system that places events at 80% probability when the true rate is closer to 65%. Underconfidence is the mirror problem: probabilities clustered near 50% for events that have much cleaner resolution signals. Both patterns make the output less useful to analysts who need to make decisions based on the numbers.
The Brier score provides a summary calibration metric. It is the mean squared difference between predicted probability and binary outcome (1 for occurrence, 0 for non-occurrence), averaged across events. Lower Brier scores indicate better calibration. A Brier score of 0.25 corresponds roughly to random guessing on binary events. Anything consistently below 0.20 for a large event sample suggests the scoring system is adding genuine predictive information.
Brier scores alone are insufficient, though. Resolution, a separate component of forecasting quality, measures whether the probability estimates are actually different for events that occur versus events that do not. A system with good calibration but poor resolution is placing events at a well-calibrated 55% when it should be placing them at 30% or 75% based on available evidence. Calibration and resolution are related but distinct properties, and a complete evaluation of a probability tool requires examining both.
Why AI Systems Miss Calibration
AI language models and many AI-adjacent scoring systems are optimized primarily for coherence, relevance, and fluency in their outputs, not for calibration against observed event frequencies. When an AI system produces a probability estimate, it is typically drawing on pattern associations in training data rather than applying a formal scoring model with documented calibration properties.
This matters in a particular way for geopolitical and political event scoring. The training data for most large language models is heavily weighted toward analysis and commentary written after events occurred, meaning the models have extensive exposure to post-hoc reasoning but limited exposure to the forward-looking probabilistic reasoning that produces calibrated estimates before outcomes are known. The result is a tendency toward confident-sounding outputs that do not reflect genuine probabilistic calibration.
There is also a structural issue with resolution windows. Calibration requires tracking predictions against outcomes, which requires that the event resolution period be complete. If a scoring system claims to assess near-term events but does not systematically track outcomes and update its calibration weights accordingly, it is not using feedback to improve. It is perpetually generating first-draft numbers without the editorial process that would make those numbers reliable.
What Calibration Tracking Actually Requires
To build a calibrated probability scoring system for near-term geopolitical events, several elements have to be in place simultaneously.
First, events must be defined with precise resolution criteria before scoring begins. An event defined as "political instability in region X" cannot be evaluated for calibration because there is no agreed standard for resolution. An event defined as "a specific ministerial position changes hands within 60 days" can be evaluated. The precision of the event definition determines whether calibration measurement is even possible.
Second, the resolution outcome must be recorded consistently and without post-hoc reinterpretation. This is harder than it sounds. There is a natural tendency to reinterpret event boundaries after the fact when outcomes are ambiguous, and that reinterpretation can corrupt a calibration dataset over time. The resolution judgment has to be made against the original event definition without access to hindsight.
Third, calibration analysis must be stratified by event type. A system that is well-calibrated for electoral outcomes may be poorly calibrated for regulatory events, because the information environment and update dynamics differ substantially. Aggregate Brier scores across event types can mask poor performance in specific categories. Useful calibration reporting distinguishes performance by event domain and time horizon.
What Good Calibration Reporting Looks Like for Users
From an analyst's perspective, calibration reporting needs to answer specific questions: how many events has this system scored and resolved in this event category? What is the Brier score for that sample? How does the calibration plot look across the probability range? Are there systematic biases (consistent overconfidence at high probabilities, or underconfidence in the middle)?
A system that reports these numbers transparently, including sample sizes and confidence intervals around the Brier score itself, is making a fundamentally different claim than a system that says "our AI accurately predicts geopolitical events." The first claim is falsifiable and auditable. The second is marketing language.
The sample size caveat matters more than it is usually acknowledged. Calibration analysis on fewer than 100 resolved events per category produces wide uncertainty intervals around Brier scores, meaning the apparent calibration level may not be reliably distinguishable from random chance. Analysts using calibration data to justify reliance on a probability tool should understand what sample sizes underlie the reported metrics.
Calibration in Multi-Source Aggregation Systems
When probability scores are produced by aggregating multiple source types, rather than by a single model, calibration becomes more complex and more interesting. Each source category may have a distinct calibration profile for a given event type: government document feeds may be well-calibrated for regulatory events but lag on electoral outcomes; expert analyst networks may be well-calibrated for policy directional events but noisy on timing.
A well-designed aggregation system uses these distinct calibration profiles to weight source contributions dynamically by event type. If news wire coverage has historically produced overconfident readings for political transition events in a particular region, the aggregation model should discount wire contributions for that event category, even when those contributions are consistent with other source types. This is the mechanism by which calibration history feeds back into the scoring engine itself rather than remaining a static audit metric.
The implication for analysts is that a probability reading from a well-designed aggregation system carries embedded calibration information even when that information is not surfaced explicitly. The number reflects both the raw signal from sources and the adjustment for how reliable those sources have historically been for this event type. Understanding this layered structure changes how analysts should interpret the outputs rather than treating them as raw machine-generated probability estimates.
The Trust Question
We are sometimes asked whether Cade Market can provide its own calibration statistics. The honest answer is that calibration analysis requires a sufficient resolved event sample per category, and we are not going to report numbers based on a sample size that cannot support reliable inference. We track this data continuously and will publish calibration analysis when the sample is large enough to say something meaningful.
What we can say now is that we built the resolution tracking infrastructure before we built the scoring engine. Event definitions are recorded before scoring begins. Outcomes are recorded against those definitions after resolution without retroactive adjustment. The calibration dataset is accumulating in a form that will support analysis. That is the prerequisite work. The calibration numbers will follow from it, and we will publish them when they are defensible rather than when they are merely available.