The core output of Cade Market is a probability score: a number between 0 and 1 that represents our best current estimate of the likelihood that a specific defined event occurs within a specified timeframe. This post explains how that number is produced, what it reflects, and what happens when the sources contributing to it disagree.
We are not trying to produce a comprehensive technical specification here. The goal is to explain the mechanics at a level of detail that helps users understand what the number means, why it moves when it moves, and what the limits of the method are. Users who treat the probability as a black-box output miss most of its value as an analytic instrument.
Starting with a Prior
Every Cade Market event scoring run begins with a base rate prior. The prior is derived from historical event frequencies in the same category and region: how often have comparable events resolved positively, conditional on similar observable conditions at the start of the scoring window?
For regulatory events, the prior draws on historical base rates for formal review initiation in the relevant regulatory jurisdiction and sector. For political transition events, it draws on base rates for leadership changes in comparable political system types under comparable observable conditions. The prior is not a guess. It is an empirical starting point based on what has historically happened in similar circumstances.
The prior matters because it prevents the scoring engine from overreacting to initial signals. An early wire story about a possible regulatory review should not move the probability estimate from a realistic base rate to 80% before any procedural evidence has emerged. The Bayesian update on the prior is proportional to both the strength of the new signal and the weight of the prior evidence. This dampening property keeps the estimates from being volatile in response to early, low-weight signals.
Bayesian Updating Across Source Types
As signals arrive from different source types, the engine applies Bayesian updates that move the posterior probability from the prior toward a new estimate. The update magnitude for each signal depends on three factors: the signal's evidence strength (how specifically does this signal bear on the defined event?), the source type's calibration weight for this event category (how reliably have signals from this source type predicted outcomes in comparable events?), and the signal recency (more recent signals carry higher weight than older signals in the same direction, because they reflect the current state of the information environment rather than an earlier state).
The update formula is applied sequentially as signals arrive, so the current probability reading at any point in time represents the posterior after all signals ingested to date, weighted by the three factors above. When a new signal arrives overnight, the probability moves. The size and direction of the movement reflect the signal's properties against the current posterior, not against the original prior.
This means that the same signal can have a larger or smaller effect depending on where the probability currently sits. A moderately positive signal arriving when the estimate is at 35% has a different update magnitude than the same signal arriving when the estimate is at 75%, because the Bayesian update is proportional to the information gap between the signal and the current state of belief.
How Source Weights Are Determined
Source weights are the most consequential parameter in the scoring model. A source type with a high calibration weight for a given event category will move the probability substantially when it signals. A source type with a low calibration weight will move it little.
Weights are determined by retrospective calibration analysis: for each source type and event category combination, we compare the historical signal pattern at the time events occurred with the signal pattern at the time similar events did not occur. A source type that consistently produced stronger signals in the weeks before events occurred than in the weeks before non-events is assigned a higher weight for that category. A source type with similar signal strength in both cases gets a lower weight.
This calibration is not static. As more events resolve in our dataset, the calibration weights are updated to reflect the growing evidence base. A source type that has performed well in the most recent resolution sample will see its weight rise. One that has underperformed will see its weight fall. The scoring engine is continuously incorporating the calibration feedback from resolved events, not operating on fixed parameters set at initialization.
What Happens When Sources Diverge
Source divergence is one of the most analytically interesting situations the scoring engine encounters. When the majority of source types are signaling in one direction and a minority are signaling in the opposite direction, the model has to handle a genuine disagreement in the evidence.
The first step is to determine whether the diverging signals represent different information or the same information processed differently. A government document feed and a wire service pointing in opposite directions is potential genuine divergence: one is detecting procedural activity that contradicts the public narrative, or vice versa. Wire coverage from multiple outlets pointing in the opposite direction from field intelligence may reflect a different structural situation: the wire coverage may be tracking public statements that diverge from observable ground conditions.
The model handles divergence through weighted aggregation rather than simple majority rule. The direction with higher total weight from participating source types dominates the update, but the update magnitude is damped proportionally to the degree of disagreement. High divergence produces a smaller probability shift than high agreement would produce from the same volume of signals, and it surfaces a divergence flag on the estimate. The flag tells the analyst that the current reading reflects a situation where source types are in disagreement, which warrants direct investigation before acting on the estimate.
Signal Decay and Probability Drift
Signals age. A wire story from 11 days ago that was informative at the time is less informative today because it reflects a state of the world that may have changed. The scoring engine applies a recency decay function to signals as they age, reducing their contribution to the posterior as time passes without confirmation from newer signals in the same direction.
This decay function is one reason why probability estimates sometimes drift toward the prior when no new signals arrive. It is not that the engine is becoming less confident about the event. It is that older evidence is systematically discounted, and without new evidence to replace it, the estimate has less to hold it away from the base rate. An estimate that was at 61% three weeks ago based on heavy signal volume may drift back toward a 48% base rate if the signal volume subsides, reflecting the genuine reduction in evidence supporting the elevated estimate.
Users sometimes interpret this drift as the engine "changing its mind." The more accurate interpretation is that the estimate is tracking the information environment, which has gone quiet. If the event is still live, the quiet period may mean either that the situation has stabilized or that the relevant signals are temporarily absent. The analyst's job is to assess which interpretation is more plausible given contextual knowledge that the engine does not have.
What the Score Does Not Represent
The probability score does not represent a Cade Market opinion about whether the event should occur or is desirable. It is a calibrated estimate of likelihood, not a recommendation. Some events we track have low estimated probabilities that clients hope remain low. Others have high probabilities that clients are hoping to avoid. The engine has no view on desirability; it has a current best estimate of likelihood given available evidence.
The score also does not represent a claim about the outcome. A 78% probability estimate for an event that does not occur is not a failed prediction. It is a statement that, given the information available at the time the score was generated, occurrence was substantially more likely than non-occurrence. The outcome was in the less-likely range. That happens, and over a large sample of events, it should happen about 22% of the time for events scored at 78%. What calibration tracking measures is whether the observed rate of occurrence in that range of estimates is actually close to 22% over time, which is the meaningful performance question rather than any single outcome.