The central problem in intelligence aggregation for probability scoring is not information collection. It is source weighting. When multiple sources are reporting on the same event with different signal directions, or when some sources are silent while others are active, the scoring engine has to decide how much to trust each input. That decision either improves calibration or degrades it, and the difference is not always visible until you look at resolution history.
There is a tempting shortcut: treat all sources equally, or weight by volume. More reporting from more sources drives the score up. Silence drives it down. This approach is implementable and produces outputs that look reasonable. It also produces systematically miscalibrated scores for any event type where source coverage quality is uneven, which is most event types.
Why Equal Weighting Fails
Equal source weighting fails because sources have different base rates of relevance and different calibration histories by event category. A news wire that covers global financial markets very well may be nearly useless as a leading indicator for second-tier regulatory events in emerging markets. A regional government document repository may be highly predictive for those same regulatory events and unreliable for anything involving private corporate decisions.
When you apply equal weights across these sources on a regulatory event question, the financial wire's volume of background coverage effectively dilutes the relevant signal from the government repository. The score converges toward the noisy broad-coverage source rather than the precise but narrow coverage source. The result is a probability estimate that underweights the most relevant evidence.
The failure mode is not random. It is systematic in the direction of regressing toward whatever sources have the most coverage density. For most event categories, the sources with the highest coverage density are the general-purpose news wires, which are designed to cover everything and are therefore not calibrated for any specific event type. Equal weighting effectively maximizes the contribution of the least specialized source category.
Calibration Weights and What They Represent
A calibration weight is a learned parameter that encodes a source's historical predictive accuracy for a specific event category. A source that has historically been a reliable leading indicator for legislative regulatory outcomes in a given region receives a higher weight on events in that category. A source with poor historical calibration on that category receives a lower weight, regardless of its coverage volume or general reputation.
Calibration weights are not fixed at source setup. They are updated as events resolve. Every resolution of an event that a source contributed signal to provides a data point on that source's calibration for that event category. The weight adjusts accordingly, with the adjustment rate governed by the learning rate parameter for the relevant category.
This means a new source enters the scoring system with a conservative starting weight, reflecting genuine uncertainty about its calibration, and earns higher weighting through demonstrated predictive accuracy on resolved events. A source that performs well in its first months of coverage gains influence; one that performs poorly does not. The system is self-correcting in a way that static weights are not.
A necessary condition for calibration weights to function is that events actually resolve. An event that never closes, because its resolution condition was too vague or its timeframe was too long, provides no calibration signal. This is one of the reasons why event definition discipline matters for the scoring system: vague events that never clearly resolve prevent the system from learning what worked and what did not.
Source Divergence as Signal
When sources agree, the scoring is relatively straightforward: convergent signals are aggregated with their respective calibration weights. When sources diverge significantly, the right response is not to average the signals; it is to widen the confidence interval and flag the divergence.
Source divergence is itself informative. If government document signals are pointing toward low probability of a regulatory action while field intelligence signals are pointing toward high probability, the divergence likely indicates that there is genuine uncertainty in the environment, not a technical aggregation problem. One of two things is probably true: either the formal process indicators are lagging behind private deliberations that field sources are observing, or the field sources are receiving misleading signals about actual regulatory intent. Both interpretations are relevant to the analyst; neither is captured by a simple average.
Displaying divergent source signals without resolving them into a single point estimate is a more honest output than forcing convergence. The analyst can look at the source breakdown, see that government documents are at 30% and field intelligence is at 68%, and make their own judgment about which source category they have more reason to trust in this specific context. The scoring system should not be making that judgment for them when the sources are genuinely split.
Source Silence as Negative Evidence
One of the less obvious aspects of source weighting is that source silence is not neutral. When a source that is usually active on a given event category is silent during an observation window, that silence carries negative evidential weight. A government document repository that would normally have published consultation period announcements by a certain date in the regulatory cycle, and has not, updates the probability downward relative to base rate.
Managing source silence as evidence requires knowing what activity level to expect from each source in each event category at each stage of the event lifecycle. That expectation is built from historical coverage patterns, not from prior probability alone. A source that is structurally silent in the early stages of a regulatory process and then active in the formal announcement stage should not have its early silence penalized; that is normal behavior for that source in that event phase. Only deviations from expected coverage patterns carry evidential weight.
This is methodologically more demanding than treating silence as neutral or treating it as uniformly negative. It requires maintaining source-level base rate models for coverage behavior by event category and event phase. The overhead is real. The payoff is that probability estimates are not systematically skewed by source coverage patterns that have nothing to do with underlying event probability.
Correlated Sources and Independence Weighting
A common problem in multi-source aggregation is source correlation: when multiple sources are drawing on the same underlying information, their signals are not independent, and aggregating them as if they were independent overestimates signal strength.
News wire sources are particularly prone to this. A regional wire story that gets picked up by international wire services generates multiple coverage events from a single underlying information source. If all three wire events are treated as independent signals, the score moves substantially on what is actually a single data point. Wire correlation detection is an important component of source weighting architecture, especially during periods of high news volume when amplification effects are strongest.
Government document sources are generally less correlated with each other because they represent distinct institutional outputs. A parliamentary committee publication and a regulatory agency consultation notice are unlikely to be drawing on the same underlying source, even when they concern the same event. The independence assumption is more defensible for documents than for wire coverage, which shapes how we apply independence corrections across source categories.
What Source Weighting Cannot Fix
Source weighting improves calibration but does not eliminate it. There are event categories where no available source has demonstrated historical calibration because the events are rare, because the source environment is thin, or because the resolution conditions are genuinely ambiguous. In these cases, conservative default weights and wide confidence intervals are the appropriate response, not aggressive weighting adjustments that create false precision.
The goal is a scoring system that knows what it knows and represents uncertainty where it exists. Source weighting is the mechanism for encoding what the system has learned about source reliability. It is not a mechanism for compensating for sources that are simply not informative for the question being asked. When the source environment is inadequate for a particular event category, the honest output is a wide interval, not a confidently wrong point estimate.