Scoring
How much to trust an individual score
Overall model performance is the same for every repository. To say how much to trust an individual score, each prediction also carries an evidence-support value built from three repository-specific factors — one weak factor pulls the whole thing down. It measures how much observable signal backs the score, not statistical confidence.
Data coverage
Share of expected maintenance signals actually observed; missing ones are filled with the training-cohort average and add no evidence.
In-distribution fit
Share of observed signals that sit within the normal range the model was trained on, rather than an extreme it never saw.
Calibration support
How many past examples backed the band this prediction falls in — more support means a steadier estimate.
A separate margin shows how decisive a call is — how far the score sits from the boundary between buckets — kept apart from evidence support because a well-supported score can still land close to the line.