OSS Risk Radar

Scoring

How much to trust an individual score

Overall model performance is the same for every repository. To say how much to trust an individual score, each prediction also carries an evidence-support value built from three repository-specific factors — one weak factor pulls the whole thing down. It measures how much observable signal backs the score, not statistical confidence.

Data coverage

Share of expected maintenance signals actually observed; missing ones are filled with the training-cohort average and add no evidence.

In-distribution fit

Share of observed signals that sit within the normal range the model was trained on, rather than an extreme it never saw.

Calibration support

How many past examples backed the band this prediction falls in — more support means a steadier estimate.

A separate margin shows how decisive a call is — how far the score sits from the boundary between buckets — kept apart from evidence support because a well-supported score can still land close to the line.