Brier Score
Install and import#
npm install fintech-algorithmsimport { brierScore } from "fintech-algorithms/model-validation-and-backtesting/classification-and-score-validation/brier-score";Signature#
brierScore(inputs)Averages the weighted squared gap between each predicted probability and its realised label, and compares that average against the prevalence-only baseline to give a skill score.
Parameters#
| Name | Type | Notes |
|---|---|---|
inputs | { records: Array<{ id: string; label: 0 | 1; probability: number; weight?: number; score_available_at?: string; label_available_at?: string }>; evaluation_cutoff?: string } | The forecast population. Every record needs a unique nonempty id, a label that is exactly the number 0 or 1, and a finite probability in [0,1]; weight defaults to 1 and must be positive. If evaluation_cutoff is supplied, any score_available_at or label_available_at on a record is compared against it as a string and must not sort after it.records: nonempty · probability: between 0 and 1 inclusive · weight: positive, default 1 |
Returns#
{ brier_score: number; event_rate: number; baseline_brier: number; brier_skill: number | null; weight_sum: number; record_count: number; state: string }
brier_score is the weighted mean squared error, event_rate the weighted positive share, baseline_brier is event_rate * (1 - event_rate), and brier_skill is 1 - brier_score / baseline_brier or null when the baseline is 0 because the population is single-class. state is probability-evaluated.
Errors#
- When
recordsis absent, not an array, or empty — throws RangeError - When a record
idis missing, empty, or repeats an earlier one — throws RangeError - When a
labelis anything other than the number 0 or 1 — throws RangeError - When a
probabilityfalls outside [0,1] — throws RangeError - When a
weightis zero or negative — throws RangeError - When a record
probabilityorweightis not a finite number — throws TypeError - When
score_available_atorlabel_available_atsorts afterevaluation_cutoff— throws RangeError
Complexity: time O(n),
space O(n).
Worked example#
verified This is the worked example published in the article, replayed by the test suite on every run. The output cannot drift.
Input#
{
"records": [
{
"id": "R01",
"label": 1,
"score": 0.95,
"probability": 0.92,
"weight": 1,
"sector": "Banking",
"country": "Egypt",
"regime": "Expansion",
"score_available_at": "2025-01-01T00:00:00Z",
"label_available_at": "2026-01-01T00:00:00Z"
},
{
"id": "R02",
"label": 0,
"score": 0.9,
"probability": 0.88,
"weight": 1,
"sector": "Insurance",
"country": "Egypt",
"regime": "Expansion",
"score_available_at": "2025-01-01T00:00:00Z",
"label_available_at": "2026-01-01T00:00:00Z"
},
{
"id": "R03",
"label": 1,
"score": 0.9,
"probability": 0.84,
"weight": 1,
"sector": "Markets",
"country": "Saudi Arabia",
"regime": "Expansion",
"score_available_at": "2025-01-01T00:00:00Z",
"label_available_at": "2026-01-01T00:00:00Z"
}
],
"evaluation_cutoff": "2026-06-30T00:00:00Z"
}Call#
brierScore(inputs)Returns#
object with 7 fields: brier_score, event_rate, baseline_brier, brier_skill, weight_sum, record_count, state
{
"brier_score": 0.24828333333333338,
"event_rate": 0.4166666666666667,
"baseline_brier": 0.24305555555555552,
"brier_skill": -0.021508571428571654,
"weight_sum": 24,
"record_count": 24,
"state": "probability-evaluated"
}Other exports#
This module also exports
rocCurveAndRocAuc, precisionRecallCurveAndPrAuc, logLoss, reliabilityDiagramAndExpectedCalibrationError, gainsLiftAndDecileCapture, costSensitiveThresholdOptimization, scoreStabilityAndMigrationMatrix, sliceBasedValidationBySectorCountryAndRegime, rareEventBacktestAndConfidenceBounds, calculate. Every module additionally exports run as an alias of its
primary function, and a meta object carrying its catalog id, domain, family,
shape and article URL.
Diagrams#
Calculation flow#
Brier Score calculation flow
flowchart LR
S1["Validate probability label weight and cutoff"]
S2["Compute weighted event rate"]
S3["Compute each squared probability error"]
S4["Average errors by total weight"]
S5["Compute constantrate baseline"]
S1 --> S2
S2 --> S3
S3 --> S4
S4 --> S5
S5 --> D{"inputs must be probabilities rather than arbitrary scores"}
D --> O["brier_score + diagnostics"]
O --> A["Audit: For binary outcomes and probabilities in 01 Brier score is"]
How it works#
This page states the contract — how to call it correctly. The article explains the concept: why it works, and where it breaks.
References#
- Revised Guidance on Model Risk Management — Board of Governors of the Federal Reserve System, OCC, and FDIC
- Verification of Forecasts Expressed in Terms of Probability — Glenn W. Brier
- Probability calibration — scikit-learn maintainers
- Evidence boundary