fintech-algorithms
Using a coding agent? Give it the skill: npx skills add IslamBaraka90/Fintech-Algorithms-Library What it does →

Reliability Diagram and Expected Calibration Error

Install and import#

bash
npm install fintech-algorithms
ts
import { reliabilityDiagramAndExpectedCalibrationError } from "fintech-algorithms/model-validation-and-backtesting/classification-and-score-validation/reliability-diagram-and-expected-calibration-error";

Signature#

reliabilityDiagramAndExpectedCalibrationError(inputs)

Drops probability forecasts into equal-width bins, contrasts each bin's mean forecast with its realised event rate, and aggregates those gaps into the expected, maximum, and signed calibration errors.

Parameters#

NameTypeNotes
inputs{ records: Array<{ id: string; label: 0 | 1; probability: number; weight?: number; score_available_at?: string; label_available_at?: string }>; evaluation_cutoff?: string; bins?: number; bin_strategy?: "uniform" }The forecast population plus the binning choice. Every record needs a unique nonempty id, a label that is exactly the number 0 or 1, and a finite probability in [0,1]; weight defaults to 1 and must be positive. bins defaults to 5 and must be an integer from 2 to 20. bin_strategy defaults to and only accepts uniform. If evaluation_cutoff is supplied, any score_available_at or label_available_at on a record is compared against it as a string and must not sort after it.
bins: integer between 2 and 20, default 5 · bin_strategy: uniform only · probability: between 0 and 1 inclusive

Returns#

{ bins: Array<{ index: number; lower: number; upper: number; right_inclusive: boolean; record_count: number; weight_sum: number; mean_probability: number | null; event_rate: number | null; signed_gap: number | null; absolute_gap: number | null }>; expected_calibration_error: number; maximum_calibration_error: number; signed_calibration_error: number; bin_count: number; weight_sum: number; state: string }

One entry per bin, in ascending probability order, with its edges and the forecast-versus-outcome gap; an empty bin reports zero counts and null for the four statistics. expected_calibration_error is the weight-share average of the absolute gaps, maximum_calibration_error the largest absolute gap, signed_calibration_error the weight-share average of the signed gaps, and state is calibration-evaluated.

Errors#

  • When records is absent, not an array, or empty — throws RangeError
  • When a record id is missing, empty, or repeats an earlier one — throws RangeError
  • When a label is anything other than the number 0 or 1 — throws RangeError
  • When a probability falls outside [0,1] — throws RangeError
  • When a weight is zero or negative — throws RangeError
  • When bins is not an integer — throws RangeError
  • When bins is below 2 or above 20 — throws RangeError
  • When bin_strategy is anything other than uniform — throws RangeError
  • When score_available_at or label_available_at sorts after evaluation_cutoff — throws RangeError

Complexity: time O(n + bins), space O(n).

Worked example#

verified This is the worked example published in the article, replayed by the test suite on every run. The output cannot drift.

Input#

inputs
{
  "records": [
    {
      "id": "R01",
      "label": 1,
      "score": 0.95,
      "probability": 0.92,
      "weight": 1,
      "sector": "Banking",
      "country": "Egypt",
      "regime": "Expansion",
      "score_available_at": "2025-01-01T00:00:00Z",
      "label_available_at": "2026-01-01T00:00:00Z"
    },
    {
      "id": "R02",
      "label": 0,
      "score": 0.9,
      "probability": 0.88,
      "weight": 1,
      "sector": "Insurance",
      "country": "Egypt",
      "regime": "Expansion",
      "score_available_at": "2025-01-01T00:00:00Z",
      "label_available_at": "2026-01-01T00:00:00Z"
    },
    {
      "id": "R03",
      "label": 1,
      "score": 0.9,
      "probability": 0.84,
      "weight": 1,
      "sector": "Markets",
      "country": "Saudi Arabia",
      "regime": "Expansion",
      "score_available_at": "2025-01-01T00:00:00Z",
      "label_available_at": "2026-01-01T00:00:00Z"
    }
  ],
  "evaluation_cutoff": "2026-06-30T00:00:00Z",
  "bins": 5,
  "bin_strategy": "uniform"
}

Call#

reliabilityDiagramAndExpectedCalibrationError(inputs)

Returns#

object with 7 fields: bins, expected_calibration_error, maximum_calibration_error, signed_calibration_error, bin_count, weight_sum, state

{
  "bins": [
    {
      "index": 1,
      "lower": 0,
      "upper": 0.2,
      "right_inclusive": false,
      "record_count": 5,
      "weight_sum": 5,
      "mean_probability": 0.084,
      "event_rate": 0.2,
      "signed_gap": 0.116,
      "absolute_gap": 0.116
    },
    {
      "index": 2,
      "lower": 0.2,
      "upper": 0.4,
      "right_inclusive": false,
      "record_count": 5,
      "weight_sum": 5,
      "mean_probability": 0.27999999999999997,
      "event_rate": 0.4,
      "signed_gap": 0.12000000000000005,
      "absolute_gap": 0.12000000000000005
    },
    {
      "index": 3,
      "lower": 0.4,
      "upper": 0.6,
      "right_inclusive": false,
      "record_count": 5,
      "weight_sum": 5,
      "mean_probability": 0.48,
      "event_rate": 0.4,
      "signed_gap": -0.07999999999999996,
      "absolute_gap": 0.07999999999999996
    }
  ],
  "expected_calibration_error": 0.14250000000000002,
  "maximum_calibration_error": 0.28,
  "signed_calibration_error": -0.044166666666666674,
  "bin_count": 5,
  "weight_sum": 24,
  "state": "calibration-evaluated"
}

Other exports#

This module also exports rocCurveAndRocAuc, precisionRecallCurveAndPrAuc, brierScore, logLoss, gainsLiftAndDecileCapture, costSensitiveThresholdOptimization, scoreStabilityAndMigrationMatrix, sliceBasedValidationBySectorCountryAndRegime, rareEventBacktestAndConfidenceBounds, calculate. Every module additionally exports run as an alias of its primary function, and a meta object carrying its catalog id, domain, family, shape and article URL.

Diagrams#

Reliability Diagram and Expected Calibration Error — article hero
Reliability Diagram and Expected Calibration Error — decision boundaries
Reliability Diagram and Expected Calibration Error — method selection
Reliability Diagram and Expected Calibration Error — system map

Calculation flow#

Reliability Diagram and Expected Calibration Error calculation flow
flowchart LR
    S1["Validate probabilities labels weights bins strategy an"]
    S2["Create equalwidth bin boundaries"]
    S3["Assign each probability with finalbin endpoint handlin"]
    S4["Calculate support mean probability and event rate"]
    S5["Retain empty bins with null diagnostics"]
    S1 --> S2
    S2 --> S3
    S3 --> S4
    S4 --> S5
    S5 --> D{"bin count and endpoint assignment must remain fixed for co"}
    D --> O["expected_calibration_error + diagnostics"]
    O --> A["Audit: Nonempty bin weights sum to total weight and ECE lies in 0"]

How it works#

This page states the contract — how to call it correctly. The article explains the concept: why it works, and where it breaks.

Read the article →

References#

The rest of the Classification and Score Validation family#