fintech-algorithms
Using a coding agent? Give it the skill: npx skills add IslamBaraka90/Fintech-Algorithms-Library What it does →

Rare-Event Backtest and Confidence Bounds

Install and import#

bash
npm install fintech-algorithms
ts
import { rareEventBacktestAndConfidenceBounds } from "fintech-algorithms/model-validation-and-backtesting/classification-and-score-validation/rare-event-backtest-and-confidence-bounds";

Signature#

rareEventBacktestAndConfidenceBounds(inputs)

Builds a Wilson score interval around an observed event count out of a number of trials and reports whether the model's expected probability falls inside that interval, below it, or above it.

Parameters#

NameTypeNotes
inputs{ observed_events: number; trials: number; expected_probability: number; confidence_level: 0.90 | 0.95 | 0.99; sampling_assumption: "independent-bernoulli-approximation"; horizon_start: string; horizon_end: string; knowledge_cutoff: string }The backtest tally and the interval settings. observed_events and trials must both be integers with 0 <= observed_events <= trials and trials > 0. expected_probability is the model's predicted event rate and must be a finite number in [0,1]. confidence_level must be exactly 0.90, 0.95, or 0.99, each mapping to a fixed z value. sampling_assumption must be the string independent-bernoulli-approximation, which is the only assumption this implementation supports. horizon_start, horizon_end and knowledge_cutoff are required ISO-8601 UTC instants of the form YYYY-MM-DDTHH:MM:SSZ; the horizon must have positive length and must close no later than the knowledge cutoff.
observed_events: integer, 0 or more, no larger than trials · trials: integer greater than 0 · expected_probability: between 0 and 1 inclusive · confidence_level: 0.90, 0.95, or 0.99

Returns#

{ observed_events: number; trials: number; observed_rate: number; expected_probability: number; expected_count: number; confidence_level: number; z_value: number; wilson_center: number; wilson_half_width: number; wilson_lower: number; wilson_upper: number; consistency: "expected-below-interval" | "inside-interval" | "expected-above-interval"; sampling_assumption: string; state: string }

The tally is echoed back with observed_rate and expected_count alongside the interval: z_value, wilson_center, wilson_half_width, and the clamped wilson_lower and wilson_upper bounds, with the lower bound pinned to 0 when no event was observed and the upper bound pinned to 1 when every trial was an event. consistency places expected_probability against those bounds, and state is rare-event-evaluated.

Errors#

  • When observed_events or trials is not a finite number — throws TypeError
  • When observed_events or trials is not an integer — throws RangeError
  • When trials is 0 or negative, or observed_events is negative or larger than trials — throws RangeError
  • When expected_probability is not a finite number — throws TypeError
  • When expected_probability falls outside [0,1] — throws RangeError
  • When confidence_level is not one of 0.90, 0.95, or 0.99 — throws RangeError
  • When sampling_assumption is not independent-bernoulli-approximation — throws RangeError
  • When horizon_start, horizon_end or knowledge_cutoff is missing or not a string — throws TypeError
  • When a horizon or cutoff value is not an ISO-8601 UTC instant of the form YYYY-MM-DDTHH:MM:SSZ — throws RangeError
  • When horizon_start is not strictly before horizon_end — throws RangeError
  • When horizon_end is later than knowledge_cutoff, so the tally would include future evidence — throws RangeError

Complexity: time O(1), space O(1).

Worked example#

verified This is the worked example published in the article, replayed by the test suite on every run. The output cannot drift.

Input#

inputs
{
  "observed_events": 3,
  "trials": 1000,
  "expected_probability": 0.002,
  "confidence_level": 0.95,
  "horizon_start": "2025-01-01T00:00:00Z",
  "horizon_end": "2026-01-01T00:00:00Z",
  "knowledge_cutoff": "2026-06-30T00:00:00Z",
  "sampling_assumption": "independent-bernoulli-approximation"
}

Call#

rareEventBacktestAndConfidenceBounds(inputs)

Returns#

object with 14 fields: observed_events, trials, observed_rate, expected_probability, expected_count, confidence_level, z_value, wilson_center, …

{
  "observed_events": 3,
  "trials": 1000,
  "observed_rate": 0.003,
  "expected_probability": 0.002,
  "expected_count": 2,
  "confidence_level": 0.95,
  "z_value": 1.959963984540054,
  "wilson_center": 0.004901898967320896,
  "wilson_half_width": 0.0038811150861822775,
  "wilson_lower": 0.0010207838811386186,
  "wilson_upper": 0.008783014053503173,
  "consistency": "inside-interval",
  "sampling_assumption": "independent-bernoulli-approximation",
  "state": "rare-event-evaluated"
}

Other exports#

This module also exports rocCurveAndRocAuc, precisionRecallCurveAndPrAuc, brierScore, logLoss, reliabilityDiagramAndExpectedCalibrationError, gainsLiftAndDecileCapture, costSensitiveThresholdOptimization, scoreStabilityAndMigrationMatrix, sliceBasedValidationBySectorCountryAndRegime, calculate. Every module additionally exports run as an alias of its primary function, and a meta object carrying its catalog id, domain, family, shape and article URL.

Diagrams#

Rare-Event Backtest and Confidence Bounds — article hero
Rare-Event Backtest and Confidence Bounds — decision boundaries
Rare-Event Backtest and Confidence Bounds — method selection
Rare-Event Backtest and Confidence Bounds — system map

Calculation flow#

Rare-Event Backtest and Confidence Bounds calculation flow
flowchart LR
    S1["Validate count trials expected probability confidence "]
    S2["Calculate observed rate and expected count"]
    S3["Map confidence level to z"]
    S4["Calculate Wilson center and halfwidth"]
    S5["Clamp bounds to 01"]
    S1 --> S2
    S2 --> S3
    S3 --> S4
    S4 --> S5
    S5 --> D{"matured trials and the independentBernoulli approximation "}
    D --> O["observed_rate + diagnostics"]
    O --> A["Audit: Wilson bounds lie in 01 lower  observed rate  upper includ"]

How it works#

This page states the contract — how to call it correctly. The article explains the concept: why it works, and where it breaks.

Read the article →

References#

The rest of the Classification and Score Validation family#