fintech-algorithms
Using a coding agent? Give it the skill: npx skills add IslamBaraka90/Fintech-Algorithms-Library What it does →

Precision-Recall Curve and PR-AUC

Install and import#

bash
npm install fintech-algorithms
ts
import { precisionRecallCurveAndPrAuc } from "fintech-algorithms/model-validation-and-backtesting/classification-and-score-validation/precision-recall-curve-and-pr-auc";

Signature#

precisionRecallCurveAndPrAuc(inputs)

Sweeps a weighted record ledger from the highest score down and reports precision and recall at every distinct score, together with the average precision obtained by summing precision over each recall increment.

Parameters#

NameTypeNotes
inputs{ records: Array<{ id: string; label: 0 | 1; score: number; weight?: number; score_available_at?: string; label_available_at?: string }>; evaluation_cutoff?: string }The scored population. Every record needs a unique nonempty id, a label that is exactly the number 0 or 1, and a finite score; weight defaults to 1 and must be positive. If evaluation_cutoff is supplied, any score_available_at or label_available_at on a record is compared against it as a string and must not sort after it.
records: nonempty · label: 0 or 1 · weight: positive, default 1

Returns#

{ points: Array<{ threshold: number | null; true_positive: number; false_positive: number; precision: number; recall: number }>; pr_auc_average_precision: number; baseline_prevalence: number; positive_weight: number; negative_weight: number; tie_group_count: number; integration_rule: string; state: string }

points opens with the classify-nothing point (threshold: null, precision: 1, recall: 0) and then carries one point per distinct score. pr_auc_average_precision is the recall-increment sum, baseline_prevalence is the positive share of total weight, integration_rule names the rule used, and state is ranking-evaluated.

Errors#

  • When records is absent, not an array, or empty — throws RangeError
  • When a record id is missing, empty, or repeats an earlier one — throws RangeError
  • When a label is anything other than the number 0 or 1 — throws RangeError
  • When a weight is zero or negative — throws RangeError
  • When a record score or weight is not a finite number — throws TypeError
  • When score_available_at or label_available_at sorts after evaluation_cutoff — throws RangeError
  • When no record carries positive weight on the positive class — throws RangeError

Complexity: time O(n log n), space O(n).

Worked example#

verified This is the worked example published in the article, replayed by the test suite on every run. The output cannot drift.

Input#

inputs
{
  "records": [
    {
      "id": "R01",
      "label": 1,
      "score": 0.95,
      "probability": 0.92,
      "weight": 1,
      "sector": "Banking",
      "country": "Egypt",
      "regime": "Expansion",
      "score_available_at": "2025-01-01T00:00:00Z",
      "label_available_at": "2026-01-01T00:00:00Z"
    },
    {
      "id": "R02",
      "label": 0,
      "score": 0.9,
      "probability": 0.88,
      "weight": 1,
      "sector": "Insurance",
      "country": "Egypt",
      "regime": "Expansion",
      "score_available_at": "2025-01-01T00:00:00Z",
      "label_available_at": "2026-01-01T00:00:00Z"
    },
    {
      "id": "R03",
      "label": 1,
      "score": 0.9,
      "probability": 0.84,
      "weight": 1,
      "sector": "Markets",
      "country": "Saudi Arabia",
      "regime": "Expansion",
      "score_available_at": "2025-01-01T00:00:00Z",
      "label_available_at": "2026-01-01T00:00:00Z"
    }
  ],
  "evaluation_cutoff": "2026-06-30T00:00:00Z"
}

Call#

precisionRecallCurveAndPrAuc(inputs)

Returns#

object with 8 fields: points, pr_auc_average_precision, baseline_prevalence, positive_weight, negative_weight, tie_group_count, integration_rule, state

{
  "points": [
    {
      "threshold": null,
      "true_positive": 0,
      "false_positive": 0,
      "precision": 1,
      "recall": 0
    },
    {
      "threshold": 0.95,
      "true_positive": 1,
      "false_positive": 0,
      "precision": 1,
      "recall": 0.1
    },
    {
      "threshold": 0.9,
      "true_positive": 2,
      "false_positive": 1,
      "precision": 0.6666666666666666,
      "recall": 0.2
    }
  ],
  "pr_auc_average_precision": 0.6220479082321188,
  "baseline_prevalence": 0.4166666666666667,
  "positive_weight": 10,
  "negative_weight": 14,
  "tie_group_count": 22,
  "integration_rule": "average-precision-right-step",
  "state": "ranking-evaluated"
}

Other exports#

This module also exports rocCurveAndRocAuc, brierScore, logLoss, reliabilityDiagramAndExpectedCalibrationError, gainsLiftAndDecileCapture, costSensitiveThresholdOptimization, scoreStabilityAndMigrationMatrix, sliceBasedValidationBySectorCountryAndRegime, rareEventBacktestAndConfidenceBounds, calculate. Every module additionally exports run as an alias of its primary function, and a meta object carrying its catalog id, domain, family, shape and article URL.

Diagrams#

Precision-Recall Curve and PR-AUC — article hero
Precision-Recall Curve and PR-AUC — decision boundaries
Precision-Recall Curve and PR-AUC — method selection
Precision-Recall Curve and PR-AUC — system map

Calculation flow#

Precision-Recall Curve and PR-AUC calculation flow
flowchart LR
    S1["Validate labels scores weights and cutoff"]
    S2["Group equal scores descending"]
    S3["Start at recall zero and precision one"]
    S4["Update TP and FP by group"]
    S5["Calculate precision and recall"]
    S1 --> S2
    S2 --> S3
    S3 --> S4
    S4 --> S5
    S5 --> D{"PR area must retain the declared averageprecision integrat"}
    D --> O["pr_auc_average_precision + diagnostics"]
    O --> A["Audit: Recall starts at 0 ends at 1 and never decreases"]

How it works#

This page states the contract — how to call it correctly. The article explains the concept: why it works, and where it breaks.

Read the article →

References#

The rest of the Classification and Score Validation family#