Skip to content

Expose instance-level metric contributions - #452

Open
jakbia wants to merge 1 commit into
NannyML:mainfrom
jakbia:instance_lev_metrics
Open

jakbia wants to merge 1 commit into
NannyML:mainfrom
jakbia:instance_lev_metrics

Conversation

@jakbia

@jakbia jakbia commented Aug 27, 2026

Copy link
Copy Markdown

Summary

  • add opt-in instance-level contributions to supported OSS calculators and estimators
  • preserve the existing chunk result, thresholds, alerts, confidence bounds, sampling errors, and default execution path
  • expose a separate, merge-friendly instance dataframe containing only row/chunk identifiers and contribution columns
  • support calculating all configured contributions or a selective subset

Interface

result = calculator.calculate(
    analysis_data,
    include_instance_metrics=['accuracy', 'roc_auc'],
    identifier_column_name='prediction_id',
)

chunk_results = result.to_df()
instance_results = result.to_instance_df()

The same arguments are available on estimate. include_instance_metrics accepts False (the default), True, or a list of metric names. to_df() is unchanged. to_instance_df() contains the identifier, chunk metadata, and one <component>_contribution column per selected result component. The original input can be joined back using the identifier.

Per-metric values

Contribution direction follows the metric itself: high accuracy contributions are favourable, while high loss, data-quality, reconstruction-error, or domain-classifier contributions are unfavourable.

Realized performance

Metric Instance value and aggregation
MAE Absolute error; chunk value is the mean.
MAPE Absolute percentage error using the existing scikit-learn epsilon denominator; chunk value is the mean.
MSE Squared error; chunk value is the mean.
MSLE Squared log1p error; chunk value is the mean.
RMSE Squared error; chunk value is sqrt(mean).
RMSLE Squared log1p error; chunk value is sqrt(mean).
Accuracy Correctness indicator; chunk value is the mean.
Precision Correctness on predicted-positive rows and NaN elsewhere; nanmean is precision.
Recall Correctness on actual-positive rows and NaN elsewhere; nanmean is recall.
Specificity Correctness on actual-negative rows and NaN elsewhere; nanmean is specificity.
F1 TP = 1, FP/FN = 0, TN = NaN; nanmean gives Jaccard J, with F1 = 2J / (1 + J).
ROC AUC Per-row Mann-Whitney concordance against the opposite class, with half credit for ties; the mean is ROC AUC.
Average precision Positive baseline credit plus ranking-inversion penalties shared by their endpoints; contributions sum to AP and badly ranked negatives can be negative.
Confusion matrix One contribution column per cell, adjusted for the configured normalization; each column sums to its chunk cell.
Business value Row payoff, adjusted for per_prediction normalization when configured; contributions sum to chunk business value.

For multiclass classification, accuracy keeps the 0/1 mean. ROC AUC and average precision are calculated one-vs-rest and macro-averaged. Precision, recall, specificity, and F1 allocate each class's macro term to participating rows and sum to the chunk metric. Confusion-matrix cells and business value use the same additive allocation as binary classification.

Confidence-based performance estimation

CBPE uses calibrated probabilities for expected class membership and the existing uncalibrated scores for ranking metrics.

Metric Instance value and aggregation
Accuracy Expected correctness; the mean is estimated accuracy.
Precision Calibrated positive probability on predicted-positive rows and NaN elsewhere; the mean is estimated precision. A no-predicted-positive chunk contributes zeros to match current behavior.
Recall Signed positive probability: positive for predicted positives and negative for predicted negatives. This ranks false-negative pressure but is not a direct reducer for the ratio.
Specificity Signed expected-negative probability: positive for predicted negatives and negative for predicted positives. This is diagnostic rather than a direct reducer.
F1 Signed positive probability by predicted class. This preserves useful row ordering but is not a direct reducer for the nonlinear ratio.
ROC AUC Probabilistic Mann-Whitney concordance. A constant reconciliation offset preserves row ordering while making the mean match the existing rounded, tie-stable estimate.
Average precision Expected positive credit plus symmetric expected inversion penalties, reconciled to the existing cumulative-count rounding; contributions sum to the estimate.
Confusion matrix Expected cell mass adjusted for normalization; each component sums to its estimated cell.
Business value Expected payoff under calibrated class probabilities; contributions sum to the estimate.

Multiclass CBPE applies the same calculations one-vs-rest and uses the existing macro reduction. Accuracy and ROC AUC aggregate by mean; precision, recall, specificity, F1, average precision, confusion-matrix cells, and business value use additive macro allocations.

Direct loss estimation

DLE exposes its existing per-row predicted loss after clipping negative predictions to zero. MAE, MAPE, MSE, and MSLE use the mean; RMSE and RMSLE use the square root of the mean. No additional loss models are fitted.

Data quality and summary statistics

Metric Instance value and aggregation
Missing values 0/1 missing indicator; mean when normalized, otherwise sum.
Unseen values 0/1 unseen indicator; mean when normalized, otherwise sum.
Out-of-range values 0/1 out-of-range indicator; mean when normalized, otherwise sum.
Average Raw row value; chunk reducer remains the average.
Sum Raw row value; chunk reducer remains the sum.
Median Raw row value; chunk reducer remains the median.
Standard deviation Squared deviation from the chunk mean; chunk value is sqrt(sum / (valid_count - 1)).
Row count One per row; contributions sum to the count.

Multivariate drift

Metric Instance value and aggregation
Data reconstruction Per-row Euclidean PCA reconstruction norm; the mean is chunk reconstruction error.
Domain classifier Analysis-row out-of-fold classifier score converted to Mann-Whitney concordance against the out-of-fold reference scores; the mean is domain-classifier AUROC. Existing CV fits are reused.

Scope and compatibility

  • instance contributions are disabled by default, so the existing schema and calculation cost remain unchanged
  • validation, chunking, thresholds, alerts, sampling errors, and failure behavior follow the existing calculator lifecycle
  • unsupported calculators reject instance metrics explicitly
  • KS, Jensen-Shannon, Hellinger, Wasserstein, L-infinity, chi-square, and continuous/categorical distribution calculators remain chunk-only because meaningful row influence would require expensive leave-one-out work
  • the implementation covers OSS only; no Pro changes are included

Verification

  • focused instance-metric suite: 39 passed
  • full suite on this commit: 1001 passed, 40 skipped, with 7 CBPE timestamp fixture failures reproduced unchanged on origin/main
  • selected lint, compile, and documentation builds completed successfully

@jakbia
jakbia requested review from nikml and nnansters as code owners August 27, 2026 07:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants