Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Interface
The same arguments are available on
estimate.include_instance_metricsacceptsFalse(the default),True, or a list of metric names.to_df()is unchanged.to_instance_df()contains the identifier, chunk metadata, and one<component>_contributioncolumn per selected result component. The original input can be joined back using the identifier.Per-metric values
Contribution direction follows the metric itself: high accuracy contributions are favourable, while high loss, data-quality, reconstruction-error, or domain-classifier contributions are unfavourable.
Realized performance
log1perror; chunk value is the mean.sqrt(mean).log1perror; chunk value issqrt(mean).NaNelsewhere;nanmeanis precision.NaNelsewhere;nanmeanis recall.NaNelsewhere;nanmeanis specificity.NaN;nanmeangives JaccardJ, with F1 =2J / (1 + J).per_predictionnormalization when configured; contributions sum to chunk business value.For multiclass classification, accuracy keeps the 0/1 mean. ROC AUC and average precision are calculated one-vs-rest and macro-averaged. Precision, recall, specificity, and F1 allocate each class's macro term to participating rows and sum to the chunk metric. Confusion-matrix cells and business value use the same additive allocation as binary classification.
Confidence-based performance estimation
CBPE uses calibrated probabilities for expected class membership and the existing uncalibrated scores for ranking metrics.
NaNelsewhere; the mean is estimated precision. A no-predicted-positive chunk contributes zeros to match current behavior.Multiclass CBPE applies the same calculations one-vs-rest and uses the existing macro reduction. Accuracy and ROC AUC aggregate by mean; precision, recall, specificity, F1, average precision, confusion-matrix cells, and business value use additive macro allocations.
Direct loss estimation
DLE exposes its existing per-row predicted loss after clipping negative predictions to zero. MAE, MAPE, MSE, and MSLE use the mean; RMSE and RMSLE use the square root of the mean. No additional loss models are fitted.
Data quality and summary statistics
sqrt(sum / (valid_count - 1)).Multivariate drift
Scope and compatibility
Verification
origin/main