Skip to content

Multiclass API

The generated items below use NumPy-style parsing because this evaluator's public docstrings follow that convention.

Evaluator

imvpy.multi_imv.MulticlassIMV

Multinomial IMV for multi-class classification problems.

This class extends IMV to classification tasks with multiple classes, providing both one-vs-all IMV scores and pairwise IMV confusion matrices.

Parameters:

  • data (DataFrame) –

    The dataset containing features and outcome variable

  • outcome_variable (str) –

    Name of the outcome/target column

  • model_creator (callable) –

    Zero-argument function returning a fresh classifier with fit and predict_proba; fitted models must expose aligned classes_ arrays.

  • n_splits (int, default: 10 ) –

    Number of folds for k-fold cross-validation

  • optional_explanatory_variables (list, default: None ) –

    List of feature column names. If None, uses all columns except outcome

  • random_state (int, default: None ) –

    Random seed for reproducibility

  • stratified (bool, default: False ) –

    False preserves the legacy shuffled KFold behavior. True selects StratifiedKFold, which is the right choice for a new analysis on imbalanced classes and guarantees no fold omits a class. Switching changes the result and must be reported.

  • verbose (bool, default: False ) –

    Print per-fold progress and results.

Source code in src/imvpy/multi_imv/evaluator.py
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
class MulticlassIMV:
    """
    Multinomial IMV for multi-class classification problems.

    This class extends IMV to classification tasks with multiple classes,
    providing both one-vs-all IMV scores and pairwise IMV confusion matrices.

    Parameters
    ----------
    data : pd.DataFrame
        The dataset containing features and outcome variable
    outcome_variable : str
        Name of the outcome/target column
    model_creator : callable
        Zero-argument function returning a fresh classifier with ``fit`` and
        ``predict_proba``; fitted models must expose aligned ``classes_`` arrays.
    n_splits : int, default=10
        Number of folds for k-fold cross-validation
    optional_explanatory_variables : list, optional
        List of feature column names. If None, uses all columns except outcome
    random_state : int, optional
        Random seed for reproducibility
    stratified : bool, default=False
        False preserves the legacy shuffled ``KFold`` behavior. True selects
        ``StratifiedKFold``, which is the right choice for a new analysis on
        imbalanced classes and guarantees no fold omits a class. Switching
        changes the result and must be reported.
    verbose : bool, default=False
        Print per-fold progress and results.
    """

    # Legacy API compatibility while retaining one canonical implementation.
    # These must stay below the docstring: a class body statement placed above a
    # string literal turns that literal into a no-op expression, leaving
    # ``MulticlassIMV.__doc__`` as None.
    ll = staticmethod(ll)
    get_w = staticmethod(get_w)

    def __init__(self, data, outcome_variable, model_creator, n_splits=10, 
                 optional_explanatory_variables=None, random_state=None,
                 stratified=False, verbose=False):
        if not isinstance(data, pd.DataFrame):
            raise TypeError("data must be a pandas DataFrame")
        if outcome_variable not in data:
            raise ValueError(f"outcome variable {outcome_variable!r} is not in data")
        self.data = data.copy()
        self.outcome_variable = outcome_variable
        self.model_creator = model_creator
        self.n_splits = n_splits
        self.optional_explanatory_variables = (
            optional_explanatory_variables 
            if optional_explanatory_variables is not None 
            else data.columns.drop(outcome_variable).tolist()
        )
        self.random_state = random_state
        self.stratified = bool(stratified)
        self.verbose = bool(verbose)

    def _split_indices(self):
        """Return original KFold indices or explicit production stratification."""
        if self.stratified:
            splitter = StratifiedKFold(
                n_splits=self.n_splits, shuffle=True, random_state=self.random_state
            )
            return splitter.split(self.data, self.data[self.outcome_variable])
        splitter = KFold(
            n_splits=self.n_splits, shuffle=True, random_state=self.random_state
        )
        return splitter.split(self.data)

    # Note: ll() and get_w() are imported from imvpy.core.
    # No need to redefine them here - this eliminates code duplication!

    @staticmethod
    def _fold_classes(model_basic, model_enhanced):
        """Column order shared by both fitted models' predict_proba output."""
        basic = np.asarray(model_basic.classes_)
        enhanced = np.asarray(model_enhanced.classes_)
        if not np.array_equal(basic, enhanced):
            raise ValueError(
                "null and enhanced models disagree on class order; their "
                "probability columns cannot be compared"
            )
        return enhanced

    @staticmethod
    def _resolve_classes(data, outcome_variable, p_base, p_enhanced, classes):
        """Label each probability column, validating shapes.

        ``classes`` names the class each column of *p_base*/*p_enhanced* belongs
        to and must be given in column order (i.e. ``model.classes_``). When it
        is omitted the columns are assumed to follow the sorted classes present
        in *data*, which is only correct if this fold contains every class the
        model was trained on.
        """
        p_base = np.asarray(p_base)
        p_enhanced = np.asarray(p_enhanced)
        if p_base.ndim != 2 or p_enhanced.shape != p_base.shape:
            raise ValueError("probability arrays must have matching (samples, classes) shapes")
        if p_base.shape[0] != len(data):
            raise ValueError("probability rows must match the number of observations")
        if classes is None:
            classes = np.sort(data[outcome_variable].unique())
            if p_base.shape[1] != len(classes):
                raise ValueError(
                    f"probability arrays have {p_base.shape[1]} columns but this fold "
                    f"contains {len(classes)} classes; pass classes=model.classes_ so "
                    "columns can be matched to labels, or use stratified folds"
                )
        else:
            classes = np.asarray(classes)
            if p_base.shape[1] != len(classes):
                raise ValueError("classes must name exactly one label per probability column")
        return classes, p_base, p_enhanced

    def multinominal_imv_matrix(self, data, outcome_variable, p_base, p_enhanced, classes=None):
        """
        Calculate pairwise IMV confusion matrix for all class combinations.

        Creates a matrix showing information gain for each pair of classes.
        Element (i,j) represents IMV when discriminating class i from class j.
        Diagonal elements are zero (no discrimination within same class).

        Parameters
        ----------
        data : pd.DataFrame
            Test data containing outcome variable
        outcome_variable : str
            Name of outcome/target column
        p_base : array-like, shape (n_samples, n_classes)
            Predicted probabilities from null model (intercept only)
        p_enhanced : array-like, shape (n_samples, n_classes)
            Predicted probabilities from model with features
        classes : array-like, optional
            Label of each probability column, in column order (``model.classes_``).
            Required whenever *data* may not contain every class the model was
            trained on; otherwise columns are matched to the wrong labels.

        Returns
        -------
        pd.DataFrame, shape (n_classes, n_classes)
            Pairwise IMV matrix indexed by *classes*. Pairs for which this fold
            holds no samples of one or both classes are NaN.


        Process:
            1. For each class pair (i, j) where i ≠ j:
            2. Filter data to only samples of class i or j
            3. Normalize probabilities for binary comparison
            4. Compute IMV comparing base vs enhanced models
            5. Store IMV(i vs j) at position [i, j]

        Interpretation:
            - High IMV(i,j): Features help distinguish class i from class j
            - Low IMV(i,j): Little information gain for this class pair
            - Matrix is exactly symmetric: IMV(i,j) == IMV(j,i). Pairwise
              renormalization gives p_j = 1 - p_i, and swapping i and j also
              flips the label, so ll() is unchanged because it is invariant
              under (y, p) -> (1-y, 1-p). This is unlike the ablation matrix,
              where the two models have independent likelihoods.
        """
        outcomes, p_base, p_enhanced = self._resolve_classes(
            data, outcome_variable, p_base, p_enhanced, classes
        )
        imv_mat = np.zeros((len(outcomes), len(outcomes)))
        labels = data[outcome_variable].to_numpy()

        for i, outcome_i in enumerate(outcomes):
            for j, outcome_j in enumerate(outcomes):
                if i == j:
                    continue
                # Filter data to only include these two classes. Column i always
                # belongs to outcome_i because both are positions in *outcomes*.
                mask = (labels == outcome_i) | (labels == outcome_j)
                y_tmp = np.where(labels[mask] == outcome_i, 1, 0)
                if y_tmp.size == 0 or y_tmp.min() == y_tmp.max():
                    # One of the two classes is absent from this fold, so the
                    # pair carries no discrimination signal here.
                    imv_mat[i, j] = np.nan
                    continue

                # Normalize probabilities for binary comparison
                denom_b = np.sum(p_base[mask][:, [i, j]], axis=1)
                denom_e = np.sum(p_enhanced[mask][:, [i, j]], axis=1)
                if np.any(denom_b <= 0) or np.any(denom_e <= 0):
                    raise ValueError(
                        f"pair ({outcome_i!r}, {outcome_j!r}) has zero combined "
                        "probability mass; cannot renormalize"
                    )
                p_b = p_base[mask, i] / denom_b
                p_e = p_enhanced[mask, i] / denom_e

                # Use shared ll() and get_w() from core module
                a0 = ll(y_tmp, p_b)
                a1 = ll(y_tmp, p_e)

                p0 = get_w(a0)
                p1 = get_w(a1)
                ew = (p1 - p0) / p0

                imv_mat[i, j] = ew

        imv_mat = pd.DataFrame(imv_mat, index=outcomes, columns=outcomes)
        return imv_mat

    def one_vs_all_single_fold(self, data, outcome_variable, p_base, p_enhanced, classes=None):
        """
        Calculate one-vs-all IMV for a single fold.

        For each class, calculates IMV treating it as positive class vs all others.

        Parameters
        ----------
        data : pd.DataFrame
            Data with outcome variable
        outcome_variable : str
            Name of outcome column
        p_base : array-like, shape (n_samples, n_classes)
            Predicted probabilities from base model
        p_enhanced : array-like, shape (n_samples, n_classes)
            Predicted probabilities from enhanced model
        classes : array-like, optional
            Label of each probability column, in column order (``model.classes_``).
            Required whenever *data* may not contain every class the model was
            trained on; otherwise columns are matched to the wrong labels.

        Returns
        -------
        pd.DataFrame
            DataFrame with class labels and their IMV scores. Classes absent
            from this fold score NaN, since one-vs-rest is unmeasurable without
            positives.
        """
        outcomes, p_base, p_enhanced = self._resolve_classes(
            data, outcome_variable, p_base, p_enhanced, classes
        )
        labels = data[outcome_variable].to_numpy()
        imv = []

        for class_index, outcome in enumerate(outcomes):
            # Binary encoding: current class vs all others. class_index indexes
            # both *outcomes* and the probability columns, so they stay aligned.
            y_tmp = np.where(labels == outcome, 1, 0)
            if y_tmp.max() == 0:
                imv.append(np.nan)
                continue

            p_b = p_base[:, class_index]
            p_e = p_enhanced[:, class_index]

            # Use shared ll() and get_w() from core module
            a0 = ll(y_tmp, p_b)
            a1 = ll(y_tmp, p_e)

            p0 = get_w(a0)
            p1 = get_w(a1)

            ew = (p1 - p0) / p0
            imv.append(ew)

        ova_df = pd.DataFrame({'class': outcomes, 'imv': imv})
        return ova_df

    def k_fold_one_vs_all(self):
        """
        Perform k-fold cross-validation for one-vs-all IMV evaluation.

        Trains null and enhanced models across k folds and computes IMV for each
        class treated as positive vs all other classes combined as negative.

        Returns
        -------
        tuple of (imv_results, imv_average)
            imv_results : list of numpy arrays
                IMV scores for each fold, shape (n_folds, n_classes)
            imv_average : numpy array
                Mean IMV scores across all folds, shape (n_classes,)

        Process:
            1. Split data into k folds
            2. For each fold:
               - Train null model (constant only) on train set
               - Train enhanced model (with features) on train set
               - Compute one-vs-all IMV on test set
            3. Average IMV scores across all folds

        Side Effects:
            Prints IMV results for all folds when ``verbose=True``.

        Example Output:
            IMV results across folds: [[0.15, 0.23, 0.18], [0.14, 0.21, 0.19], ...]

        Note:
            Uses random_state for reproducible fold splits if specified.
        """
        imv_results = []

        for train_index, test_index in self._split_indices():
            X_train = self.data[self.optional_explanatory_variables].iloc[train_index]
            X_test = self.data[self.optional_explanatory_variables].iloc[test_index]
            y_train = self.data[self.outcome_variable].iloc[train_index]
            y_test = self.data[self.outcome_variable].iloc[test_index]

            # Base model (null model with only constant)
            X_train_constant = np.ones((X_train.shape[0], 1))
            X_test_constant = np.ones((X_test.shape[0], 1))
            model_basic = self.model_creator()
            model_basic.fit(X_train_constant, y_train)
            p_base = model_basic.predict_proba(X_test_constant)

            # Enhanced model with features
            model_enhanced = self.model_creator()
            model_enhanced.fit(X_train, y_train)
            p_enhanced = model_enhanced.predict_proba(X_test)

            # Calculate IMV for this fold
            test_data = X_test.copy()
            test_data[self.outcome_variable] = y_test
            ova_df = self.one_vs_all_single_fold(
                test_data, self.outcome_variable, p_base, p_enhanced,
                classes=self._fold_classes(model_basic, model_enhanced),
            )
            imv_results.append(ova_df['imv'].values)

        # Folds that lack a class contribute NaN for it rather than a wrong value.
        imv_average = _nanmean(np.array(imv_results), axis=0)
        if self.verbose:
            print(f"IMV results across folds: {imv_results}")

        return imv_results, imv_average

    def k_fold_imv_matrix(self):
        """
        Perform k-fold cross-validation for pairwise IMV confusion matrix.

        Trains models across k folds and computes pairwise IMV matrices,
        then averages to get stable estimates of class discrimination ability.

        Returns
        -------
        tuple of (imv_matrices_list, imv_matrices_average)
            imv_matrices_list : list of numpy arrays
                IMV confusion matrix for each fold, shape (n_folds, n_classes, n_classes)
            imv_matrices_average : pd.DataFrame
                Average IMV matrix across folds, shape (n_classes, n_classes)
                with class labels as index/columns

        Process:
            1. Split data into k folds
            2. For each fold:
               - Train null model (constant only) on train set
               - Train enhanced model (with features) on train set  
               - Compute pairwise IMV matrix on test set
            3. Average matrices element-wise across all folds

        Side Effects:
            Prints the averaged IMV matrix when ``verbose=True``.

        Example Output:
            Average IMV Matrix:
                   0      1      2
            0  0.000  0.145  0.123
            1  0.145  0.000  0.098
            2  0.123  0.098  0.000

        Note:
            - Diagonal elements are always 0 (no self-discrimination)
            - Matrix is exactly symmetric (see multinominal_imv_matrix)
            - Uses random_state for reproducible splits
        """
        imv_matrices_list = []

        for train_index, test_index in self._split_indices():
            X_train = self.data[self.optional_explanatory_variables].iloc[train_index]
            X_test = self.data[self.optional_explanatory_variables].iloc[test_index]
            y_train = self.data[self.outcome_variable].iloc[train_index]
            y_test = self.data[self.outcome_variable].iloc[test_index]

            # Base model (null model with only constant)
            X_train_constant = np.ones((X_train.shape[0], 1))
            X_test_constant = np.ones((X_test.shape[0], 1))
            model_basic = self.model_creator()
            model_basic.fit(X_train_constant, y_train)
            p_base = model_basic.predict_proba(X_test_constant)

            # Enhanced model with features
            model_enhanced = self.model_creator()
            model_enhanced.fit(X_train, y_train)
            p_enhanced = model_enhanced.predict_proba(X_test)

            # Calculate IMV matrix for this fold
            test_data = self.data.iloc[test_index].copy()
            test_data[self.outcome_variable] = y_test
            imv_matrix = self.multinominal_imv_matrix(
                test_data, self.outcome_variable, p_base, p_enhanced,
                classes=self._fold_classes(model_basic, model_enhanced),
            )
            imv_matrices_list.append(imv_matrix.values)

        # Folds that lack a class contribute NaN for it rather than a wrong value.
        imv_matrices_average = _nanmean(np.stack(imv_matrices_list), axis=0)
        imv_matrices_average = pd.DataFrame(
            imv_matrices_average, 
            index=imv_matrix.index, 
            columns=imv_matrix.columns
        )
        if self.verbose:
            print("Average IMV Matrix:")
            print(imv_matrices_average)

        return imv_matrices_list, imv_matrices_average

    def multinomial_IMV_heatmap(self, imv_matrix, ax=None, figsize=(6, 6)):
        """
        Create heatmap visualization of pairwise IMV confusion matrix.

        Visualizes the IMV matrix as a colored heatmap with annotations.
        Useful for identifying which class pairs are most distinguishable.

        Parameters
        ----------
        imv_matrix : pd.DataFrame or array-like, shape (n_classes, n_classes)
            Pairwise IMV matrix to visualize (from k_fold_imv_matrix)
        ax : matplotlib.axes.Axes, optional
            Existing axis to plot on. If None, creates new figure.
        figsize : tuple, default=(6, 6)
            Figure size (width, height) if creating new figure

        Returns
        -------
        tuple or matplotlib.axes.Axes
            - If ax=None: Returns (fig, ax) tuple
            - If ax provided: Returns ax

        Visualization Details:
            - Color scheme: canonical IMV navy-to-red publication palette
            - Annotations: IMV values displayed in cells (3 decimal places)
            - Labels: "Outcome1", "Outcome2", etc. for rows and columns
            - Diagonal: Always 0 (no self-discrimination)

        Example:
            >>> imv_matrices, imv_avg = evaluator.k_fold_imv_matrix()
            >>> fig, ax = evaluator.multinomial_IMV_heatmap(imv_avg)
            >>> plt.tight_layout()
            >>> plt.show()
        """
        data = np.asarray(imv_matrix, dtype=float)
        if data.ndim != 2:
            raise ValueError("imv_matrix must be two-dimensional")
        labels = [f"Outcome{index + 1}" for index in range(data.shape[1])]
        return plot_imv_heatmap(
            data,
            ax=ax,
            figsize=figsize,
            title="IMV Confusion Matrix",
            labels=labels,
            colorbar_label="Pairwise IMV",
        )

    def multinomial_IMV_boxplot(self, imv_results, figsize=(6, 6), ax=None):
        """
        Create boxplot visualization of one-vs-all IMV distribution across folds.

        Shows the distribution and variability of IMV scores for each class
        across k-fold cross-validation. Useful for assessing stability and
        comparing class-wise information gain.

        Parameters
        ----------
        imv_results : list of arrays, shape (n_folds, n_classes)
            One-vs-all IMV results from k_fold_one_vs_all()
        figsize : tuple, default=(6, 6)
            Figure size (width, height) if creating new figure
        ax : matplotlib.axes.Axes, optional
            Existing axis to plot on. If None, creates new figure.

        Returns
        -------
        tuple or matplotlib.axes.Axes
            - If ax=None: Returns (fig, ax) tuple
            - If ax provided: Returns ax

        Visualization Details:
            - One boxplot per class showing distribution across folds
            - Box: Interquartile range (IQR) Q1-Q3
            - Whiskers: Extend to 1.5*IQR or data extremes
            - Median line: Shown within each box
            - Labels: "Outcome1", "Outcome2", etc. for each class

        Example:
            >>> imv_results, imv_avg = evaluator.k_fold_one_vs_all()
            >>> fig, ax = evaluator.multinomial_IMV_boxplot(imv_results)
            >>> plt.tight_layout()
            >>> plt.show()

        Note:
            Narrow boxes indicate stable IMV across folds.
            Wide boxes suggest fold-dependent performance.
        """
        data_matrix = np.asarray(imv_results, dtype=float)
        if data_matrix.ndim != 2:
            raise ValueError("imv_results must be a folds-by-classes matrix")
        labels = [f"Outcome{index + 1}" for index in range(data_matrix.shape[1])]
        return plot_ova_boxplot(
            data_matrix,
            ax=ax,
            figsize=figsize,
            labels=labels,
            title="Multinomial IMV across Different Outcomes",
            ylabel="IMV Value",
        )

Pairwise single-fold matrix

imvpy.multi_imv.MulticlassIMV.multinominal_imv_matrix

multinominal_imv_matrix(data, outcome_variable, p_base, p_enhanced, classes=None)

Calculate pairwise IMV confusion matrix for all class combinations.

Creates a matrix showing information gain for each pair of classes. Element (i,j) represents IMV when discriminating class i from class j. Diagonal elements are zero (no discrimination within same class).

Parameters:

  • data (DataFrame) –

    Test data containing outcome variable

  • outcome_variable (str) –

    Name of outcome/target column

  • p_base ((array - like, shape(n_samples, n_classes))) –

    Predicted probabilities from null model (intercept only)

  • p_enhanced ((array - like, shape(n_samples, n_classes))) –

    Predicted probabilities from model with features

  • classes (array - like, default: None ) –

    Label of each probability column, in column order (model.classes_). Required whenever data may not contain every class the model was trained on; otherwise columns are matched to the wrong labels.

Returns:

  • (DataFrame, shape(n_classes, n_classes)) –

    Pairwise IMV matrix indexed by classes. Pairs for which this fold holds no samples of one or both classes are NaN.

  • Process –
    1. For each class pair (i, j) where i ≠ j:
    2. Filter data to only samples of class i or j
    3. Normalize probabilities for binary comparison
    4. Compute IMV comparing base vs enhanced models
    5. Store IMV(i vs j) at position [i, j]
  • Interpretation –
    • High IMV(i,j): Features help distinguish class i from class j
    • Low IMV(i,j): Little information gain for this class pair
    • Matrix is exactly symmetric: IMV(i,j) == IMV(j,i). Pairwise renormalization gives p_j = 1 - p_i, and swapping i and j also flips the label, so ll() is unchanged because it is invariant under (y, p) -> (1-y, 1-p). This is unlike the ablation matrix, where the two models have independent likelihoods.
Source code in src/imvpy/multi_imv/evaluator.py
def multinominal_imv_matrix(self, data, outcome_variable, p_base, p_enhanced, classes=None):
    """
    Calculate pairwise IMV confusion matrix for all class combinations.

    Creates a matrix showing information gain for each pair of classes.
    Element (i,j) represents IMV when discriminating class i from class j.
    Diagonal elements are zero (no discrimination within same class).

    Parameters
    ----------
    data : pd.DataFrame
        Test data containing outcome variable
    outcome_variable : str
        Name of outcome/target column
    p_base : array-like, shape (n_samples, n_classes)
        Predicted probabilities from null model (intercept only)
    p_enhanced : array-like, shape (n_samples, n_classes)
        Predicted probabilities from model with features
    classes : array-like, optional
        Label of each probability column, in column order (``model.classes_``).
        Required whenever *data* may not contain every class the model was
        trained on; otherwise columns are matched to the wrong labels.

    Returns
    -------
    pd.DataFrame, shape (n_classes, n_classes)
        Pairwise IMV matrix indexed by *classes*. Pairs for which this fold
        holds no samples of one or both classes are NaN.


    Process:
        1. For each class pair (i, j) where i ≠ j:
        2. Filter data to only samples of class i or j
        3. Normalize probabilities for binary comparison
        4. Compute IMV comparing base vs enhanced models
        5. Store IMV(i vs j) at position [i, j]

    Interpretation:
        - High IMV(i,j): Features help distinguish class i from class j
        - Low IMV(i,j): Little information gain for this class pair
        - Matrix is exactly symmetric: IMV(i,j) == IMV(j,i). Pairwise
          renormalization gives p_j = 1 - p_i, and swapping i and j also
          flips the label, so ll() is unchanged because it is invariant
          under (y, p) -> (1-y, 1-p). This is unlike the ablation matrix,
          where the two models have independent likelihoods.
    """
    outcomes, p_base, p_enhanced = self._resolve_classes(
        data, outcome_variable, p_base, p_enhanced, classes
    )
    imv_mat = np.zeros((len(outcomes), len(outcomes)))
    labels = data[outcome_variable].to_numpy()

    for i, outcome_i in enumerate(outcomes):
        for j, outcome_j in enumerate(outcomes):
            if i == j:
                continue
            # Filter data to only include these two classes. Column i always
            # belongs to outcome_i because both are positions in *outcomes*.
            mask = (labels == outcome_i) | (labels == outcome_j)
            y_tmp = np.where(labels[mask] == outcome_i, 1, 0)
            if y_tmp.size == 0 or y_tmp.min() == y_tmp.max():
                # One of the two classes is absent from this fold, so the
                # pair carries no discrimination signal here.
                imv_mat[i, j] = np.nan
                continue

            # Normalize probabilities for binary comparison
            denom_b = np.sum(p_base[mask][:, [i, j]], axis=1)
            denom_e = np.sum(p_enhanced[mask][:, [i, j]], axis=1)
            if np.any(denom_b <= 0) or np.any(denom_e <= 0):
                raise ValueError(
                    f"pair ({outcome_i!r}, {outcome_j!r}) has zero combined "
                    "probability mass; cannot renormalize"
                )
            p_b = p_base[mask, i] / denom_b
            p_e = p_enhanced[mask, i] / denom_e

            # Use shared ll() and get_w() from core module
            a0 = ll(y_tmp, p_b)
            a1 = ll(y_tmp, p_e)

            p0 = get_w(a0)
            p1 = get_w(a1)
            ew = (p1 - p0) / p0

            imv_mat[i, j] = ew

    imv_mat = pd.DataFrame(imv_mat, index=outcomes, columns=outcomes)
    return imv_mat

One-vs-rest single fold

imvpy.multi_imv.MulticlassIMV.one_vs_all_single_fold

one_vs_all_single_fold(data, outcome_variable, p_base, p_enhanced, classes=None)

Calculate one-vs-all IMV for a single fold.

For each class, calculates IMV treating it as positive class vs all others.

Parameters:

  • data (DataFrame) –

    Data with outcome variable

  • outcome_variable (str) –

    Name of outcome column

  • p_base ((array - like, shape(n_samples, n_classes))) –

    Predicted probabilities from base model

  • p_enhanced ((array - like, shape(n_samples, n_classes))) –

    Predicted probabilities from enhanced model

  • classes (array - like, default: None ) –

    Label of each probability column, in column order (model.classes_). Required whenever data may not contain every class the model was trained on; otherwise columns are matched to the wrong labels.

Returns:

  • DataFrame –

    DataFrame with class labels and their IMV scores. Classes absent from this fold score NaN, since one-vs-rest is unmeasurable without positives.

Source code in src/imvpy/multi_imv/evaluator.py
def one_vs_all_single_fold(self, data, outcome_variable, p_base, p_enhanced, classes=None):
    """
    Calculate one-vs-all IMV for a single fold.

    For each class, calculates IMV treating it as positive class vs all others.

    Parameters
    ----------
    data : pd.DataFrame
        Data with outcome variable
    outcome_variable : str
        Name of outcome column
    p_base : array-like, shape (n_samples, n_classes)
        Predicted probabilities from base model
    p_enhanced : array-like, shape (n_samples, n_classes)
        Predicted probabilities from enhanced model
    classes : array-like, optional
        Label of each probability column, in column order (``model.classes_``).
        Required whenever *data* may not contain every class the model was
        trained on; otherwise columns are matched to the wrong labels.

    Returns
    -------
    pd.DataFrame
        DataFrame with class labels and their IMV scores. Classes absent
        from this fold score NaN, since one-vs-rest is unmeasurable without
        positives.
    """
    outcomes, p_base, p_enhanced = self._resolve_classes(
        data, outcome_variable, p_base, p_enhanced, classes
    )
    labels = data[outcome_variable].to_numpy()
    imv = []

    for class_index, outcome in enumerate(outcomes):
        # Binary encoding: current class vs all others. class_index indexes
        # both *outcomes* and the probability columns, so they stay aligned.
        y_tmp = np.where(labels == outcome, 1, 0)
        if y_tmp.max() == 0:
            imv.append(np.nan)
            continue

        p_b = p_base[:, class_index]
        p_e = p_enhanced[:, class_index]

        # Use shared ll() and get_w() from core module
        a0 = ll(y_tmp, p_b)
        a1 = ll(y_tmp, p_e)

        p0 = get_w(a0)
        p1 = get_w(a1)

        ew = (p1 - p0) / p0
        imv.append(ew)

    ova_df = pd.DataFrame({'class': outcomes, 'imv': imv})
    return ova_df

Cross-validated one-vs-rest

imvpy.multi_imv.MulticlassIMV.k_fold_one_vs_all

k_fold_one_vs_all()

Perform k-fold cross-validation for one-vs-all IMV evaluation.

Trains null and enhanced models across k folds and computes IMV for each class treated as positive vs all other classes combined as negative.

Returns

tuple of (imv_results, imv_average) imv_results : list of numpy arrays IMV scores for each fold, shape (n_folds, n_classes) imv_average : numpy array Mean IMV scores across all folds, shape (n_classes,)

Process
  1. Split data into k folds
  2. For each fold:
  3. Train null model (constant only) on train set
  4. Train enhanced model (with features) on train set
  5. Compute one-vs-all IMV on test set
  6. Average IMV scores across all folds
Side Effects

Prints IMV results for all folds when verbose=True.

Example Output

IMV results across folds: [[0.15, 0.23, 0.18], [0.14, 0.21, 0.19], ...]

Note

Uses random_state for reproducible fold splits if specified.

Source code in src/imvpy/multi_imv/evaluator.py
def k_fold_one_vs_all(self):
    """
    Perform k-fold cross-validation for one-vs-all IMV evaluation.

    Trains null and enhanced models across k folds and computes IMV for each
    class treated as positive vs all other classes combined as negative.

    Returns
    -------
    tuple of (imv_results, imv_average)
        imv_results : list of numpy arrays
            IMV scores for each fold, shape (n_folds, n_classes)
        imv_average : numpy array
            Mean IMV scores across all folds, shape (n_classes,)

    Process:
        1. Split data into k folds
        2. For each fold:
           - Train null model (constant only) on train set
           - Train enhanced model (with features) on train set
           - Compute one-vs-all IMV on test set
        3. Average IMV scores across all folds

    Side Effects:
        Prints IMV results for all folds when ``verbose=True``.

    Example Output:
        IMV results across folds: [[0.15, 0.23, 0.18], [0.14, 0.21, 0.19], ...]

    Note:
        Uses random_state for reproducible fold splits if specified.
    """
    imv_results = []

    for train_index, test_index in self._split_indices():
        X_train = self.data[self.optional_explanatory_variables].iloc[train_index]
        X_test = self.data[self.optional_explanatory_variables].iloc[test_index]
        y_train = self.data[self.outcome_variable].iloc[train_index]
        y_test = self.data[self.outcome_variable].iloc[test_index]

        # Base model (null model with only constant)
        X_train_constant = np.ones((X_train.shape[0], 1))
        X_test_constant = np.ones((X_test.shape[0], 1))
        model_basic = self.model_creator()
        model_basic.fit(X_train_constant, y_train)
        p_base = model_basic.predict_proba(X_test_constant)

        # Enhanced model with features
        model_enhanced = self.model_creator()
        model_enhanced.fit(X_train, y_train)
        p_enhanced = model_enhanced.predict_proba(X_test)

        # Calculate IMV for this fold
        test_data = X_test.copy()
        test_data[self.outcome_variable] = y_test
        ova_df = self.one_vs_all_single_fold(
            test_data, self.outcome_variable, p_base, p_enhanced,
            classes=self._fold_classes(model_basic, model_enhanced),
        )
        imv_results.append(ova_df['imv'].values)

    # Folds that lack a class contribute NaN for it rather than a wrong value.
    imv_average = _nanmean(np.array(imv_results), axis=0)
    if self.verbose:
        print(f"IMV results across folds: {imv_results}")

    return imv_results, imv_average

Cross-validated pairwise matrix

imvpy.multi_imv.MulticlassIMV.k_fold_imv_matrix

k_fold_imv_matrix()

Perform k-fold cross-validation for pairwise IMV confusion matrix.

Trains models across k folds and computes pairwise IMV matrices, then averages to get stable estimates of class discrimination ability.

Returns

tuple of (imv_matrices_list, imv_matrices_average) imv_matrices_list : list of numpy arrays IMV confusion matrix for each fold, shape (n_folds, n_classes, n_classes) imv_matrices_average : pd.DataFrame Average IMV matrix across folds, shape (n_classes, n_classes) with class labels as index/columns

Process
  1. Split data into k folds
  2. For each fold:
  3. Train null model (constant only) on train set
  4. Train enhanced model (with features) on train set
  5. Compute pairwise IMV matrix on test set
  6. Average matrices element-wise across all folds
Side Effects

Prints the averaged IMV matrix when verbose=True.

Example Output

Average IMV Matrix: 0 1 2 0 0.000 0.145 0.123 1 0.145 0.000 0.098 2 0.123 0.098 0.000

Note
  • Diagonal elements are always 0 (no self-discrimination)
  • Matrix is exactly symmetric (see multinominal_imv_matrix)
  • Uses random_state for reproducible splits
Source code in src/imvpy/multi_imv/evaluator.py
def k_fold_imv_matrix(self):
    """
    Perform k-fold cross-validation for pairwise IMV confusion matrix.

    Trains models across k folds and computes pairwise IMV matrices,
    then averages to get stable estimates of class discrimination ability.

    Returns
    -------
    tuple of (imv_matrices_list, imv_matrices_average)
        imv_matrices_list : list of numpy arrays
            IMV confusion matrix for each fold, shape (n_folds, n_classes, n_classes)
        imv_matrices_average : pd.DataFrame
            Average IMV matrix across folds, shape (n_classes, n_classes)
            with class labels as index/columns

    Process:
        1. Split data into k folds
        2. For each fold:
           - Train null model (constant only) on train set
           - Train enhanced model (with features) on train set  
           - Compute pairwise IMV matrix on test set
        3. Average matrices element-wise across all folds

    Side Effects:
        Prints the averaged IMV matrix when ``verbose=True``.

    Example Output:
        Average IMV Matrix:
               0      1      2
        0  0.000  0.145  0.123
        1  0.145  0.000  0.098
        2  0.123  0.098  0.000

    Note:
        - Diagonal elements are always 0 (no self-discrimination)
        - Matrix is exactly symmetric (see multinominal_imv_matrix)
        - Uses random_state for reproducible splits
    """
    imv_matrices_list = []

    for train_index, test_index in self._split_indices():
        X_train = self.data[self.optional_explanatory_variables].iloc[train_index]
        X_test = self.data[self.optional_explanatory_variables].iloc[test_index]
        y_train = self.data[self.outcome_variable].iloc[train_index]
        y_test = self.data[self.outcome_variable].iloc[test_index]

        # Base model (null model with only constant)
        X_train_constant = np.ones((X_train.shape[0], 1))
        X_test_constant = np.ones((X_test.shape[0], 1))
        model_basic = self.model_creator()
        model_basic.fit(X_train_constant, y_train)
        p_base = model_basic.predict_proba(X_test_constant)

        # Enhanced model with features
        model_enhanced = self.model_creator()
        model_enhanced.fit(X_train, y_train)
        p_enhanced = model_enhanced.predict_proba(X_test)

        # Calculate IMV matrix for this fold
        test_data = self.data.iloc[test_index].copy()
        test_data[self.outcome_variable] = y_test
        imv_matrix = self.multinominal_imv_matrix(
            test_data, self.outcome_variable, p_base, p_enhanced,
            classes=self._fold_classes(model_basic, model_enhanced),
        )
        imv_matrices_list.append(imv_matrix.values)

    # Folds that lack a class contribute NaN for it rather than a wrong value.
    imv_matrices_average = _nanmean(np.stack(imv_matrices_list), axis=0)
    imv_matrices_average = pd.DataFrame(
        imv_matrices_average, 
        index=imv_matrix.index, 
        columns=imv_matrix.columns
    )
    if self.verbose:
        print("Average IMV Matrix:")
        print(imv_matrices_average)

    return imv_matrices_list, imv_matrices_average

Evaluator plotting methods

imvpy.multi_imv.MulticlassIMV.multinomial_IMV_heatmap

multinomial_IMV_heatmap(imv_matrix, ax=None, figsize=(6, 6))

Create heatmap visualization of pairwise IMV confusion matrix.

Visualizes the IMV matrix as a colored heatmap with annotations. Useful for identifying which class pairs are most distinguishable.

Parameters

imv_matrix : pd.DataFrame or array-like, shape (n_classes, n_classes) Pairwise IMV matrix to visualize (from k_fold_imv_matrix) ax : matplotlib.axes.Axes, optional Existing axis to plot on. If None, creates new figure. figsize : tuple, default=(6, 6) Figure size (width, height) if creating new figure

Returns

tuple or matplotlib.axes.Axes - If ax=None: Returns (fig, ax) tuple - If ax provided: Returns ax

Visualization Details
  • Color scheme: canonical IMV navy-to-red publication palette
  • Annotations: IMV values displayed in cells (3 decimal places)
  • Labels: "Outcome1", "Outcome2", etc. for rows and columns
  • Diagonal: Always 0 (no self-discrimination)
Example

imv_matrices, imv_avg = evaluator.k_fold_imv_matrix() fig, ax = evaluator.multinomial_IMV_heatmap(imv_avg) plt.tight_layout() plt.show()

Source code in src/imvpy/multi_imv/evaluator.py
def multinomial_IMV_heatmap(self, imv_matrix, ax=None, figsize=(6, 6)):
    """
    Create heatmap visualization of pairwise IMV confusion matrix.

    Visualizes the IMV matrix as a colored heatmap with annotations.
    Useful for identifying which class pairs are most distinguishable.

    Parameters
    ----------
    imv_matrix : pd.DataFrame or array-like, shape (n_classes, n_classes)
        Pairwise IMV matrix to visualize (from k_fold_imv_matrix)
    ax : matplotlib.axes.Axes, optional
        Existing axis to plot on. If None, creates new figure.
    figsize : tuple, default=(6, 6)
        Figure size (width, height) if creating new figure

    Returns
    -------
    tuple or matplotlib.axes.Axes
        - If ax=None: Returns (fig, ax) tuple
        - If ax provided: Returns ax

    Visualization Details:
        - Color scheme: canonical IMV navy-to-red publication palette
        - Annotations: IMV values displayed in cells (3 decimal places)
        - Labels: "Outcome1", "Outcome2", etc. for rows and columns
        - Diagonal: Always 0 (no self-discrimination)

    Example:
        >>> imv_matrices, imv_avg = evaluator.k_fold_imv_matrix()
        >>> fig, ax = evaluator.multinomial_IMV_heatmap(imv_avg)
        >>> plt.tight_layout()
        >>> plt.show()
    """
    data = np.asarray(imv_matrix, dtype=float)
    if data.ndim != 2:
        raise ValueError("imv_matrix must be two-dimensional")
    labels = [f"Outcome{index + 1}" for index in range(data.shape[1])]
    return plot_imv_heatmap(
        data,
        ax=ax,
        figsize=figsize,
        title="IMV Confusion Matrix",
        labels=labels,
        colorbar_label="Pairwise IMV",
    )

imvpy.multi_imv.MulticlassIMV.multinomial_IMV_boxplot

multinomial_IMV_boxplot(imv_results, figsize=(6, 6), ax=None)

Create boxplot visualization of one-vs-all IMV distribution across folds.

Shows the distribution and variability of IMV scores for each class across k-fold cross-validation. Useful for assessing stability and comparing class-wise information gain.

Parameters

imv_results : list of arrays, shape (n_folds, n_classes) One-vs-all IMV results from k_fold_one_vs_all() figsize : tuple, default=(6, 6) Figure size (width, height) if creating new figure ax : matplotlib.axes.Axes, optional Existing axis to plot on. If None, creates new figure.

Returns

tuple or matplotlib.axes.Axes - If ax=None: Returns (fig, ax) tuple - If ax provided: Returns ax

Visualization Details
  • One boxplot per class showing distribution across folds
  • Box: Interquartile range (IQR) Q1-Q3
  • Whiskers: Extend to 1.5*IQR or data extremes
  • Median line: Shown within each box
  • Labels: "Outcome1", "Outcome2", etc. for each class
Example

imv_results, imv_avg = evaluator.k_fold_one_vs_all() fig, ax = evaluator.multinomial_IMV_boxplot(imv_results) plt.tight_layout() plt.show()

Note

Narrow boxes indicate stable IMV across folds. Wide boxes suggest fold-dependent performance.

Source code in src/imvpy/multi_imv/evaluator.py
def multinomial_IMV_boxplot(self, imv_results, figsize=(6, 6), ax=None):
    """
    Create boxplot visualization of one-vs-all IMV distribution across folds.

    Shows the distribution and variability of IMV scores for each class
    across k-fold cross-validation. Useful for assessing stability and
    comparing class-wise information gain.

    Parameters
    ----------
    imv_results : list of arrays, shape (n_folds, n_classes)
        One-vs-all IMV results from k_fold_one_vs_all()
    figsize : tuple, default=(6, 6)
        Figure size (width, height) if creating new figure
    ax : matplotlib.axes.Axes, optional
        Existing axis to plot on. If None, creates new figure.

    Returns
    -------
    tuple or matplotlib.axes.Axes
        - If ax=None: Returns (fig, ax) tuple
        - If ax provided: Returns ax

    Visualization Details:
        - One boxplot per class showing distribution across folds
        - Box: Interquartile range (IQR) Q1-Q3
        - Whiskers: Extend to 1.5*IQR or data extremes
        - Median line: Shown within each box
        - Labels: "Outcome1", "Outcome2", etc. for each class

    Example:
        >>> imv_results, imv_avg = evaluator.k_fold_one_vs_all()
        >>> fig, ax = evaluator.multinomial_IMV_boxplot(imv_results)
        >>> plt.tight_layout()
        >>> plt.show()

    Note:
        Narrow boxes indicate stable IMV across folds.
        Wide boxes suggest fold-dependent performance.
    """
    data_matrix = np.asarray(imv_results, dtype=float)
    if data_matrix.ndim != 2:
        raise ValueError("imv_results must be a folds-by-classes matrix")
    labels = [f"Outcome{index + 1}" for index in range(data_matrix.shape[1])]
    return plot_ova_boxplot(
        data_matrix,
        ax=ax,
        figsize=figsize,
        labels=labels,
        title="Multinomial IMV across Different Outcomes",
        ylabel="IMV Value",
    )

MultinomialIMV is an identity alias of MulticlassIMV. multinominal_imv_matrix intentionally retains its historical misspelling.