Multiple options to load predictions in ValidMind Datasets

This notebook guides you through loading predictions in ValidMind dataset objects using the assign_predictions() function. The function is designed to enable developers to support various way to load predictions in the dataset object so that tests can make use of it.

This guide includes the code required to:

Install the client library

The client library provides Python support for the ValidMind Developer Framework. To install it:

%pip install -q validmind

Initialize the client library

ValidMind generates a unique code snippet for each registered model to connect with your developer environment. You initialize the client library with this code snippet, which ensures that your documentation and tests are uploaded to the correct model when you run the notebook.

Get your code snippet:

  1. In a browser, log into the Platform UI.

  2. In the left sidebar, navigate to Model Inventory and click + Register new model.

  3. Enter the model details and click Continue. (Need more help?)

    For example, to register a model for use with this notebook, select:

    • Documentation template: Binary classification
    • Use case: Marketing/Sales - Attrition/Churn Management

    You can fill in other options according to your preference.

  4. Go to Getting Started and click Copy snippet to clipboard.

Next, replace this placeholder with your own code snippet:

# Replace with your code snippet

import validmind as vm

vm.init(
    api_host="https://api.prod.validmind.ai/api/v1/tracking",
    api_key="...",
    api_secret="...",
    project="...",
)

Preview the documentation template

A template predefines sections for your documentation project and provides a general outline to follow, making the documentation process much easier.

You will upload documentation and test results into this template later on. For now, take a look at the structure that the template provides with the vm.preview_template() function from the ValidMind library and note the empty sections:

vm.preview_template()

Load the sample dataset

The sample dataset used here is provided by the ValidMind library. To be able to use it, you need to import the dataset and load it into a pandas DataFrame, a two-dimensional tabular data structure that makes use of rows and columns:

# Import the sample dataset from the library

from validmind.datasets.classification import customer_churn as demo_dataset

print(
    f"Loaded demo dataset with: \n\n\t• Target column: '{demo_dataset.target_column}' \n\t• Class labels: {demo_dataset.class_labels}"
)

raw_df = demo_dataset.load_data()
raw_df.head()

Prepocess the raw dataset

Preprocessing performs a number of operations to get ready for the subsequent steps:

  • Preprocess the data: Splits the DataFrame (df) into multiple datasets (train_df, validation_df, and test_df) using demo_dataset.preprocess to simplify preprocessing.
  • Separate features and targets: Drops the target column to create feature sets (x_train, x_val) and target sets (y_train, y_val).
train_df, validation_df, test_df = demo_dataset.preprocess(raw_df)
x_train = train_df.drop(demo_dataset.target_column, axis=1)
y_train = train_df[demo_dataset.target_column]
x_val = validation_df.drop(demo_dataset.target_column, axis=1)
y_val = validation_df[demo_dataset.target_column]

Train models for testing

  • Initialize XGBoost and Logistic Regression Classifiers
from sklearn.linear_model import LogisticRegression
import xgboost

%matplotlib inline

xgb = xgboost.XGBClassifier(early_stopping_rounds=10)
xgb.set_params(
    eval_metric=["error", "logloss", "auc"],
)
xgb.fit(
    x_train,
    y_train,
    eval_set=[(x_val, y_val)],
    verbose=False,
)

lr = LogisticRegression(random_state=0)
lr.fit(
    x_train,
    y_train,
)

Initialize ValidMind objects

Initialize the ValidMind models

vm_model_xgb = vm.init_model(
    xgb,
    input_id="xgb",
)
vm_model_lr = vm.init_model(
    lr,
    input_id="lr",
)

Initialize the ValidMind datasets

Before you can run tests, you must first initialize a ValidMind dataset object using the init_dataset function from the ValidMind (vm) module.

This function takes a number of arguments:

  • dataset — the raw dataset that you want to provide as input to tests
  • input_id - a unique identifier that allows tracking what inputs are used when running each individual test
  • target_column — a required argument if tests require access to true values. This is the name of the target column in the dataset
  • class_labels — an optional value to map predicted classes to class labels

With all datasets ready, you can now initialize the raw, training and test datasets (raw_df, train_df and test_df) created earlier into their own dataset objects using vm.init_dataset():

vm_raw_ds = vm.init_dataset(
    input_id="raw_dataset",
    dataset=raw_df,
    target_column=demo_dataset.target_column,
)

vm_train_ds = vm.init_dataset(
    input_id="train_dataset",
    dataset=train_df,
    target_column=demo_dataset.target_column,
)
vm_test_ds = vm.init_dataset(
    input_id="test_dataset", dataset=test_df, target_column=demo_dataset.target_column
)

Options to load predictions using the developer frameworks

1. Load predictions from a file

This creates a new column called <model_id>_prediction in the dataset and assigns metadata to track that the <model_id>_prediction column is linked to the model <model_id>

Predictions calculated outside of VM

import pandas as pd

train_xgb_prediction = pd.DataFrame(xgb.predict(x_train), columns=["xgb_prediction"])
test__xgb_prediction = pd.DataFrame(xgb.predict(x_val), columns=["xgb_prediction"])

train_lr_prediction = pd.DataFrame(lr.predict(x_train), columns=["lr_prediction"])
test_lr_prediction = pd.DataFrame(lr.predict(x_val), columns=["lr_prediction"])

Assign predictions to the training dataset

We can now use the assign_predictions() method from the Dataset object to link existing predictions to any model:

vm_train_ds.assign_predictions(
    model=vm_model_xgb, prediction_values=train_xgb_prediction.xgb_prediction.values
)
vm_train_ds.assign_predictions(
    model=vm_model_lr, prediction_values=train_lr_prediction.lr_prediction.values
)

Run an example test

Now, let’s run an example test such as MinimumAccuracy twice to show how we’re able to load the correct model predictions by using the model input parameter, even though we’re passing the same train_ds dataset instance to the test:

full_suite = vm.tests.run_test(
    "validmind.model_validation.sklearn.MinimumAccuracy",
    inputs={"dataset": vm_train_ds, "model": vm_model_xgb},
)
full_suite = vm.tests.run_test(
    "validmind.model_validation.sklearn.MinimumAccuracy",
    inputs={
        "dataset": vm_train_ds,
        "model": vm_model_lr,
    },
)

Run an example test

Now, let’s run an example test such as MinimumAccuracy twice to show how we’re able to load the correct model predictions by using the model input parameter, even though we’re passing the same train_ds dataset instance to the test:

full_suite = vm.tests.run_test(
    "validmind.model_validation.sklearn.MinimumAccuracy",
    inputs={"dataset": vm_train_ds, "model": vm_model_xgb},
)
full_suite = vm.tests.run_test(
    "validmind.model_validation.sklearn.MinimumAccuracy",
    inputs={
        "dataset": vm_train_ds,
        "model": vm_model_lr,
    },
)