%pip install -q validmindGuide to Running Tests with Multiple Datasets
This notebook guides you through a mechanism of running a test that requires more than one dataset.
This guide includes the code required to:
- Load the demo dataset
- Prepocess the raw dataset and Train a model for testing
- Initialize ValidMind objects
- Run a test that requires multiple datasets
Install the client library
The client library provides Python support for the ValidMind Developer Framework. To install it:
Initialize the client library
ValidMind generates a unique code snippet for each registered model to connect with your developer environment. You initialize the client library with this code snippet, which ensures that your documentation and tests are uploaded to the correct model when you run the notebook.
Get your code snippet:
In a browser, log into the Platform UI.
In the left sidebar, navigate to Model Inventory and click + Register new model.
Enter the model details and click Continue. (Need more help?)
For example, to register a model for use with this notebook, select:
- Documentation template:
Binary classification - Use case:
Marketing/Sales - Attrition/Churn Management
You can fill in other options according to your preference.
- Documentation template:
Go to Getting Started and click Copy snippet to clipboard.
Next, replace this placeholder with your own code snippet:
# Replace with your code snippet
import validmind as vm
vm.init(
api_host="https://api.prod.validmind.ai/api/v1/tracking",
api_key="...",
api_secret="...",
project="...",
)Preview the documentation template
A template predefines sections for your documentation project and provides a general outline to follow, making the documentation process much easier.
You will upload documentation and test results into this template later on. For now, take a look at the structure that the template provides with the vm.preview_template() function from the ValidMind library and note the empty sections:
vm.preview_template()Load the sample dataset
The sample dataset used here is provided by the ValidMind library. To be able to use it, you need to import the dataset and load it into a pandas DataFrame, a two-dimensional tabular data structure that makes use of rows and columns:
# Import the sample dataset from the library
from validmind.datasets.classification import customer_churn as demo_dataset
print(
f"Loaded demo dataset with: \n\n\t• Target column: '{demo_dataset.target_column}' \n\t• Class labels: {demo_dataset.class_labels}"
)
raw_df = demo_dataset.load_data()
raw_df.head()Prepocess the raw dataset
Preprocessing performs a number of operations to get ready for the subsequent steps:
- Preprocess the data: Splits the DataFrame (
df) into multiple datasets (train_df,validation_df, andtest_df) usingdemo_dataset.preprocessto simplify preprocessing. - Separate features and targets: Drops the target column to create feature sets (
x_train,x_val) and target sets (y_train,y_val).
train_df, validation_df, test_df = demo_dataset.preprocess(raw_df)
x_train = train_df.drop(demo_dataset.target_column, axis=1)
y_train = train_df[demo_dataset.target_column]
x_val = validation_df.drop(demo_dataset.target_column, axis=1)
y_val = validation_df[demo_dataset.target_column]Train models for testing
Initialize XGBoost and Logistic Regression Classifiers
from sklearn.linear_model import LogisticRegression
import xgboost
%matplotlib inline
xgb = xgboost.XGBClassifier(early_stopping_rounds=10)
xgb.set_params(
eval_metric=["error", "logloss", "auc"],
)
xgb.fit(
x_train,
y_train,
eval_set=[(x_val, y_val)],
verbose=False,
)Initialize ValidMind objects
Initialize the ValidMind model
vm_model_xgb = vm.init_model(
xgb,
input_id="xgb",
)Initialize the ValidMind datasets
Before you can run tests, you must first initialize a ValidMind dataset object using the init_dataset function from the ValidMind (vm) module.
This function takes a number of arguments:
dataset— the raw dataset that you want to provide as input to testsinput_id- a unique identifier that allows tracking what inputs are used when running each individual testtarget_column— a required argument if tests require access to true values. This is the name of the target column in the datasetclass_labels— an optional value to map predicted classes to class labels
With all datasets ready, you can now initialize the raw, training and test datasets (raw_df, train_df and test_df) created earlier into their own dataset objects using vm.init_dataset():
vm_train_ds = vm.init_dataset(
input_id="train_dataset",
dataset=train_df,
target_column=demo_dataset.target_column,
)
vm_test_ds = vm.init_dataset(
input_id="test_dataset", dataset=test_df, target_column=demo_dataset.target_column
)Run a test that requires multiple datasets
We are going to show the following in next two blocks:
- Assign predictions for
vm_train_dsandvm_test_ds - Run
RobustnessDiagnosiswhich is one example test that takes two input datasets
Run predictions and link with the model
vm_train_ds.assign_predictions(model=vm_model_xgb)
vm_test_ds.assign_predictions(model=vm_model_xgb)Run test
vm.tests.run_test(
"validmind.model_validation.sklearn.RobustnessDiagnosis",
inputs={"datasets": (vm_train_ds, vm_test_ds), "model": vm_model_xgb},
)