%pip install -q validmindRunning Individual Documentation Sections
This notebook guides you through using the run_documentation_tests() function within the ValidMind Developer Framework for targeted testing. The function is designed to enable developers to run tests on individual sections or specific groups of sections in their model documentation.
As a model developer, running individual documentation sections is useful in various development scenarios. For instance, when updates are made to a model, often only certain parts of the documentation require revision. The run_documentation_tests() function allows you to directly test only these affected sections, thus saving you time and ensuring that the documentation accurately reflects the latest changes.
This guide includes the code required to:
- Load the demo dataset
- Prepocess the raw dataset
- Train a model for testing
- Initialize ValidMind objects
- Run the data preparation documentation section
- Run the model development documentation section
- Run multiple documentation sections
Before you begin
For access to all features available in this notebook, create a free ValidMind account.
Signing up is FREE — Sign up now
If you encounter errors due to missing modules in your Python environment, install the modules with pip install, and then re-run the notebook. For more help, refer to Installing Python Modules.
Install the client library
Initialize the client library
ValidMind generates a unique code snippet for each registered model to connect with your developer environment. You initialize the client library with this code snippet, which ensures that your documentation and tests are uploaded to the correct model when you run the notebook.
Get your code snippet:
In a browser, log into the Platform UI.
In the left sidebar, navigate to Model Inventory and click + Register new model.
Enter the model details, making sure to select Binary classification as the template and Marketing/Sales - Attrition/Churn Management as the use case, and click Continue. (Need more help?)
Go to Getting Started and click Copy snippet to clipboard.
Next, replace this placeholder with your own code snippet:
# Replace with your code snippet
import validmind as vm
vm.init(
api_host="https://api.prod.validmind.ai/api/v1/tracking",
api_key="...",
api_secret="...",
project="...",
)%matplotlib inline
import xgboost as xgbPreview the documentation template
A template predefines sections for your model documentation and provides a general outline to follow, making the documentation process much easier.
You will upload documentation and test results into this template later on. For now, take a look at the structure that the template provides with the vm.preview_template() function from the ValidMind library and note the empty sections:
vm.preview_template()Load the Demo Dataset
# You can also import taiwan_credit like this:
# from validmind.datasets.classification import taiwan_credit as demo_dataset
from validmind.datasets.classification import customer_churn as demo_dataset
df = demo_dataset.load_data()Prepocess the raw dataset
train_df, validation_df, test_df = demo_dataset.preprocess(df)Train a model for testing
We train a simple customer churn model for our test.
x_train = train_df.drop(demo_dataset.target_column, axis=1)
y_train = train_df[demo_dataset.target_column]
x_val = validation_df.drop(demo_dataset.target_column, axis=1)
y_val = validation_df[demo_dataset.target_column]
model = xgb.XGBClassifier(early_stopping_rounds=10)
model.set_params(
eval_metric=["error", "logloss", "auc"],
)
model.fit(
x_train,
y_train,
eval_set=[(x_val, y_val)],
verbose=False,
)Initialize ValidMind objects
We initize the objects required to run test suites using the ValidMind framework.
vm_dataset = vm.init_dataset(
input_id="raw_dataset",
dataset=df,
target_column=demo_dataset.target_column,
class_labels=demo_dataset.class_labels,
)
vm_train_ds = vm.init_dataset(
input_id="train_dataset",
dataset=train_df,
type="generic",
target_column=demo_dataset.target_column,
)
vm_test_ds = vm.init_dataset(
input_id="test_dataset",
dataset=test_df,
type="generic",
target_column=demo_dataset.target_column,
)
vm_model = vm.init_model(model, input_id="model")Assign predictions to the datasets
We can now use the assign_predictions() method from the Dataset object to link existing predictions to any model. If no prediction values are passed, the method will compute predictions automatically:
vm_train_ds.assign_predictions(
model=vm_model,
)
vm_test_ds.assign_predictions(
model=vm_model,
)Run the data preparation section
In this section, we focus on running the tests within the data preparation section of the model documentation. After running this function, only the tests associated with this section will be executed, and the corresponding section in the model documentation will be updated.
results = vm.run_documentation_tests(
section="data_preparation",
inputs={
"dataset": vm_dataset,
},
)Run the model development section
In this section, we focus on running the tests within the model development section of the model documentation. After running this function, only the tests associated with this section will be executed, and the corresponding section in the model documentation will be updated.
results = vm.run_documentation_tests(
section="model_development",
inputs={
"dataset": vm_train_ds,
"model": vm_model,
"datasets": (vm_train_ds, vm_test_ds),
},
)Run multiple model documentation sections
This section demonstrates how you can execute both the data preparation and model development sections using run_documentation_tests(). After running this function, the tests associated with both sections will be executed, and their corresponding model documentation sections updated.
results = vm.run_documentation_tests(
section=["model_development", "model_diagnosis"],
inputs={
"dataset": vm_test_ds,
"model": vm_model,
"datasets": (vm_train_ds, vm_test_ds),
},
)