Multi-model Example

A Vollo program can contain more than one model. The accelerator holds all the models of a program at the same time, and each inference selects a model by its index. This lets one accelerator serve several models without a program reload.

This example compiles two scikit-learn models into one program. The models can have different numbers of trees, depths and input features.

import numpy as np
import vollo_trees_compiler as vtc
from sklearn.ensemble import RandomForestRegressor
from skl2onnx.common.data_types import FloatTensorType
from skl2onnx import convert_sklearn

def train_and_export(name, n_estimators, max_depth, n_features):
    X = np.random.rand(2**max_depth, n_features)
    y = np.random.rand(2**max_depth)
    random_forest = RandomForestRegressor(
        n_estimators=n_estimators, max_depth=max_depth
    )
    random_forest.fit(X, y)

    initial_type = [("input", FloatTensorType([1, n_features]))]
    onnx_model = convert_sklearn(
        random_forest,
        initial_types=initial_type,
        target_opset=12
    )
    with open(f"{name}.onnx", "wb") as f:
        f.write(onnx_model.SerializeToString())

train_and_export("model_a", n_estimators=128, max_depth=8, n_features=64)
train_and_export("model_b", n_estimators=256, max_depth=6, n_features=256)

Lower each ONNX model to a vollo_trees_compiler.Forest, as in the single-model example.

forest_a = vtc.Forest.from_onnx("model_a.onnx")
forest_b = vtc.Forest.from_onnx("model_b.onnx")

All the models in one program must give their outputs in the same precision. The precision of a model is set by the leaf type in its ONNX file (see Supported Models), and Forest.output_precision reports it.

assert forest_a.output_precision() == forest_b.output_precision()

Compile the list of forests into one program with vollo_trees_compiler.Forest.forests_to_program_f32. The position of a forest in the list is its model index in the program.

config = vtc.Config.amd_v80_u256()
program = vtc.Forest.forests_to_program_f32([forest_a, forest_b], config)
program.save("multimodel.vollo")

Simulation

Evaluation and cycle estimates take a model_ix argument, which defaults to model 0. The input must have the number of features of the selected model.

input_a = np.random.rand(forest_a.num_input_features())
input_b = np.random.rand(forest_b.num_input_features())

print(f"Model 0 output: {program.eval(input_a, model_ix=0)}")
print(f"Model 1 output: {program.eval(input_b, model_ix=1)}")

print(f"Model 0 pessimistic cycle estimate: {program.pessimistic_cycle_estimate(model_ix=0)}")
print(f"Model 1 pessimistic cycle estimate: {program.pessimistic_cycle_estimate(model_ix=1)}")

The models in a program share the tree units of the accelerator. The cycle estimate of a model in a multi-model program can therefore differ from the estimate of the same model compiled on its own.

Inference

The Vollo runtime loads a multi-model program in the same way as a single-model program. The model index given to vollo_rt_add_job selects the model to run. This snippet extends the C example to run one inference on model 1:

// The program holds both models
assert(vollo_rt_num_models(ctx) == 2);

// Model 1 is the second forest in the list given to the compiler
size_t model_index = 1;
assert(vollo_rt_model_input_num_elements(ctx, model_index, 0) == 256);

float input_tensor[256] = {0};
float output_tensor[1];

const void* inputs[1] = {input_tensor};
void* outputs[1] = {output_tensor};
const number_format input_formats[1] = {number_format_fp32};
const number_format output_formats[1] = {number_format_fp32};

EXIT_ON_ERROR(vollo_rt_add_job(
  ctx,
  model_index,
  0,  // user_ctx
  input_formats,
  (const void* const*)inputs,
  output_formats,
  (void* const*)outputs));

// Poll for completion as in the C example

Jobs for one model complete in order, but jobs for different models do not. Use the user_ctx argument to tell the completions apart when several models run at the same time.