First Compile#
This chapter shows the shortest path from an original model to an NPU binary (MXQ).
Original model -> MBLT -> MXQ
The examples use a ResNet-50 ONNX model, but the same flow applies to other model families.
Compile with CLI#
Quickest first run#
Specify just the model and target device to compile right away.
python -m qbcompiler compile \
--model /workspace/resnet50.onnx \
--backend onnx \
--target-device regulus-rb \
--output /workspace/resnet50.mxq
On success, the MXQ file is created.
ls -lh /workspace/resnet50.mxq
Add calibration data for better accuracy#
The result above was quantized without calibration, so accuracy may be significantly degraded. For production builds, provide representative calibration data with --calib-data-path. If you haven’t prepared calibration data yet, see Prepare Calibration Data.
python -m qbcompiler compile \
--model /workspace/resnet50.onnx \
--backend onnx \
--calib-data-path /workspace/calibration/resnet50 \
--target-device regulus-rb \
--output /workspace/resnet50.mxq
The CLI mirrors the main pipeline stages. You can also run individual stages directly:
# Original model -> MBLT (parse only)
python -m qbcompiler parse --model model.onnx --backend onnx --output model.mblt
# MBLT -> MXQ (quantize only)
python -m qbcompiler quantize --mblt model.mblt --calib-data-path calib --output model.mxq
Use python -m qbcompiler <command> --help to see the exact options available in the installed version.
Compile with Python#
The same operation is available from Python.
from qbcompiler import mxq_compile
mxq_compile(
model="/workspace/resnet50.onnx",
backend="onnx",
calib_data_path="/workspace/calibration/resnet50",
target_device="regulus-rb",
save_path="/workspace/resnet50.mxq",
)
For PyTorch models, provide a feed_dict whose keys match the model’s forward parameter names:
import torch
import torchvision
from qbcompiler import mxq_compile
model = torchvision.models.resnet50(pretrained=True).eval().cpu()
feed_dict = {"x": torch.randn(1, 3, 224, 224)}
mxq_compile(
model=model,
backend="torch",
feed_dict=feed_dict,
calib_data_path="/workspace/calibration/resnet50",
target_device="regulus-rb",
save_path="/workspace/resnet50.mxq",
)
The tensor shapes in feed_dict must be valid for the original model. For HuggingFace/transformers LLMs, use the torch backend and the LLM-specific settings described in the Transformer chapter.
Key arguments at a glance:
Purpose |
Python argument |
|---|---|
Original model |
|
Input framework |
|
Calibration samples |
|
Target NPU |
|
Output MXQ path |
|
PyTorch example inputs |
|
The same compile configuration concepts are available from CLI and Python. Use a config file when the build has many options or must be reproduced exactly.
qb Compiler provides built-in presets for common model families such as classification, yolo_640, llm, and vision_transformer. List available presets with python -m qbcompiler presets. See Vision and Transformer for model-family specific compilation guides.
Prepare Calibration Data#
Calibration data should match the images or tensors the compiled model will see during real inference. For image models, you can either pass a directory of raw image files directly to calib_data_path and define the model preprocessing with PreprocessingConfig, or use prepared NumPy tensors.
Use Raw Images with a Preprocessing Pipeline#
For the standard TorchVision ResNet-50 recipe, keep representative JPEG or PNG images in a directory such as /workspace/calibration/imagenet-1k-selected. Do not convert them to .npy files. Pass that directory to calib_data_path and configure the preprocessing pipeline in the Python compile call:
from qbcompiler import (
CalibrationConfig,
PreprocessingConfig,
Uint8InputConfig,
mxq_compile,
)
preprocess_pipeline = [
{"op": "resize", "height": 256, "width": 256, "mode": "bilinear"},
{"op": "centerCrop", "height": 224, "width": 224},
{
"op": "normalize",
"scaleToUint8": True, # [0, 255] -> [0.0, 1.0]
"mean": [0.485, 0.456, 0.406],
"std": [0.229, 0.224, 0.225],
"fuseIntoFirstLayer": True,
},
]
preprocessing_config = PreprocessingConfig(
apply=True,
auto_convert_format=True,
pipeline=preprocess_pipeline,
input_configs={},
)
calibration_config = CalibrationConfig(
method=1,
output=0,
mode=1,
max_percentile={"percentile": 0.9999, "topk_ratio": 0.01},
)
mxq_compile(
model="/workspace/resnet50.onnx",
backend="onnx",
calib_data_path="/workspace/calibration/imagenet-1k-selected",
save_path="/workspace/resnet50.mxq",
target_device="regulus-rb",
device="cpu",
inference_scheme="single",
image_channels=3,
preprocessing_config=preprocessing_config,
uint8_input_config=Uint8InputConfig(apply=True, inputs=[]),
calibration_config=calibration_config,
)
image_channels=3 converts grayscale calibration images to RGB when necessary. auto_convert_format=True handles the input format conversion. The pipeline uses bilinear resize, center crop, and ImageNet-style scaling and normalization. fuseIntoFirstLayer=True and Uint8InputConfig let the MXQ model accept uint8 image input while fusing normalization into the first layer.
inference_scheme="single" compiles the MXQ for Single core mode. Other values such as multi, global4, and global8 select different ARIES runtime core modes; see Reference — inferenceScheme before changing this value.
Use this method only when the pipeline matches the model’s training and inference recipe. If the model uses a different image size, color order, resize rule, crop rule, scaling, or normalization, update the pipeline accordingly.
Use Prepared NumPy Tensors#
qb Compiler can also create preprocessed NumPy calibration samples with either a YAML preprocessing description or a Python preprocessing function.
YAML-based preprocessing example:
from qbcompiler.calibration import make_calib
make_calib(
args_pre="/workspace/resnet50.yaml",
data_dir="/workspace/calibration/cali_1000",
save_dir="/workspace/calibration",
save_name="resnet50",
max_size=100,
)
Example preprocessing YAML:
Datatype: Image
GetImage:
to_float32: false
channel_order: RGB
Pre-Order: [ResizeTorch, CenterCrop, Normalize, SetOrder]
Pre-processing:
ResizeTorch:
size: 256
interpolation: bilinear
CenterCrop:
size: [224, 224]
Normalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
to_float_div255: true
SetOrder:
shape: HWC
Manual preprocessing example:
import numpy as np
import torch
import torchvision.transforms.functional as F
from PIL import Image
from torchvision.transforms import InterpolationMode
from qbcompiler.calibration import make_calib_man
def preprocess_resnet50(img_path: str):
img = Image.open(img_path)
out = F.pil_to_tensor(img)
out = F.resize(out, size=256, interpolation=InterpolationMode.BILINEAR)
out = F.center_crop(out, output_size=(224, 224))
out = out.to(torch.float, copy=False) / 255.0
out = F.normalize(out, [0.485, 0.456, 0.406], [0.229, 0.224, 0.225])
return np.transpose(out.numpy(), axes=[1, 2, 0])
make_calib_man(
pre_ftn=preprocess_resnet50,
data_dir="/workspace/calibration/cali_1000",
save_dir="/workspace/calibration",
save_name="resnet50",
max_size=100,
)
Both examples create:
/workspace/calibration/resnet50: a directory of preprocessed.npysamples/workspace/calibration/resnet50.txt: a metadata file listing the samples
Either path can be used as calibration input.
Verify the Compiled Model#
An MXQ file that builds without errors does not guarantee correct inference. Run the compiled model on the NPU and compare the results against the original floating-point model.
Run inference with qbruntime#
Install the Mobilint runtime library (pip install mobilint-qb-runtime) and run the MXQ on the target device. See the Runtime Installation Guide for setup details.
import qbruntime
import numpy as np
import torch
from PIL import Image
from torchvision.transforms import functional as F, InterpolationMode
acc = qbruntime.Accelerator(0)
mc = qbruntime.ModelConfig()
mxq_model = qbruntime.Model("resnet50.mxq", mc)
mxq_model.launch(acc)
img = Image.open("test.jpg").convert("RGB")
out = F.pil_to_tensor(img)
out = F.resize(out, size=256, interpolation=InterpolationMode.BILINEAR)
out = F.center_crop(out, output_size=(224, 224))
out = out.to(torch.float) / 255.0
out = F.normalize(out, [0.485, 0.456, 0.406], [0.229, 0.224, 0.225])
image = np.transpose(out.numpy(), axes=[1, 2, 0]) # HWC float32
output = mxq_model.infer(image)
The preprocessing must match how the model was calibrated. The example above uses float32 normalized input because the default compile path uses float32 calibration data. For uint8 input workflows that fuse normalization into the model, see the SDK Tutorial — Image Classification.
For complete runtime examples including model loading, postprocessing, and top-k output, see the SDK Tutorial — Runtime.
Check the output#
Run a few representative inputs and confirm the output is reasonable:
Classification: top-5 predictions should include the expected class
Detection: bounding boxes should appear at sensible positions and sizes
LLM: generated text should be coherent and follow the prompt
If the output looks wrong — misclassified images, empty detections, garbled text — revisit the calibration data quality, try a different preset, or adjust quantization options. See Model Quantization for tuning guidance.
Inspect Compile Outputs#
For repeatable builds, keep these files together:
original model or model revision
calibration data directory or
.txtmetadata filecompile config or preset name
target device string
generated
.mxq