Reference#
This chapter is a reference for the current qbcompiler command line, Python API, configuration schema, supported inputs, and related terms. The source of truth is the current qbcompiler implementation.
CLI Reference#
Run the CLI as:
python -m qbcompiler <command> [options]
The current command set is compile, parse, quantize, dump-config, info, check, presets, validate, and extract-body. The table below covers the five this manual builds its examples on; presets is described under Compile Configuration.
Command |
Purpose |
Main options |
|---|---|---|
|
End-to-end model to |
|
|
Convert an input model to |
|
|
Convert a runnable |
|
|
Write a default or preset-derived |
|
|
Print the provenance recorded in a |
|
--backend defaults to onnx. --calib-data-path is repeatable. --compile-config accepts JSON, YAML, or YML and is the extension point for settings that do not have dedicated CLI flags. dump-config chooses JSON or YAML from the output file suffix.
For subprocess integrations, the CLI has an opt-in JSONL protocol. Set QBCOMPILER_JSONL=1, true, yes, or on to emit structured status, progress, result, log, and error messages on stdout. Without that environment variable, normal logs are written to stderr.
Python API Reference#
Public functions are available from qbcompiler.
API |
Purpose |
|---|---|
|
Main Python entry point. Compiles a model or runnable |
|
Parses a source model (ONNX / PyTorch / TensorFlow / TF Lite / TorchScript), then compiles it to |
|
Compiles an existing |
|
|
|
|
|
Returns the provenance record of a |
|
Exports a model to |
|
Returns preset metadata dictionaries. |
|
Writes a |
|
Writes a |
|
Sets the log level of every |
The main mxq_compile parameters are:
from qbcompiler import mxq_compile
mxq_compile(
model="model.onnx",
target_device="aries-rb",
calib_data_path="calib",
save_path="model.mxq",
backend="onnx",
device="gpu",
config_preset="classification",
compile_config=None,
)
Important parameters include model, target_device, calib_data_path, save_path, backend, feed_dict, dynamic_axes, multi_shape, yolo_decode_include, exclude_first_subgraph, device, inference_scheme, use_random_calib, cpu_offload, optimize_option, bias_correction, buffer_mode, force_npu_input_reposition, force_npu_output_reposition, image_channels, split_blocks, split_parts, config_preset, compile_config, config_save_path, model_part, model_part_options, and sub-config objects such as calibration_config, bit_config, llm_config, hessian_quant_config, bias_correction_config, mod_config, equivalent_transformation_config, search_weight_scale_config, uint8_input_config, preprocessing_config, save_sample_config, resource_management_config, and the keyword-only extra_output_config.
multi_shape and extra_output_config are new in qbcompiler 1.4:
multi_shapecompiles one model at several sizes of one input dimension in a single call. It is keyed by input name likefeed_dict, and each entry names the axis and the sizes it takes, for example{"x": {"axis": 3, "values": [100, 200, 300]}}or{"x": {"axis": (2, 3), "values": [(224, 224), (448, 448)]}}for axes that move together. The input axis named bymulti_shape.axismust already be marked as a dynamic axis when the model is exported to ONNX.widths = (100, 200, 300) mxq_compile( model="model.onnx", backend="onnx", target_device="aries-rb", feed_dict={"x": example}, multi_shape={"x": {"axis": 3, "values": widths}}, calib_data_path=[f"calib/w{w}" for w in widths], save_path="model.mxq", )
multi_shaperequiresfeed_dict. It resizes the named axis of its tensors to create an example tensor for each input size, and parses the model once per size. Several inputs are walked together by index, so theirvalueslists must be the same length.multi_shapecannot be combined with thedynamic_axesargument.mblt_compile()writes one.mbltper size: pass a list ofmblt_save_pathvalues, or one path to get<stem>_<i>.mblt. When an original model is passed withmulti_shape,mxq_compile()creates one.mbltper input size in a temporary directory, runs quantization once, and packs the results into one.mxq. Alternatively, pass the paths of existing.mbltfiles asmodel. Because those files are already parsed, parser arguments such asfeed_dictandmulti_shapedo not apply and produce a warning if passed. Before quantization, the compiler instead compares the source-model information and graph structure recorded in the files, rejecting a set that does not represent the same model and reporting the difference. In both multi-shape cases,calib_data_pathtakes one calibration directory per size, paired by index withvalues.extra_output_configadds named intermediate layers to the MXQ’s outputs. See Compile Configuration — Exporting intermediate activations.
qbcompiler 1.4 also removes the legacy Model_Dict parser and the arguments only it read. qbcompiler.Model_Dict, qbcompiler.compiler.LegacyCompiler and qbcompiler.compiler.compiler_legacy no longer exist; use mxq_compile() or mblt_compile(). Passing in_dformats or input_shape_dict raises ValueError. Input layouts are inferred from the model, and multi_shape replaces input_shape_dict. layer_bias_correction and layer_bias_correction_config are renamed to bias_correction and bias_correction_config.
config_save_path, model_part and model_part_options are new in qbcompiler 1.3:
config_save_pathwrites the fully resolvedCompileConfig— the normalized configuration after every layer of the precedence order below has been applied — before compilation starts. A.yamlor.ymlsuffix writes YAML, any other suffix writes JSON, and parent directories are created. Because it is written up front, the record survives a compile that fails after it starts, and the file can be passed straight back ascompile_config=to reproduce the compile. It is written by the quantize phase, so acompilethat fails while parsing produces nothing. The CLI exposes it as--config-save-pathonquantizeandcompile.model_partnames which part of a multi-part torch model to parse ("vision","language","encoder", and so on), andmodel_part_optionspasses extra arguments that part needs, such as{"mel_frames": 100}. A model declaring exactly one part resolves it fromNone; a model declaring several requires a name. List them withavailable_parts(model)fromqbcompiler.model_dict.parser.patcher.parts. These replace the removedhf_configargument, which now raisesValueError.
Configuration resolution priority (highest wins): explicit arguments → sub-config objects → compile_config file or config_preset, whichever is given → defaults. Passing both uses the preset and ignores the file. For details and usage examples, see Compile Configuration — Config Resolution Priority.
CompileConfig Schema#
CompileConfig is a Pydantic model with extra="forbid" and aliases enabled. Use alias names in JSON/YAML files. A full template can be generated with:
python -m qbcompiler dump-config --output compile_config.yaml
python -m qbcompiler dump-config --preset yolo_640 --output yolo_640.yaml
Top-level fields:
Alias |
Python field |
Default |
Meaning |
|---|---|---|---|
|
|
|
Model path list. Usually supplied by API/CLI instead. |
|
|
|
Calibration dataset path list. |
|
|
|
MXQ output path list. |
|
|
|
Generate random calibration data. |
|
|
|
Compile-time NPU core-assignment scheme used by the generated MXQ. See inferenceScheme. |
|
|
|
Enable CPU offloading for unsupported groups. |
|
|
|
Force input reposition operations onto NPU. |
|
|
|
Force output reposition operations onto NPU. |
|
|
|
Back-end optimization selector. Accepted values are |
|
|
|
Buffer serialization mode. |
|
|
|
Compilation compute device: |
|
|
|
Computation dtype metadata. |
|
|
|
Enable debug mode. |
|
|
|
Enable trace mode. |
|
|
|
Number of image channels; |
|
|
|
Config schema version. |
|
|
|
LLM multi-MXQ split points by block index. |
|
|
|
Split LLM transformer blocks into N parts. |
Nested sections:
Alias |
Type |
Purpose |
|---|---|---|
|
|
Treat model inputs as uint8, optionally by input name. |
|
|
Input preprocessing pipeline such as resize, letterbox, crop, normalize, and format conversion. |
|
|
Weight dtype, the GPU memory budget for weight quantization ( |
|
|
Quantization method, output quantization, calibration mode, clipping/statistics, and layer overrides. |
|
|
Quantization precision settings. |
|
|
HessianQuant optimization settings, including |
|
|
Corrects the per-channel bias that quantization introduces. On by default. Replaces |
|
|
MOD training/optimization settings. |
|
|
LLM sequence/cache/runtime/debug settings. If |
|
|
Sparse MoE calibration expert selection. |
|
|
SmoothQuant-like, rotation, and FFN equivalent transformations. |
|
|
Weight-scale search. |
|
|
Load externally supplied scale entries. |
|
|
Internal runtime metadata. |
|
|
Save sample data for inspection/debugging. |
|
|
Add named intermediate layers to the MXQ’s outputs ( |
CompileConfig.from_file() accepts .json, .yaml, and .yml. For compatibility, it also flattens grouped keys such as quantization.calibration, quantization.bit, advancedQuantization.hessianQuant, advancedQuantization.mod, advancedQuantization.EquivalentTransformation, advancedQuantization.searchWeightScale, advancedQuantization.loadScale, and advancedQuantization.biasCorrection, which replaces 1.3’s advancedQuantization.layerBiasCorrection from 1.4. That grouped spelling is also the one the tool writes — dump-config and --config-save-path both emit it. Flattening overwrites rather than merges, so a top-level calibration block added beside a quantization.calibration one is discarded with no message: do not mix the two spellings in one file.
inferenceScheme#
inferenceScheme selects how NPU core work is assigned when qb Runtime uses the MXQ for inference. Because the compiled MXQ is prepared for the selected core-assignment scheme, an MXQ built only for one mode cannot later be switched arbitrarily to another mode at runtime. One built with all carries every mode the model and target support, but the caller must name one: qb Runtime’s default CoreMode::Auto accepts only an MXQ with exactly one mode, so an all build needs an explicit setSingleCoreMode() or setGlobal4CoreMode() — see ARIES ModelConfig Configuration.
Available choices are single, multi, global4, global8, and all. single uses independent Local Cores, multi is the cluster-level 4-batch mode, and global4/global8 use 4 or 8 Local Cores together for one input. all prepares the MXQ for every core mode supported by the selected model and target. Support can vary by model, target, and compiler version. REGULUS targets have a single NPU core, so MXQ files compiled for REGULUS can use only single mode. From qbcompiler 1.3 this is enforced rather than left to fail later: multi, global4 and global8 are rejected outright on a single-core target, and all narrows to single there. The check runs before quantization regardless of whether the scheme or the target device was set first. global is a legacy/compatibility value; use global4 or global8 for new settings.
If llm.attributes.runtime.batchSize is greater than 1, LLM/KV-cache transformer compilation accepts single, global4 and global8, and rejects multi and all rather than narrowing them — each segment compiles under exactly one scenario, and guessing which one was meant would build for a core set nobody asked for. Vision batch inference and LLM batch inference do not use the same core-mode rule.
For what each mode means, see ARIES Core Mode. For runtime core/cluster selection, see ARIES ModelConfig Configuration.
RuntimeOptions#
RuntimeOptions currently contains only:
Field |
Default |
Meaning |
|---|---|---|
|
|
Compiler/runtime version metadata. |
This section is runtime metadata rather than a normal user tuning surface. Prefer the compile, calibration, preprocessing, LLM, and optimization sections for user-controlled behavior.
Preset List#
Preset |
Extends |
Behavior |
|---|---|---|
|
none |
Image classification defaults: |
|
none |
Object detection defaults: |
|
|
Enables uint8 input and standard Torchvision preprocessing: resize shortest side to 256 with the Pillow backend ( |
|
|
Enables uint8 input and 640x640 letterbox preprocessing with pad value 114 and the OpenCV backend ( |
|
|
Enables uint8 input and 1280x1280 letterbox preprocessing with pad value 114 and the OpenCV backend ( |
From qbcompiler 1.4 the three image presets use the Pillow and OpenCV backends to match their reference evaluators, so their calibration tensors and quantization results can differ from 1.3.
| llm | none | Enables LLM config with maxSequenceLength=4096, maxCacheLength=4096, calibration.mode=0, calibration.output=0, full-sequence-length LLM calibration, and the QK, UD, VO, SpinR1, SpinR2 and OptimizeFFN equivalent transformations. |
| llm_fast | llm | Turns off those six equivalent transformations and full-sequence-length calibration, trading accuracy for compile time. |
| vision_transformer | none | Transformer-oriented calibration/bit settings. |
| multimodal | none | Enables LLM handling and uses calibration.method=3. |
Supported Frameworks#
The public backend strings are below. Backend names are case-insensitive. From qbcompiler 1.4 every backend uses the current parser; tf, tflite and torchscript no longer go through the legacy parser.
Backend |
Input |
|---|---|
|
ONNX model path. This is the primary path and is used by |
|
TensorFlow SavedModel directory, Keras |
|
TensorFlow Lite |
|
Path of a |
|
PyTorch model or HuggingFace/transformers model path/object flows. HuggingFace LLMs use this backend path, not a separate |
Supported Target Device#
Use these target device strings in CLI and Python API calls:
target device |
Meaning |
|---|---|
|
REGULUS RA target. |
|
ARIES RB target. |
|
REGULUS RB target. |
|
REGULUS RB target connected over USB. From qbcompiler 1.3. |
The validator accepts the public strings above. Internal enum names use underscores, but user-facing values use hyphens.
Supported Operators#
Mobilint IR Operations List lists the operations Mobilint hardware runs natively. It is not the whole answer: an operation outside that list may still compile through a graph-level transformation, and whether any given one runs natively also depends on the model, backend, shapes and target device, which the parser and target-device allocation logic settle per compile.
Available reference paths:
For non-ONNX backends, parse/compile the model and inspect unsupported groups with
.mbltand CPU offloading workflows.If a model contains unsupported groups,
cpuOffload/cpu_offloadcan partition those groups for CPU execution where the runtime flow supports it.
Glossary#
Term |
Meaning |
|---|---|
|
Mobilint executable package produced by quantization/compilation and run on Mobilint NPUs. |
|
Mobilint intermediate model format used between parsing and MXQ generation. |
|
Convert a source model into |
|
Convert a runnable |
|
End-to-end parse plus quantize flow. |
|
Source model framework identifier such as |
|
Mobilint NPU target device string such as |
NPU Chip |
Mobilint NPU product name such as ARIES or REGULUS. |
target device |
Compile target such as |
|
Representative input tensors used to derive quantization scales/statistics. |
|
Partitioning unsupported graph groups for CPU execution while supported body subgraphs run on NPU. |
|
The largest NPU-runnable supported subgraph extracted from a partitioned |
|
Built-in partial |
License#
See License for the open-source license notices included with this manual.