API Reference#
|
Mobilint SDK qb Compiler v1.4
MCS002-EN
|
Mobilint qb Compiler API provides functions for compiling models from various framework(onnx, pytorch, transformers, etc.) to MXQ format. More...
Classes | |
| class | qbcompiler.calibration.preprocess.GetNpy |
| Load NumPy data from an in-memory array or a .npy file. More... | |
| class | qbcompiler.calibration.preprocess.NpyNormalize |
| class | qbcompiler.calibration.preprocess.NpyPad |
| Pad NumPy arrays to a desired shape. More... | |
| class | qbcompiler.calibration.preprocess.GetImage |
| Retrieve an image tensor from a file path or NumPy array using OpenCV. More... | |
| class | qbcompiler.calibration.preprocess.Pad |
| Pad images to a target shape or to meet a size divisor. More... | |
| class | qbcompiler.calibration.preprocess.Normalize |
| Apply mean/std normalization to images. More... | |
| class | qbcompiler.calibration.preprocess.ResizeTorch |
| class | qbcompiler.calibration.preprocess.Resize |
| Resize images with optional aspect-ratio preservation. More... | |
| class | qbcompiler.calibration.preprocess.CenterCrop |
| Perform a centered crop on the image. More... | |
| class | qbcompiler.calibration.preprocess.SetOrder |
| class | qbcompiler.calibration.preprocess.preproc_builder |
| Build a calibration preprocessing pipeline from a YAML or dict configuration. More... | |
| class | GetText |
| Extract text strings from raw inputs. More... | |
| class | TokenizeToEmbedding |
| Convert tokenized text into embeddings using a Hugging Face model. More... | |
| class | YoloPre |
| Apply YOLO-style resize, padding, and normalization to images. More... | |
Functions | |
| qbcompiler.frontend.mxq_compile (model, str target_device, Union[str, List[str]] calib_data_path=UNSET, int save_subgraph_type=0, output_subgraph_path="", Union[str, List[str]] save_path=UNSET, backend="onnx", feed_dict=None, dynamic_axes=None, Optional[dict] multi_shape=None, yolo_decode_include=False, exclude_first_subgraph=False, device=UNSET, inference_scheme=UNSET, use_random_calib=UNSET, cpu_offload=UNSET, optimize_option=UNSET, buffer_mode=UNSET, force_npu_input_reposition=UNSET, force_npu_output_reposition=UNSET, image_channels=UNSET, split_blocks=UNSET, split_parts=UNSET, Optional[str] config_preset=None, Optional[CompileConfig] compile_config=None, Optional[ResourceManagementConfig] resource_management_config=None, Optional[CalibrationConfig] calibration_config=None, Optional[BitConfig] bit_config=None, Optional[LlmConfig] llm_config=None, Optional[HessianQuantConfig] hessian_quant_config=None, Optional[ModConfig] mod_config=None, Optional[EquivalentTransformationConfig] equivalent_transformation_config=None, Optional[SearchWeightScaleConfig] search_weight_scale_config=None, Optional[SaveSampleConfig] save_sample_config=None, Optional[Uint8InputConfig] uint8_input_config=None, Optional[PreprocessingConfig] preprocessing_config=None, bias_correction=UNSET, Optional[BiasCorrectionConfig] bias_correction_config=None, Optional[str] model_part=None, Optional[dict] model_part_options=None, Optional[str] config_save_path=None, *, Optional[ExtraOutputConfig] extra_output_config=None, **kwargs) | |
| Compile a model into a Mobilint eXeCUtable (MXQ) package for execution on Mobilint NPUs. | |
| qbcompiler.frontend.mxq_compile_from_source (model, str target_device, Union[str, List[str]] calib_data_path=UNSET, int save_subgraph_type=0, output_subgraph_path="", Union[str, List[str]] save_path=UNSET, backend="onnx", feed_dict=None, dynamic_axes=None, Optional[dict] multi_shape=None, yolo_decode_include=False, exclude_first_subgraph=False, device=UNSET, inference_scheme=UNSET, use_random_calib=UNSET, cpu_offload=UNSET, optimize_option=UNSET, buffer_mode=UNSET, force_npu_input_reposition=UNSET, force_npu_output_reposition=UNSET, image_channels=UNSET, split_blocks=UNSET, split_parts=UNSET, Optional[str] config_preset=None, Optional[CompileConfig] compile_config=None, Optional[ResourceManagementConfig] resource_management_config=None, Optional[CalibrationConfig] calibration_config=None, Optional[BitConfig] bit_config=None, Optional[LlmConfig] llm_config=None, Optional[HessianQuantConfig] hessian_quant_config=None, Optional[ModConfig] mod_config=None, Optional[EquivalentTransformationConfig] equivalent_transformation_config=None, Optional[SearchWeightScaleConfig] search_weight_scale_config=None, Optional[SaveSampleConfig] save_sample_config=None, Optional[Uint8InputConfig] uint8_input_config=None, Optional[PreprocessingConfig] preprocessing_config=None, bias_correction=UNSET, Optional[BiasCorrectionConfig] bias_correction_config=None, Optional[str] model_part=None, Optional[dict] model_part_options=None, Optional[str] config_save_path=None, *, Optional[ExtraOutputConfig] extra_output_config=None, **kwargs) | |
| Compile a raw framework model (ONNX / PyTorch / TensorFlow / TF-Lite) into an MXQ package. | |
| qbcompiler.frontend.mxq_compile_from_mblt (str|list[str] mblt, str target_device, Union[str, List[str]] calib_data_path=UNSET, Union[str, List[str]] save_path=UNSET, backend="onnx", device=UNSET, inference_scheme=UNSET, use_random_calib=UNSET, cpu_offload=UNSET, optimize_option=UNSET, buffer_mode=UNSET, force_npu_input_reposition=UNSET, force_npu_output_reposition=UNSET, image_channels=UNSET, split_blocks=UNSET, split_parts=UNSET, Optional[str] config_preset=None, Optional[CompileConfig] compile_config=None, Optional[ResourceManagementConfig] resource_management_config=None, Optional[CalibrationConfig] calibration_config=None, Optional[BitConfig] bit_config=None, Optional[LlmConfig] llm_config=None, Optional[HessianQuantConfig] hessian_quant_config=None, Optional[ModConfig] mod_config=None, Optional[EquivalentTransformationConfig] equivalent_transformation_config=None, Optional[SearchWeightScaleConfig] search_weight_scale_config=None, Optional[SaveSampleConfig] save_sample_config=None, Optional[Uint8InputConfig] uint8_input_config=None, Optional[PreprocessingConfig] preprocessing_config=None, bias_correction=UNSET, Optional[BiasCorrectionConfig] bias_correction_config=None, Optional[str] config_save_path=None, *, Optional[ExtraOutputConfig] extra_output_config=None, **kwargs) | |
Compile an existing Mobilint IR (.mblt) file into an MXQ package. | |
| qbcompiler.frontend.mblt_compile (str|Any model, str|list[str] mblt_save_path, str target_device, backend="onnx", device="cpu", feed_dict=None, dynamic_axes=None, Optional[dict] multi_shape=None, yolo_decode_include=False, cpu_offload=False, exclude_first_subgraph=False, Optional[str] model_part=None, Optional[dict] model_part_options=None, **kwargs) | |
| Export a model to the Mobilint .mblt format without producing an MXQ package. | |
| None | qbcompiler.frontend._mblt_compile_dispatch (str|Any model, str|list[str] mblt_save_path, str target_device, backend="onnx", device=UNSET, feed_dict=None, dynamic_axes=None, Optional[dict] multi_shape=None, yolo_decode_include=False, cpu_offload=UNSET, exclude_first_subgraph=False, Optional[str] model_part=None, Optional[dict] model_part_options=None, **kwargs) |
| Route a compile-to-mblt call to the compile pipeline. | |
| None | qbcompiler.frontend._mxq_compile_pipeline (model, str target_device, Union[str, List[str]] calib_data_path=UNSET, int save_subgraph_type=0, output_subgraph_path="", Union[str, List[str]] save_path=UNSET, backend="onnx", feed_dict=None, dynamic_axes=None, Optional[dict] multi_shape=None, yolo_decode_include=False, exclude_first_subgraph=False, device=UNSET, inference_scheme=UNSET, use_random_calib=UNSET, cpu_offload=UNSET, optimize_option=UNSET, buffer_mode=UNSET, force_npu_input_reposition=UNSET, force_npu_output_reposition=UNSET, image_channels=UNSET, split_blocks=UNSET, split_parts=UNSET, Optional[str] config_preset=None, Optional[CompileConfig|str] compile_config=None, Optional[ResourceManagementConfig] resource_management_config=None, Optional[CalibrationConfig] calibration_config=None, Optional[BitConfig] bit_config=None, Optional[LlmConfig] llm_config=None, Optional[HessianQuantConfig] hessian_quant_config=None, Optional[ModConfig] mod_config=None, Optional[EquivalentTransformationConfig] equivalent_transformation_config=None, Optional[SearchWeightScaleConfig] search_weight_scale_config=None, Optional[SaveSampleConfig] save_sample_config=None, Optional[Uint8InputConfig] uint8_input_config=None, Optional[PreprocessingConfig] preprocessing_config=None, bias_correction=UNSET, Optional[BiasCorrectionConfig] bias_correction_config=None, Optional[str] model_part=None, Optional[dict] model_part_options=None, Optional[str] config_save_path=None, *, Optional[ExtraOutputConfig] extra_output_config=None, **kwargs) |
| Quantize and compile through the ConfigManager-resolved pipeline. | |
| None | qbcompiler.frontend._mblt_compile_pipeline (str|object model, str target_device, str|list[str] mblt_save_path, backend="onnx", object device=UNSET, feed_dict=None, dynamic_axes=None, Optional[dict] multi_shape=None, yolo_decode_include=False, object cpu_offload=UNSET, exclude_first_subgraph=False, Optional[str] model_part=None, Optional[dict] model_part_options=None, Optional[str] config_preset=None, Optional[CompileConfig|str] compile_config=None, Optional[str] compilation_mode=None, **kwargs) |
| Export to .mblt through the ConfigManager-resolved pipeline. | |
| None | qbcompiler.frontend.mblt_compile_with_callback (str model, str mblt_save_path, str target_device, str backend="onnx", Optional[str] device=None, Optional[bool] cpu_offload=None, Optional[str] config_preset=None, Optional[CompileConfig|str] compile_config=None, Optional[str] compilation_mode=None, *, Callable[[int, str], None] progress_callback) |
| Compile-to-mblt entry point with progress callbacks. | |
| None | qbcompiler.frontend.mxq_compile_with_callback (str model, str target_device, str save_path, str backend="onnx", Optional[str] device=None, Optional[Union[str, List[str]]] calib_data_path=None, Optional[bool] use_random_calib=None, Optional[str] config_preset=None, Optional[CompileConfig|str] compile_config=None, Optional[str] config_save_path=None, *, Callable[[int, str], None] progress_callback) |
| Compile/quantize entry point with progress callbacks. | |
| qbcompiler.calibration.make_calib_data.get_yaml (str yaml_path) | |
| qbcompiler.calibration.make_calib_data.make_calib_man (pre_ftn, str data_dir, str save_dir, str save_name, str anno_json=None, str file_format="%012d.jpg", int max_size=-1, bool remove_npy=False, int seed=2023, bool save_calib_msg=False, str msg_path=None) | |
| From images and a user-provided preprocessing function, generate calibration NumPy files and a manifest. | |
| str | qbcompiler.calibration.make_calib_data.make_calib (Union[str, dict] args_pre, str data_dir, str save_dir, str save_name=None, str anno_json=None, str file_format="%012d.jpg", int max_size=-1, bool remove_npy=False, int seed=2023, bool save_calib_msg=False, str msg_path=None) |
| Create calibration NumPy files and a manifest using a YAML or dict preprocessing configuration. | |
| List[str] | qbcompiler.calibration.make_calib_data.make_calib_llm (Union[str, dict] args_pre, str dataset_name, Union[List[str], str] subset_list, str text_field, str split, str save_dir, int max_calib=512, Optional[str] save_prefix=None) |
| Generate calibration embeddings for LLM models using Hugging Face datasets. | |
| qbcompiler.calibration.preprocess.preprocess_fnc (fnc, Datatype="Image") | |
Variables | |
| dict | qbcompiler.calibration.preprocess.DTYPE_GETDATA_MAP = {"Image": "GetImage", "Npy": "GetNpy", "LLM": "GetText"} |
Detailed Description
Mobilint qb Compiler API provides functions for compiling models from various framework(onnx, pytorch, transformers, etc.) to MXQ format.
Function Documentation
◆ mxq_compile()
| qbcompiler.frontend.mxq_compile | ( | model, | |
| str | target_device, | ||
| Union[str, List[str]] | calib_data_path = UNSET, | ||
| int | save_subgraph_type = 0, | ||
| output_subgraph_path = "", | |||
| Union[str, List[str]] | save_path = UNSET, | ||
| backend = "onnx", | |||
| feed_dict = None, | |||
| dynamic_axes = None, | |||
| Optional[dict] | multi_shape = None, | ||
| yolo_decode_include = False, | |||
| exclude_first_subgraph = False, | |||
| device = UNSET, | |||
| inference_scheme = UNSET, | |||
| use_random_calib = UNSET, | |||
| cpu_offload = UNSET, | |||
| optimize_option = UNSET, | |||
| buffer_mode = UNSET, | |||
| force_npu_input_reposition = UNSET, | |||
| force_npu_output_reposition = UNSET, | |||
| image_channels = UNSET, | |||
| split_blocks = UNSET, | |||
| split_parts = UNSET, | |||
| Optional[str] | config_preset = None, | ||
| Optional[CompileConfig] | compile_config = None, | ||
| Optional[ResourceManagementConfig] | resource_management_config = None, | ||
| Optional[CalibrationConfig] | calibration_config = None, | ||
| Optional[BitConfig] | bit_config = None, | ||
| Optional[LlmConfig] | llm_config = None, | ||
| Optional[HessianQuantConfig] | hessian_quant_config = None, | ||
| Optional[ModConfig] | mod_config = None, | ||
| Optional[EquivalentTransformationConfig] | equivalent_transformation_config = None, | ||
| Optional[SearchWeightScaleConfig] | search_weight_scale_config = None, | ||
| Optional[SaveSampleConfig] | save_sample_config = None, | ||
| Optional[Uint8InputConfig] | uint8_input_config = None, | ||
| Optional[PreprocessingConfig] | preprocessing_config = None, | ||
| bias_correction = UNSET, | |||
| Optional[BiasCorrectionConfig] | bias_correction_config = None, | ||
| Optional[str] | model_part = None, | ||
| Optional[dict] | model_part_options = None, | ||
| Optional[str] | config_save_path = None, | ||
| * | , | ||
| Optional[ExtraOutputConfig] | extra_output_config = None, | ||
| ** | kwargs ) |
Compile a model into a Mobilint eXeCUtable (MXQ) package for execution on Mobilint NPUs.
When no explicit value is provided for a parameter marked with UNSET, the default from CompileConfig is used. This allows a compile_config file or object to supply the value without being overridden by function-level defaults.
Configuration is resolved in priority order (highest to lowest):
- Explicitly passed function arguments
- Individual sub-config objects (calibration_config, llm_config, etc.)
- kwargs partial overrides (quantization_method, weight_dtype, etc.)
- compile_config (CompileConfig object or JSON/YAML file) or config_preset
- CompileConfig field defaults
- Parameters
-
model string or model instance. Model path. When using backend="onnx", this should be the path to an ONNX model file. Whenbackendis "torchscript", pass the path of atorch.jit.savearchive (it is exported to ONNX first, sofeed_dictis required); when "torch", provide a standard PyTorch model. Forbackend="tf", pass a TensorFlow SavedModel directory, a Keras .keras/.h5 file or a frozen GraphDef .pb. Forbackend="tflite", pass the .tflite model path. The path of an existing.mblt skips parsing and only quantizes. A list or tuple is accepted for one purpose only: multi-shape compilation. Its entries must be existing.mblt files that are the sizes of one model – written by amulti_shapecompile, or by onemblt_compile()call per size – and they are quantized in one call into a single multi-shape.mxq (thencalib_data_pathtakes one directory per file, in the same order). A list of source models, or of unrelated.mblt files, is not a valid input: before quantization the files are checked to be one model at different sizes and the set is refused with the first difference otherwise. Two things are compared. The source model each file records in its provenance (the source file's sha256, the HuggingFace revision, ...): files made from different checkpoints are rejected even when their graphs match. Only fields that pin the weights are grounds for refusal; a backend or class name that disagrees is reported instead, since those also change with where the compile ran. And the graph itself: same NPU/CPU partition, same layer types in graph order (options and weights are not compared, since they legitimately differ per size). A file that records no such identity – an in-memory torch model leaves only its class name – is logged as unverified, because for it "same graph" is all that was checked. To compile a source model at several sizes in one call, pass one model withmulti_shapeinstead.calib_data_path string or list of strings. Path(s) to the calibration dataset. Accepts either a text/json file that lists NumPy files or a directory that contains the pre-processed NumPy files. save_subgraph_type int. Controls optional MBLT subgraph exports: 0 disables exports; 1 saves only the graph structure; 2 saves graph structure plus weights; 3 saves the graph structure split into multiple subgraphs; 4 saves both structure and weights split into multiple subgraphs. Defaults to 0. output_subgraph_path string. Destination path for the exported .mblt file when save_subgraph_typeis 1–4. The resulting file can be used for visualization. Defaults to "".save_path string or list of strings. Output MXQ filename(s). When omitted, defaults to "{model_name}.mxq" derived from the model path basename. backend string. Framework used to generate the Mobilint IR. Case-insensitive, so "ONNX" and "onnx" are the same backend. "onnx", "torch", "tf" (also spelled "tensorflow" or "keras"), "tflite" and "torchscript" are all parsed by the current parser. Any other value raises ValueError. Defaults to "onnx". Whenmodelis a path to an already-compiled .mblt file (the output ofmblt_compile()), it is compiled directly with MXQ regardless of thebackendvalue, and the file is validated beforehand; it must be a runnable artifact produced bymblt_compile()rather than asave_subgraph_typepreview export.target_device string. Target NPU device for parser/compiler configuration (for example "aries-rb"). Required by execution paths that parse or quantize a model. device string. Compilation and inference device: "cpu" or "gpu". When omitted, uses CompileConfig default. feed_dict dict. Example input tensors for shape inference and inference validation. multi_shape dict. Compile the model at several sizes of a dynamic dimension in one call. Keyed by input name like feed_dict; each entry names the axis (an int, or a tuple of axes that move together) and the sizes it takes: {"x": {"axis": 3, "values": [100, 200, 300]}} or {"x": {"axis": (2, 3), "values": [(224, 224), (448, 448)]}}. Requiresfeed_dict:each value yields its own example tensor derived from it, and the model is parsed once per value into its own .mblt. Several inputs are walked together by index, so their "values" lists must be equally long. Cannot be combined withdynamic_axes. Withmodel_partthe part is built once per size from that size's example tensors, so a part that sizes itself from its feed (Qwen3-ASR audio's "mel_chunk") compiles at every listed size.dynamic_axes dict. Marks model axes as dynamic. Keys are input names and values map dimension indices to aliases (for example {"input": {2: "seq_len"}}). inference_scheme string. NPU inference scheme. One of "single", "multi", "global", "global4", or "global8". When omitted, uses CompileConfig default. yolo_decode_include bool. Determines whether YOLO decode runs on NPU. Defaults to False.use_random_calib bool. Generates random calibration data to validate model compilability. When omitted, uses CompileConfig default. cpu_offload bool. Enables CPU offloading during NPU inference. When omitted, uses CompileConfig default. optimize_option int. Compiler optimization strategy selector. When omitted, uses CompileConfig default. buffer_mode int. Buffer serialization mode: 0 uses a naive buffer, 1 uses an mmap-backed buffer. When omitted, uses CompileConfig default. force_npu_input_reposition bool. Force input reposition operations to run on NPU instead of CPU. When omitted, uses CompileConfig default. force_npu_output_reposition bool. Force output reposition operations to run on NPU instead of CPU. When omitted, uses CompileConfig default. image_channels int. Number of image channels (0 for auto-detect). When omitted, uses CompileConfig default. split_blocks list of int. Multi-MXQ split points by transformer block index. Only supported for LLM models. When omitted, uses CompileConfig default. split_parts int. Evenly split transformer blocks into N MXQ parts. Only supported for LLM models. When omitted, uses CompileConfig default. bias_correction bool. Enables in-schedule integer bias correction during weight quantization. When omitted, uses bias_correction_configor the CompileConfig value.exclude_first_subgraph bool. Applies only when CPU offloading is enabled: exclude the first subgraph from the final graph if it is unsupported. Defaults to False.config_preset string. Name of a built-in configuration preset. When provided, loads the preset via CompileConfig.from_preset(). Defaults to None (no preset). Available presets: - "classification": Image classification models (ResNet, EfficientNet, ViT, etc.)
- "detection": Object detection models (YOLO, SSD, DETR, etc.)
- "classification_torchvision": Torchvision classification models with standard preprocessing
- "yolo_640": YOLO detection models with 640x640 letterbox preprocessing
- "yolo_1280": YOLO detection models with 1280x1280 letterbox preprocessing
- "llm": Large Language Models (LLaMA, Qwen, Gemma, etc.)
- "llm_fast": LLM with faster compilation (less accuracy optimization)
- "vision_transformer": Vision Transformer models (ViT, DeiT, Swin, etc.)
- "multimodal": Multimodal models (CLIP, BLIP, LLaVA, etc.)
compile_config CompileConfig or string. CompileConfig object or path to JSON/YAML configuration file containing all compilation settings (resourceManagement, calibration, bit, hessianQuant, mod, llm, etc.). config_save_path string. When provided, the fully resolved configuration - the normalized CompileConfig after every layer of the precedence order above has been applied, including all sub-configurations - is written to this path before compilation starts. The format follows the path suffix: ".yaml"/".yml" produce YAML, anything else produces JSON. Parent directories are created as needed. The saved file can be fed back as compile_configto reproduce the same compilation. Defaults to None (no config file is written).resource_management_config ResourceManagementConfig. Resource management configuration object. calibration_config CalibrationConfig. Calibration configuration object. bit_config BitConfig. Bit configuration object for quantization precision settings. llm_config LlmConfig. LLM configuration object. hessian_quant_config HessianQuantConfig. HessianQuant (Hessian-based Quantization) configuration object. bias_correction_config BiasCorrectionConfig. In-schedule integer bias correction settings. model_part string. Selects a specific part of a torch model to parse ("vision", "language", "encoder", ...). Typically used to support models with complex architectures (e.g. Qwen3-VL) by compiling one part at a time. Parts a model declares: parser.patcher.parts.available_parts(model). A model declaring exactly one part resolves it from None; one declaring several requires a name. model_part_options dict. Extra arguments for the named part (e.g. {"mel_frames": 100} for Qwen3-ASR audio). mod_config ModConfig. MOD (Metric-based Optimization and Distillation) configuration object. equivalent_transformation_config EquivalentTransformationConfig. Configuration for equivalent transformations like SmoothQuant. search_weight_scale_config SearchWeightScaleConfig. Configuration for weight scale search. save_sample_config SaveSampleConfig. Configuration for sample data generation and saving. uint8_input_config Uint8InputConfig. Configuration for uint8 input handling. preprocessing_config PreprocessingConfig. Preprocessing pipeline configuration. extra_output_config ExtraOutputConfig. Promote named intermediate layers to extra model outputs. Keyword-only. Names are matched against the post-fusion graph, after quantization, so a layer the graph optimizer folded away is rejected by name. kwargs dict. Additional compiler arguments. Supports partial config overrides such as quantization_method,quantization_mode,percentile,weight_dtype,ram_usage,max_sequence_length, etc. Also accepts deprecated parameters (quantization_config,advanced_quantization_config,input_process_config,save_sample,sample_dtype) which will emit DeprecationWarning.
- Returns
- None.
- Using configuration
- There are three ways to configure quantization settings:
- Load all settings from a JSON/YAML config file: from qbcompiler import mxq_compilemxq_compile(model="path/to/model.onnx",target_device="aries-rb",calib_data_path="path/to/calib",compile_config="path/to/config.json", # or config.yamldevice="gpu",)
- Pass individual sub-config objects: from qbcompiler import mxq_compileResourceManagementConfig,CalibrationConfig,BitConfig,HessianQuantConfig,ModConfig,LlmConfig,)resource_mgmt = ResourceManagementConfig(weight_dtype="float32")calib_cfg = CalibrationConfig(method=1, mode=1)bit_cfg = BitConfig(...)hessian_quant_cfg = HessianQuantConfig(apply=True)mod_cfg = ModConfig(apply=False)llm_cfg = LlmConfig(apply=True)mxq_compile(model="path/to/model.onnx",target_device="aries-rb",calib_data_path="path/to/calib",resource_management_config=resource_mgmt,calibration_config=calib_cfg,bit_config=bit_cfg,hessian_quant_config=hessian_quant_cfg,mod_config=mod_cfg,llm_config=llm_cfg,device="gpu",)
- Automatically applied partial configuration overrides: Please refer to the mxq_compile function and the quantization configuration section for the meaning of the quantization-related numeric values.from qbcompiler import mxq_compilemxq_compile(model="path/to/model.onnx",target_device="aries-rb",calib_data_path="path/to/calib",quantization_method=1, # per channel quantizationquantization_mode=1, # max percentile quantizationpercentile=0.999, # percentile value for max percentile quantizationquantization_output=0, # per layer quantization for the output layerdevice="gpu",)
- Compiling models with custom inputs (ONNX/Torch/TensorFlow)
- Provide NumPy inputs whenever the model omits shape information so that qbcompiler can infer unknown dimensions and data formats.
Definition at line 63 of file frontend.py.
◆ mxq_compile_from_source()
| qbcompiler.frontend.mxq_compile_from_source | ( | model, | |
| str | target_device, | ||
| Union[str, List[str]] | calib_data_path = UNSET, | ||
| int | save_subgraph_type = 0, | ||
| output_subgraph_path = "", | |||
| Union[str, List[str]] | save_path = UNSET, | ||
| backend = "onnx", | |||
| feed_dict = None, | |||
| dynamic_axes = None, | |||
| Optional[dict] | multi_shape = None, | ||
| yolo_decode_include = False, | |||
| exclude_first_subgraph = False, | |||
| device = UNSET, | |||
| inference_scheme = UNSET, | |||
| use_random_calib = UNSET, | |||
| cpu_offload = UNSET, | |||
| optimize_option = UNSET, | |||
| buffer_mode = UNSET, | |||
| force_npu_input_reposition = UNSET, | |||
| force_npu_output_reposition = UNSET, | |||
| image_channels = UNSET, | |||
| split_blocks = UNSET, | |||
| split_parts = UNSET, | |||
| Optional[str] | config_preset = None, | ||
| Optional[CompileConfig] | compile_config = None, | ||
| Optional[ResourceManagementConfig] | resource_management_config = None, | ||
| Optional[CalibrationConfig] | calibration_config = None, | ||
| Optional[BitConfig] | bit_config = None, | ||
| Optional[LlmConfig] | llm_config = None, | ||
| Optional[HessianQuantConfig] | hessian_quant_config = None, | ||
| Optional[ModConfig] | mod_config = None, | ||
| Optional[EquivalentTransformationConfig] | equivalent_transformation_config = None, | ||
| Optional[SearchWeightScaleConfig] | search_weight_scale_config = None, | ||
| Optional[SaveSampleConfig] | save_sample_config = None, | ||
| Optional[Uint8InputConfig] | uint8_input_config = None, | ||
| Optional[PreprocessingConfig] | preprocessing_config = None, | ||
| bias_correction = UNSET, | |||
| Optional[BiasCorrectionConfig] | bias_correction_config = None, | ||
| Optional[str] | model_part = None, | ||
| Optional[dict] | model_part_options = None, | ||
| Optional[str] | config_save_path = None, | ||
| * | , | ||
| Optional[ExtraOutputConfig] | extra_output_config = None, | ||
| ** | kwargs ) |
Compile a raw framework model (ONNX / PyTorch / TensorFlow / TF-Lite) into an MXQ package.
This is the raw-model half of mxq_compile(): it parses the model into Mobilint IR and then quantizes and compiles it in one pass. The parameters, their defaults, and the configuration precedence rules are identical to mxq_compile() — see that function for the full reference and usage examples.
Passing the path of an existing .mblt file here raises ValueError; use mxq_compile_from_mblt() for that input instead. mxq_compile() routes between the two automatically.
- Parameters
-
model string or model instance. Raw model or path to one. Must not be an existing .mblt file.target_device string. Target NPU device (for example "aries-rb").
- Returns
- None.
Definition at line 442 of file frontend.py.
◆ mxq_compile_from_mblt()
| qbcompiler.frontend.mxq_compile_from_mblt | ( | str | list[str] | mblt, |
| str | target_device, | ||
| Union[str, List[str]] | calib_data_path = UNSET, | ||
| Union[str, List[str]] | save_path = UNSET, | ||
| backend = "onnx", | |||
| device = UNSET, | |||
| inference_scheme = UNSET, | |||
| use_random_calib = UNSET, | |||
| cpu_offload = UNSET, | |||
| optimize_option = UNSET, | |||
| buffer_mode = UNSET, | |||
| force_npu_input_reposition = UNSET, | |||
| force_npu_output_reposition = UNSET, | |||
| image_channels = UNSET, | |||
| split_blocks = UNSET, | |||
| split_parts = UNSET, | |||
| Optional[str] | config_preset = None, | ||
| Optional[CompileConfig] | compile_config = None, | ||
| Optional[ResourceManagementConfig] | resource_management_config = None, | ||
| Optional[CalibrationConfig] | calibration_config = None, | ||
| Optional[BitConfig] | bit_config = None, | ||
| Optional[LlmConfig] | llm_config = None, | ||
| Optional[HessianQuantConfig] | hessian_quant_config = None, | ||
| Optional[ModConfig] | mod_config = None, | ||
| Optional[EquivalentTransformationConfig] | equivalent_transformation_config = None, | ||
| Optional[SearchWeightScaleConfig] | search_weight_scale_config = None, | ||
| Optional[SaveSampleConfig] | save_sample_config = None, | ||
| Optional[Uint8InputConfig] | uint8_input_config = None, | ||
| Optional[PreprocessingConfig] | preprocessing_config = None, | ||
| bias_correction = UNSET, | |||
| Optional[BiasCorrectionConfig] | bias_correction_config = None, | ||
| Optional[str] | config_save_path = None, | ||
| * | , | ||
| Optional[ExtraOutputConfig] | extra_output_config = None, | ||
| ** | kwargs ) |
Compile an existing Mobilint IR (.mblt) file into an MXQ package.
This is the pre-parsed half of mxq_compile(): parsing already happened (via mblt_compile()), so this function only quantizes and compiles. Quantization parameters, their defaults, and the configuration precedence rules are identical to mxq_compile() — see that function for the full reference and usage examples.
Parser-only arguments (save_subgraph_type, output_subgraph_path, feed_dict, dynamic_axes, multi_shape, yolo_decode_include, exclude_first_subgraph) are absent because the graph is already parsed. mxq_compile() accepts them for backward compatibility and drops them when routing here; either entry point logs a warning naming the ones that were given, so a multi_shape or feed_dict that cannot apply is not mistaken for one that did.
- Parameters
-
mblt string or list of strings. Path to a runnable .mblt produced bymblt_compile(); the file is validated up front and asave_subgraph_typepreview export is rejected. Or a list of such paths: every file is quantized in one call and packed into a single.mxq, as amulti_shapecompile does with the files it writes per size, andcalib_data_pathtakes one calibration directory per file, in the same order. The files must be one graph at different sizes; before quantization the pipeline checks that every file has the same NPU/CPU partition and the same layer types in graph order, and refuses the set with the first difference otherwise. Options and weights are not compared – shape-dependent options and constant folding legitimately differ per size.target_device string. Target NPU device (for example "aries-rb"). backend string. Retained for backward compatibility and ignored: the graph is already parsed, so no parser backend is selected.
- Returns
- None.
- Exceptions
-
ValueError mbltis not an existing.mblt file, or a list has an entry that is not one; a file is not a runnable artifact for the requestedcpu_offloadsetting; or the files of a list are not the same graph.
Definition at line 565 of file frontend.py.
◆ mblt_compile()
| qbcompiler.frontend.mblt_compile | ( | str | Any | model, |
| str | list[str] | mblt_save_path, | ||
| str | target_device, | ||
| backend = "onnx", | |||
| device = "cpu", | |||
| feed_dict = None, | |||
| dynamic_axes = None, | |||
| Optional[dict] | multi_shape = None, | ||
| yolo_decode_include = False, | |||
| cpu_offload = False, | |||
| exclude_first_subgraph = False, | |||
| Optional[str] | model_part = None, | ||
| Optional[dict] | model_part_options = None, | ||
| ** | kwargs ) |
Export a model to the Mobilint .mblt format without producing an MXQ package.
- Parameters
-
model string or model instance. Source model or path to compile. mblt_save_path string or list of strings. Output path for the .mblt artifact; with multi_shape, one path per value.backend string. Framework identifier, case-insensitive. "onnx", "torch", "tf" (also spelled "tensorflow" or "keras"), "tflite" and "torchscript" are all parsed by the current parser; a "torchscript" archive is exported to ONNX first and needs feed_dict. Any other value raisesValueError. Defaults to "onnx".device string. Compilation device ("cpu" or "gpu"). Defaults to "cpu". feed_dict dict. Example inputs used for shape inference. multi_shape dict. Compile the model at several sizes of a dynamic dimension in one call. Keyed by input name like feed_dict; each entry names the axis (an int, or a tuple of axes that move together) and the sizes it takes: {"x": {"axis": 3, "values": [100, 200, 300]}} or {"x": {"axis": (2, 3), "values": [(224, 224), (448, 448)]}}. Requiresfeed_dict:each value yields its own example tensor derived from it, and the model is parsed once per value into its own .mblt. Several inputs are walked together by index, so their "values" lists must be equally long. Pass onemblt_save_pathper value (a list), or a single path to get <stem>_<i>.mblt. Cannot be combined withdynamic_axes. Withmodel_partthe part is built once per size from that size's example tensors, so a part that sizes itself from its feed (Qwen3-ASR audio's "mel_chunk") compiles at every listed size.dynamic_axes dict. Declares dynamic axes per input name. yolo_decode_include bool. Runs YOLO decode on NPU when True. cpu_offload bool. Enables CPU offloading for unsupported groups. exclude_first_subgraph bool. Applies only when CPU offloading is enabled: exclude the first subgraph from the final graph if it is unsupported. model_part string. Selects a specific part of a torch model to parse ("vision", "language", "encoder", ...). Typically used to support models with complex architectures (e.g. Qwen3-VL) by compiling one part at a time. Parts a model declares: parser.patcher.parts.available_parts(model). A model declaring exactly one part resolves it from None; one declaring several requires a name. model_part_options dict. Extra arguments for the named part (e.g. {"mel_frames": 100} for Qwen3-ASR audio). kwargs dict. Additional arguments forwarded to the compiler.
- Returns
- None.
Definition at line 694 of file frontend.py.
◆ _mblt_compile_dispatch()
|
protected |
Route a compile-to-mblt call to the compile pipeline.
Shared by :func:mblt_compile and :func:mblt_compile_with_callback so both validate target_device and normalize backend the same way. device / cpu_offload default to UNSET here rather than to concrete values: that is what lets the callback wrapper forward only the flags its caller actually set and leave the rest to the merged CompileConfig. :func:mblt_compile keeps its own historical "cpu" / False defaults and passes them explicitly.
Definition at line 761 of file frontend.py.
◆ _mxq_compile_pipeline()
|
protected |
Quantize and compile through the ConfigManager-resolved pipeline.
Serves the backends in MODEL_DICT_BACKENDS; the callers (:func:mxq_compile_from_source, :func:mxq_compile_from_mblt) have already validated target_device and normalized backend.
Definition at line 808 of file frontend.py.
◆ _mblt_compile_pipeline()
|
protected |
Export to .mblt through the ConfigManager-resolved pipeline.
compilation_mode is an internal knob: one of {"release", "dev", "debug"} or None (default) to defer to the MBLT_APP_ENV environment variable. When set it overrides the env and drives the parser's inference_validation / device_alloc / log_level cascade — the same preset as MBLT_APP_ENV. Not surfaced in the CLI on purpose; intended for in-process scripts and tests that want scoped dev-mode validation without mutating process-wide environment.
Definition at line 915 of file frontend.py.
◆ mblt_compile_with_callback()
| None qbcompiler.frontend.mblt_compile_with_callback | ( | str | model, |
| str | mblt_save_path, | ||
| str | target_device, | ||
| str | backend = "onnx", | ||
| Optional[str] | device = None, | ||
| Optional[bool] | cpu_offload = None, | ||
| Optional[str] | config_preset = None, | ||
| Optional[CompileConfig | str] | compile_config = None, | ||
| Optional[str] | compilation_mode = None, | ||
| * | , | ||
| Callable[[int, str], None] | progress_callback ) |
Compile-to-mblt entry point with progress callbacks.
Only forwards explicitly-set values so the UNSET defaults (and the merged CompileConfig behind them) apply when the caller omits a flag – which is why this goes through _mblt_compile_dispatch rather than :func:mblt_compile, whose own device / cpu_offload defaults would override the config. target_device is required and has no config fallback — callers (e.g. the CLI) validate it via validate_target_device before invoking (QC-228).
Definition at line 968 of file frontend.py.
◆ mxq_compile_with_callback()
| None qbcompiler.frontend.mxq_compile_with_callback | ( | str | model, |
| str | target_device, | ||
| str | save_path, | ||
| str | backend = "onnx", | ||
| Optional[str] | device = None, | ||
| Optional[Union[str, List[str]]] | calib_data_path = None, | ||
| Optional[bool] | use_random_calib = None, | ||
| Optional[str] | config_preset = None, | ||
| Optional[CompileConfig | str] | compile_config = None, | ||
| Optional[str] | config_save_path = None, | ||
| * | , | ||
| Callable[[int, str], None] | progress_callback ) |
Compile/quantize entry point with progress callbacks.
target_device is required and has no config fallback — callers (e.g. the CLI) validate it via validate_target_device before invoking, and the wrapper forwards it unconditionally (QC-228).
Like :func:mblt_compile_with_callback, only explicitly-set values are forwarded, so :func:mxq_compile's UNSET defaults (and the merged CompileConfig behind them) still apply to whatever the caller omits. config_save_path passes straight through, and the resolved config is written there before compilation starts; this is what backs the CLI's --config-save-path on quantize / compile.
Definition at line 1014 of file frontend.py.
◆ get_yaml()
| qbcompiler.calibration.make_calib_data.get_yaml | ( | str | yaml_path | ) |
Definition at line 29 of file make_calib_data.py.
◆ make_calib_man()
| qbcompiler.calibration.make_calib_data.make_calib_man | ( | pre_ftn, | |
| str | data_dir, | ||
| str | save_dir, | ||
| str | save_name, | ||
| str | anno_json = None, | ||
| str | file_format = "%012d.jpg", | ||
| int | max_size = -1, | ||
| bool | remove_npy = False, | ||
| int | seed = 2023, | ||
| bool | save_calib_msg = False, | ||
| str | msg_path = None ) |
From images and a user-provided preprocessing function, generate calibration NumPy files and a manifest.
- Parameters
-
pre_ftn Callable. Pre-processing function that accepts an image path and returns a NumPy array. data_dir string. Directory containing the raw calibration images. save_dir string. Directory in which to store the generated NumPy files and metadata text file. save_name string. Basename for the saved assets. NumPy files are written to {save_dir}/{save_name}_npyand the manifest to{save_dir}/{save_name}.txt. Defaults to the basename ofdata_dirwhen omitted.anno_json string. Optional COCO-style annotation file. When provided, samples are randomly selected while maintaining class balance. file_format string. Filename format that uses image_idx(for example "%012d.jpg").max_size int. Maximum number of calibration samples. A value of -1 disables the limit. remove_npy bool. Remove existing NumPy outputs before generating new data. seed int. Random seed applied to sampling and class-balanced selection. Defaults to 2023. save_calib_msg bool. Persist the calibration metadata as an MSGpack file in addition to the manifest. msg_path string. Explicit MSGpack output path. When omitted, the path is auto-generated by dataset name and sample count.
- Returns
- string. Path to the generated text file that lists the preprocessed NumPy samples.
- Note
- Calibration samples are written as float32 NumPy arrays derived from the
pre_ftnoutput. Whenmax_sizeis positive and the dataset is larger, a random subset is used. COCO annotations override this sampling to balance categories.
Definition at line 40 of file make_calib_data.py.
◆ make_calib()
| str qbcompiler.calibration.make_calib_data.make_calib | ( | Union[str, dict] | args_pre, |
| str | data_dir, | ||
| str | save_dir, | ||
| str | save_name = None, | ||
| str | anno_json = None, | ||
| str | file_format = "%012d.jpg", | ||
| int | max_size = -1, | ||
| bool | remove_npy = False, | ||
| int | seed = 2023, | ||
| bool | save_calib_msg = False, | ||
| str | msg_path = None ) |
Create calibration NumPy files and a manifest using a YAML or dict preprocessing configuration.
- Parameters
-
args_pre string or dict. Path to a YAML file or an in-memory dictionary describing the preprocessing pipeline. See the preproc_builder class for all supported operators and parameters. data_dir string. Directory of calibration images. save_dir string. Directory where the generated NumPy files and manifest text file are stored. save_name string. Optional basename for outputs. Defaults to the basename of data_dir.anno_json string. Optional COCO-format annotation file used to sample images with class balance. file_format string. Filename template built from image_idx(defaults to "%012d.jpg").max_size int. Maximum number of calibration samples to generate; -1 keeps all available items. remove_npy bool. Remove existing NumPy outputs before running preprocessing. seed int. Random seed used for sampling. Defaults to 2023. save_calib_msg bool. Save calibration metadata as MSGpack alongside NumPy and text outputs. msg_path string. Explicit MSGpack output path. Auto-generated from dataset name and sample count when omitted.
- Returns
- string. Path to the manifest text file that lists the generated NumPy samples.
- Note
- The preprocessing pipeline is constructed through
preproc_builderand executed viamake_calib_man. Use the YAML schema to combine primitives such as GetImage, Resize, Normalize, Pad, CenterCrop, and SetOrder in a Pre-Order list.
Definition at line 139 of file make_calib_data.py.
◆ make_calib_llm()
| List[str] qbcompiler.calibration.make_calib_data.make_calib_llm | ( | Union[str, dict] | args_pre, |
| str | dataset_name, | ||
| Union[List[str], str] | subset_list, | ||
| str | text_field, | ||
| str | split, | ||
| str | save_dir, | ||
| int | max_calib = 512, | ||
| Optional[str] | save_prefix = None ) |
Generate calibration embeddings for LLM models using Hugging Face datasets.
Args: args_pre: YAML path or preprocessing configuration dictionary. Must set Datatype to "LLM". dataset_name: Name of the Hugging Face dataset to load. subset_list: List of dataset subsets (e.g. language splits) to iterate. split: HF dataset split to use. text_field: Key containing text in dataset entries. save_dir: Root directory where calibration samples are stored. max_calib: Maximum number of calibration samples per subset. save_prefix: Override prefix used in output directory naming. If None, derives from tokenizer directory name when available.
Returns: List of paths to generated calibration txt files, one per subset.
Definition at line 200 of file make_calib_data.py.
◆ preprocess_fnc()
| qbcompiler.calibration.preprocess.preprocess_fnc | ( | fnc, | |
| Datatype = "Image" ) |
Definition at line 28 of file preprocess.py.
Variable Documentation
◆ DTYPE_GETDATA_MAP
| dict qbcompiler.calibration.preprocess.DTYPE_GETDATA_MAP = {"Image": "GetImage", "Npy": "GetNpy", "LLM": "GetText"} |
Definition at line 24 of file preprocess.py.
Generated by