qbcompiler.configs.models.HessianQuantConfig Class Reference

qbcompiler.configs.models.HessianQuantConfig Class Reference#

Mobilint SDK qb Compiler: qbcompiler.configs.models.HessianQuantConfig Class Reference
Mobilint SDK qb Compiler v1.4
MCS002-EN
qbcompiler.configs.models.HessianQuantConfig Class Reference

Configuration for HessianQuant algorithm. More...

Inheritance diagram for qbcompiler.configs.models.HessianQuantConfig:

Classes

class  Attributes

Public Member Functions

"HessianQuantConfig" with_updates (self, **kwargs)
 Return a copy with updated fields.

Static Public Attributes

 model_config
bool apply = Field(default=False, description="If true, apply HessianQuant")
str hessian_dtype
str solver
bool rescomp
 Attributes
 default_factory
 alias

Detailed Description

Configuration for HessianQuant algorithm.

Defines parameters controlling whether and how HessianQuant is applied during quantization, including layer-level inclusion/exclusion lists.

Parameters
applybool. If true, apply HessianQuant
hessian_dtypestr. Storage dtype for the accumulated HessianQuant Hessian. bf16 halves its memory footprint (host RAM when the Hessian lives on CPU, VRAM when on GPU) — critical for large models such as MoE with thousands of expert FFN Hessians. Compute stays float32 regardless: the per-batch matmul and accumulation run in float32 and the solve upcasts back to float32; only the persistent accumulator is bfloat16. Accepted values:
"fp32" - Store the Hessian in float32. The default.
"bf16" - Store the Hessian in bfloat16 (half the memory).
solverstr. HessianQuant solve algorithm used to compute per-layer integer weights. symmetric targets the output on the quantized input; the asymmetric solvers target the original-weight float output and differ in how they absorb the input mismatch. Accepted values:
"symmetric" - The default. min ||(W - Q) X_q||^2: sequential greedy rounding via a Cholesky-based Hessian solve.
"asymmetric_causal" - min ||W X_fp - Q X_q||^2: each column's input-mismatch residual is fed only to the columns after it, scaled by alpha.
"asymmetric_refit" - min ||W X_fp - Q X_q||^2: least-squares refit W*^T = H^-1 G W^T of every column, then a symmetric solve around W*.
rescompbool. Enable residual-compensation correction. When true, each solver additionally applies an R correction term that compensates for accumulated weight drift from inter-block error propagation. Requires crossGram collection even for solver=symmetric.
attributesAttributes. HessianQuant algorithm attributes

Definition at line 770 of file models.py.

Member Function Documentation

◆ with_updates()

"HessianQuantConfig" qbcompiler.configs.models.HessianQuantConfig.with_updates ( self,
** kwargs )

Return a copy with updated fields.

Definition at line 850 of file models.py.

Member Data Documentation

◆ model_config

qbcompiler.configs.models.HessianQuantConfig.model_config
static
Initial value:
= ConfigDict(
populate_by_name=True,
extra="forbid",
)

Definition at line 789 of file models.py.

◆ apply

bool qbcompiler.configs.models.HessianQuantConfig.apply = Field(default=False, description="If true, apply HessianQuant")
static

Definition at line 794 of file models.py.

◆ hessian_dtype

str qbcompiler.configs.models.HessianQuantConfig.hessian_dtype
static
Initial value:
= Field(
default="fp32",
alias="hessianDtype",
description="Storage dtype for the accumulated HessianQuant Hessian. bf16 halves its memory footprint (host RAM when the Hessian lives on CPU, VRAM when on GPU) — critical for large models such as MoE with thousands of expert FFN Hessians. Compute stays float32 regardless: the per-batch matmul and accumulation run in float32 and the solve upcasts back to float32; only the persistent accumulator is bfloat16.",
)

Definition at line 795 of file models.py.

◆ solver

str qbcompiler.configs.models.HessianQuantConfig.solver
static
Initial value:
= Field(
default="symmetric",
description="HessianQuant solve algorithm used to compute per-layer integer weights. symmetric targets the output on the quantized input; the asymmetric solvers target the original-weight float output and differ in how they absorb the input mismatch.",
)

Definition at line 800 of file models.py.

◆ rescomp

bool qbcompiler.configs.models.HessianQuantConfig.rescomp
static
Initial value:
= Field(
default=False,
description="Enable residual-compensation correction. When true, each solver additionally applies an R correction term that compensates for accumulated weight drift from inter-block error propagation. Requires crossGram collection even for solver=symmetric.",
)

Definition at line 804 of file models.py.

◆ Attributes

qbcompiler.configs.models.HessianQuantConfig.Attributes
static

Definition at line 848 of file models.py.

◆ default_factory

qbcompiler.configs.models.HessianQuantConfig.default_factory
static

Definition at line 848 of file models.py.

◆ alias

qbcompiler.configs.models.HessianQuantConfig.alias
static

Definition at line 848 of file models.py.


The documentation for this class was generated from the following file: