qbcompiler.configs.models.HessianQuantConfig Class Reference#
|
Mobilint SDK qb Compiler v1.4
MCS002-KR
|
qbcompiler.configs.models.HessianQuantConfig Class Reference
Configuration for HessianQuant algorithm. More...
Inheritance diagram for qbcompiler.configs.models.HessianQuantConfig:
Classes | |
| class | Attributes |
Public Member Functions | |
| "HessianQuantConfig" | with_updates (self, **kwargs) |
| Return a copy with updated fields. | |
Static Public Attributes | |
| model_config | |
| bool | apply = Field(default=False, description="If true, apply HessianQuant") |
| str | hessian_dtype |
| str | solver |
| bool | rescomp |
| Attributes | |
| default_factory | |
| alias | |
Detailed Description
Configuration for HessianQuant algorithm.
Defines parameters controlling whether and how HessianQuant is applied during quantization, including layer-level inclusion/exclusion lists.
- Parameters
-
apply bool. If true, apply HessianQuant hessian_dtype str. Storage dtype for the accumulated HessianQuant Hessian. bf16 halves its memory footprint (host RAM when the Hessian lives on CPU, VRAM when on GPU) — critical for large models such as MoE with thousands of expert FFN Hessians. Compute stays float32 regardless: the per-batch matmul and accumulation run in float32 and the solve upcasts back to float32; only the persistent accumulator is bfloat16. Accepted values:
"fp32" - Store the Hessian in float32. The default.
"bf16" - Store the Hessian in bfloat16 (half the memory).
solver str. HessianQuant solve algorithm used to compute per-layer integer weights. symmetric targets the output on the quantized input; the asymmetric solvers target the original-weight float output and differ in how they absorb the input mismatch. Accepted values:
"symmetric" - The default. min ||(W - Q) X_q||^2: sequential greedy rounding via a Cholesky-based Hessian solve.
"asymmetric_causal" - min ||W X_fp - Q X_q||^2: each column's input-mismatch residual is fed only to the columns after it, scaled by alpha.
"asymmetric_refit" - min ||W X_fp - Q X_q||^2: least-squares refit W*^T = H^-1 G W^T of every column, then a symmetric solve around W*.
rescomp bool. Enable residual-compensation correction. When true, each solver additionally applies an R correction term that compensates for accumulated weight drift from inter-block error propagation. Requires crossGram collection even for solver=symmetric. attributes Attributes. HessianQuant algorithm attributes
Member Function Documentation
◆ with_updates()
| "HessianQuantConfig" qbcompiler.configs.models.HessianQuantConfig.with_updates | ( | self, | |
| ** | kwargs ) |
Member Data Documentation
◆ model_config
|
static |
◆ apply
|
static |
◆ hessian_dtype
|
static |
Initial value:
= Field(
default="fp32",
alias="hessianDtype",
description="Storage dtype for the accumulated HessianQuant Hessian. bf16 halves its memory footprint (host RAM when the Hessian lives on CPU, VRAM when on GPU) — critical for large models such as MoE with thousands of expert FFN Hessians. Compute stays float32 regardless: the per-batch matmul and accumulation run in float32 and the solve upcasts back to float32; only the persistent accumulator is bfloat16.",
)
◆ solver
|
static |
Initial value:
= Field(
default="symmetric",
description="HessianQuant solve algorithm used to compute per-layer integer weights. symmetric targets the output on the quantized input; the asymmetric solvers target the original-weight float output and differ in how they absorb the input mismatch.",
)
◆ rescomp
|
static |
Initial value:
= Field(
default=False,
description="Enable residual-compensation correction. When true, each solver additionally applies an R correction term that compensates for accumulated weight drift from inter-block error propagation. Requires crossGram collection even for solver=symmetric.",
)
◆ Attributes
|
static |
◆ default_factory
|
static |
◆ alias
The documentation for this class was generated from the following file:
Generated by