SketchSSM calibration: Qwen3.8 Flash-Next (NVFP4 weights)
calibration.pt is a portable SketchSSM
calibration file for Qwen3.8 Flash-Next. It contains the group-shared sketch basis and
the Full-Gram allocation scores from which the per-head rank table and the
ordered frames for any mean rank are derived. It contains no model weights.
Calibration weights
Collected with NVFP4 weights: RadixArk/Qwen3.8-Flash-Next-NVFP4 at revision 7b719225242aacd3dbd3f9407468c2ee9a9d2594 (an NVFP4 checkpoint quantized with Model Optimizer). The basis and the allocation scores depend on the weights, so use this file with these weights; for another precision or checkpoint, calibrate with that checkpoint.
Contents
| Field | Value |
|---|---|
| Base model | Qwen3.8 Flash-Next (Gated DeltaNet) |
| Recurrent layers | 36 |
| State heads per layer | 48 |
| Key dim K / value dim V | 128 / 128 |
| Basis groups per layer | 16 |
| Window W | 16 |
| Erase factor | yes |
| Allocation rank cap | 60 |
Basis omega |
float32, shape (36, 16, 78, 128) |
| File size | 24,098,949 bytes |
Layers are stored in the order of the model's recurrent layers. The file loads with torch.load(..., weights_only=True).
The offline calibration guide
documents its keys.
Verified mean ranks
For these mean ranks, the derived tables equal the published bundle tables and the exported frames equal those of the bundle export path:
| Mean rank | Dense heads | Sketch heads |
|---|---|---|
| 3 | 0 | 1728 |
| 4 | 0 | 1728 |
| 7 | 1 | 1727 |
| 11 | 2 | 1726 |
| 26 | 287 | 1441 |
Other mean ranks are allocated with the same rule but have no stored table to
compare with. manifest.json lists the SHA-256 of calibration.pt and these
results.
Usage
With the SketchSSM repository, export the frames for a mean rank:
hf download SketchSSM/Qwen3.8-Flash-Next-NVFP4 calibration.pt --local-dir calibration
python -m offline_calibration export --calibration calibration/calibration.pt \
--mean-rank 7 --out frames.pt
With vLLM (requires the SketchSSM vLLM fork with calibration-file support):
vllm serve RadixArk/Qwen3.8-Flash-Next-NVFP4 --sketchssm SketchSSM/Qwen3.8-Flash-Next-NVFP4 --sketchssm-mean-rank 7
How it was produced
From the public calibration bundle offline_calibration/example/qwen_flash_next in the
SketchSSM repository:
python -m offline_calibration package --bundle offline_calibration/example/qwen_flash_next --out calibration.pt
package re-allocates every configured mean rank from the packaged contents
and fails unless each table equals the bundle table.
License
This calibration file is released under the Apache License 2.0, like the SketchSSM repository. It is derived from the base model's weights, so use it under the base model's license as well.
Model tree for SketchSSM/Qwen3.8-Flash-Next-NVFP4
Base model
Qwen/Qwen3.8-Flash-Next