FOUR-BIT VALUES ARE TINY. THE SCALING STRATEGY MUST BE FINER.
Rack-scale inference first meets the storage and movement cost of weights and KV Cache. E2M1 has only four bits; if an entire tensor shares one coarse range, local outliers squeeze every other value.
LEVEL-1 MICRO-BLOCK SCALENONE×LEVEL-2 GLOBAL SCALENONE
CHOOSE A FOUR-BIT SCALING STRATEGY0 / 6
The fewer the bits, the more scaling granularity matters. Compression is a measurable tradeoff among capacity, bandwidth, and acceptable error—not a contest for the smallest dtype.
The code and 96 values teach structure; they are not a model-specific API or quality result. Real deployment requires a supported quantizer, calibration data, and target-task validation.