BACK TO MAIN PATHPARALLEL HORIZONS
LESSON 07RACK-SCALE INFERENCE
READ
  1. 01COMPRESS
  2. 02PLACE
  3. 03RECOVER
  4. 04COMPLETE

STEP 01 · COMPRESS

FOUR-BIT VALUES ARE TINY. THE SCALING STRATEGY MUST BE FINER.

Rack-scale inference first meets the storage and movement cost of weights and KV Cache. E2M1 has only four bits; if an entire tensor shares one coarse range, local outliers squeeze every other value.
NVFP4 QUANTIZATION RECIPETEACHING FLOW · VALIDATE QUALITY
01recipe = Float4()
02with te.autocast(enabled=True, recipe=recipe):
03logits = model(tokens)
04quality = validate(logits, reference)
CHOOSE A FOUR-BIT SCALING STRATEGY
QUANTIZED WEIGHT MICRO-BLOCKS6 × 16 TEACHING VALUES
CHOOSE A FOUR-BIT SCALING STRATEGY0 / 6

The fewer the bits, the more scaling granularity matters. Compression is a measurable tradeoff among capacity, bandwidth, and acceptable error—not a contest for the smallest dtype.

The code and 96 values teach structure; they are not a model-specific API or quality result. Real deployment requires a supported quantizer, calibration data, and target-task validation.