BACK TO MAIN PATHPARALLEL HORIZONS
LESSON 06TRANSFORMER SCALING
READ
  1. 01CALIBRATE
  2. 02SHARD
  3. 03OVERLAP
  4. 04COMPLETE

STEP 01 · CALIBRATE

BEFORE REMOVING BITS, MAKE THE VALUES FIT THE FORMAT.

Transformer matrix operations already use Tensor Cores, but forward activations and backward gradients do not share one distribution. Casting every tensor to the same FP8 format clips large values or erases small ones.
TRANSFORMER ENGINE · FP8 RECIPEE4M3 FORWARD · E5M2 BACKWARD
01recipe = DelayedScaling(
02fp8_format=Format.E4M3,
03amax_history_len=0)
04with te.autocast(enabled=True, recipe=recipe):
05loss = transformer(tokens)
06loss.backward()
CHOOSE A NUMERICAL RECIPE
FORWARD / BACKWARD VALUE TRACECONCEPTUAL TENSOR DISTRIBUTION
CHOOSE AN FP8 RECIPE0 / 6

Low precision is not one dtype change. Observe tensor distributions, then choose formats, Scale factors, and which state remains at higher precision.

The bars illustrate range and scaling, not real model tensors or an accuracy measurement. Production training must validate loss, convergence, and downstream quality.