BACK TO MAIN PATHPARALLEL HORIZONS
LESSON 05RESOURCE ISOLATION
READ
  1. 01CONTEND
  2. 02ISOLATE
  3. 03FEED
  4. 04COMPLETE

STEP 01 · CONTEND

THE ONLINE REQUEST ARRIVED. A LARGE BATCH ALREADY OWNS THE GPU.

The offline job wants throughput while online inference must return inside a latency budget. They already use two CUDA Streams, yet still share the same compute and memory resources.
SHARED GPU · TWO WORKLOADS2 STREAMS · 1 RESOURCE POOL
01cudaStream_t batch, online;
02batchKernel<<<largeGrid, block, 0, batch>>>(archive);
03// request arrives while batchKernel is active
04infer<<<smallGrid, block, 0, online>>>(request);
05cudaStreamSynchronize(online);
BOTH WORKLOADS READY TO SUBMIT0 / 8 TICKS
RESOURCE OCCUPANCY TRACECONCEPTUAL SCHEDULE
ONLINE LATENCY BUDGET≤ 4 TICKSOBSERVED COMPLETION

TICK and the 24 cells are teaching units. They show ordering and contention, not an SM count, millisecond latency, or performance measurement.