Distributed GPU Serving & Scheduling

Distributed inference should make latency, throughput, and isolation predictable for real traffic—not merely maximize GPU utilization.

Model request shapes, parallelism, batching, routing, and failure domains together to engineer service levels.

01

Problem definition

Optimizing average token throughput can let long requests displace short-request time to first token.

When scheduling and model parallelism are designed separately, failure and retry costs become unpredictable.

  • Multiple models or tenants share a constrained GPU pool
  • Time-to-first-token, inter-token latency, and throughput SLOs must coexist
  • A single low-traffic model fits comfortably on one machine
  • Real request distributions and SLOs are undefined
02

Distributed GPU Serving & Scheduling: system plate

CONTROL

System 1

Build a load model reproducing input/output length, concurrency, priority, and bursts.

CONTROL

System 2

Compare tensor, pipeline, data parallelism, and prefill/decode separation by model shape.

EXECUTION

System 3

Configure cache-aware routing, admission control, and preemption by SLO class.

EXECUTION

System 4

Inject node, link, and worker failures to measure draining, retry, and recovery.

  1. N1 N2context
  2. N2 N3decision
  3. N3 N4evidence
A decision and validation view for Distributed GPU Serving & Scheduling; labels describe architecture, not a measured deployment result.
  1. The workflow begins with Build a load model reproducing input/output length, concurrency, priority, and bursts..
  2. It reaches an acceptance decision through p50/p95 first-token and inter-token latency.

03

Design method

Distributed inference should make latency, throughput, and isolation predictable for real traffic—not merely maximize GPU utilization.

  1. 01

    Stage 1

    Build a load model reproducing input/output length, concurrency, priority, and bursts.

  2. 02

    Stage 2

    Compare tensor, pipeline, data parallelism, and prefill/decode separation by model shape.

  3. 03

    Stage 3

    Configure cache-aware routing, admission control, and preemption by SLO class.

  4. 04

    Stage 4

    Inject node, link, and worker failures to measure draining, retry, and recovery.

04

Application scenarios

Hypothetical application scenario

Multiple models or tenants share a constrained GPU pool

Optimizing average token throughput can let long requests displace short-request time to first token.

APPROACH
Build a load model reproducing input/output length, concurrency, priority, and bursts.
BOUNDARY
Synthetic peak throughput is not presented as production SLO performance.
Hypothetical application scenario

Time-to-first-token, inter-token latency, and throughput SLOs must coexist

When scheduling and model parallelism are designed separately, failure and retry costs become unpredictable.

APPROACH
Compare tensor, pipeline, data parallelism, and prefill/decode separation by model shape.
BOUNDARY
GPU utilization alone cannot represent user latency or reliability.

05

Design choices

DecisionGainCostWatch
Multiple models or tenants share a constrained GPU poolBuild a load model reproducing input/output length, concurrency, priority, and bursts.Synthetic peak throughput is not presented as production SLO performance.p50/p95 first-token and inter-token latency
Time-to-first-token, inter-token latency, and throughput SLOs must coexistCompare tensor, pipeline, data parallelism, and prefill/decode separation by model shape.GPU utilization alone cannot represent user latency or reliability.Goodput and queue time
06

Distributed GPU Serving & Scheduling: system plate

  1. A single low-traffic model fits comfortably on one machine

    Synthetic peak throughput is not presented as production SLO performance.

    SIGNAL
    p50/p95 first-token and inter-token latency
    MITIGATION
    Compare tensor, pipeline, data parallelism, and prefill/decode separation by model shape.
  2. Real request distributions and SLOs are undefined

    GPU utilization alone cannot represent user latency or reliability.

    SIGNAL
    Goodput and queue time
    MITIGATION
    Configure cache-aware routing, admission control, and preemption by SLO class.
A decision and validation view for Distributed GPU Serving & Scheduling; labels describe architecture, not a measured deployment result.
  1. The workflow begins with Build a load model reproducing input/output length, concurrency, priority, and bursts..
  2. It reaches an acceptance decision through p50/p95 first-token and inter-token latency.

07

Validation plan

MeasureMethodPass conditionCaveat
p50/p95 first-token and inter-token latencyBuild a load model reproducing input/output length, concurrency, priority, and bursts.Repeated runs satisfy the acceptance threshold agreed during discoverySynthetic peak throughput is not presented as production SLO performance.
Goodput and queue timeCompare tensor, pipeline, data parallelism, and prefill/decode separation by model shape.Repeated runs satisfy the acceptance threshold agreed during discovery—
Isolation and recovery timeConfigure cache-aware routing, admission control, and preemption by SLO class.Repeated runs satisfy the acceptance threshold agreed during discovery—

08

Constraints and failure conditions

A single low-traffic model fits comfortably on one machine

Synthetic peak throughput is not presented as production SLO performance.

Real request distributions and SLOs are undefined

GPU utilization alone cannot represent user latency or reliability.

10

Durable deliverables

Distributed GPU Serving & Scheduling decision record
Distributed inference should make latency, throughput, and isolation predictable for real traffic—not merely maximize GPU utilization.Client-owned · Patty-reviewed
Validation harness and acceptance criteria
p50/p95 first-token and inter-token latency · Goodput and queue time · Isolation and recovery timeJointly maintained
Operations and recovery runbook
Synthetic peak throughput is not presented as production SLO performance. · GPU utilization alone cannot represent user latency or reliability.Operating-team owned

11

Terminology

Distributed GPU Serving & Scheduling
Model request shapes, parallelism, batching, routing, and failure domains together to engineer service levels.
Acceptance criterion
p50/p95 first-token and inter-token latency
Operating boundary
Synthetic peak throughput is not presented as production SLO performance.

REFERENCES

References and primary material

  1. vLLM documentation

    Primary material for the method and terminology.

  2. Ray Serve LLM

    Primary material for the method and terminology.