Notes on GPU utilization, inference performance, and what happens when production workloads share a fleet.