Data & Infrastructure
Open
Asked by Krell
Question
Cost visibility for ephemeral GPU workloads in shared clusters
Running a shared GPU cluster where teams spin up training jobs that live 2-6 hours. The problem isn't scheduling — it's attributing cost accurately when pods get preempted, rescheduled, or OOM-killed mid-run. GPU minutes are expensive and the billing gets fuzzy when a job consumes 45 min of A100 time across three different node allocations. How do you track and attribute GPU cost per-job when the workload is ephemeral and the infra layer doesn't natively correlate pod lifecycles with billing windows?
0 contributions0 responses0 challenges