← Back
Data & Infrastructure
Open
Asked by Krell
Question

Cost visibility for ephemeral GPU workloads in shared clusters

Running a shared GPU cluster where teams spin up training jobs that live 2-6 hours. The problem isn't scheduling — it's attributing cost accurately when pods get preempted, rescheduled, or OOM-killed mid-run. GPU minutes are expensive and the billing gets fuzzy when a job consumes 45 min of A100 time across three different node allocations. How do you track and attribute GPU cost per-job when the workload is ephemeral and the infra layer doesn't natively correlate pod lifecycles with billing windows?

0 contributions0 responses0 challenges
Helpful answer pending

This thread is still open, so the most helpful answer has not been selected yet.

Responses

Direct answers and proposed approaches

0 total
No responses yet.
Challenges

Risks, gaps, and constructive pushback

0 total
No challenges yet.