Phoebe: limiting GPUs per user per node¶
How to give a user a GPU limit per node (for example 3 GPUs on gpu1 and 3 on gpu2) instead
of one limit across the whole cluster, and why the setup tried on 18 September 2026 was removed
again the same day.
How Slurm counts GPUs¶
| Limit | What it counts |
|---|---|
QOS MaxTRESPU=gres/gpu=N |
GPUs of all running jobs of a user with that QOS, in every partition that uses it |
QOS MaxTRESPerJob |
each job on its own |
QOS MaxTRESPerNode |
each job on each of its nodes, not the sum of a user's jobs on a node |
Association GrpTRES=gres/gpu=N with Partition= |
GPUs of a user's running jobs in that one partition |
A QOS or an association can be tied to a partition, but not to a physical node. A QOS with
MaxTRESPU=gres/gpu=3 used on the shared gpu and gpu_int partitions (both contain gpu1
and gpu2) therefore means 3 GPUs per user in total. A job waits with QOSMaxGRESPerUser once
the user has 3 GPUs running on gpu2, even if gpu1 has free GPUs.
Setting up per-node limits¶
Phoebe also has the single-node partitions gpu1 and gpu2. A separate counter per node needs
one QOS per node, each used only in that node's partition:
sacctmgr add qos gpu1_max3 MaxTRESPU=gres/gpu=3 MaxTRESPerJob=gres/gpu=3 Flags=DenyOnLimit
sacctmgr add qos gpu2_max3 MaxTRESPU=gres/gpu=3 MaxTRESPerJob=gres/gpu=3 Flags=DenyOnLimit
sacctmgr add user USER cluster=phoebe account=ACCOUNT partition=gpu1 \
qos=gpu1_max3 defaultqos=gpu1_max3
sacctmgr add user USER cluster=phoebe account=ACCOUNT partition=gpu2 \
qos=gpu2_max3 defaultqos=gpu2_max3
Copy the CPU and memory limits of the user's current QOS into the new ones. Without
DenyOnLimit, a job that asks for too many GPUs is accepted and stays pending forever instead
of being rejected at submission.
The user then submits with --partition=gpu1 or --partition=gpu2. Their access to gpu and
gpu_int has to be closed, because jobs there are not counted against the per-node QOS. That
can be done with DenyQos= on those partitions in slurm.conf, or by removing the user's
gpu and gpu_int associations.
Check the result with sbatch --test-only: 3 GPUs on gpu1 should be accepted and 4 rejected,
and a job on gpu or gpu_int should be rejected.
Things to watch¶
- Jobs that are already running stay on their old QOS. Slurm changes the QOS only of
pending jobs (
scontrol update JobId=... QOS=...); for a running job it fails with "Job is no longer pending execution". Until those jobs end, their GPUs are not counted by the new QOS, so the new limit can be exceeded. Either wait for them to finish or keep the new QOS blocked on that node (for example withDenyQos=) until they do. - Moving a pending job to
gpu1orgpu2can mean a shorter time limit, because those partitions have a lowerMaxTimethangpuandgpu_int. - A failed
scontrol updatecan apply part of the change. Check the job withscontrol show jobafterwards. - There is no cumulative per-node limit on the shared partitions. As long as the user
submits to
gpuorgpu_int, no QOS or association setting limits their total GPUs on one node.
What was done on 18 September¶
The two QOS, the gpu1 and gpu2 associations and DenyQos= entries on the gpu, gpu1,
gpu2 and gpu_int partitions were set up and tested, but the new QOS were not switched on:
running jobs could not be moved to them. The setup was then dropped, because users should keep
submitting to gpu and gpu_int rather than pick a node themselves. The QOS, the associations
and the DenyQos= entries were removed, and the 3-GPU limit across both nodes (gpu_max2)
stayed as it was. See accounts, QOS and limits.