Phoebe: MIG trial on gpu2 and GPU node topology¶
On 6 May 2026 two of gpu2's eight A100s were split into small MIG instances. The trial was rolled back on 13–14 May. The CPU topology changes made at the same time stayed.
CPU topology of the GPU nodes¶
Several attempts on 6 May changed how Slurm sees the CPUs of gpu1 and gpu2:
GetEnvTimeout=2removed, andParameters=l3cache_as_sockettried ongpu2.- Replaced by the cluster-wide
SlurmdParameters=numa_node_as_socket. gpu1andgpu2redefined fromSockets=2 CoresPerSocket=32toSockets=8 CoresPerSocket=8: each NUMA node is now a socket.
With 8 "sockets" of 8 cores, each GPU can be bound to the cores closest to it.
MIG on gpu2¶
MIG was set up for one specific job that needed small GPU slices, and removed again once the test was over. It was never meant as a permanent configuration.
gpu2 was changed from Gres=gpu:a100:8 to Gres=gpu:a100:6,gpu:nvidia_a100_1g.10gb:14:
/dev/nvidia0–3and/dev/nvidia6–7stayed as full A100s./dev/nvidia4and/dev/nvidia5were each split into seven1g.10gbMIG instances (14 in total).gres.conflisted every device by hand (AutoDetect=off), with the cores closest to each GPU.
On 13–14 May gpu2 went back to Gres=gpu:a100:8, with the original one-line gres.conf
entry. MIG is disabled on all eight cards today.
How it was set up¶
Nothing was scripted and no systemd unit recreates the instances at boot, so MIG did not
survive a reboot. The steps below come from root's shell history on gpu2 and from
/etc/slurm on slurm1.
Two GPU numberings
nvidia-smi -i N uses the PCI bus order. /dev/nvidiaN in gres.conf uses the device
minor number. They differ on gpu2 and the minor numbers change between reboots: in May the
MIG cards were nvidia-smi -i 6 and -i 7 (= /dev/nvidia4 and /dev/nvidia5); today
-i 6 is /dev/nvidia0. Look up the current mapping before changing anything:
-
Drain the node and make sure nothing uses the two cards:
-
Enable MIG mode on the two cards (by
nvidia-smiindex). If the card is busy it needssystemctl stop nvidia-persistencedandnvidia-smi --gpu-reset -i N, or a reboot. -
Create seven
1g.10gbGPU instances (profile 19, seenvidia-smi mig -lgip) on each card, each with its compute instance (-C): -
On slurm1, change the
gpu2node line inslurm.conftoGres=gpu:a100:6,gpu:nvidia_a100_1g.10gb:14and list every device ingres.conf. Each MIG instance is the parent/dev/nvidiaNplus its two capability files; their numbers are in/proc/driver/nvidia-caps/mig-minorson gpu2 (gpu4/gi7/access 606,gpu4/gi7/ci0/access 607, ...):NodeName=gpu2 AutoDetect=off Name=gpu Type=a100 File=/dev/nvidia[0-1] Cores=8-15 NodeName=gpu2 AutoDetect=off Name=gpu Type=a100 File=/dev/nvidia[2-3] Cores=24-31 NodeName=gpu2 AutoDetect=off Name=gpu Type=a100 File=/dev/nvidia[6-7] Cores=56-63 NodeName=gpu2 AutoDetect=off Name=gpu Type=nvidia_a100_1g.10gb MultipleFiles=/dev/nvidia4,/dev/nvidia-caps/nvidia-cap606,/dev/nvidia-caps/nvidia-cap607 Cores=40-47 # ... one line per MIG instance, 7 on /dev/nvidia4 and 7 on /dev/nvidia5 -
Restart
slurmctldon slurm1 andslurmdon gpu2, then resume the node.
The full MIG gres.conf is kept on slurm1 as
/etc/slurm/gres.conf.bak-20260513190911-gpu2-physical (the matching slurm.conf has the same
suffix). It is also in the /etc/slurm git history (commits de315ae and 5dd77a1), and the
lines are still in the current gres.conf and slurm.conf, commented out.
How it was rolled back¶
Then the gpu2 lines in slurm.conf and gres.conf were restored to Gres=gpu:a100:8 and
the single gres.conf line, and slurmd was restarted.