Why Is My Job Not Starting?¶
On this Page
- Finding your job in the queue and reading its state
- What each pending reason means, and what to do about it
- Getting an estimated start time
- Making your next job start sooner
A job that is not running is not necessarily stuck. Slurm holds it until a node with the resources you asked for is free, and until it is your job's turn. The queue tells you which of those two is happening.
Where Is My Job?¶
Terminal
squeue --me
JOBID PARTITION NAME USER ST TIME NODES NODELIST(REASON)
669120 aoraki analysis abcde01 R 1:12:33 1 aoraki07
669121 aoraki_gpu train abcde01 PD 0:00 1 (Priority)
Two columns matter:
ST— the state.Ris running,PDis pending,CGis completing.NODELIST(REASON)— for a running job, the node it is on. For a pending job, the reason it is not running yet, in brackets.
What Does the Reason in Brackets Mean?¶
| Reason | What it means | What to do |
|---|---|---|
(Priority) |
Other jobs are ahead of yours | Wait, or ask for less so backfill can slot you in |
(Resources) |
Your job is next, but no node has enough free right now | Wait — it will start as jobs ahead of it finish |
(QOSMaxJobsPerUserLimit) |
You are already running the maximum number of jobs allowed in that partition. On GPU partitions that is 2 | Wait for one to finish, or cancel an idle session |
(QOSMaxGRESPerUser) |
You already hold the maximum GPUs allowed | As above |
(AssocMaxJobsLimit) |
An account-wide job limit is in force | Wait, or email rtis.support@otago.ac.nz |
(Dependency) |
The job it depends on has not finished yet | Nothing — unless the other job failed, in which case cancel this one |
(DependencyNeverSatisfied) |
The job it depended on failed or was cancelled, so this one will never run | scancel it and resubmit |
(ReqNodeNotAvail) |
A node you asked for is down, draining or reserved — often ahead of scheduled maintenance | Drop the explicit node request, or wait for the maintenance window to pass |
(PartitionTimeLimit) |
Your --time is longer than the partition allows |
Reduce --time, or use aoraki_long |
(PartitionConfig) |
Your request cannot be satisfied by any node in that partition — usually too many cores or too much memory | Reduce the request, or choose a partition that has nodes that size |
(JobHeldUser) / (JobHeldAdmin) |
The job is held. You can release your own with scontrol release <jobid> |
Release it, or contact rtis.support@otago.ac.nz for an admin hold |
(BeginTime) |
You asked for it to start later with --begin |
Nothing |
The current partition limits are in the Cluster Overview.
When Will My Job Start?¶
Terminal
squeue --me --start
This adds a START_TIME column with Slurm's estimate. scontrol shows the same thing for
one job, along with everything else Slurm knows about it:
Terminal
scontrol show job <jobid>
The estimate moves
It is a projection based on the jobs currently queued, and it assumes every one of them runs for its full requested wall time. Most finish early, so jobs usually start sooner than the estimate. New submissions and higher-priority work can push it the other way. Treat it as a rough guide, not a booking.
How Do I Make My Job Start Sooner?¶
In order of how much difference it makes:
- Ask for less time. Slurm backfills short jobs into gaps in the schedule, so a job asking for 2 hours has far more places to fit than the same job asking for 3 days. This is the single most effective change — see Why Asking for Less Starts Sooner.
- Ask for less memory and fewer cores. Both make the gap your job needs smaller. Use
seffon a previous run to find out what it really needed. - Use a different partition. If your work does not need the general-purpose nodes,
aoraki_shortandaoraki_smalluse cores that would otherwise sit idle on the GPU nodes, and are often much quicker to start. - Do not name a specific node with
--nodelistunless you genuinely need it. You are then waiting for one machine instead of dozens.
How Busy Is the Cluster?¶
Terminal
sinfo -o "%20P %5D %14F %10m %11l"
PARTITION NODES NODES(A/I/O/T) MEMORY TIMELIMIT
aoraki* 27 23/4/0/27 1030000+ 7-00:00:00
aoraki_bigcpu 10 8/2/0/10 1500000 14-00:00:00
aoraki_gpu 10 10/0/0/10 770000+ 7-00:00:00
The NODES(A/I/O/T) column is Allocated / Idle / Other / Total. A partition
showing 10/0/0/10 has nothing free.
To see which GPUs each node has, and how many are already taken:
Terminal
sinfo -p aoraki_gpu -o "%20P %10N %20G %30C"
Queue and Availability shows the same picture as graphs over the last week, including how long jobs on each partition have actually been waiting before they start — a better guide to when the cluster is usually quiet, and to whether a different partition would get you running sooner.
Why Did a Job I Submitted Later Start First?¶
Two mechanisms, both working as intended:
- Backfill. While Slurm holds nodes free for a large job at the front of the queue, it looks down the queue for smaller jobs that can start and finish before the large job is due to begin. A short job jumps ahead because it fits in the gap.
- Fairshare. Priority takes account of how much of the cluster you and your group have used recently. Heavy recent use lowers your priority relative to someone who has been idle.
Can I Ask for More Than a Partition Allows?¶
No. Aoraki enforces partition limits at submission, so a request over the limit is either
rejected immediately by sbatch or sits pending with (PartitionConfig) or
(PartitionTimeLimit) and never runs.
Check your request against the partition limits before submitting. If your work genuinely needs more than a partition allows, email rtis.support@otago.ac.nz — limits can be raised for a specific piece of work.