Supported Slurm settings in Cluster Director

This document describes the Slurm settings that you can configure when you create or modify a cluster in Cluster Director. Use these settings to customize the behavior of the Slurm scheduler and other Slurm components for the entire cluster, specific partitions, or specific nodesets to support your workloads.

To learn more about how Cluster Director uses Slurm to handle cluster management, job scheduling, and workload orchestration, see Slurm orchestration in Cluster Director.

Specify custom Slurm settings

When you create or modify a cluster, you can customize Slurm settings to adapt your cluster to your workload requirements. If you don't specify a value for a setting, then Cluster Director uses the default value listed in this document.

To specify custom Slurm settings, use one of the following methods:

  • REST API: specify settings in the orchestrator.slurm.config, orchestrator.slurm.partitions[].config, or orchestrator.slurm.nodeSets[].config fields in your request body.

The following sections list the supported Slurm settings for clusters, partitions, and compute nodes.

Cluster-level settings

The following table lists the supported cluster-level settings (orchestrator.slurm.config) for Slurm:

API field Slurm parameter Data type Description Default value
defMemPerCpu DefMemPerCPU String The default memory in megabytes (MB) that Slurm allocates for each CPU in your cluster. 0
enforcePartLimits EnforcePartLimits String Whether Slurm enforces partition limits when you submit or run jobs. NO
firstJobId FirstJobId String The job ID that Slurm assigns to the first job after the cluster starts. The value of each subsequent job ID increases by 1. 1
jobRequeue JobRequeue String Whether batch jobs return to the queue if a compute node encounters a host error. 1
overTimeLimit OverTimeLimit String The number of minutes that a job can run past its time limit before Slurm stops the job. 0
requeueExitCodes RequeueExit Array of integers The exit codes that cause Slurm to return a batch job to the queue. null (unset)
requeueHoldExitCodes RequeueExitHold Array of integers The exit codes that cause Slurm to return a batch job to the queue in a held state, where the job waits until you submit it again or delete it. null (unset)
schedulerParameters SchedulerParameters JSON object (Message) The behavior of the Slurm job scheduler. salloc_wait_nodes,ignore_prefer_validation,bf_continue,nohold_on_prolog_fail
selectTypeParameters SelectTypeParameters String How Slurm allocates compute resources to jobs. CR_Core_Memory
accountingStorageTres AccountingStorageTRES String The trackable resources (TRES), such as GPUs, that Slurm records in job usage logs. gres/gpu,gres/tpu
fairShareDampeningFactor FairShareDampeningFactor String A numerical factor that sets a priority on how nested accounts maintain fair access to cluster resources. 1
priorityCalcPeriod PriorityCalcPeriod String The time interval, in seconds, at which Slurm recalculates job priority. 300 (5 minutes)
priorityDecayHalfLife PriorityDecayHalfLife String The time it takes for past resource usage to decrease by half. 7-0 (seven days)
priorityFavorSmall PriorityFavorSmall String Whether Slurm gives small jobs higher priority in the queue than large jobs. NO
priorityFlags PriorityFlags String The configuration options that modify how Slurm calculates job priority. None (unset)
priorityMaxAge PriorityMaxAge String The time a job must wait in the queue to reach maximum age priority. 7-0 (seven days)
priorityType PriorityType String The plugin that Slurm uses to calculate job priority, such as priority/multifactor. priority/multifactor
priorityUsageResetPeriod PriorityUsageResetPeriod String The time interval at which Slurm resets historical resource usage records to zero. NONE (disabled)
priorityWeightAge PriorityWeightAge String The weight that Slurm assigns to queue wait time when it calculates job priority. 0
priorityWeightAssoc PriorityWeightAssoc String The weight that Slurm assigns to account associations when it calculates job priority. 0
priorityWeightFairshare PriorityWeightFairshare String The weight that Slurm assigns to fair-share resource usage when it calculates job priority. 0
priorityWeightJobSize PriorityWeightJobSize String The weight that Slurm assigns to requested job size when it calculates job priority. 0
priorityWeightPartition PriorityWeightPartition String The weight that Slurm assigns to the job partition when it calculates job priority. 0
priorityWeightQos PriorityWeightQOS String The weight that Slurm assigns to quality of service (QoS) when it calculates job priority. 0
priorityWeightTres PriorityWeightTRES String The weight that Slurm assigns to trackable resources (TRES), such as GPUs, when it calculates job priority. 0 (unset)
preemptExemptTime PreemptExemptTime String The minimum time a job must run before a higher-priority job can preempt the job. 0 (disabled)
preemptMode PreemptMode Array of strings The action that Slurm takes when it preempts a job, such as CANCEL or REQUEUE. OFF
preemptParameters PreemptParameters JSON object (Message) The configuration options that control how Slurm preempts jobs. None (unset)
preemptType PreemptType String The plugin that Slurm uses to select which jobs to preempt. preempt/none (unset)

Partition-level settings

The following table lists the supported partition-level settings (orchestrator.slurm.partitions[].config) for Slurm:

API field Slurm parameter Data type Description Default value
allowAccounts AllowAccounts String A comma-separated list of user accounts that can run jobs in your cluster partition. ALL
allowQos AllowQos String A comma-separated list of quality of service (QoS) levels that can run jobs in your cluster partition. ALL
defaultTime DefaultTime String The default runtime limit for jobs that don't specify a time limit in your cluster partition. None (unset)
defMemPerCpu DefMemPerCPU String The default memory in megabytes (MB) that Slurm allocates for each CPU in your cluster partition. 0 (unlimited)
denyAccounts DenyAccounts String A comma-separated list of user accounts that can't run jobs in your cluster partition. None (unset)
denyQos DenyQos String A comma-separated list of quality of service (QoS) levels that can't run jobs in your cluster partition. None (unset)
exclusiveUser ExclusiveUser String Whether Slurm reserves entire compute nodes for jobs from a single user. NO
graceTime GraceTime String The time, in seconds, that Slurm grants to a job before Slurm stops it during preemption. 0
maxTime MaxTime String The maximum runtime limit for jobs in your cluster partition. UNLIMITED
maxNodes MaxNodes String The maximum number of nodes that a job can request in your cluster partition. UNLIMITED
overSubscribe OverSubscribe String Whether multiple jobs can share compute resources on the same node. NO
overTimeLimit OverTimeLimit String The number of minutes that a job can run past its time limit in your cluster partition before Slurm stops the job. 0
preemptMode PreemptMode String The preemption mechanism for jobs in your cluster partition, which overrides the cluster-level setting. Inherits the value set in orchestrator.slurm.config.preemptMode
priorityJobFactor PriorityJobFactor String A numerical value that Slurm adds to the priority score of jobs in your cluster partition. 1
priorityTier PriorityTier String The priority tier that Slurm uses to schedule and preempt jobs across partitions. 1
qos QOS String The default quality of service (QoS) that Slurm assigns to jobs in your cluster partition. None (unset)
tresBillingWeights TRESBillingWeights String The factors that Slurm uses to calculate resource usage costs from trackable resources (TRES) in your cluster partition. CPU=1

Node-level settings

The settings in the following table correspond to Slurm NodeName parameters, not NodeSet directives. In Cluster Director, you apply these settings to all compute nodes in a nodeset by using the orchestrator.slurm.nodeSets[].config field. You can't configure individual compute nodes in a nodeset separately, but different nodesets can have different node settings.

The following table lists the supported node settings for Slurm:

API field Slurm parameter Data type Description Default value
cpuSpecList CpuSpecList String A comma-separated list of CPU IDs that Slurm reserves for system processes on each node in your nodeset. None (unset)
features Features String A comma-separated list of labels that you assign to nodes in your nodeset. None (unset)
memSpecLimit MemSpecLimit String The amount of memory in megabytes (MB) that Slurm reserves for system processes on each node in your nodeset. 0 (unset)
weight Weight String A priority weight that Slurm uses to select nodes for jobs. Slurm allocates nodes with lower weights before nodes with higher weights. 1
coreSpecCount CoreSpecCount String The number of CPU cores that Slurm reserves for system processes on each node in your nodeset. 0 (unset)

What's next