.. _slurm_reference: Complete SLURM Directive Reference ================================== .. contents:: Table of Contents :local: :depth: 2 Introduction ------------ This guide provides a complete reference for the most important SLURM directives used on GridUnesp. Use it as a quick reference when creating your submission scripts. .. note:: For complete practical examples, see: :ref:`running_simulations`. Basic Script Structure ---------------------- Every SLURM script follows this structure: .. code-block:: bash #!/bin/bash #SBATCH [DIRECTIVES HERE] # Load modules module load software/version # Your commands ./your_program Essential Directives -------------------- These are the directives you will use in almost every job: Job Identification ^^^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--job-name=NAME`` - Job name (shown in squeue) - ``#SBATCH --job-name=my_simulation`` * - ``--output=FILE`` - Output file (stdout) - ``#SBATCH --output=job_%j.out`` * - ``--error=FILE`` - Error file (stderr) - ``#SBATCH --error=job_%j.err`` .. tip:: **Useful variables**: * ``%j`` = Job ID * ``%x`` = Job name * ``%u`` = Username Example: ``--output=logs/%x_%j.out`` creates ``logs/my_job_12345.out`` Compute Resources ^^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--nodes=N`` - Number of nodes - ``#SBATCH --nodes=2`` * - ``--ntasks=N`` - Total number of tasks/processes - ``#SBATCH --ntasks=16`` * - ``--ntasks-per-node=N`` - Tasks per node - ``#SBATCH --ntasks-per-node=28`` * - ``--cpus-per-task=N`` - CPUs per task (for threads) - ``#SBATCH --cpus-per-task=4`` Partition and Time ^^^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--partition=NAME`` - Partition to use - ``#SBATCH --partition=short`` * - ``--time=HH:MM:SS`` - Maximum execution time - ``#SBATCH --time=12:00:00`` * - ``--time=D-HH:MM:SS`` - Time with days - ``#SBATCH --time=2-00:00:00`` Memory ^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--mem=SIZE`` - Total memory per node - ``#SBATCH --mem=64G`` * - ``--mem-per-cpu=SIZE`` - Memory per CPU - ``#SBATCH --mem-per-cpu=4G`` .. warning:: **Do not use** ``--mem`` and ``--mem-per-cpu`` together! Choose one. Special Resources ^^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--gres=RESOURCE:N`` - Generic resources (GPUs) - ``#SBATCH --gres=gpu:1`` * - ``--gres=gpu:MODEL:N`` - Specific GPU - ``#SBATCH --gres=gpu:l40s:2`` Advanced Directives ------------------- E-mail and Notifications ^^^^^^^^^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--mail-type=EVENT`` - When to send e-mail - ``#SBATCH --mail-type=END,FAIL`` * - ``--mail-user=EMAIL`` - Your e-mail - ``#SBATCH --mail-user=you@unesp.br`` Available events for ``--mail-type``: * ``BEGIN`` - When the job starts * ``END`` - When the job finishes * ``FAIL`` - If the job fails * ``ALL`` - All events * ``NONE`` - No notifications Job Arrays ^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--array=INDICES`` - Create a job array - ``#SBATCH --array=1-100`` * - ``--array=LIST`` - Specific indices - ``#SBATCH --array=1,5,10,20`` * - ``--array=START-END:STEP`` - With increment - ``#SBATCH --array=0-100:5`` * - ``--array=INDICES%LIMIT`` - Maximum simultaneous - ``#SBATCH --array=1-1000%50`` **Environment variables in arrays**: * ``$SLURM_ARRAY_TASK_ID`` - Current job array index * ``$SLURM_ARRAY_JOB_ID`` - Job array ID * ``$SLURM_ARRAY_TASK_COUNT`` - Total number of jobs Job Dependencies ^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--dependency=TYPE:JOBID`` - Depends on another job - ``#SBATCH --dependency=afterok:12345`` Dependency types: * ``after:JOBID`` - After job starts * ``afterok:JOBID`` - After job finishes successfully * ``afternotok:JOBID`` - After job fails * ``afterany:JOBID`` - After job finishes (success or failure) * ``singleton`` - Only one job with this name at a time Resource Control ^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--exclusive`` - Exclusive node (not shared) - ``#SBATCH --exclusive`` * - ``--constraint=FEATURE`` - Requires specific feature - ``#SBATCH --constraint=avx2`` Task Distribution ^^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--distribution=TYPE`` - How to distribute tasks - ``#SBATCH --distribution=cyclic`` Distribution types: * ``block`` - Contiguous block * ``cyclic`` - Round-robin * ``plane`` - By plane Output Control ^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--open-mode=MODE`` - File opening mode - ``#SBATCH --open-mode=append`` Modes: * ``append`` - Add to existing file * ``truncate`` - Overwrite file Restart and Checkpoints ^^^^^^^^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--requeue`` - Allow automatic requeue - ``#SBATCH --requeue`` * - ``--no-requeue`` - Do not allow requeue - ``#SBATCH --no-requeue`` .. Limits and Restrictions ^^^^^^^^^^^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 30 40 30 * - Directive - Description - Example * - ``--nice=VALUE`` - Adjust priority (0-10000) - ``#SBATCH --nice=100`` * - ``--qos=QOS`` - Quality of Service - ``#SBATCH --qos=normal`` Complete Practical Examples --------------------------- Simple Serial Job ^^^^^^^^^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=serial_test #SBATCH --partition=short #SBATCH --ntasks=1 #SBATCH --cpus-per-task=1 #SBATCH --mem=4G #SBATCH --time=02:00:00 #SBATCH --output=serial_%j.out #SBATCH --error=serial_%j.err python my_script.py OpenMP Parallel Job ^^^^^^^^^^^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=openmp_job #SBATCH --partition=short #SBATCH --nodes=1 #SBATCH --ntasks=1 #SBATCH --cpus-per-task=28 #SBATCH --mem=64G #SBATCH --time=12:00:00 #SBATCH --output=openmp_%j.out export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK ./my_openmp_program MPI Parallel Job ^^^^^^^^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=mpi_job #SBATCH --partition=short #SBATCH --nodes=4 #SBATCH --ntasks-per-node=28 #SBATCH --mem-per-cpu=4G #SBATCH --time=24:00:00 #SBATCH --output=mpi_%j.out module load openmpi/4.1 mpirun -np 112 ./my_mpi_program Hybrid MPI + OpenMP Job ^^^^^^^^^^^^^^^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=hybrid_job #SBATCH --partition=medium #SBATCH --nodes=2 #SBATCH --ntasks-per-node=4 #SBATCH --cpus-per-task=7 #SBATCH --mem=64G #SBATCH --time=48:00:00 #SBATCH --output=hybrid_%j.out export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK module load openmpi/4.1 srun ./hybrid_program GPU Job ^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=gpu_job #SBATCH --partition=gpu #SBATCH --nodes=1 #SBATCH --ntasks=1 #SBATCH --cpus-per-task=8 #SBATCH --gres=gpu:1 #SBATCH --mem=32G #SBATCH --time=12:00:00 #SBATCH --output=gpu_%j.out module load cuda/11.8 ./cuda_program Multi-GPU Job ^^^^^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=multi_gpu #SBATCH --partition=gpu #SBATCH --nodes=1 #SBATCH --gres=gpu:4 #SBATCH --cpus-per-task=48 #SBATCH --mem=128G #SBATCH --time=24:00:00 #SBATCH --output=mgpu_%j.out module load cuda/11.8 ./multi_gpu_program Job Array ^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=array_job #SBATCH --partition=short #SBATCH --array=1-100 #SBATCH --ntasks=1 #SBATCH --mem=4G #SBATCH --time=01:00:00 #SBATCH --output=array_%A_%a.out # Process the file for this index INPUT_FILE="input_${SLURM_ARRAY_TASK_ID}.dat" OUTPUT_FILE="output_${SLURM_ARRAY_TASK_ID}.dat" ./process $INPUT_FILE > $OUTPUT_FILE Job Array with Concurrency Limit ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=limited_array #SBATCH --partition=short #SBATCH --array=1-1000%50 #SBATCH --ntasks=1 #SBATCH --mem=2G #SBATCH --time=00:30:00 #SBATCH --output=limited_%A_%a.out # Maximum of 50 jobs running simultaneously python process.py --id $SLURM_ARRAY_TASK_ID Jobs with Dependencies ^^^^^^^^^^^^^^^^^^^^^^ .. code-block:: bash # Job 1: Preprocessing JOBID1=$(sbatch --parsable preprocess.sh) # Job 2: Processing (depends on Job 1) JOBID2=$(sbatch --parsable --dependency=afterok:$JOBID1 process.sh) # Job 3: Post-processing (depends on Job 2) sbatch --dependency=afterok:$JOBID2 postprocess.sh Job with E-mail Notification ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=long_simulation #SBATCH --partition=medium #SBATCH --ntasks=28 #SBATCH --time=72:00:00 #SBATCH --mail-type=BEGIN,END,FAIL #SBATCH --mail-user=you@unesp.br #SBATCH --output=simulation_%j.out # Receive e-mail when the job starts, finishes or fails ./long_simulation SLURM Environment Variables --------------------------- SLURM defines several useful environment variables: Identification ^^^^^^^^^^^^^^ .. code-block:: bash $SLURM_JOB_ID # Job ID $SLURM_JOB_NAME # Job name $SLURM_JOB_USER # Submitting user $SLURM_SUBMIT_DIR # Submission directory Resources ^^^^^^^^^ .. code-block:: bash $SLURM_NTASKS # Number of tasks $SLURM_CPUS_PER_TASK # CPUs per task $SLURM_MEM_PER_NODE # Memory per node (MB) $SLURM_NNODES # Number of nodes Nodes and Location ^^^^^^^^^^^^^^^^^^ .. code-block:: bash $SLURM_NODELIST # List of allocated nodes $SLURM_NODE_ALIASES # Node aliases $SLURM_JOB_NODELIST # Job nodes Arrays ^^^^^^ .. code-block:: bash $SLURM_ARRAY_JOB_ID # Job array ID $SLURM_ARRAY_TASK_ID # Current task index $SLURM_ARRAY_TASK_MIN # Minimum index $SLURM_ARRAY_TASK_MAX # Maximum index Usage Example ^^^^^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --job-name=test #SBATCH --ntasks=4 echo "Job ID: $SLURM_JOB_ID" echo "Job Name: $SLURM_JOB_NAME" echo "Number of tasks: $SLURM_NTASKS" echo "Allocated nodes: $SLURM_NODELIST" echo "Submission directory: $SLURM_SUBMIT_DIR" Useful SLURM Commands --------------------- Submission and Control ^^^^^^^^^^^^^^^^^^^^^^ .. code-block:: bash # Submit a job sbatch script.sh # Submit with option overrides sbatch --time=10:00:00 --mem=32G script.sh # Cancel a job scancel JOBID # Cancel all your jobs scancel -u $USER # Cancel by name scancel --name=job_name Monitoring ^^^^^^^^^^ .. code-block:: bash # View queue squeue # View only your jobs squeue -u $USER # View details of a job scontrol show job JOBID scontrol show job 123456 # View partition information sinfo # View resource usage sstat --format=JobID,MaxRSS,AveCPU JOBID.batch sstat --format=JobID,MaxRSS,AveCPU 14321.batch History and Accounting ^^^^^^^^^^^^^^^^^^^^^^ .. code-block:: bash # View completed jobs sacct # View details of a specific job sacct -j JOBID --format=JobID,JobName,State,Elapsed,MaxRSS # View jobs from the last 24 hours sacct --starttime=now-1day Tips and Best Practices ----------------------- Estimating Resources ^^^^^^^^^^^^^^^^^^^^ 1. **Start small**: Test with short jobs first 2. **Monitor actual usage**: Use ``sacct`` to see how much you actually used 3. **Adjust gradually**: Increase resources as needed Optimizing Wait Time ^^^^^^^^^^^^^^^^^^^^ 1. **Request realistic time**: Short jobs start faster 2. **Only request what you need**: Do not ask for 28 cores if you need 4 3. **Consider job arrays**: Instead of one giant job Avoiding Common Mistakes ^^^^^^^^^^^^^^^^^^^^^^^^ ❌ **Don't do this**: .. code-block:: bash #SBATCH --nodes=2 #SBATCH --ntasks=1 # ERROR: wastes resources since one node will be idle ✅ **Do this instead**: .. code-block:: bash #SBATCH --nodes=2 #SBATCH --ntasks-per-node=8 Organizing Outputs ^^^^^^^^^^^^^^^^^^ .. code-block:: bash #!/bin/bash #SBATCH --output=logs/%x_%j.out #SBATCH --error=logs/%x_%j.err # Create logs directory if it does not exist mkdir -p logs Debugging ^^^^^^^^^ For debugging, add information to the output: .. code-block:: bash #!/bin/bash #SBATCH ... echo "Job started at: $(date)" echo "Running on: $(hostname)" echo "Directory: $(pwd)" echo "Allocated nodes: $SLURM_NODELIST" echo "" # Your code here echo "" echo "Job finished at: $(date)" References ---------- * `SLURM Official Documentation `_ * `SLURM Cheat Sheet `_ * :ref:`running_simulations` - Practical job guide * :ref:`partitions_limits` - Limits and policies .. seealso:: **Other resources**: * :ref:`job_array` - More about job arrays * :ref:`shared_memory` - OpenMP * :ref:`distributed_memory` - MPI * :ref:`gpu_usage` - Using GPUs --- **Last updated about SLURM Directive References**: February 2025 **Tip**: Bookmark this page for quick reference!