Skip to content

Replace the slurm script generation tool with a proper workflow #98

Description

@zonca

Most probably Nextflow, the advantage is that can track progress status, restart jobs that fail, create all the slurm scripts automatically.

https://www.zonca.dev/posts/2025-10-07-running-nextflow-on-expanse

Run at NERSC with workflow queue

By ChatGPT, needs testing

Here’s the shortest working recipe to run a Nextflow manager on NERSC’s workflow QOS (so it can stay up for weeks) while submitting tasks to Slurm.


1) One-time setup

  • Request access to the workflow QOS (quick form). ([NERSC Documentation]1)
  • Put your Nextflow work dir on $SCRATCH (Lustre) because Nextflow needs file locking there. ([NERSC Documentation]2)
  • Install Nextflow once in a scratch or home bin:
cd $SCRATCH/nextflow && curl -s https://get.nextflow.io | bash

([NERSC Documentation]2)


2) Nextflow config (tasks run on Slurm)

Create nextflow.config in your pipeline dir:

process {
  executor          = 'slurm'
  // NERSC-friendly scheduler polling/rate limits
  queueSize         = 15
  pollInterval      = '5 min'
  dumpInterval      = '6 min'
  queueStatInterval = '5 min'
  exitReadTimeout   = '13 min'
  killBatchSize     = 30
  submitRateLimit   = '20 min'

  // Default Slurm options for short tests; override per-process as needed
  clusterOptions    = '-q debug -t 00:30:00 -C cpu'
}

workDir = "${System.getenv('SCRATCH')}/nf-work"

singularity {
  enabled = true    // OR use Shifter below
}

// If you prefer Shifter at NERSC, use:
// shifter.enabled = true
// and set process.container = 'docker://your/image:tag'

The polling/ratelimit knobs and sample clusterOptions are from NERSC’s Nextflow guidance. Shifter is supported by Nextflow via the shifter scope if you’d rather not use Apptainer/Singularity. ([NERSC Documentation]2)

Override clusterOptions inside any heavy process, e.g.
clusterOptions = '-q regular -t 05:30:00 -C cpu'. ([NERSC Documentation]2)


3) Run the manager on the workflow queue via scrontab

Use NERSC scrontab with #SCRON headers to keep Nextflow running/restarting on the workflow QOS:

$SCRATCH/nf-launch.sh

#!/usr/bin/env bash
set -euo pipefail
cd /path/to/your/pipeline
$SCRATCH/nextflow/nextflow run main.nf -resume -with-report report.html -with-timeline timeline.html

scrontab -e entry (UTC times):

#SCRON -q workflow
#SCRON -C cron
#SCRON -A <your_account>
#SCRON -t 30-00:00:00
#SCRON -o $SCRATCH/nf-workflow-%j.out
#SCRON --open-mode=append
0 */1 * * * /bin/bash $SCRATCH/nf-launch.sh
  • This runs your manager under the workflow QOS (up to 90-day walltime; ¼ of a login node) and restarts hourly if interrupted. ([NERSC Documentation]1)

Notes & tips

  • For long campaigns, NERSC recommends placing the manager in workflow QOS rather than on a login node. If it ever stops (maintenance, etc.), just re-enable the scrontab entry and keep using -resume. ([NERSC Documentation]2)
  • For GPU tasks, remember to request GPUs in your per-process clusterOptions (e.g., --gpus-per-node=4) and use -C gpu. ([NERSC Documentation]3)
  • Always keep workDir on $SCRATCH. ([NERSC Documentation]2)

If you want, tell me your project account and whether you prefer Shifter or Apptainer, and I’ll drop in an exact nextflow.config + scrontab tailored to your pipeline.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions