Skip to content

Routing tasks to a different queue

A submission takes one --queue-id, plus an optional --fallback-queue-id for spot-capacity retries. Every process in the workflow runs on that queue.

That’s a problem when only one or two processes are long-running or need capacity spot never holds onto: routing the whole run to an on-demand queue is expensive for every other task, and the fallback queue doesn’t really fix it either — it only kicks in after a task has already been interrupted, and only for one retry.

What you actually want is for most processes to keep running on spot, while a specific process — by name or by label — always runs on-demand from the first attempt. Nextflow already supports this, no GENI change required.

Nextflow config wins by directive priority, lowest to highest:

  1. generic process.* settings — this is where GENI sets process.queue
  2. a directive written directly in the process body
  3. a withLabel: selector
  4. a withName: selector

Config is merged by scope path, so GENI’s generic process.queue only overwrites other generic process.queue settings — it never erases a selector block in your workflow’s own nextflow.config. A withName: or withLabel: block that sets queue always wins for the processes it matches, regardless of what GENI’s config says.

Any active queue in the same environment as the engine will do — typically an on-demand one:

job-queue/geni-queue-a1b2c3d4
geni queue list
geni queue get <queue-id>

The simplest form. Add this to the nextflow.config bundled at the root of your workflow ZIP:

process {
withName: 'BWA_MEM|GATK_HAPLOTYPECALLER' {
queue = 'arn:aws:batch:us-east-1:123456789012:job-queue/geni-queue-a1b2c3d4'
}
}

The selector only replaces queue — errorStrategy and maxRetries are still inherited from GENI’s generic settings.

Recommended: a hardcoded ARN pins the workflow to one environment. Read it from the params file instead, in a closure so it resolves per task rather than at config-parse time:

nextflow.config
params.ondemandQueue = null // default: no override
process {
withLabel: 'long_running' {
queue = { params.ondemandQueue ?: task.config.queue }
}
}
main.nf
process GATK_HAPLOTYPECALLER {
label 'long_running'
cpus 16
memory '64.GB'
...
}

Then pass the ARN as a parameter at submission time:

{
"input": "s3://my-bucket/samples.csv",
"ondemandQueue": "arn:aws:batch:us-east-1:123456789012:job-queue/geni-queue-a1b2c3d4"
}
Terminal window
geni submission create \
--workflow-name wgs-germline \
--workflow-version 1.2.0 \
--engine-id <engine-id> \
--queue-id <spot-queue-id> \
--params-file params.json \
--output-folder s3://my-bucket/output/

Omit ondemandQueue and the labelled processes fall back to the normal spot queue, so the workflow stays runnable in any environment.

Option C — a directive in the process body

Section titled “Option C — a directive in the process body”

Also beats GENI’s generic setting, at the cost of hardcoding the ARN into the .nf file itself:

process GATK_HAPLOTYPECALLER {
queue 'arn:aws:batch:us-east-1:123456789012:job-queue/geni-queue-a1b2c3d4'
...
}
Terminal window
geni submission tasks --submission-id <id>
aws batch describe-jobs --jobs <jobId> --query 'jobs[].jobQueue'

The pinned process’s job should sit on the on-demand queue’s ARN from the very first attempt; everything else stays on the queue passed via --queue-id.