TrimGalore

Description

According to the documentation of TrimGalore, this is a tool for consistent adapter trimming and removal of low-quality bases from next-generation sequencing (NGS) data. It has special handling for RRBS (Reduced Representation Bisulfite Sequencing) libraries.

It automates adapter detection, offers Phred-based quality trimming, support for paired-end data, and parallelization for efficient processing.

Available Versions

  • trim_galore/2.2.0 (default)

Note

To check the available versions:

module avail trim_galore

Loading the Module

# Load TrimGalore
module load trim_galore/2.2.0

# Verify installation
trim_galore --version

# Check help
trim_galore --help

Job Submission

Below are example scripts to use TrimGalore on GridUnesp. Always remember to define INPUT and OUTPUT for job-nanny.

Example 1: Single-end Processing

submit_trimgalore_se.sh
#!/bin/bash
#SBATCH -J trim_se
#SBATCH -N 1
#SBATCH -n 1
#SBATCH -t 12:00:00
#SBATCH --mem=8G

export INPUT="sample.fastq.gz"
export OUTPUT="sample_trimmed.fq.gz"

module load trim_galore/2.2.0

job-nanny trim_galore --fastqc sample.fastq.gz

Example 2: Paired-end Processing

submit_trimgalore_pe.sh
#!/bin/bash
#SBATCH -J trim_pe
#SBATCH -N 1
#SBATCH -c 8
#SBATCH -t 24:00:00
#SBATCH --mem=16G

export INPUT="sample_R1.fastq.gz sample_R2.fastq.gz"
export OUTPUT="sample_R1_val_1.fq.gz sample_R2_val_2.fq.gz"

module load trim_galore/2.2.0

job-nanny trim_galore --paired --cores $SLURM_CPUS_PER_TASK \
                      --fastqc sample_R1.fastq.gz sample_R2.fastq.gz

Example 3: Paired-end with Uncompressed Files

submit_trimgalore_pe_uncompressed.sh
#!/bin/bash
#SBATCH -J trim_pe_uncomp
#SBATCH -N 1
#SBATCH -c 8
#SBATCH -t 24:00:00
#SBATCH --mem=16G

export INPUT="sample_R1.fastq sample_R2.fastq"
export OUTPUT="sample_R1_val_1.fq sample_R2_val_2.fq"

module load trim_galore/2.2.0

job-nanny trim_galore --paired --cores $SLURM_CPUS_PER_TASK \
                      sample_R1.fastq sample_R2.fastq

Note

By default, output compression mirrors input compression: - Input .fastq.gz → output .fq.gz - Input .fastq → output .fq

Example 4: RRBS Mode (Bisulfite-seq)

submit_trimgalore_rrbs.sh
#!/bin/bash
#SBATCH -J trim_rrbs
#SBATCH -N 1
#SBATCH -c 4
#SBATCH -t 24:00:00
#SBATCH --mem=16G

export INPUT="bisulfite_R1.fastq.gz bisulfite_R2.fastq.gz"
export OUTPUT="bisulfite_R1_val_1.fq.gz bisulfite_R2_val_2.fq.gz"

module load trim_galore/2.2.0

job-nanny trim_galore --paired --rrbs --cores $SLURM_CPUS_PER_TASK \
                      bisulfite_R1.fastq.gz bisulfite_R2.fastq.gz

Example 5: Size Optimization with --clumpify

For data like ATAC-seq or amplicons that benefit from reordering for better compression:

submit_trimgalore_clumpify.sh
#!/bin/bash
#SBATCH -J trim_clump
#SBATCH -N 1
#SBATCH -c 8
#SBATCH -t 12:00:00
#SBATCH --mem=16G

export INPUT="amplicon_data.fastq.gz"
export OUTPUT="amplicon_trimmed.fq.gz"

module load trim_galore/2.2.0

# For intermediate files: --clumpify
# For storage: add --compression 6
job-nanny trim_galore --clumpify --cores $SLURM_CPUS_PER_TASK \
                      --compression 6 amplicon_data.fastq.gz

Example 6: Poly-A Removal

For RNA-seq libraries enriched with poly-A:

submit_trimgalore_polya.sh
#!/bin/bash
#SBATCH -J trim_polya
#SBATCH -N 1
#SBATCH -c 8
#SBATCH -t 12:00:00
#SBATCH --mem=16G

export INPUT="rna_sample_R1.fastq.gz rna_sample_R2.fastq.gz"
export OUTPUT="rna_sample_R1_val_1.fq.gz rna_sample_R2_val_2.fq.gz"

module load trim_galore/2.2.0

job-nanny trim_galore --paired --polyA --cores $SLURM_CPUS_PER_TASK \
                      --fastqc rna_sample_R1.fastq.gz rna_sample_R2.fastq.gz

Example 7: Force Uncompressed Output

submit_trimgalore_dont_gzip.sh
#!/bin/bash
#SBATCH -J trim_nogzip
#SBATCH -N 1
#SBATCH -c 4
#SBATCH -t 12:00:00
#SBATCH --mem=8G

export INPUT="sample.fastq.gz"
export OUTPUT="sample_trimmed.fq"

module load trim_galore/2.2.0

job-nanny trim_galore --dont_gzip --cores $SLURM_CPUS_PER_TASK \
                      sample.fastq.gz

Job Array for Multiple Samples

submit_trimgalore_array.sh
#!/bin/bash
#SBATCH -J trim_array
#SBATCH --array=1-10
#SBATCH -N 1
#SBATCH -c 4
#SBATCH -t 12:00:00
#SBATCH --mem=8G

SAMPLES=(
    "sample1"
    "sample2"
    "sample3"
    "sample4"
    "sample5"
    "sample6"
    "sample7"
    "sample8"
    "sample9"
    "sample10"
)

SAMPLE=${SAMPLES[$SLURM_ARRAY_TASK_ID-1]}
export INPUT="${SAMPLE}.fastq.gz"
export OUTPUT="${SAMPLE}_trimmed.fq.gz"

module load trim_galore/2.2.0

job-nanny trim_galore --fastqc ${SAMPLE}.fastq.gz

Job Array for Paired-end (Multiple Samples)

submit_trimgalore_array_pe.sh
#!/bin/bash
#SBATCH -J trim_array_pe
#SBATCH --array=1-10
#SBATCH -N 1
#SBATCH -c 8
#SBATCH -t 24:00:00
#SBATCH --mem=16G

SAMPLES=(
    "sample1"
    "sample2"
    "sample3"
    "sample4"
    "sample5"
    "sample6"
    "sample7"
    "sample8"
    "sample9"
    "sample10"
)

SAMPLE=${SAMPLES[$SLURM_ARRAY_TASK_ID-1]}
export INPUT="${SAMPLE}_R1.fastq.gz ${SAMPLE}_R2.fastq.gz"
export OUTPUT="${SAMPLE}_R1_val_1.fq.gz ${SAMPLE}_R2_val_2.fq.gz"

module load trim_galore/2.2.0

job-nanny trim_galore --paired --cores $SLURM_CPUS_PER_TASK \
                      --fastqc ${SAMPLE}_R1.fastq.gz ${SAMPLE}_R2.fastq.gz

Main Options and Features

TrimGalore Options

Option

Abbrev./Value

Description

--paired

Paired-end mode

--cores N

-j N

Number of cores for parallel processing (up to ~8)

--fastqc

Runs integrated FastQC on trimmed files

--rrbs

Special mode for RRBS libraries

--polyA

Removes poly-A tails (recommended for RNA-seq)

--nextseq N

Quality trimming specific to 2-color instruments

--clumpify

Reorders reads for better compression (requires --cores >= 2)

--compression N

gzip compression level (1 to 9). Default: 1 (fastest)

--retain_unpaired

Keeps unpaired reads in separate files

--dont_gzip

Outputs plain text (uncompressed)

--length N

Discards reads shorter than N (default: 20)

--quality N

Phred quality for trimming (default: 20)

Tip

For paired-end data, consider using --retain_unpaired to avoid losing reads that lost their pair during trimming.

Output Files

Files generated by TrimGalore

Mode

Output files

Single-end

*_trimmed.fq.gz, *_trimming_report.txt, *_trimming_report.json

Paired-end

*_val_1.fq.gz, *_val_2.fq.gz (valid pairs)

Paired-end (with --retain_unpaired)

*_unpaired_1.fq.gz, *_unpaired_2.fq.gz (unpaired reads)

FastQC

HTML report and ZIP file with post-trimming quality

Optimization Tips

  1. Parallelism: Use --cores up to 8 for near-linear speed gains

  2. Compression: --compression 1 (default) is faster; --compression 6-9 reduces file size

  3. Clumpify: Benefits ATAC-seq, amplicons, RNA-seq; requires --cores >= 2

  4. Memory: TrimGalore is lightweight, but integrated FastQC may consume more memory

References

See also