############################################################################ annotation pipeline will be started with the following parameters: general parameters library file= /faststorage/project/PAN_illumina/people/thomas/2026_03_18_RNAseq_HsBam_forGEO.txt - all 6 libraries are available type of libraries= RNAseq storage location= /faststorage/project/PAN_illumina/results//thomasgrnbk/dm6/RNAseq/2026-03-18-total_RNAseq_SLX-19644_GEO/ tmp storage location= /faststorage/project/PAN_illumina/tmp//thomasgrnbk/dm6/RNAseq/2026-03-18-total_RNAseq_SLX-19644_GEO/ processing parameters genome assembly version= dm6 genome annotation version= r6.63 Y-chromosome= excluded from analysis length filtering= 18 - 1000 trimming= Yes 6 - 200 MM allowed in mapping= 2 track normalization= track normalization to 10M uniquely mapping reads read extension= 0 filtered read classes= rRNA:tRNA:mito color for track= 0,128,0 ############################################################################ start command: /faststorage/project/PAN_illumina/backup/scripts/AnnotationPipeline/annotate_reads.sh -i /faststorage/project/PAN_illumina/people/thomas/2026_03_18_RNAseq_HsBam_forGEO.txt -F total_RNAseq_SLX-19644_GEO -t RNAseq -v dm6 -V r6.63 -m 18 -M 1000 -f 6 -l 200 -s 2 -e 0 -J 1000000 Version= 4.0 Release Branch CommitID= ############################################################################ libraries analyzed in this run: /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext03.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_piwiKD_00h_rep1 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext20.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_piwiKD_00h_rep2 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext01.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_whiteKD_00h_rep1 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext05.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_whiteKD_00h_rep2 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext02.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_rhiKD_00h_rep1 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext19.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_rhiKD_00h_rep2 all 6 libraries are available ############################################################################ libraries as supplied by the user: /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext20.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_piwiKD_00h_rep2 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext19.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_rhiKD_00h_rep2 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext05.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_whiteKD_00h_rep2 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext03.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_piwiKD_00h_rep1 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext02.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_rhiKD_00h_rep1 /faststorage/project/PAN_illumina/data/libSTORAGE/2023_02_THG_HSbam_RNAseq/ftp1.cruk.cam.ac.uk/SLX-19644.NEBNext01.HW7MTDRX2.s_2.r_1.fq.gz HsBam_totalRNAseq_whiteKD_00h_rep1 ############################################################################ help file of used version: usage: /home/thomasgrnbk/PAN_illumina/scripts/AnnotationPipeline/annotate_reads.sh options ############################################################################### This scripts analyses deep-sequencing libraries and generates several statistics and UCSC compatiple tracks. It can also convert raw-bam files located at NGS into adaptor clipped and N trimmed fasta files. Version= 4.0 Do not start the script in the background by adding "&" to the start command. to attach the screen again use the command screen -r A detailed description of the functionality and usage can be found in: "/faststorage/project/PAN_illumina/backup/scripts/AnnotationPipeline/"README.md OPTIONS: -h, --help Show this message OBLIGATORY [but only one of the two options at any time] -i, --input FILE File containing the information for all the libraries -u, --update PATH Update old AnnotationPipeline run to newest version [path to folder] OPTIONAL - available in interactive mode -F, --folder-name NAME Name attachment for the output folder -t, --type TYPE TYPE of LIBRARY -j, --slam Libraries contain SLAMseq T->C conversions -v, --genome-version VER Genome version (dm3, dm6, ASM...) -V, --version VER Annotation version -D, --dge Perform DGE analysis (RNAseq only) bam -> fasta --demux-only Only demultiplex NGS data without processing individual libraries -x, --force-preprocess Force preprocessing even if input not bam -N, --n-trimm NUM Trim N random nucleotides from each end -P, --raw-paired Also output raw paired file -p, --only-paired Only paired end file (implies -P and -Q) -Q, --fastq-out Output trimmed fastq -q, --fastq-out-raw Output raw demultiplexed fastq -A, --custom-adaptors Use custom adaptor sequences -r, --raw Only generate demultiplexed preprocessed files -R, --subsample NUM Subsample fraction of reads read-preprocessing -m, --min-length NUM Minimal read length after trimming -M, --max-length NUM Maximal read length after trimming -T, --trimm Enable trimming -f, --first NUM First base to keep -l, --last NUM Last base to keep --polya Remove polyA stretches -I, --invert Invert reads -S, --single-end Force single-end processing -2, --second-mate Only process 2nd mate -J, --max-count NUM Max sequence count to consider -y, --demux-fasta Enable sRBC demultiplexing mapping & analysis -Y, --y-chrom Include Y chromosome -s, --mismatches NUM Allowed mismatches -E, --random-multi Distribute multimappers evenly -W, --wig-norm Normalize wig to 10M uniquely mapped -w, --wig-fasta-norm Normalize wig to 10M input reads -L, --spike-in-norm Spike-in based normalization -z, --no-norm Disable normalization --extra-seq FILE fasta-File containing extra sequences to be added to the TE-histogram analysis --force-quant-unstranded Force quantification to GeTMM for unstranded libraries (e.g. ChIPseq) -5, --only-5end Wig from 5' end only -1, --ping-pong Ping‑pong/phasing analysis (sRNAseq) -n, --no-stranded Unstranded analysis -e, --extend NUM Extend reads to fragment length -b, --export-bam Export collapsed bam -B, --export-bam-uncollapsed Export uncollapsed bam -3, --export-salmon Export salmon raw files -c, --color RGB Track color (R,G,B) -a, --auto-scale Set flag to set wig tracks to auto scale in UCSC user management & data import -G, --geo Create GEO export -U, --user NAME Run as different user ############################################################################ total time = 45.45 min