• Not collapsing for disk usage and BLAST function • Summarize results by myself • Quality control will be added in near future • NGS QC Toolkit (建樂學長) may be used • Automation & Parallelization NCBI SRA-Toolkit dump to fasta/q format Original datasets on GEO SRRxxx.sra, SRRyyy.sra, ... Fastq format with QC and sequencing details SRRxxx.fastq, SRRyyy.fastq, ... FastX Toolkit clip off 3' adapter Only clipped sequences left SRRxxx_clipped.fastq, SRRyyy_clipped.fastq, ... Quality Control discard low score reads Fasta Converter file format conversion Simpler file format: fasta SRRxxx_clipped.fasta, SRRyyy_clipped.fasta, ... BLAST+ make blast database Original datasets on GEO SRRxxx.sra, SRRyyy.sra, ... BLAST+ blastn query for every candidates on every dataset novel miRNA candidates candidate01.fa, candidate02.fa, ... Handmade Script summary all queries BLAST detail results for every query candidate01_xxx.csv, candidate02_xxx.csv, … candidate01_yyy.csv, candidate02_yyy.csv, … …, …, … Summarized read count for all candidates candidate01-10_xxx-zzz.csv Excel table output Script Automation
into csv (comma separate file), which can be imported into excel directly • like usual command line tool • argument take • standard lib of Python Bioinformatics and Biostatistics Core, NTU Center of Genomic Medicine 4
tissues • Illumina Genome Analyzer IIx • 32 patients and 62 samples in total • 21 adenocarcinoma patient (肺腺癌) • 11 squamous cell carcinoma patients (犬鱗狀上皮細胞癌?) • Normalize expression rate using sample size • start from raw sequence data • use read count after 3’ adapter clipped • adapter-only, too-short, N’s, no adapter clipped reads will be discarded. Bioinformatics and Biostatistics Core, NTU Center of Genomic Medicine 6
tool logged the reads in terms of their length, • After I computed the ratio of input(output) to number of sequence, • All ratios are exactly 53. • Still don’t know why, since they run the same argument • Probably due to the different platform they used Bioinformatics and Biostatistics Core, NTU Center of Genomic Medicine 8
to 建樂學長 for great help and inspiration • information at its supplementary file • so each sample has relatively small size • Illumina Genome Analyzer IIx • 185 unique samples and 54 samples in replicate • total 245 samples • includes IDC, ILC, DCIS, Apocrine, Adenoid, Metaplastic, Atypical Medullary => treated as ‘other’ type of samples • normal: 16, other: 229 Bioinformatics and Biostatistics Core, NTU Center of Genomic Medicine 10
Clipping data made by FastX Toolkit • ratio is not const as well • need to debug its source code • planning to replace this tool • However, fastq-dump and BLAST makedb gave the correct and reasonable read count • verify by counting number of lines in fasta file Bioinformatics and Biostatistics Core, NTU Center of Genomic Medicine 11
• Average expression rate of same type of samples • However, not all samples in one type have similar behavior Bioinformatics and Biostatistics Core, NTU Center of Genomic Medicine 13 ¯ Xn = 1 n n X i=0 xi Ni ⇥ 106 ¯ Xn : average expression rate n : number of samples of same type xi : read count in sample i Ni : size of sample i
• “AVERAGE” expression rate? • define the different trend of expression across cancer type • select possible miRNA candidates • not sure whether this method is correct • will normalized expression rate differs by sample size? • Original data should be included Bioinformatics and Biostatistics Core, NTU Center of Genomic Medicine 20