Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
7
5. What is the range of the E-values?
0.50
6. What is the biological significance of these hits / is there any biological
meaning?
While the hits are real, they mean that they represent sequences that actually exist in the
database. However, we know that our query sequences are completely random, which
means that they have no evolutionary relation to the hits. Due to the pool of data being so
vast, the only reason we came across our hits was due to pure chance. Our sequence is
very small, consisting of only 25bp and so its size allows for a higher probability of matching
with other sequences in the database.
Activity 3: Protein sequences and BLASTP
Sequence Manipulation Suite: Random Protein Sequence at
http://www.bioinformatics.org/sms2/random_protein.html
Generate three protein sequences of length 25 aa :
• The distribution of amino acids will be equal (5% prob) and this is different from true
biological sequences - however this is not important for this first part of the exercise.
• The process for which BLASTP selects candidate sequences for full Smith-Waterman
alignment is different from BLASTN. (BLASTN - a single short (11 bp +) perfect match hit is
needed. For BLASTP - a pair of "near match" hits of 3 aa within a 40 aa window is needed).
1. Go to "Protein BLAST" page at NCBI and choose blastp .
2. Paste in the sequences in FASTA format, and choose the "NR" database (this is the
protein version, consisting of translated CDS'es, UniProt etc).
3. Set the parameters as follows:
a. In the "Algorithm Parameters" section select BLOSUM62 as the
b. alignment matrix to use
c. Set the "Expect threshold" to 1000 (default: 10)
d. DISABLE the "Short queries" parameters (otherwise your selected parameters will be
ignored).
4. Perform the BLAST search.
My Random Sequence:
>random sequence 1 consisting of 25 residues.
RHLSFHAPQWYQDHSQKRGAANFVS