Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
1
PAIRWISE ALIGNMENT
Pairwise Sequence Alignment is used to identify regions of similarity that may indicate functional,
structural and/or evolutionary relationships between two biological sequences (protein or nucleic acid).
By contrast, Multiple Sequence Alignment (MSA) is the alignment of three or more biological sequences
of similar length. From the output of MSA applications, homology can be inferred and the evolutionary
relationship between the sequences studied.
Pairwise alignment is performed using an algorithm known as Dynamic Programming (DP), which has 2
main variants:
- Global alignment (Needleman-Wunsch) — the sequences are forced to be aligned across their
entire length. (http://www.ebi.ac.uk/Tools/psa/emboss_needle/help/index-protein.html)
- Local alignment (Smith-Waterman) — only the best matching sub-part of the sequences are
aligned).
(http://www.ebi.ac.uk/Tools/psa/emboss_water/help/index-protein.html)
Activity 1: Write a brief description about the two proteins P29600 and
P41363
P29600:
Is a protein called Subtilisin Savinase from the organism Lederbergia lentus (Bacillus lentus). It’s an
extracellular alkaline serine protea and its main function is to catalyzes the hydrolysis of proteins and
peptide amides.
P41363:
Is a protein called Thermostable alkaline protease that is coded from the gene BH0855 on the organism
Alkalihalobacillus halodurans (strain ATCC BAA-125 / DSM 18197 / FERM 7344 / JCM 9153 / C-125)
(Bacillus halodurans). Its only known function is that it shows keratinolytic activity.
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
2
Activity 2: Global Alignment
1. Go to EBI Pairwise Sequence Alignment (PSA) at
https://www.ebi.ac.uk/Tools/psa/emboss_needle/,
2. Go to Global Alignment âž” Needle (EMBOSS) âž” Protein Needle (Short for Needleman &
Wunsch)
3. Paste in the sequences - including the header line - one sequence in each search
4. box. Set format to pair
5. Go to "STEP 2" âž”"More options" âž” Matrix and set to "BLOSUM62";
6. Set End Gap Penalty to "true" (otherwise, the algorithm ignores gaps at the very ends of the
alignments, i.e. gaps positioned at the end of one of the sequences would not cost anything).
7. Finally, start the alignment by clicking the button labeled "Submit"
RESULTS OF THE ALIGNMENT:
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
3
1. What is the Alignment score?
860.5
2. What is the Alignment length?
361
3. What is the % and fraction Identity? (The value reported for "Identity" includes
perfect matches only)
176/269 (65.4%)
4. What is the % and fraction Similarity? (The value reported for "Similarity" includes
perfect matches + "close" mismatches)
214/269 (79.6%)
Results from NCBI:
Activity 3: Local Alignment
1. Go to https://www.ebi.ac.uk/Tools/psa/ âž” Local Alignment âž” Water (EMBOSS)
2. Select protein alignment.
3. Use protein set A
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
4
1. What is the Alignment score?
916.0
2. What is the Alignment length?
269
3. What is the % and fraction Identity? (The value reported for "Identity" includes
perfect matches only)
176/269 (65.4%)
4. What is the % and fraction Similarity? (The value reported for "Similarity" includes
perfect matches + "close" mismatches)
214/269 (79.6%)
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
5
5. Do you think one of them is more biologically relevant than the other, or do both
contribute valuable information?
Since the two sequences are of different lengths makes the most sense to use the Smith-
Waterman algorithm ("local alignment"), as this will provide an analysis of differences and
similarities for that part of the sequence which is actually comparable.
The sequences are very similar so there generally are not much difference in the information
you get from using local and global alignment.
6. How were the amino acid sequences of the two proteins determined? (Hint: look
at the titles of the papers, and the Cited for fields, listed in the Reference
sections).
The P29600 sequence is derived from 3D structure.
And P41363 is translated from DNA + information from protein sequencing.
7. Subcellular localization: Where in (or outside) the cell do the enzymes function?
Both of the proteins are Secreted proteins.
8. The feature table contains details about the regions/domains of the protein - try
to do a comparison to spot the differences between the two UniProt entries (Hint:
focus on the "PTM / processing" section).
P29600 starts with the sequence of the mature protein.
P41363 starts with a signal peptide (pos: 1-24), then pro-peptide (25-93), and then comes the
mature protein.
Both signal peptide (function: signal for export of the protein) and the pro-peptide (function:
helps protein to fold correctly or make sure the protein is not active before it is in the correct
place - especially important for proteases that should not start to do its function before being
activated.
Activity 4: Exploring Gaps
We have been comparing sequences that are relatively closely related. Let’s
examine what happens when we compare human protein (P07498) with mice protein
(P06796) and, or prokaryotic Savinase (P29600) protease with a human peptidase
(P29144).
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
6
- With protein sequences sets B and, or C. Perform global alignment (Needle)
- Set End Gap Penalty to "true":
1. What is the Alignment score?
276.5
2. What is the Alignment length?
193
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
7
3. What is the % and fraction Identity? (The value reported for "Identity" includes
perfect matches only)
Identity: 76/193 (39.4%) and Similarity: 106/193 (54.9%)
4. What is the % and fraction Similarity? (The value reported for "Similarity" includes
perfect matches + "close" mismatches)
23/196 (11.96%)
- Repeat the alignment with End Gap Penalty set to "false" and report the same
results as above.
The results are the same as reported above.
- Repeat the alignment again — this time using the local alignment algorithm
(Water) — and report the same results as above.
The results are the same as reported above
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
8
Do you think local or global alignment is best for finding similar parts of distantly related
proteins? Why? Hint: Distantly related proteins typically share a core that relates to the
function of the protein
The prokaryotic protease matches only a single area in the middle of the human protease. In contrast,
this is not clear in the global alignment with end gaps, which "spreads" the short sequence over the
entire length. For distantly related sequences, it would be best to use local alignment because it would
actually provide an optimal analysis of the comparable part of the sequences.
Open the "Shuffle Protein" page
(http://www.bioinformatics.org/sms2/shuffle_protein.html) in a new window/tab, paste
in the tripeptidyl peptidase sequence and shuffle it.
Shuffle Protein results:
Version1:
>results for 1249 residue sequence "sp|P29144|TPP2_HUMAN Tripeptidyl-peptidase 2 OS=Homo sapiens OX=9606
GN=TPP2 PE=1 SV=4" starting
"MATAATEEPF"FCLKCEETKLGSQKVMGSTVNQKKKIVNSEKHGPYNVEKIIAVSLNFCNVILSGIARDGRDNHDVLARPLDHSYEKNENEQNPLYPDNKFHGSMNLAAN
AVVIHHSLALRIDQKCGSKPPTLLDAYDVYQVCKGEIGERPPVSLSELESKWPSITELKMPSGYISTGGWRLIGTEEDFEWIQLPTMTCSSTEWNNFRDLEGPKISELLKG
TELSDPWYLYFHRLLDLELLDWKLSEKGFTVIFYETYGKQLSHGKTETTYHAYGKKSLLSVSQLSDRDNRSRVDVIWIALALGAINAIVTTNSVKVAAVKALSWVVKATFG
CIYVTGPDSRKKAIRAYVGDGFRIYDVLTAQAMKNTGLTRMVLFNSAYKAGLSDNYVPNKLGDSAERHGRTDITAELYTEGLRNVSTAGETYLPDAAPGTKLLDAKNDPSG
PDQRAYSYKSLSLNKGYVSARSTPSDGVQANGRKHSLRIISAAAIVSSPPGYIKLPEEFPFYSNHCVQFITTERPNPETQVCWEITCRSTKTLGLMRQSAALLTIANLSFT
VLGQGLDGQCGTTGGLANELQSEDRKPPGILLRSCINTRKGPPEQNLEALWKTTLPVIIEHHGDESKSNFETPFYKTIQHLVEHQHHNKYIVYTTLSHEGSEAIYASRRPK
WNGAGANYELENFRYHEILGDFPQADGATGEDIEKLANVIEGYRIEESSTTQAGLVSKGIGVKAFKDLSITWMPVELPDHNARLDKKLYIASGVEGVITLDICEQGLKICV
LQQDPVPADKPFCYHSEFQPSATLIPFIDDRSEEDNIGGPDIMTVNMIAEQPVAIRQYDSSVHKTAAFFANSATAMPFAGMRVIAHRTPAGHYDLQENASSWVVSCVVWTE
EWDLRGRSSLLTGMLNTNVIVSFKSATFKHGAVSDHTDAITDCRPTVPDSKGACQKGYEEKMIEGLASIELKTHKANADVEDFPQNPPGLYVHKADLCVQNEGLAIDIKVL
HLPNFGRDNREILALPNSSVAKEPLNHHAVCKIISAISSNKDPFLEIKIEGATVVFKKTVSTEGTDKQLVGYFPLFTNPSGADGNNPMVLPLANENLIKDKGYIYDRPSDG
VTVCTAYGALGQKGACWQYIASWAMDKYSGSPDAYNKVYKLPEKVILLEIWPSSVCMRAPEKHDFEVCPEENLEGGMAAQLTKLPVNVLDQTKCYDLHNELHGVEYGFGRL
VIPIEQFSADNDKKAERVWSAQTKAPPLLVMDVCEWHAHL
Version 2:
>results for 1249 residue sequence "sp|P29144|TPP2_HUMAN Tripeptidyl-peptidase 2 OS=Homo sapiens OX=9606 GN=TPP2 PE=1
SV=4" starting "MATAATEEPF"
NHHIAIDKEPRWLVADTKEASNGVAVNVYIGKEIADLTEGAGLDVTAYCADNPELSGPGDSYEIDARPQKEKGSPMALGWAVPKWEDVDVVAPKICTIGAAIYELRINDAVQGFKELCP
PTTKIHTNPIWASDNSIWNLCDKKIGSMDISDFFVDQFFEDEYHYPMAEGLTLPDLMSGDWEQINKYMITKESSIMAPCNRAYLTWLKGTDCSVRQAMPGWDIPHSAVSMSNLCFNVSV
QVTEHSLVSIVKNKRAPAGQLQLKLSSKDTDGARPVGIYRNKAEEHSTLIGEAIVSLLGLQINHKIYSCKIAPTASDEWLPKGYERSREKGGKSYEDRTTINTKVGEISVWSDRADYDI
NFEVQITVVPEAPYPYGTLCATISNSRVVRGVALHDLLPYFEGHKQYITQVKKRDCPDMIGMLKRSTVLSEFSVSQTSIDGYSFDRSGERDMISAYAIEHNFQQGDQLEPYKGLKGNTI
VKVGGLGENQKLPCPLGNIAAQHDTYFVSYNVALLIAVQGPAPDCAPPNLLLADWVVTNIVSNEIPMKTTASLTCNLDFKHDPTGDSASPLGEGESYVVLLIAKDVEKGAYALKGDCAV
NDYSIFTLIYKLQYPYAERYSGTKAIVRSPCYKQPRLKDLATELGKFTANNDNLHHSTDEGPTVRLCPRSFELSKQPHCPSEMGWRSNPWSGLLEVIAKELGTTRTGKIPSYPALLLGD
LPQLLLVEAKLVGSRDLTFSKTFTVSGPVRDGDIEWHIQSEAYKCEKAKREAFAGIPSGRTEPKTFHLLPACSFSAAAGKDIITAKAWGGHLQSPKSTIVWDGYNSTTFPACNHAKEEL
VSPQEINDWISWIGKIRELHDDSLKLYAPMVKTPGTASGTVTNAAGDDNFGDMPTPNTKKLLEVRNMNIEKTLQTSGQCPLAEANGIRLASALLSYQVPLNLLVQVRTILGTSIVEAFG
WHCPLLHEILKKRNLDLTDVINAAITVFQSDDIKSDGAYINTQETPAAKSQATHALSWRADGLAEQHNHNSSKKNTKGSCVVGHRGFPGLACMEFEVETSKKKKVEFTMGVFLYVDEVK
EKLTGLYYDELEAFSSANIARHLLLYGNILILGEALVYEDQSRTVFEQKLKIPTPYSQFPARETHCHKQNRSGRSNKETLNHESYNVHEHVLSDYINVNEEEVDSNHHLEVVIRSGLGG
FYEPGLEKV CYRTPKNNNGALFLPNFRELSGLAEYPLFMHQRIKRVVTFVEGVNAHGK
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
9
Version 3:
>results for 1249 residue sequence "sp|P29144|TPP2_HUMAN Tripeptidyl-peptidase 2 OS=Homo sapiens OX=9606 GN=TPP2 PE=1
SV=4" starting "MATAATEEPF"
GTFTNDPDPNCIRIIYPSTAKMEGTNKHDVEGALYKEAVYRAVDPFGREGYHTDFTIGPQGQLYSPDYCQGAPTGNALLVNPEIAQVYTVLIPDTIEEKTGDSRLTVFAKVGKSTFKYL
SCINWPKARMAATFWGWLTNLLKVFDVKLNYCEPQVTDAMLPNAIIDECISGSREESRSGSESIIDHNGYNVPSSVELGETQMLKTEEVQLAAEDKAYDRGGEGISPAVIVFEKYGSLE
LSKSNIISGEYVLSFTVRHHSKPGCYYPQAPENWHSLHKNVRVSRASSLLDPALQTQLPLYPTMRVAKTWQSYNIKPEALPEATHDKIVLVKEDFIVKMLKKENAGGNEKYSGRVSLKS
APWAEIEGYATAAFVDKNRAATSEITGHSKIADVRKIQRGHRTYGGFVLPDYALPWPGTPPHVALIASADRGVSDQPKDRHESIDIVLSGIKFRYLQDTIAIAHISLIVELGKLYDVNG
AQPLQGELWCDGLLTALGKVQTANLSEAFKIVHVETYRMTIKVNEIKAEHKVSEPELNWKAPLCNDQIHGYRSSLFKSSAYDRGALLSLCISKGRNSEIKAVEGSEGNAFGTRPNTIVM
AGDGHKESVTSLDMDWLTRYPEDRAIETYVETTHFYLLLEELNIVYAVFGEGALTNDQFNDTKSFMTQAKISVLPDSNNQDKLDVNNVPQILSALAWANCGVNGDLQLLDHQPATLLLL
KLARCVKFQILMKQKALRGQQFKCPKYPYTVTVIWTRTTPYDFAENEGYIIPINAGGPPAMVLGLLNSRDCSPHVWNIVWESVDAVTLISYGHTILLLTVTGGLFVAQHTDGNTGRYAR
LSVPSPNKSWLLDKQGVTYELGAVVEEPQPEKDGEFGCVELSESWAMKESWNNAGLTLSIALSIMGNVVENSPVESKVSDSFPANFMITVTKADPHAKKDQHAHRGVSNHNPRNITGFC
RNPYEESLGTGQPGALAADENLIHTTCYTKHLWSLHRLECDGMVASPVLTNTNLKLIEIKEFKEKGKKFHRPGSPRLSCDCLAPADVDSINAWLHEETKHTLLLFTKFNGRIICAIDIS
LSDKSILCALVDSANDLCDDNYQTNQSSQPILRMLCSIARGKEENDPCASVARRPSYTGKHKVHCAPEGGYSAADKEVGPEAHYGLAEEKHFSSLFVGQKGKFEECSDEFGPDYESRVV
GAFKNGEPPEKTLNKDQDDIVQMSSIMANKTIGIYDHAKILSWYIKGLPRYTTVLPDKK
Then align Savinase and the shuffled tripeptidyl peptidase sequence using local alignment:
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
10
Repeat the above procedure two more times, so that you align Savinase with three
different shuffled versions of tripeptidyl peptidase.
For version 2:
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
11
For version 3:
1. How do the local alignments look? (What are the ranges of Alignment score,
Alignment length, Identity, Similarity, and gap percentage)?
We can see the results from each alignment in the screenshots above.
The results are different for every Alignemnt because the shuffled sequences are all different.
If we had completed the experiment 100 times instead of only 3 times, we could have done
statistics on the outcome and calculated confidence limits and therefore evaluated the degree
of statistical significance from a given alignment score (more on statistical significance when we
come to BLAST).
Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
12
2. Comparing the Savinase/shuffled alignment to the previous Savinase/Human
Peptidase alignment - how will you judge the alignment with human peptidase
now? (More/Less confidence in relation between the sequences?).
The score is higher with the Savinase/Human Peptidase alignment than what we got with the
shuffled sequences. However, only the scores show a clear difference between them.