Bioinformatics and Functional Genomics Student ID: 21485094
Lecturer: Dr. Obed Brew Beatriz Manso
1
PROTEIN STRUCTURE ANALYSIS AND MODELL ING
An important aspect of systems biology is the modelling of proteins. Such modeling can be used to
identify the genes and proteins that interact to generate a transformation, examine the instructions
or algorithms that modify the transformation or cause a cell to function in a particular way.
Modeling of proteins is part of an evolutionary cycle in which experimental biology also plays a key
role: a typical protein modelling cycle begins with the development of a model that represents the
biology of the thing that is being modelled, then the model is tested to see if it behaves the same
way as the biological protein. Alternatively, you could tweak the model until it closely resembles the
biological protein and repeat this cycle until your model accurately reflects reality.
Identify possible orthologs and or paralogs, and compare their 3D structure(s)
with that of the new protein.. Perform multiple sequence alignment,
phylogeny, and predict the protein’s structure and its function.
Part 1: Prediction of Primary Protein Structure
1. Predict the primary protein structure of the gene sequence in
“Practical08_UnnamedSequence” by uploading the FASTA sequence into the online program
AUGUSTUS (http://bioinf.uni-greifswald.de/augustus/submission.php).
2. Set the parameters in AUGUSTUS as follows:
a. Organism: ‘Homo sapiens’
b. Report genes on: both strands
c. Alternative transcripts: few.
3. Run AUGUSTUS and Click the link ‘graphical and text results’
4. View protein sequence by clicking “predicted amino acid sequences”
5. Save the predicted protein sequence as a FASTA file by copying the sequence into notepad.
In the file tab, select ‘save as’, save as type ‘all files’ and save the protein as a file with the
extension ‘.fa’ or ‘.fasta’.
6. Identify what protein FASTA file was predicted from the gene sequence by uploading the
protein FASTA file into BLASTp and assessing the matches.
>UnamedSequence
NNNGAAAAAAAANAACCTAGANAGTGGGNTGTTCANGGGGGGGGGGAGGAATCTTTGNTNCGCCAGGCCTCTTNGGCTTCAAAAGGAAGCTGCCTAA
GTACCTGCTCTTTACCAGCCCCCAGGAGAACCCCTGGGGCCACAAGCGCAGCTACCGCCTGCAGATCCACTCCATGGCCGACCAGGTGCTGCCCCCAGGC
TGGCAGGAGGAGCAGGCCATCACCTGGGCAAGGTACCCCCTGGCAGTGACCAAGTACCGGGAGTCGGAGCTGTGCAGCAGCAGCATCTACCACCAGAA
CGACCCCTGGCACCCGCCCGTGGTCTTCGAGCAGTTTCTTCACAACAACGAGAACATTGAAAATGAGGACCCGGTGGCCTGGGTGACGGTGGGCTTCCTG
CACATNCCCCACTCAGAGGACATTCCCAACACAGCCACACCTGGGAACTCCGTGGGCTTCCTGCTCCGGCCATTCAACTTCTTCCAGAGGACCCCTCCCTG
GCATCCAGAGACACTGGGGATCGTGTGGCCTCGGGACAACGGNCCCAACTACCGTTCCACGCTGGGATCCCTGANCAATCGAATCCCCGCGCCGCCATG
GCGNNCGGGANCATGCCAACTTCNNTCCAATTCCCTATAGCGAATCTNTNNTACAATTACTGCCCTCCGTTTACAACCNATCGAATNGGAAAACCTTTGG
CGTTCCCAACTTAATCGCCTTTNCANNCCAATCCCCTTCGCGTTNGCTGCAAAACAAAAAGCCCNANCCGATCCCTTCCNAAAANTTCCAACCTATGGCAT
GGACCCCCCTGATCGCGATAACCTNCGCTCTNCGTGNTACCCTCCAACNGACNTTNAATTNCANNNCTTAACNCCCNTCTNCG