- ES Español

- EN English

5.57. Bioinformatics (Mandatory)
- Semester: 9th Sem. Credits: 4
- Hour of this course: Theory: 2 hours; Practice: 4 hours;
- Syllabus:
- htmlonly

Español

English - Prerrequisites: None
5.57.1. Justification ↑ Back to top
The use of computational methods in the biological sciences has become one of the key tools for the field of molecular biology, being a fundamental part of research in this area.
In Molecular Biology, there are several applications that involve both DNA, protein analysis or sequencing of the human genome, which depend on computational methods. Many of these problems are really complex and deal with large data sets.
This course can be used to see concrete use cases of several areas of knowledge of Computer Science such as Programming Languages (PL), Algorithms and Complexity (AL), Probabilities and Statistics, Information Management (IM), Intelligent Systems (IS).
5.57.2. Generales Goals ↑ Back to top
- That the student has a solid knowledge of molecular biological problems that challenge computing.
- That the student is able to abstract the essence of the various biological problems to pose solutions using their knowledge of Computer Science
5.57.3. Contribution to Outcomes ↑ Back to top
- AG-C11) Use of Tools: Applies modern computing tools in problem solving. (Usage)
- AG-C10) Inquiry: Studies complex computing problems using information science methods. (Usage)
5.57.4. Content ↑ Back to top
5.57.4.1. Introduction to Molecular Biology (4 hours) [Skills AG-C10,AG-C11] ↑ Back to top
Bibliography: (Clote and Backofen, 2000; Setubal and Meidanis, 1997)
Topics
- Review of organic chemistry: molecules and macromolecules, sugars, nucleic acids, nucleotides, RNA, DNA, proteins, amino acids and levels of structure in proteins.
- The Dogma of Life: From DNA to Proteins, Transcription, Translation, Protein Synthesis.
- Genome study: Maps and sequences, specific techniques
Learning Outcomes
- Achive a general knowledge of the most important topics in Molecular Biology. [Familiarity]
- Understand that biological problems are a challenge to the computational world. [Assessment]
5.57.4.2. Sequence Comparison (4 hours) [Skills AG-C10,AG-C11] ↑ Back to top
Bibliography: (Clote and Backofen, 2000; Setubal and Meidanis, 1997; Pevzner, 2000)
Topics
- Sequences of nucleotides and amino acid sequences.
- Sequence alignment, paired alignment problem, exhaustive search, Dynamic programming, global alignment, local alignment, gaps penalty
- Comparison of multiple sequences: sum of pairs, complexity analysis by dynamic programming, alignment heuristics, star algorithm, progressive alignment algorithms.
Learning Outcomes
- Understand and solve the problem of aligning a pair of sequences. [Usage]
- Understand and solve the problem of multiple sequence alignment. [Usage]
- Know the various algorithms for aligning existing sequences in the literature . [Familiarity]
5.57.4.3. Phylogenetic Trees (4 hours) [Skills AG-C10,AG-C11] ↑ Back to top
Bibliography: (Clote and Backofen, 2000; Setubal and Meidanis, 1997; Pevzner, 2000)
Topics
- Phylogeny: Introduction and phylogenetic relations
- Phylogenetic trees: definition, type of trees, problem of search and reconstruction of trees
- Reconstruction methods: parsimony methods, distance methods, maximum likelihood methods, confidence of reconstructed trees
Learning Outcomes
- Understand the concept of phylogeny, phylogenetic trees and the methodological difference between biology and molecular biology. [Familiarity]
- Understand the problem of the reconstruction of phylogenetic trees, to know and apply the main algorithms for the reconstruction of phylogenetic trees. [Assessment]
5.57.4.4. DNA Sequence Assembling (4 hours) [Skills AG-C10,AG-C11] ↑ Back to top
Bibliography: (Setubal and Meidanis, 1997; Aluru, 2006)
Topics
- Biological basis: ideal case, difficulties, alternative methods for DNA sequencing
- Formal Assembly Models: Shortest Common Superstring, Reconstruction, Multicontig
- Algorithms for sequence assembly: representation of overlaps, paths to create superstrings, voracious algorithm, acyclic graphs.
- Assembly heuristics: search for overlays, ordering fragments, alignments and consensus.
Learning Outcomes
- Understand the computational challenge of the Sequence Assembly problem. [Familiarity]
- Understand the principle of formal model for assembly. [Assessment]
- Know the main heuristics for the problem of assembjale of DNA sequences[Usage]
5.57.4.5. Secondary and tertiary structures (4 hours) [Skills AG-C10,AG-C11] ↑ Back to top
Bibliography: (Setubal and Meidanis, 1997; Clote and Backofen, 2000; Aluru, 2006)
Topics
- Molecular structures: primary, secondary, tertiary, quaternary.
- Prediction of secondary structures of RNA: formal model, pair energy, structures with independent bases, solution with Dynamic Programming, structures with loops.
- Protein folding: Estructuras en proteinas, problema de protein folding.
- Protein Threading: Definitions, Branch \ Bound Algorithm, Branch \ Bound for protein threading.
- Structural Alignment: Definitions, DALI algorithm
Learning Outcomes
- Know the protein structures and the necessity of computational methods for the prediction of the geometry. [Familiarity]
- Know the algorithms for solving prediction problems of secondary structures RNA, and structures in proteins. [Assessment]
5.57.4.6. Probabilistic Models in Molecular Biology (4 hours) [Skills AG-C10,AG-C11] ↑ Back to top
Bibliography: (R et al., 1998; Clote and Backofen, 2000; Aluru, 2006; Krogh et al., 1994)
Topics
- Probability: Random Variables, Markov Chains, Metropoli-Hasting Algorithm, Markov Random Fields, and Gibbs Sampler, Maximum Likelihood.
- Hidden Markov Models (HMM), parameter estimation, Viterbi algorithm and Baul-Welch method, Application in paired and multiple alignments, Motifs detection in proteins, in eukaryotic DNA, in sequences families.
- Probabilistic phylogeny: probabilistic models of evolution, likelihood of alignments, likelihood for inference, comparison of probailistic and non-probabilistic methods
Learning Outcomes
- Review concepts of Probabilistic Models and understand their importance in Computational Molecular Biology. [Assessment]
- Know and apply Hidden Markov Models for various analyzes in Molecular Biology.. [Usage]
- Know the application of probabilistic models in Phylogeny and to compare them with non-probabilistic models[Assessment]
5.57.5. Bibliography ↑ Back to top
Clote, P. and Backofen, R. (2000). Computational Molecular Biology: An Introduction. John Wiley & Sons Ltd. 279 pages.
Setubal, J. C. and Meidanis, J. (1997). Introduction to computational molecular biology. Boston: PWS Publishing Company.
Pevzner, P. A. (2000). Computational Molecular Biology: an Algorithmic Approach. The MIT Press, Cambridge, Massachusetts.
Aluru, S., editor (2006). Handbook of Computational Molecular Biology. Computer and Information Science Series. Chapman & Hall, CRC, Boca Raton, FL.
R, D., S.R, E., A, K., and G, M. (1998). Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids. Cambridge University Press.
Krogh, A., Brown, M., Mian, I. S., Sjölander, K., and Haussler, D. (1994). Hidden markov models in computational biology, applications to protein modeling. J Mol. Biol, 235:1501–1531.