Introduction to Bioinformatics

Biology has traditionally been an observational rather than a deductive science. Although recent developments have not altered this basic orientation, the nature of the data has radically changed. It is arguable that until recently all biological observations were fundamentally anecdotal - admittedly with varying degrees of precision, some very high indeed. However, in the last generation the data have become not only much more quantitative and precise, but, in the case of nucleotide and amino acid sequences, they have become discrete. It is possible to determine the genome sequence of an individual organism or clone not only completely, but in principle exactly. Experimental error can never be avoided entirely, but for modern genomic sequencing it is extremely low.
Not that this has converted biology into a deductive science. Life does obey principles of physics and chemistry, but for now life is too complex, and too dependent on historical contingency, for us to deduce its detailed properties from basic principles.
A second obvious property of the data of bioinformatics is their very very large amount. Currently the nucleotide sequence databanks contain 16 10 9 bases (abbreviated 16 Gbp). If we use the approximate size of the human genome - 3.2 10 9 letters - as a unit, this amounts to five HUman Genome Equivalents (or 2 huges, an apt name). For a comprehensible standard of comparison, 1 huge is comparable to the number of characters appearing in six complete years of...