Text-to-Speech Synthesis

The speech-production process was qualitatively described in Chapter 7. There we showed that speech is produced by a source, such as the glottis, which is subsequently modified by the vocal tract acting as a filter. In this chapter, we turn our attention to developing a more-formal quantitative model of speech production, using the techniques of signals and filters described in Chapter 10.
Such models often come under the heading of the acoustic theory of speech production, which refers both to the general field of research in mathematical speech-production models and to the book of that title by Fant [158]. Although considerable previous work in this field had been done prior to its publication, this book was the first to bring together various strands of work and describe the whole process in a unified manner. Furthermore, Fant backed his study up with extensive empirical studies with X-rays and mechanical models to test and verify the speech-production models being proposed. Since then, many refinements to the model have been made, as researchers have investigated trying to improve the accuracy and practicalities of these models. Here we focus on the single most widely accepted model, but conclude the chapter with a discussion on variations on this.
As with any modelling process, we have to reach a compromise between a model that accurately describes the phenomena in question and one that is simple, effective and suited to practical needs. If we tried to capture every aspect...