Programming with Hyper-Threading Technology: How to Write Multithreaded Software for Intel IA-32 Processors

Until the mid-1990s, most programmers had it easy. They could write desktop programs that consisted of a single stream of instructions, which the processor executed sequentially in a fairly straightforward manner. Things were simple. Programs could rely on this linear execution model, because user interaction had not yet reached its current level of sophistication. There were few graphical interfaces and networks. Most PCs were standalone devices that ran a small slate of applications that had text interfaces.
In this simpler time, chip vendors' efforts to improve performance consisted mostly of gently accelerating the clock that is, increasing the megahertz at which the processor worked. As more-advanced operating systems and applications came to market, designers of microprocessors began looking for additional ways to boost performance.
Semiconductor architects soon came upon the idea of adding capabilities to the processor by which it could look ahead and execute certain upcoming instructions out of order. This out-of-order execution, later called dynamic execution, ushered in numerous advances in semiconductor technology, which will be discussed later. By executing some instructions out of order, more work could be done by the processor without disturbing the fundamental architecture of the single instruction queue. Succeeding generations of processors, such as the Intel Pentium II and the Intel Pentium III processors, greatly extended this capability. These processors could pre-execute very long sequences of instructions, thereby achieving remarkable time savings.
To increase performance even more, Intel started researching methods to accomplish more work on each processor clock tick. On servers, the...