Flight time is not propagation delay, a statistical budget saves far less at one error in a million million than at three sigma, and the arithmetic that ended the wide parallel bus can be written down in a page.
Propagation delay is a property of the line: length divided by velocity, about a hundred and seventy picoseconds per inch in stripline on a dielectric constant of four. Flight time is the interval between the driver crossing its own switching threshold and the receiver crossing its. These are different quantities and on a badly matched net they differ by a lot.
The reason is the reflection arithmetic of deck 01. A driver whose output impedance is below the line impedance launches more than half the swing, the open receiver doubles it, and the threshold is crossed on the first incident wave — flight time is then propagation delay plus a little loading. A driver whose impedance is well above the line impedance launches too little, the receiver does not cross, and it has to wait for the wave to go to the source, reflect, and come back.
The staircase below is the voltage at the receiver of a line driven by a source three hundred ohms above a fifty-ohm line — a weak driver, a long net, or a bus with more loads on it than it was designed for. Each step is one round trip, and the receiver's threshold is not crossed until the third.
This is the failure mode that made multi-drop buses hard and eventually killed them. Loads distributed along a line lower its effective impedance and load it capacitively, so a driver adequate for the unloaded net becomes inadequate for the populated one, and the flight time depends on how many cards are plugged in. A point-to-point serial link has exactly one load, a matched termination and a flight time equal to the propagation delay, which is one of several reasons the industry went that way.
A synchronous transfer requires the data to be stable for a setup time before the capturing clock edge and a hold time after it. Written out, with the clock arriving at the receiver later than at the transmitter by a skew:
$$T > t_{co} + t_{\text{flight,max}} + t_{su} + t_{\text{skew}},$$
$$t_{co} + t_{\text{flight,min}} > t_{h} + t_{\text{skew}}.$$
Notice what is absent from the second. There is no period in it. A hold violation cannot be fixed by slowing the clock down, because slowing the clock does not change how soon after an edge the new data arrives. Setup failures are frequency problems and hold failures are layout or silicon problems, and that asymmetry is the single most useful thing to know about timing closure.
A hold violation also gets worse when the line gets shorter, which is counter-intuitive enough that it is worth stating plainly: shortening a trace reduces the minimum flight time and can break a transfer that worked. It is the one case in this entire series where more copper is the fix.
On a parallel bus the skew between lanes comes directly out of the budget, so it is worth knowing which contributions dominate. The term everybody controls is rarely the largest.
Length matching is the term the layout tool measures and the reviewer checks. The laminate's own variation — the glass weave of deck 03 — is invisible to the tool, scales with board length rather than with routing tolerance, and on a long bus is comparable or larger. Tightening the matching constraint from twenty-five to five thousandths of an inch buys very little if the weave is contributing four picoseconds regardless.
A budget can be added two ways. The worst-case sum takes every term at its extreme and assumes they all go wrong together. The statistical sum treats the independent ones as independent and adds them in quadrature, then multiplies by however many standard deviations the required confidence demands.
The statistical result is smaller, and the question worth asking is whether the independence is earned. It usually is for manufacturing spread across many boards: two boards from different panels genuinely have uncorrelated laminate thickness. It usually is not for terms that share a cause: two lanes on the same board see the same laminate lot, the same temperature and the same supply rail, so their variations move together and quadrature addition understates the total.
Here is the part that is specific to high-speed links and is frequently carried over incorrectly from processor timing closure, where statistical methods transformed the discipline.
Conventional statistical timing works at three standard deviations, which for a handful of terms is a large saving over a worst-case sum. A serial link has to work at an error ratio of one in a million million, which is slightly past seven standard deviations. The quadrature sum is the same; the multiplier is more than twice as large.
With only a few terms the statistical budget at $10^{-12}$ can be larger than the worst-case sum, which sounds absurd and is correct: if four terms are each treated as a three-sigma spread and you then demand seven sigma of the combination, you are asking for more than their arithmetic sum. The technique pays only when there are many independent terms, and how many is a computation rather than a matter of taste.
The historical argument can be written as arithmetic. Three clocking schemes are budgeted against the same eight-inch board, and each spends its cycle on something different.
One clock distributed to both ends. The data crosses the board and is captured against a clock that has separately crossed the board, so the whole flight time and all of the clock skew come out of one cycle.
The clock travels with the data and sees the same flight time, so the bulk of it cancels. What remains is how well the lanes can be matched to each other, which is the previous slide's skew budget.
Timing is recovered from the data on each lane separately. Neither flight time nor lane-to-lane skew appears in the budget at all; what remains is jitter and whatever eye the channel leaves.
The model on the previous slide is crude — a handful of terms, one board length, no attempt at a particular technology. It is worth checking against history anyway, because if it is even approximately right it should reproduce the rates at which each scheme actually ran out.
A source-synchronous bus thirty-two bits wide at its limit carries a respectable aggregate. A single serial lane carries a fraction of that — and costs two conductors rather than thirty-two, needs no length matching between lanes, and can be added one lane at a time. The parallel bus did not lose on bandwidth per pin, it lost on bandwidth per pin per unit of design effort, and then it lost on bandwidth per pin as well.
| Quantity | Expression | Worth remembering |
|---|---|---|
| Propagation delay | $\ell\sqrt{\varepsilon_{\text{eff}}}/c$ | A property of the line |
| Flight time | Threshold to threshold | A property of the line, the driver and the load together |
| Setup | $T > t_{co}+t_{\text{flight}}+t_{su}+t_{\text{skew}}$ | Has a period in it; slowing down fixes it |
| Hold | $t_{co}+t_{\text{flight,min}} > t_h+t_{\text{skew}}$ | No period; slowing down does nothing |
| Worst-case sum | $\sum |a_i|$ | Safe, pessimistic, and independent of confidence level |
| Statistical sum | $Q\sqrt{\sum \sigma_i^2}$ | Depends on the confidence demanded; at $10^{-12}$, $Q=7.03$ |
| Common clock | — | Spends the whole flight time; ran out first |
| Source synchronous | — | Cancels flight time; limited by lane matching |
| Embedded clock | — | Spends neither; limited by the channel and by jitter |
Deck 15 of eleven in Signal Integrity & High-Speed Digital Design. Every figure on this page is computed by si_models/deck15 and embedded as data; nothing is typed in by hand.
Single-page HTML · KaTeX-rendered maths · no build step. Source on GitHub