PI-based CDR

Wang, Zhaowen. Efficient and High-Performance Clocking Circuits for High-Speed Data Links. 2022. Columbia University, PhD dissertation. Academic Commons,[https://academiccommons.columbia.edu/doi/10.7916/g3f1-4e71]
The PI-based architecture decouples the high-frequency clock synthesis and local clock deskew, allowing to optimize the power consumption and circuit area at a system level
Phase Interpolator (PI)
!!! Clock Edges
And for a phase interpolator, you need those reference clocks to be completely the opposite. Ideally they would be triangular shaped

four input clocks given by the cyan, black, magenta, red
John T. Stonick, ISSCC 2011 tutorial. "DPLL Based Clock and Data Recovery"
kink problem

B. Razavi, "The Design of a Phase Interpolator [The Analog Mind]," IEEE Solid-State Circuits Magazine, Volume. 15, Issue. 4, pp. 6-10, Fall 2023.(https://www.seas.ucla.edu/brweb/papers/Journals/BR_SSCM_4_2023.pdf)
Predistortion - sinusoidal
the interpolating inverters near the midscale can be made weaker so as to obtain more uniform phase increments. Alternatively, those at the top and bottom of the array can be made stronger
Single-Quadrant PI
\[ V_o(t) = m \cdot \sin(\omega t + \frac{\pi}{2}) + p\cdot \sin(\omega t) = m\cdot \cos(\omega t) + p\cdot \sin(\omega t) = \sqrt{m^2+p^2} \sin(\omega t + \phi) \]
where \(\tan \phi = \frac{m}{p} = \frac{1-p}{p}\) and \(p = \frac{1}{1+\tan \phi}\)

1 | phi = (32:-1:0)./32*pi/2; |
Eight-Quadrant PI
\[\begin{align}
V_o(t) &= m \cdot \sin(\omega t + \frac{\pi}{4}) + p\cdot
\sin(\omega t) = \frac{\sqrt{2}}{2}m\cdot \cos(\omega t) + \left(
\frac{\sqrt{2}}{2}m + p\right)\cdot \sin(\omega t) \\
&= \sqrt{m^2 + p^2 +\sqrt{2}pm}\cdot \sin(\omega t + \phi)
\end{align}\]
where \(\tan\phi = \frac{\sqrt{2}m}{\sqrt{2}m+2p} = \frac{\sqrt{2}-\sqrt{2}p}{\sqrt{2}+(2-\sqrt{2})p}\)

1 | phi = (16:-1:0)./16*pi/4; |
Predistortion - square wave
Weinlader, Daniel, Thomas H. Lee and James A. Gasbarro. "Precision CMOS receivers for VLSI testing applications." (2001). [https://www-vlsi.stanford.edu/people/alum/pdf/0111_Weinlader_Precision_CMOS_Receivers_.pdf]

Suppose \(V_i(t) = 1- e^{-\frac{t}{\tau}}\) and \(V_q(t) = 1-e^{-\frac{t-\Delta t}{\tau}}\) with \(t\ge \Delta t\) \[ \frac{1}{2} = (1-\alpha)\cdot V_i(t) + \alpha \cdot V_q(t) \] yield triggering time \[ t = \tau \ln\left[ 1 + \alpha \left(e^{\frac{\Delta t}{\tau}}-1\right)\right] + \tau \ln 2 \] Then \[\begin{align} \frac{\partial t}{\partial \alpha} &= \tau \frac{e^{\frac{\Delta t}{\tau }}-1}{1+\alpha(e^{\frac{\Delta t}{t}}-1)} \gt 0 \\ \frac{\partial^2 t}{\partial \alpha^2} &= -\tau \frac{\left(e^{\frac{\Delta t}{\tau }}-1\right)^2}{\left(1+\alpha(e^{\frac{\Delta t}{t}}-1)\right)^2} \lt 0 \end{align}\]
As a conclusion, heavier weight while \(\alpha\) approaching to 1 in order to improve linearity

1 | import numpy as np |
Input/Output amplitude
A constant Output amplitude is desired because the swing-dependent delay characteristic of the CML-to-CMOS (C2C) circuit results in AM–PM distortion which eventually manifests as phase nonlinearity
Current-Mode Phase Interpolator
Current-mode PIs (CMPIs) can achieve high linearity but at the cost of digital overhead to generate sinusoidal weights
Voltage-Mode Phase Interpolator
Integrating-Mode Phase Interpolator
PI Nonlinearity Effect
Wang, Zhaowen. Efficient and High-Performance Clocking Circuits for High-Speed Data Links. 2022. Columbia University, PhD dissertation. Academic Commons,[https://academiccommons.columbia.edu/doi/10.7916/g3f1-4e71]
Sampling Offset due to PI Nonlinearity

Let \(T=T_{\mathrm{LSB}}\), and write a PI’s output time as
\[ t(k)=t_{\mathrm{ref}}+[k+I(k)]T \]
where \(I(k)=\mathrm{INL}(k)\), expressed in LSBs. Then
\[ \mathrm{DNL}(k)=\frac{t(k+1)-t(k)}{T}-1 =I(k+1)-I(k) \]
Thus INL describes the error at a code, while DNL describes the error in one code-to-code step.
1. Why \((0.5+|\mathrm{DNL}_p|)T\)?
For an ideal PI, available phases are spaced by \(T\). If the desired transition lies halfway between two available phases, selecting the nearest phase leaves an error of
\[ |E_e|\le \frac{T}{2} \]
That explains the \(0.5\)
With DNL, a particular step has width
\[ t(k+1)-t(k)=[1+\mathrm{DNL}(k)]T \]
A larger step means a larger gap in which the desired phase might lie.
The excerpt can be read as budgeting
\[ \underbrace{0.5T}_{\text{nominal quantization}} + \underbrace{|\mathrm{DNL}_p|T}_{\text{nonlinearity allowance}} \]
But for a monotonic PI that actually selects the nearest available phase, the tighter bound is half the largest actual step:
\[ \boxed{|E_e| \le \frac{1+\max_k\mathrm{DNL}(k)}{2}T \le \left(0.5+\frac{|\mathrm{DNL}_p|}{2}\right)T} \]
So the paper's \(0.5+|\mathrm{DNL}_p|\) is a looser bound under this model
2. Why does the data-clock bound become \((1+|\mathrm{DNL}_p|+|\mathrm{INL}_{pp}|)T\)?
\(E_e\) and \(E_d\) are timing errors, measured in seconds—not the clock times themselves
\(E_e\): edge-sampling clock error relative to the ideal data-transition time: \[ E_e=t_e-t_{\text{transition}}. \]
\(E_d\): data-sampling clock error relative to the ideal data-sampling time, assumed here to be half a UI after that transition: \[ E_d=t_d-\left(t_{\text{transition}}+\frac{\mathrm{UI}}{2}\right). \]
Here \(t_e\) and \(t_d\) are the actual sampling times. A positive error means the clock samples late; a negative error means it samples early.
Writing \(H=\mathrm{UI}/2\), these definitions give
\[ \boxed{E_d=E_e+\underbrace{(t_d-t_e-H)}_{\text{error in edge-to-data spacing}}.} \]
Let the desired edge-to-data spacing be
\[ H=\frac{\mathrm{UI}}{2}, \]
and let the data-clock code be \(k+m\) when the edge-clock code is \(k\). Then
\[ t_d-t_e=mT+[I(k+m)-I(k)]T. \]
Consequently, the data-clock error relative to its ideal sampling position is
\[ \boxed{ E_d = E_e +\underbrace{(mT-H)}_{\text{spacing quantization}} +\underbrace{[I(k+m)-I(k)]T}_{\text{relative INL error}}. } \]
If \(m\) is chosen by rounding \(H/T\), then
\[ |mT-H|\le 0.5T. \]
Combining this with the paper’s edge-clock allowance gives
\[ |E_d| \le \underbrace{(0.5+|\mathrm{DNL}_p|)T}_{\text{edge-clock error}} +\underbrace{0.5T}_{\text{spacing quantization}} +\underbrace{\mathrm{INL}_{pp}T}_{\text{relative INL error}}, \]
which produces the quoted expression.
Using tighter edge-clock bound
\[ |E_d| \le \underbrace{\left(0.5+\frac{|\mathrm{DNL}_p|}{2}\right)T_{\mathrm{LSB}}}_{\text{edge-clock error}} +\underbrace{0.5T_{\mathrm{LSB}}}_{\text{spacing quantization}} +\underbrace{\mathrm{INL}_{pp}T_{\mathrm{LSB}}}_{\text{relative INL error}}, \]
so
\[ \boxed{|E_d|\le \left(1+\frac{|\mathrm{DNL}_p|}{2}+\mathrm{INL}_{pp}\right)T_{\mathrm{LSB}}} \]
If \(H/T\) is an integer—for example, a full-period PI with a number of steps divisible by eight can represent \(45^\circ\) exactly—then \(mT-H=0\). That extra \(0.5T\) is unnecessary. (If the desired edge-to-data spacing is exactly representable by an integer number of PI steps, the spacing-quantization term vanishes, and the constant \(1\) becomes \(0.5\))
The CDR finds an edge-clock code, and the data-clock code is obtained by adding a fixed code offset
With \(T=T_{\mathrm{LSB}}\) and \(H=\mathrm{UI}/2\):
\[ k_e=k,\qquad m=\operatorname{round}\left(\frac{H}{T}\right),\qquad k_d=k_e+m \]
Here, \(m\) is a number of PI steps, not a time. Thus, for an ideal PI, the timing relationship is
\[ \boxed{ t_d=t_e+\operatorname{round}\left(\frac{\mathrm{UI}}{2T}\right)T} \]
The circuit operates continuously: the CDR adjusts \(k_e\), and the data-clock code follows as \(k_d=k_e+m\). It does not need to measure a numerical value of \(t_e\) before generating the data clock.
For nonlinear PIs, adding \(m\) codes does not necessarily add exactly \(mT\) in time. Assuming a common timing reference,
\[ \boxed{ t_d=t_e+mT+ \left[I_d(k_e+m)-I_e(k_e)\right]T} \]
That last term is precisely the relative INL error we discussed. The subscripts allow the edge and data clocks to come from different PIs.
Deterministic Jitter due to PI Nonlinearity

\[
K_{f,PI} = \frac{1/2^N}{T_m/T_o}\cdot \frac{1}{T_o} = \frac{f_m}{2^N}
\]
For DCO \[ K_{f,DCO} = \frac{K_T}{T_o}\cdot \frac{1}{T_o} = \frac{K_T}{T_o^2} \]
PI vs. PLL based CDR
PCI Express Jitter Modeling Revision 1.0RD July 14, 2004

\[
H_1 - \left[H_1(1-H_3) + H_2H_3\right] = (H_1-H_2)H_3
\]
reference
A. K. Mishra, Y. Li, P. Agarwal and S. Shekhar, "Improving Linearity in CMOS Phase Interpolators," in IEEE Journal of Solid-State Circuits, vol. 58, no. 6, pp. 1623-1635, June 2023 [pdf]
Cortiula A, Menin D, Bandiziol A, Driussi F, Palestri P. Modeling of Phase-Interpolator-Based Clock and Data Recovery for High-Speed PAM-4 Serial Interfaces. Electronics. 2025; [https://www.mdpi.com/2079-9292/14/10/1979]
G. Souliotis, A. Tsimpos and S. Vlassis, "Phase Interpolator-Based Clock and Data Recovery With Jitter Optimization," in IEEE Open Journal of Circuits and Systems, vol. 4, pp. 203-217, 2023 [https://ieeexplore.ieee.org/document/10184121]
B. Razavi, "The Design of a Phase Interpolator [The Analog Mind]," in IEEE Solid-State Circuits Magazine, vol. 15, no. 4, pp. 6-10, Fall 2023 [https://www.seas.ucla.edu/brweb/papers/Journals/BR_SSCM_4_2023.pdf]
T. Chan Carusone, T. O. Dickson, S. Palermo, S. Shekhar and M. Mansuri, "Modern Wireline Transceivers," in IEEE Journal of Solid-State Circuits, vol. 61, no. 2, pp. 395-422, Feb. 2026 [https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=11311714]