Sections15
  1. Quick note:
  2. Lorentz transformation: general structure
  3. 1D boost
  4. Solution
  5. Rapidity and velocity
  6. Physical consequences
  7. Time dilation
  8. Length contraction (the wrong way to see the concept)
  9. Relativity of simultaneity
  10. Length contraction, done properly
  11. Boost of a moving object
  12. Boost of light
  13. Doppler effect
  14. Aberration
  15. 2D boost

Quick note:§

Remembers that in previous articles, we defined the lorentz transformation using

ΛTηΛ=η\Lambda^T\eta\Lambda = \eta

This immidieatly hint us to have η\eta in the calculation of invariant, as the definition above is already about invariant of Λ\Lambda no matter what it is it should obey the definition. Let's sandwich it with XX and note that X′=ΛXX' = \Lambda X.

XTΛTηΛX=XTηX(ΛX)Tη(ΛX)=XTηX(X′)Tη(X′)=XTηX\begin{align*} X^T\Lambda^T\eta\Lambda X &= X^T\eta X \\ (\Lambda X)^T\eta(\Lambda X) &= X^T\eta X \\ (X')^T\eta(X') &= X^T\eta X \\ \end{align*}

Yeah, we see that indeed XTηXX^T\eta X is invariant, write it in index notation it is ημνXνXμ\eta_{\mu\nu}X^\nu X^\mu being invariant. And it is

 ds2:=−c2 dt2+ dx2+ dy2+ dz2\dd s^2 := -c^2\dd t^2 + \dd x^2 + \dd y^2 + \dd z^2

This invariant, we call it proper length squared, whos actual meaning will be clear later. Although from first principle this is not how we defined lorentz transformation. But since this invariant in universally true, we might as well treat it as an alternative definition of lorentz transformation. A transformation that keep this value invariant.

And note that we can form an invariant of another unit ( of time ) rather than lenght by dividing through −c2-c^2. We will get

dτ2:=dt2−1c2( dx2+ dy2+ dz2)d\tau^2 := dt^2 - \frac{1}{c^2}(\dd x^2 + \dd y^2 + \dd z^2)

Again, we will see why is this the proper time later.

Invariant in general

Note that the derivation of invariant works for any four vectors, not only for four positions

Lorentz transformation: general structure§

Let's remind ourselves that we have a generator GG for the Lorentz transformation

Λ=e−iG\Lambda = e^{-iG}

(using a different convention from previous articles, but the structure stays the same). We have the generator

G=i(0ϕxϕyϕzϕx0−θzθyϕyθz0−θxϕz−θyθx0)G=i\begin{pmatrix} 0 & \phi_x & \phi_y & \phi_z \\ \phi_x & 0 & -\theta_z & \theta_y \\ \phi_y & \theta_z & 0 & -\theta_x \\ \phi_z & -\theta_y & \theta_x & 0\\ \end{pmatrix}

This is the general so(1,3)\mathfrak{so}(1,3), the Lie algebra of SO+(1,3)SO^+(1,3) — the proper, orthochronous Lorentz group, the part reachable continuously from the identity by exponentiating a generator like GG. (The full Lorentz group O(1,3)O(1,3) also contains parity and time reversal, sitting in disconnected pieces that no real GG can reach this way — we won't need them here.) We split GG into six basis generators, exactly matching the six degrees of freedom: three boosts, three rotations.

Kx=(0i00i00000000000)Ky=(00i00000i0000000)Kz=(000i00000000i000)K_x = \begin{pmatrix} 0 & i & 0 & 0 \\ i & 0 & 0 & 0 \\ 0 & 0 & 0 & 0 \\ 0 & 0 & 0 & 0 \end{pmatrix} \qquad K_y = \begin{pmatrix} 0 & 0 & i & 0 \\ 0 & 0 & 0 & 0 \\ i & 0 & 0 & 0 \\ 0 & 0 & 0 & 0 \end{pmatrix} \qquad K_z = \begin{pmatrix} 0 & 0 & 0 & i \\ 0 & 0 & 0 & 0 \\ 0 & 0 & 0 & 0 \\ i & 0 & 0 & 0 \end{pmatrix} Jx=(00000000000−i00i0)Jy=(0000000i00000−i00)Jz=(000000−i00i000000)J_x = \begin{pmatrix} 0 & 0 & 0 & 0 \\ 0 & 0 & 0 & 0 \\ 0 & 0 & 0 & -i \\ 0 & 0 & i & 0 \end{pmatrix} \qquad J_y = \begin{pmatrix} 0 & 0 & 0 & 0 \\ 0 & 0 & 0 & i \\ 0 & 0 & 0 & 0 \\ 0 & -i & 0 & 0 \end{pmatrix} \qquad J_z = \begin{pmatrix} 0 & 0 & 0 & 0 \\ 0 & 0 & -i & 0 \\ 0 & i & 0 & 0 \\ 0 & 0 & 0 & 0 \end{pmatrix}

so that

G=ϕxKx+ϕyKy+ϕzKz+θxJx+θyJy+θzJz=ϕ⋅K+θ⋅JG=\phi_xK_x + \phi_yK_y + \phi_zK_z + \theta_xJ_x + \theta_yJ_y + \theta_zJ_z = \boldsymbol\phi\cdot \mathbf K + \boldsymbol\theta\cdot \mathbf J

with the meaning of ϕ\boldsymbol\phi and θ\boldsymbol\theta now immediate: one parametrizes boosts, the other rotations.

These six generators don't just sit next to each other — they talk to each other:

[Ji,Jj]=iϵijkJk,[Ji,Kj]=iϵijkKk,[Ki,Kj]=−iϵijkJk[J_i,J_j]=i\epsilon_{ijk}J_k, \qquad [J_i,K_j]=i\epsilon_{ijk}K_k, \qquad [K_i,K_j]=-i\epsilon_{ijk}J_k

Keep that last relation in your pocket. We won't need it for a while, but it is not decorative.

A fully general six-parameter Lorentz transformation can be written down by exponentiating GG directly, but the closed form carries no physical insight on its own. So instead of solving the general case, we'll build up from the physically transparent one. Homogeneity and isotropy of space mean we only need to understand a boost with no rotation — so set θ=0\boldsymbol\theta = 0 — and by isotropy, any single boost direction is as good as any other. We start with one dimension.

1D boost§

Solution§

Let only ϕx\phi_x be nonzero. Call it simply ϕ\phi, and its generator simply KK, so that in one dimension

K=i(0110)K = i\begin{pmatrix} 0 & 1 \\ 1 & 0 \end{pmatrix}

Notice

K2=−IK3=−KK4=IK5=K…K^2 = -I \qquad K^3 = -K \qquad K^4=I \qquad K^5=K \quad\ldots

which turns the exponential into a familiar pair of series:

Λ=e−iϕK=I−iϕK+12!ϕ2I−i13!ϕ3K+14!ϕ4I+⋯=Icosh⁡ϕ−iKsinh⁡ϕ\Lambda=e^{-i\phi K}=I-i\phi K + \frac{1}{2!}\phi^2 I - i\frac{1}{3!}\phi^3K + \frac{1}{4!}\phi^4 I+ \dots = I\cosh\phi-iK\sinh\phi

so that

Λ=(cosh⁡ϕsinh⁡ϕsinh⁡ϕcosh⁡ϕ)\Lambda = \begin{pmatrix}\cosh\phi & \sinh\phi \\ \sinh\phi & \cosh\phi\end{pmatrix}

Rapidity and velocity§

ϕ\phi is just a parameter so far — let's find out what it physically means. Boost into the frame of an object moving as x=vtx=vt, and demand that in the new frame the object sits still: x′=0x' = 0.

ctsinh⁡ϕ+vtcosh⁡ϕ=0β:=vc=−tanh⁡ϕ\begin{align*} ct\sinh\phi + vt\cosh\phi &= 0 \\ \beta := \frac{v}{c} &= -\tanh\phi \end{align*}

Set γ:=cosh⁡ϕ=(1−β2)−1/2\gamma := \cosh\phi = (1-\beta^2)^{-1/2}, so sinh⁡ϕ=−γβ\sinh\phi = -\gamma\beta, giving

Λ=γ(1−β−β1)\Lambda = \gamma\begin{pmatrix} 1 & -\beta \\ -\beta & 1 \end{pmatrix}

with γ≥1\gamma \ge 1, ∣β∣<1|\beta| < 1. Notice γ\gamma doesn't care about the sign of vv — the transformation matrix has the same structure whether you boost forward or backward, which is just isotropy of space showing up again. That also tells you Λ−1(β)=Λ(−β)\Lambda^{-1}(\beta)=\Lambda(-\beta): undoing a boost is the same as boosting the other way.

The parameter ϕ\phi itself has a name — rapidity — and unlike velocity, rapidities for collinear boosts simply add. We won't need that fact yet, but it's worth knowing it's there.

Physical consequences§

Everything that follows — time dilation, simultaneity, length contraction — comes out of exactly the same two equations:

c dt′=γ(c dt−β dx), dx′=γ( dx−βc dt)c\dd t' = \gamma(c\dd t - \beta \dd x), \qquad \dd x' = \gamma(\dd x - \beta c\dd t)

What changes from one effect to the next is only which pair of events we plug in, and which simultaneity condition we impose. Watch how much mileage we get out of the same formula.

Time dilation§

Take a clock sitting still in SS, so its worldline has  dx=0\dd x = 0. Then

c dt′=cΛμt  dXμ=γ(c dt−β dx)=γc dt\begin{align*} c\dd t' &= c\Lambda^t_\mu\,\dd X^\mu \\ &= \gamma(c\dd t-\beta\dd x) \\ &= \gamma c \dd t \end{align*}

Since SS is the clock's own rest frame,  dt\dd t here is its proper time  dτ\dd\tau ( defined as the time measured in the rest frame ), so really

 dt′=γ  dτ,i.e.Δt=γ Δτ\dd t' = \gamma\, \dd\tau, \qquad\text{i.e.}\qquad \Delta t = \gamma\,\Delta\tau

The moving observer watches the stationary clock accumulate less proper time than their own coordinate time ticks off — a moving clock runs slow. It's worth stressing that  dt′=γ dt\dd t' = \gamma \dd t is not the universal statement of time dilation; it only holds because we chose  dx=0\dd x = 0. The frame-independent statement is the one with Δτ\Delta\tau in it.

And let's we call we call − ds2/c2= dτ2-\dd s^2/c^2 = \dd\tau^2 and we can see it why here, since it is reference frame independant ( invariant ), we use the rest reference frame (( dxi)2=0(\dd x^i)^2 = 0). Then we find

− ds2/c2= dt2= dτ2-\dd s^2/c^2 = \dd t^2 = \dd\tau^2

Length contraction (the wrong way to see the concept)§

Now for length — and here we need to slow down, because length is a trickier thing to define than time. Time can be read off along a single worldline. Length can't: it needs the positions of two endpoints, at the same time, in whichever frame is doing the measuring.

If you note down where the front of a train is in the morning, and where its back is at night after it's crossed an entire country, that difference is not the train's length — you've just measured the gap between two unrelated events. So the measurement has to be simultaneous. Let's see what happens if we're careless about whose simultaneity we mean.

Take two endpoint events simultaneous in SS, so  dt=0\dd t = 0, and push them through the spatial transformation:

 dx′=Λμx dXμ=γ( dx−βc dt)=γ dx\begin{align*} \dd x' &= \Lambda^x_\mu \dd X^\mu \\ &= \gamma(\dd x-\beta c\dd t) \\ &= \gamma \dd x \end{align*}

So the moving object comes out longer by a factor of γ\gamma. That should make you suspicious — length contraction is supposed to shrink things, not stretch them. Something in this calculation isn't measuring what we think it's measuring.

Relativity of simultaneity§

Here's the catch. We fixed  dt=0\dd t = 0 — simultaneous in SS — but we never checked whether those same two events are simultaneous in S′S'. Let's check:

c dt′=cΛμt  dXμ=γ(c dt−β dx)=−γβ dx\begin{align*} c\dd t' &= c\Lambda^t_\mu\,\dd X^\mu \\ &= \gamma(c\dd t - \beta \dd x)\\ &= -\gamma\beta\dd x \end{align*}

Since  dx≠0\dd x \neq 0 for two distinct endpoints,  dt′≠0\dd t' \neq 0. The two events we used — simultaneous in SS — are not simultaneous in S′S'. That's exactly the mistake in the train story above: we measured one end and the other end at different times, just dressed up in Lorentz-transformation language instead of morning-and-night language. What we computed as "γ dx\gamma \dd x" was never the train's length in S′S' at all.

The same logic runs in reverse. Force  dt′=0\dd t' = 0 instead — simultaneous in the moving frame — and see what that requires of  dt\dd t:

c dt′=0=cΛμt  dXμ=γ(c dt−β dx)c dt=β dx\begin{align*} c\dd t' = 0 &= c\Lambda^t_\mu\,\dd X^\mu \\ &= \gamma(c\dd t-\beta\dd x) \\ c\dd t &= \beta\dd x \end{align*}

So events simultaneous in S′S' are generally not simultaneous in SS either. Simultaneity isn't something the two frames agree on — it's frame-dependent, full stop, and it was hiding inside the Lorentz transformation the whole time.

Length contraction, done properly§

Armed with that, let's redo the measurement honestly. If we want the train's proper length, we need to measure both ends at the same time in the train's own rest frame S′S' — that means  dt′=0\dd t' = 0, not  dt=0\dd t = 0. Using the relation we just derived, c dt=β dxc\dd t = \beta \dd x, substitute into the spatial transformation:

 dx′=Λμx dXμ=γ( dx−βc dt)=γ( dx−β2 dx)=γ(1−β2) dx=1γ dx\begin{align*} \dd x' &= \Lambda^x_\mu \dd X^\mu \\ &= \gamma(\dd x-\beta c\dd t) \\ &= \gamma(\dd x - \beta^2 \dd x)\\ &= \gamma(1-\beta^2 )\dd x \\ &= \frac{1}{\gamma}\dd x \end{align*}

And we write  dx\dd x as the proper length L0L_0 and  dx′\dd x' as the measured length. Where the proper length is deifned as the length measreud in the rest frame.

L=L0γL = \frac{L_0}{\gamma}

This is length contraction. The earlier factor of γ\gamma wasn't wrong arithmetic — it was the right arithmetic applied to the wrong pair of events. The deeper lesson: a Lorentz transformation doesn't just rescale numbers, it also reshuffles which events count as simultaneous, and length contraction is really a simultaneity effect wearing a geometry costume.

And since, now proper length is defined, let's see why the invariant  ds2\dd s^2 is the proper length ( squared ). Well, since it is invariant ( it's value doesn't depend on the reference frame ), we can evaluate it's value in any reference frame that makes the meaning trivial. We chose the frame that makes dt=0dt = 0, Which measure the length of the object in the rest frame, which by definition of proper length.

 ds2=−c2 dt2+ dx2+ dy2+ dz2= dx2+ dy2+ dz2=L02\dd s^2 = -c^2\dd t^2 + \dd x^2 + \dd y^2 + \dd z^2 = \dd x^2 + \dd y^2 + \dd z^2 = L_0^2

Boost of a moving object§

Now let's keep the velocity explicit instead of setting it to zero. Suppose an object moves at velocity uu in SS, so  dx=u dt\dd x = u\dd t, and we boost by vv:

 dx′=Λμx dXμ=γ( dx−βc dt)=γ(u−v) dt dx′ dt′=γ(u−v) dt dt′u′=γ(u−v) dtγ( dt−vc2 dx)u′=u−v1−uvc2\begin{align*} \dd x' &= \Lambda^x_\mu \dd X^\mu \\ &= \gamma(\dd x-\beta c\dd t) \\ &= \gamma(u-v)\dd t \\ \frac{\dd x'}{\dd t'} &= \gamma(u-v)\frac{\dd t}{\dd t'} \\ u' &= \gamma (u-v) \frac{\dd t}{\gamma\left(\dd t - \dfrac{v}{c^2}\dd x\right)} \\ u' &= \frac{u-v}{1 - \dfrac{uv}{c^2}} \end{align*}

If u,v≪cu,v \ll c, the denominator collapses to 11 and we recover the Galilean u′≈u−vu' \approx u - v, exactly as it should.

Boost of light§

Now push it to the extreme: set u=cu = c.

u′=c−v1−vc=cu'=\frac{c - v}{1-\dfrac{v}{c}}=c

Light stays light speed no matter how you boost. That's not an extra rule bolted on afterward — it falls straight out of the same transformation we've been using all along.

Doppler effect§

A source at rest in SS emits with period T0T_0 — its proper period. Time dilation alone tells us the observed period lengthens to T=γT0T = \gamma T_0. But that's not the whole story: while the wave takes time to cross the widening gap, the observer (moving away, say) keeps retreating further, so each successive crest has a little farther to travel than the last. That extra travel time is t=s/c=(v/c)Tt = s/c = (v/c)T, and folding it in gives

Tobs=(1+β)γT0=T01+β1−βT_\text{obs}=(1 + \beta)\gamma T_0 = T_0\sqrt\frac{1+\beta}{1-\beta}

so the observed frequency is

fobs=1−β1+β f0f_\text{obs}=\sqrt\frac{1-\beta}{1+\beta}\,f_0

This is the longitudinal Doppler effect — receding along the line of sight gives you both the time-dilation piece and the widening-gap piece together. If instead the motion is purely transverse — perpendicular to the line of sight, so the separation isn't changing — that second piece vanishes and only the dilation survives:

fobs=1γf0f_\text{obs}=\frac{1}{\gamma}f_0

Aberration§

Doppler changes a wave's frequency. The same boost, applied to a different object, also changes the direction it appears to arrive from — that's aberration.

Let a photon travel in SS at angle θ\theta to the xx-axis, so its momentum components are px=pcos⁡θp_x=p\cos\theta, py=psin⁡θp_y=p\sin\theta, with E=pcE=pc. Photon momentum transforms as a four-vector, exactly the way (ct,x)(ct,x) did — we haven't shown why (E/c, px, py, pz)(E/c,\,p_x,\,p_y,\,p_z) forms a four-vector yet, that's a later article, so take it on credit for now:

px′=γ(px−βE/c),E′/c=γ(E/c−βpx)p_x' = \gamma(p_x - \beta E/c), \qquad E'/c = \gamma(E/c - \beta p_x)

Divide the first by the second, and use px/(E/c)=cos⁡θp_x/(E/c)=\cos\theta:

cos⁡θ′=cos⁡θ−β1−βcos⁡θ\cos\theta' = \frac{\cos\theta - \beta}{1-\beta\cos\theta}

A source moving toward you doesn't just blueshift — its light also gets funneled forward, toward θ′=0\theta'=0, no matter which direction it was actually emitted in. Push β→1\beta\to1 and nearly everything a fast-moving source radiates ends up crammed into a narrow forward cone. That's the headlight effect, and it's the same Λ\Lambda we've had since the first page.

2D boost§

By isotropy, a boost pointed in some other single direction is no more interesting than the 1D case — the math gets uglier, the physics doesn't. The only reason 2D boosts deserve their own section is what happens when you chain two boosts in different, non-collinear directions. (Chain two boosts in the same direction and you just get one bigger boost — nothing new there.)

So look back at the generator, and consider composing two boosts:

Λ=e−iϕxKxe−iϕyKy\Lambda = e^{-i\phi_x K_x}e^{-i\phi_y K_y}

For small ϕx,ϕy\phi_x, \phi_y, expand to leading order:

Λ=(I−iϕxKx)(I−iϕyKy)=I−i(ϕxKx+ϕyKy)−ϕxϕyKxKy+O(ϕ3)\Lambda = (I-i\phi_xK_x)(I-i\phi_yK_y) = I - i(\phi_xK_x+\phi_yK_y) - \phi_x\phi_y K_xK_y + O(\phi^3)

This is where that commutation relation we set aside earlier finally gets used. Recall [Kx,Ky]=−iJz[K_x,K_y] = -iJ_z, so KxKy=KyKx−iJzK_xK_y = K_yK_x - iJ_z, and substituting,

Λ=I−i(ϕxKx+ϕyKy)−ϕxϕyKyKx+iϕxϕyJz+O(ϕ3)\Lambda = I - i(\phi_xK_x+\phi_yK_y) - \phi_x\phi_y K_yK_x + i\phi_x\phi_y J_z + O(\phi^3)

There it is: a JzJ_z term. Two pure boosts, composed, have produced a piece of a rotation.

(That's only the leading term. For finite rapidities, the exact rotation angle θW\theta_W between two perpendicular boosts of rapidity ϕ\phi and χ\chi works out to tan⁡(θW/2)=tanh⁡(ϕ/2)tanh⁡(χ/2)\tan(\theta_W/2)=\tanh(\phi/2)\tanh(\chi/2), which reduces to our ϕxϕy\phi_x\phi_y result once both are small. We won't derive the finite case here, but it's worth knowing the infinitesimal calculation isn't hiding anything qualitatively different from the exact one.)

Here's the physical picture behind that algebra. Take a rod already moving along xx, and in its own instantaneous rest frame, give both of its ends a simultaneous sideways kick in yy — a "rigid" perpendicular push. Simultaneous in that frame, that is. That frame is reached from the lab by an xx-boost, and an xx-boost's relativity of simultaneity depends on xx-position — so two events that share a yy or zz coordinate but differ in xx, simultaneous in the boosted frame, are not simultaneous back in the lab. Since the rod's two ends are separated along xx — the same direction as the boost it's already carrying — the "simultaneous" kick lands on one end before the other, as clocked in the lab. By the time you look at the whole rod at a single lab instant, one end has had a head start accelerating sideways and the other hasn't caught up — so the rod appears to have tilted, not because anything physically twisted it, but because the two kicks were smeared apart in time. And this only shows up when the boost you add is genuinely non-collinear with the one already there — a further kick purely along xx, collinear with the first boost, never desynchronizes anything.

This composition of non-collinear boosts producing a rotation is called a Wigner rotation. Let an object's velocity change continuously — think of it as an unbroken chain of these infinitesimal non-collinear boosts — and the accumulated rotation of its comoving frame is what we call Thomas precession.

non-commuting boosts  ⟶  Wigner rotation  ⟶  Thomas precession (continuous acceleration)\text{non-commuting boosts} \;\longrightarrow\; \text{Wigner rotation} \;\longrightarrow\; \text{Thomas precession (continuous acceleration)}

It's worth stating the group-theoretic version of what just happened: pure boosts do not close under composition. Compose two non-collinear ones and you land outside the set of pure boosts entirely, in something with a rotation baked in. Boosts alone aren't a subgroup of the Lorentz group — which is exactly why so(1,3)\mathfrak{so}(1,3) needed both K\mathbf K and J\mathbf J from the very first page. K\mathbf K by itself doesn't close under commutation; [Ki,Kj]∝Jk[K_i,K_j]\propto J_k kicks you straight back out into rotation territory.

So that commutation relation we asked you to remember, [Ki,Kj]=−iϵijkJk[K_i, K_j] = -i\epsilon_{ijk}J_k, was never just bookkeeping. At the infinitesimal level it's already saying: change the direction you're boosting in, and a rotation is unavoidable.

This isn't just a geometric curiosity. The textbook non-relativistic calculation of spin-orbit coupling in hydrogen — the electron's spin interacting with the effective magnetic field produced by its own orbit — comes out a factor of 2 too large compared to experiment, until you account for the fact that the electron's rest frame is itself undergoing exactly this kind of precession as it orbits the nucleus. Thomas precession supplies the missing factor of 1/21/2. A purely kinematic fact about composing boosts turns out to matter for the fine structure of every atom you're made of.

Discussion

no comments
Commenting as a guest — sign in to comment as yourself.

No comments yet — yours could open the discussion.