Seidel Aberrations

← Back to Knowledge Share

A perfect lens would turn every point of an object into a point in the image. A real one does not. The image of a star spreads into a disc, a comet, a short line or a smear, and which of those you get depends on where in the field the star sits and how wide the aperture is. Seidel’s answer, now more than a century and a half old, is still the first language for describing that: five named ways a rotationally symmetric system fails.

The five are spherical aberration, coma, astigmatism, field curvature and distortion. They are not five separate physical effects that happen to occur together. They are the five terms that survive when the wavefront error of any rotationally symmetric system is expanded to the lowest interesting order, which is why every lens, telescope and microscope objective is discussed in the same vocabulary. This note builds that expansion, shows what each term does to an image, and works out how large each has to be before it matters. Their orthogonal cousins, the Zernike polynomials, get their own note.


1. The Wave Aberration Function

Everything is measured at the exit pupil, the last surface where the light is still a converging wave rather than an image. If the system were perfect, the wave leaving the pupil would be a sphere centered on the ideal image point. That sphere is the reference sphere, and the wave aberration function \(W\) is the gap between it and the actual wavefront, measured along the ray as an optical path difference. A length, usually quoted in waves as \(W/\lambda\).

\(W\) depends on where the object point is and on which part of the pupil the light goes through, so it takes three dimensionless arguments:

  • \(H\), the field coordinate: the object point’s distance from the axis divided by the largest distance the system is asked to image, so \(H = 1\) at the corner of the frame;
  • \(\rho\), the radial pupil coordinate: distance from the pupil center divided by the pupil radius \(a\), so \(\rho = 1\) at the rim;
  • \(\theta\), the azimuth around the pupil, measured from the plane containing the axis and the object point, so \(\theta = 0\) points toward the field and \(\theta = 90^\circ\) is perpendicular to it.
Upper half of the pupil: the wave aberration W is the gap along the ray between the reference sphere centred on the ideal image point and the actual wavefront, and its slope moves the ray by epsilon in the image plane. exit pupil image plane reference sphere actual wavefront W ε ideal image point ρ
Figure 1. The upper half of the pupil, with the curvature and the error exaggerated. The reference sphere (dashed) is centered on the ideal image point; the actual wavefront (solid) departs from it by \(W\), measured along the ray. Because the ray runs perpendicular to the wavefront, that departure lands the ray a distance \(\varepsilon\) away from the ideal point.

The wavefront shape matters because rays travel perpendicular to it: tilt a patch of wavefront and the ray through it lands somewhere else. To first order the transverse ray error in the image plane is the pupil gradient of \(W\),

\[\varepsilon_x = \frac{R}{n'a}\,\frac{\partial W}{\partial \rho_x}, \qquad \varepsilon_y = \frac{R}{n'a}\,\frac{\partial W}{\partial \rho_y},\]

where \(\rho_x = \rho\cos\theta\) and \(\rho_y = \rho\sin\theta\) are Cartesian pupil coordinates, \(R\) is the distance from the exit pupil to the image plane, and \(n'\) is the refractive index of image space. Signs depend on the convention chosen for \(W\); the magnitudes are what matter here. The useful consequence is that a term in \(W\) growing as \(\rho^4\) produces ray errors growing as \(\rho^3\), one power lower, which is where the older name for these five, the third-order aberrations, comes from.

One more thing is hiding in the definition. The reference sphere is centered on a point that you get to choose. Move that center sideways and \(W\) picks up a term proportional to \(\rho\cos\theta\); move it along the axis and \(W\) picks up a term proportional to \(\rho^2\). So tilt and defocus are not really aberrations, they are statements about where you decided to put the image, and any aberration can be partly hidden by refocusing. Section 4 makes that precise.

2. Why There Are Exactly Five

A rotationally symmetric system cannot tell one azimuth from another. Rotate the object point about the axis and the entire aberration pattern rotates rigidly with it. That one fact fixes the possible form of \(W\). Treating the field and pupil positions as vectors, the only combinations that survive an arbitrary rotation of both are

\[H^2, \qquad \rho^2, \qquad H\rho\cos\theta,\]

the squared lengths of the two vectors and their dot product. So \(W\) is a function of those three quantities and of nothing else.

Expand that function as a power series and sort the terms by their total order in \(H\) and \(\rho\). The lowest group is second order: a constant \(W_{000}\) (piston, an overall delay that no detector sees), \(W_{111}H\rho\cos\theta\) (tilt, which moves the image point) and \(W_{020}\rho^2\) (defocus, which moves the focus along the axis). None of them blurs anything; they say where the image is, which is why they belong to the system’s first-order, or Gaussian, description. The first group that actually degrades an image is fourth order, and it has exactly five members:

\[W(H, \rho, \theta) = W_{040}\,\rho^4 + W_{131}\,H\rho^3\cos\theta + W_{222}\,H^2\rho^2\cos^2\theta + W_{220}\,H^2\rho^2 + W_{311}\,H^3\rho\cos\theta.\]

The subscripts are a counting scheme, not magic: in \(W_{klm}\), \(k\) is the power of \(H\), \(l\) the power of \(\rho\) and \(m\) the power of \(\cos\theta\). Every Seidel term has \(k + l = 4\), and the five combinations above are the only ones that can be built from the three invariants at that order. Each \(W_{klm}\) is a coefficient with units of length, fixed by the glasses, curvatures, thicknesses and stop position of the particular system.

Two practical consequences follow from the expansion itself, before any optics.

  • The coefficients add. Each surface in a system contributes its own five numbers and they sum, which is why a design program reports the five sums surface by surface and why a designer can cancel a positive contribution at one surface against a negative one at another. Lens design, to a first approximation, is the art of making five sums small at once.
  • The expansion is a truncation, not a law. A system that is fast or wide enough will show sixth-order and higher terms, and then the Seidel numbers stop predicting the image. They stay useful as the leading behavior and as the language designers think in.

3. What Each Term Does

The gradient rule turns each term into a shape. Figure 2 traces a grid of rays, evenly spread over the pupil, from each aberration to the paraxial image plane; the cross marks where a perfect system would put them.

Spot diagrams at the paraxial focus for the five Seidel aberrations: a symmetric blur for spherical, a comet for coma, a line for astigmatism, a uniform disc for field curvature and a shifted point for distortion. spherical coma astigmatism field curvature distortion
Figure 2. Spot diagrams for the five Seidel terms at the paraxial focus, each panel scaled to its own size. The rays are a hexapolar grid on the pupil, so rings of dots are rings of the pupil. Relative to a unit coefficient the spreads differ: spherical reaches \(4\), coma \(3\), astigmatism and field curvature \(2\), distortion \(1\).

Spherical Aberration

\(W_{040}\rho^4\) is the only term with no \(H\) in it. It is the same everywhere in the field, on axis included, and it is the reason a fast lens can be soft even at the center of the frame. The ray error is radial and grows as \(\rho^3\), so each pupil zone focuses at its own distance and the image of a point becomes a bright core inside a halo. The cures are the classic ones: bend or split the element, add an asphere, or stop down.

Coma

\(W_{131}H\rho^3\cos\theta\) is linear in field, so it vanishes on axis and grows steadily toward the corner. Working out its gradient, the ring of radius \(\rho\) in the pupil maps to a circle in the image of radius proportional to \(\rho^2\), whose center is displaced by twice its own radius. The circles therefore nest into a one-sided flare with a bright head at the ideal point, bounded by two straight lines at \(30^\circ\) either side of the axis, which is the comet the aberration is named after. Its asymmetry is what makes it nastier than its size suggests: the centroid of the blur is not the ideal image point, so coma shifts apparent positions as well as blurring them.

Astigmatism and Field Curvature

These two travel together, both quadratic in field. Split the astigmatism term with a double-angle identity,

\[W_{222}H^2\rho^2\cos^2\theta = \tfrac{1}{2}W_{222}H^2\rho^2 + \tfrac{1}{2}W_{222}H^2\rho^2\cos 2\theta,\]

and the first piece is plain defocus while the second is the part that cannot be focused away. The consequence is two line foci: the fan of rays in the meridional plane and the fan perpendicular to it come to focus at different distances, giving a line at one focus, a line at right angles at the other, and a small circle of least confusion between them. The two are separated along the axis by \(\delta z = 8N^2W_{222}H^2\), with \(N\) the working f-number.

\(W_{220}H^2\rho^2\) is the same \(\rho^2\) shape without the azimuthal part: defocus that grows with field, that is, a curved image surface. Even a system with every other aberration removed has one, set by the Petzval sum of the glasses and curvatures, and a flat sensor can only cut through it. When astigmatism is also present there are two such surfaces, one for each fan, sitting either side of the Petzval surface with the tangential one three times as far from it as the sagittal.

Distortion

\(W_{311}H^3\rho\cos\theta\) is linear in \(\rho\), and a wavefront error linear in pupil coordinate is a tilt. It does not blur anything: the image of a point stays sharp and simply lands in the wrong place, displaced by an amount growing as \(H^3\). Straight lines bow outward (barrel) or inward (pincushion) depending on the sign. Because the information is all still there, it is the one Seidel aberration a computer can undo after the fact by resampling, which is why phone cameras ship with heavy distortion and correct it in software.

aberrationwave termfieldpupil radiusimage of a point
spherical\(W_{040}\,\rho^4\)none\(a^4\)core plus halo
coma\(W_{131}\,H\rho^3\cos\theta\)\(H\)\(a^3\)one-sided comet
astigmatism\(W_{222}\,H^2\rho^2\cos^2\theta\)\(H^2\)\(a^2\)two line foci
field curvature\(W_{220}\,H^2\rho^2\)\(H^2\)\(a^2\)disc off a curved surface
distortion\(W_{311}\,H^3\rho\cos\theta\)\(H^3\)\(a\)sharp point, wrong place

The last two columns are how each coefficient itself scales, with \(a\) the pupil radius, once the normalized coordinates are unwrapped. They are the reason the five behave so differently when a lens is stopped down or a sensor is made larger, which is section 4.

4. How Much Does Each One Cost?

Peak wavefront error is the wrong currency, because an error confined to the very rim of the pupil matters less than one spread over half of it. The standard measure is the root-mean-square wavefront error over the pupil,

\[\sigma^2 = \big\langle W^2 \big\rangle - \big\langle W \big\rangle^2, \qquad \big\langle f \big\rangle = \frac{1}{\pi}\int_0^{2\pi}\!\!\int_0^1 f(\rho, \theta)\,\rho\,d\rho\,d\theta,\]

where \(\langle f\rangle\) is the average of \(f\) over the unit pupil disc and \(\sigma\) is in the same units as \(W\), usually waves. Subtracting the mean discards piston, which no measurement sees. Two properties make \(\sigma\) the number to quote: independent contributions add in quadrature, and for small errors it predicts how much light stays in the core of the image of a point through the Maréchal approximation to the Strehl ratio,

\[S \approx \exp\!\left[-\left(\frac{2\pi\sigma}{\lambda}\right)^{2}\right],\]

with \(S\) the peak intensity of the real image of a point divided by the peak a perfect system of the same aperture would give, and \(\lambda\) the wavelength. The approximation holds while \(\sigma \lesssim \lambda/10\). It also supplies the usual meaning of “diffraction limited”: \(\sigma = \lambda/14\) gives \(S \approx 0.8\), the Maréchal criterion.

Since refocusing and recentering the image cost nothing, each aberration should be charged at its best focus rather than at the paraxial one. The difference is large, and not the same for every term:

aberrationRMS at paraxial focusRMS at best focuswhat was subtracted
spherical\(0.298\,W_{040}\)\(0.075\,W_{040}\)defocus \(-W_{040}\rho^2\)
coma\(0.354\,W_{131}H\)\(0.118\,W_{131}H\)tilt \(-\tfrac{2}{3}W_{131}H\rho\cos\theta\)
astigmatism\(0.250\,W_{222}H^2\)\(0.204\,W_{222}H^2\)defocus \(-\tfrac{1}{2}W_{222}H^2\rho^2\)
field curvature\(0.289\,W_{220}H^2\)\(0\)defocus: it is defocus
distortion\(0.500\,W_{311}H^3\)\(0\)tilt: it is tilt

Three things worth reading off it. Refocusing buys a factor of four on spherical and three on coma, but only a fifth on astigmatism, because the two line foci cannot both sit in the image plane. Field curvature and distortion cost nothing in this accounting: one is defocus that varies with field and the other is tilt that varies with field, and they are problems only because sensors are flat and mappings are supposed to be linear. And the coefficients are per unit of \(W\), so one wave of \(W_{040}\) is \(0.075\) waves RMS at best focus, that is \(\lambda/13.4\): one wave of Seidel spherical aberration is almost exactly the most a system can carry and still pass the Maréchal criterion.

Aperture and Field

Because \(\rho\) and \(H\) are normalized, all the dependence on how fast and how wide the system is sits inside the coefficients. Unwrapping it for a fixed focal length, where the pupil radius \(a\) is inversely proportional to the working f-number \(N\), gives

\[W_{040} \propto N^{-4}, \qquad W_{131} \propto N^{-3}H, \qquad W_{222},\, W_{220} \propto N^{-2}H^2, \qquad W_{311} \propto N^{-1}H^3.\]

Stopping down one stop, which multiplies \(N\) by \(\sqrt{2}\), therefore cuts spherical aberration by four, coma by 2.8 and astigmatism by two, and barely touches distortion. Going from f/8 to f/4 runs it the other way: sixteen times the spherical aberration.

Defocus obeys a useful companion relation. Shifting the image plane by \(\delta z\) adds \(W_{020} = \delta z / (8N^2)\) waves of defocus, so the classical quarter-wave depth of focus is \(\delta z = \pm 2\lambda N^2\), which at f/4 and 550 nm is \(\pm 17.6\) µm. The same relation converts any of the focus shifts above into millimetres on a focusing ring.

5. A Worked Example: Stopping Down

Take a fast lens that leaves two waves of Seidel spherical aberration wide open at f/2.8, at 550 nm. Spherical is the term that matters here because it does not care about field: this is what the center of the frame looks like. Everything needed to follow the iris closing is already above. The coefficient falls as \(N^{-4}\), the RMS at best focus is \(W_{040}/(6\sqrt{5})\), the Strehl ratio follows from the Maréchal approximation, the best focus sits \(\delta z = 8N^2W_{040}\) from the paraxial one, and the Airy radius of a perfect system of the same aperture is \(1.22\lambda N\).

aperture\(W_{040}\) (waves)RMS at best focusStrehlfocus shiftAiry radius
f/2.82.0000.149 waves0.4269.0 µm1.88 µm
f/40.4800.036 waves0.9533.8 µm2.68 µm
f/5.60.1250.009 waves0.99717.2 µm3.76 µm
f/80.0300.002 waves1.0008.5 µm5.37 µm

Read across the rows and the whole behavior of a fast lens falls out. Wide open it is nowhere near diffraction limited: 0.149 waves RMS is twice the Maréchal allowance and well under half the light stays in the core. One stop down the aberration has dropped by a factor of four, the Strehl is 0.95, and the lens has become, for practical purposes, perfect. Two stops down the wavefront error is \(\lambda/107\), which buys nothing, because by then the Airy radius has grown to 3.8 µm: the image is softer than at f/4 for a reason that has nothing to do with aberration. The best working aperture is the crossover, around f/4 here, the first stop at which aberration is no longer what limits the image.

The focus-shift column is the practical trap in the table. Best focus for spherical aberration is not the paraxial focus, and as the aperture closes it slides back toward it: 69 µm at f/2.8, half that at f/4, half again at f/5.6. A camera focused wide open and then stopped down to shoot is focused in the wrong place by tens of microns, which is the familiar focus shift of fast lenses, and it is why a lens is tested at the aperture it will be used at.

One caveat before generalizing: this is the center of the frame. At the corner, coma and astigmatism take over, and they fall only as \(N^{-3}\) and \(N^{-2}\). Stopping down two stops divides spherical aberration by sixteen but astigmatism only by four, so the corner keeps improving long after the center has stopped, which is exactly the pattern lens tests show.

6. Where the Five Stop Working

The Seidel expansion earns its place by being the lowest-order truth about a symmetric system. Each of those words is a limit.

  • Rotationally symmetric. Tilt or decenter one element and the derivation collapses, because the aberration pattern no longer rotates with the object point. Such systems are described by nodal aberration theory, in which each Seidel term acquires a node that wanders around the field, and freeform optics lives outside the list entirely.
  • Lowest order. A fast or wide system shows sixth-order terms next, among them zonal spherical aberration, which leaves a residual ring after the fourth-order part is corrected. A design with its five Seidel sums driven to zero can still have a visibly imperfect image.
  • Monochromatic. Color is not in the list. Axial and lateral chromatic aberration are their own pair of first-order defects and usually dominate a simple lens long before the five do.
  • Geometric. The expansion counts optical path, not diffraction. Once the aberrations get small the image of a point stops being a ray pattern and becomes an Airy disc, which is why section 4 switched from spot sizes to RMS and Strehl.

There is one more limitation, subtler and more practical. The five terms, together with defocus and tilt, are not orthogonal over the pupil. Section 4 already showed the symptom: how much spherical aberration a wavefront “has” depends on where the focus is put, because \(\rho^4\) and \(\rho^2\) overlap. Fit a measured wavefront to Seidel terms and every coefficient shifts when a term is added or dropped, which makes the numbers hard to compare between instruments or across a design’s history. Rebuilding the same information on an orthogonal basis fixes that, and the standard choice is the subject of the companion note on Zernike polynomials.

Intuitively: Symmetry Writes the List

The five aberrations are not five separate discoveries that happened to fill the list. They are the complete inventory of ways a rotationally symmetric system is allowed to fail at the lowest order, and the field dependence of each one counts how much of the symmetry has been given up.

On axis, the system looks identical from every direction around the pupil. An error that depends on the azimuth would have to point somewhere, and there is nowhere to point, so the only defect available depends on \(\rho\) alone: spherical aberration, the same across the whole field. Move the object off axis and the pupil acquires a direction, namely toward the field, and an error linear in that direction becomes possible: coma, and it grows in proportion to how far off axis you have gone. Go further and the pupil is seen obliquely, foreshortened along one axis and not the other, and the error can differ between those two axes: astigmatism, quadratic in field, with field curvature as the part of the same quadratic that does not care about azimuth. Furthest along, the mapping from object to image can stretch without blurring at all: distortion, cubic in field.

That is the whole story of the list. One term for the symmetric case, then one more for each step away from symmetry, each carrying one more power of the field. Seidel's contribution was not to notice five phenomena but to prove that there are no others to notice, at least until the next order.