Depth from Focus & Defocus
- 1. Overview
- 2. Point Spread Function (PSF)
- 3. Depth from Focus (DFF)
- 4. Depth from Defocus (DFD)
- 5. Technical Comparison Summary
In computer vision, depth and shape recovery methods are generally divided into two main categories: active methods (laser scanners, structured light, etc.) and passive methods (stereo vision, shape from motion, etc.). Based on optical focus constraints, Depth from Focus (DFF) and Depth from Defocus (DFD) are passive and powerful depth sensing techniques that leverage the finite depth of field of single-lens cameras as a physical depth cue.
1. Overview
In images captured with a camera having a shallow depth of field, only objects located at the plane of focus appear sharp and crisp; objects in front of or behind this plane become defocused and blurred. According to optical physics, the amount and structure of blur are directly related to the physical distance of the object from the focus plane.
However, estimating local blur amount from a single image is mathematically an under-constrained problem. Given a single image patch, it is impossible to distinguish whether it appears blurry because it was captured out of focus or because the object’s original surface texture is inherently smooth/blurry. For example, a sharp photo of a smooth white wall looks identical locally to an out-of-focus photo of the same wall.
To overcome this ambiguity, multiple images taken under different focus settings or camera parameters are required. Two primary paradigms have been developed:
- Depth from Focus (DFF): Sweeps the focus plane step-by-step across the scene to collect a large focal stack. For each pixel coordinate, it searches for the image slice where contrast and sharpness are maximized.
- Depth from Defocus (DFD): Typically captures only two or three images with different focus or aperture settings. It calculates scene depth directly using analytical formulas or optimization techniques by analyzing relative blur ratios between the images.
2. Point Spread Function (PSF)
To mathematically model defocus blur, the spatial energy distribution formed on the sensor by an ideal point light source (impulse) must be defined. This distribution is called the Point Spread Function (PSF).
2.1 Circle of Confusion Geometry
According to the Gaussian Lens Law, a scene point at distance $u$ (or $o$) from a lens with focal length $f$ focuses perfectly at distance $v$ (or $i$) behind the lens:
$$\frac{1}{f} = \frac{1}{u} + \frac{1}{v}$$
If the sensor (image plane) is positioned at distance $s$ instead of the ideal focus distance $v$, focused rays intersect the sensor plane forming a circular light patch. Assuming a circular aperture, this base of the light cone is called the Blur Circle or Circle of Confusion.
Using similar triangles, the diameter of the blur circle ($b$) is related to aperture diameter ($D$) as follows:
$$\frac{b}{D} = \frac{|v - s|}{v} \implies b = D \cdot s \left| \frac{1}{s} - \frac{1}{v} \right|$$
This equation demonstrates two physical ways to control defocus blur amount:
- Vary Sensor Position ($s$): Translating the focal plane back and forth across the scene.
- Vary Aperture Size ($D$): Stopping down the lens (reducing $D$) narrows the light cone, shrinking blur diameter ($b$) and increasing depth of field.
2.2 Pillbox vs. Gaussian PSF Models
In an ideal, diffraction-free optical system, light distribution across the blur circle can be modeled as a uniform circular disk. This is termed the Pillbox Function:
$$h_{\text{pillbox}}(x, y) = \begin{cases} \frac{4}{\pi b^2}, & x^2 + y^2 \leq \frac{b^2}{4} \ 0, & \text{otherwise} \end{cases}$$
The normalization factor $\frac{4}{\pi b^2}$ enforces conservation of optical energy across expanding blur circles.
In real-world optical systems, diffraction at aperture edges, optical aberrations, surface roughness, and spatial pixel integration prevent sharp-edged pillbox distributions. Consequently, practical PSFs are realistically modeled as smooth Gaussian Functions:
$$h_{\text{Gaussian}}(x, y) = \frac{1}{2\pi \sigma^2} e^{-\frac{x^2+y^2}{2\sigma^2}}$$
Empirically, Gaussian standard deviation ($\sigma$) relates to blur circle diameter ($b$) as:
$$\sigma \approx \frac{b}{2} \propto D \cdot s \left| \frac{1}{s} - \frac{1}{v} \right|$$
2.3 Convolution and Low-Pass Filter Equivalence
Assuming depth is locally constant over small patches, defocus imaging behaves as a Linear Shift-Invariant (LSI) system. Under LSI assumptions, the captured blurry image $g(x,y)$ equals the focused image $f(x,y)$ convolved with the PSF $h(x,y)$:
$$g(x, y) = f(x, y) * h(x, y)$$
In the frequency (Fourier) domain, convolution converts to pointwise multiplication:
$$G(u, v) = F(u, v) \cdot H(u, v)$$
Because the Fourier transform of a Gaussian is also a Gaussian, an expanding PSF in spatial space ($\sigma$ growth) corresponds to a narrower Gaussian filter in frequency space.
Optically, defocus acts as a Low-Pass Filter. It preserves low-frequency macro structure while attenuating high-frequency textures, sharp edges, and fine details. Depth algorithms evaluate this high-frequency loss to infer distance.
3. Depth from Focus (DFF)
Depth from Focus (DFF) sweeps the focus plane across the scene in step increments to collect a focal stack, identifying the focal plane slice where high-frequency content peaks for each pixel.
3.1 Focus Measure and Modified Laplacian
To evaluate sharpness across a focal stack, a local Focus Measure operator is defined. Since defocus suppresses high frequencies, local brightness variations (second derivatives) quantify sharpness.
In standard Laplacian operators, horizontal and vertical second derivatives can have opposite signs and cancel out. To prevent cancellation, the Modified Laplacian ($\nabla_M^2$) sums absolute partial second derivatives:
$$\nabla_M^2 I = \left| \frac{\partial^2 I}{\partial x^2} \right| + \left| \frac{\partial^2 I}{\partial y^2} \right|$$
On discrete pixel grids, partial derivatives are approximated by finite difference kernels:
$$\frac{\partial^2 I}{\partial x^2} = I(x+1, y) - 2I(x, y) + I(x-1, y)$$
$$\frac{\partial^2 I}{\partial y^2} = I(x, y+1) - 2I(x, y) + I(x, y-1)$$
The focus measure score $M(x,y)$ is computed by accumulating Modified Laplacian values within a local window (typically $3 \times 3$ or $5 \times 5$):
$$M(x, y) = \sum_{i=x-K}^{x+K} \sum_{j=y-K}^{y+K} \nabla_M^2 I(i, j)$$
3.2 Gaussian Interpolation for Smooth Reconstruction
Assigning depth directly to discrete focal stack layer indices causes depth resolution to be constrained by stack size $N$, creating staircase/contouring artifacts on 3D models. Increasing $N$ significantly increases capture time and memory footprint.
To achieve sub-stack precision, the focus measure distribution $M(s)$ near its peak is modeled as a Gaussian bell curve:
$$M(s) = M_p e^{-\frac{(s - \bar{s})^2}{2\sigma_m^2}}$$
Taking the natural logarithm linearizes the Gaussian model:
$$\ln M(s) = \ln M_p - \frac{(s - \bar{s})^2}{2\sigma_m^2}$$
Selecting the three highest discrete focus scores ($M_1, M_2, M_3$) and equal step size $\Delta s = s_2 - s_1 = s_3 - s_2$, an analytical closed-form solution yields continuous sensor location $\bar{s}$:
$$\bar{s} = s_2 + \frac{\Delta s \left( \ln M_3 - \ln M_1 \right)}{2 \left( 2 \ln M_2 - \ln M_1 - \ln M_3 \right)}$$
Substituting continuous $\bar{s}$ into the Gaussian lens law produces smooth, high-precision 3D depth maps.
DFF is extensively applied in microscopy and industrial quality inspection where shallow depth of field lens systems operate at micron scales.
Key Limitation: DFF relies strictly on high-frequency surface texture; smooth, untextured regions cannot produce differential contrast variations across focus steps.
4. Depth from Defocus (DFD)
While DFF offers high accuracy, collecting tens of images is impractical for real-time video capture (30 FPS). Depth from Defocus (DFD) estimates depth rapidly by analyzing relative blur differences between as few as two images.
4.1 Naive DFD Solution (Ratio of Fourier Transforms)
Consider two images ($g_1, g_2$) of scene $f(x,y)$ taken with aperture sizes $D_1, D_2$ resulting in PSF widths $\sigma_1, \sigma_2$. Since aperture settings are hardware-controlled, the ratio of PSF widths is known:
$$\frac{\sigma_1}{\sigma_2} = \frac{D_1}{D_2} \implies \sigma_2 = \sigma_1 \frac{D_2}{D_1}$$
Writing spatial convolution equations in Fourier space:
$$G_1(u, v) = F(u, v) \cdot H_{\sigma_1}(u, v)$$
$$G_2(u, v) = F(u, v) \cdot H_{\sigma_2}(u, v)$$
Taking the ratio of Fourier transforms cancels the unknown true focused image $F(u,v)$:
$$\frac{G_1(u, v)}{G_2(u, v)} = \frac{H_{\sigma_1}(u, v)}{H_{\sigma_2}(u, v)}$$
Substituting Gaussian PSF formulas and taking logarithms yields an explicit expression for $\sigma_1$:
$$\sigma_1^2 - \sigma_2^2 = \frac{\ln G_2(u, v) - \ln G_1(u, v)}{2 \pi^2 (u^2 + v^2)}$$
Solving for $\sigma_1$ gives blur diameter $b_1 = 2\sigma_1$, from which scene depth $u$ is directly computed.
Warning: The naive Fourier ratio approach is unstable under sensor noise due to high-frequency division ($u^2+v^2$).
4.2 Reconstruction-Based Stable DFD
To handle noise robustly, Favaro (2003) and Pentland (1987) proposed an optimization-based formulation. The focused image $f$ and blur parameter $\sigma_1$ are estimated jointly by minimizing reconstruction error $E$:
$$E = \iint \left( g_1(x, y) - h_{\sigma_1} * f(x, y) \right)^2 dx dy + \iint \left( g_2(x, y) - h_{\sigma_1 \frac{D_2}{D_1}} * f(x, y) \right)^2 dx dy$$
Setting partial derivatives with respect to parameters to zero ($\frac{\partial E}{\partial \sigma_1} = 0, \frac{\partial E}{\partial f} = 0$) provides stable iterative solutions resilient to image noise.
4.3 Real-Time (Video-Rate) DFD System Architecture (Nayar 1996)
Nayar designed a dual-sensor optical setup utilizing a beam-splitter prism behind a single lens to split light onto two CCD sensors placed at different optical path lengths.
flowchart LR
Scene["Scene"] --> Lens["Single Lens"]
Lens --> BeamSplitter["Prism / Beam-Splitter"]
BeamSplitter --> CCD1["CCD1 (Near Focused Image)"]
BeamSplitter --> CCD2["CCD2 (Far Focused Image)"]
style Scene fill:#1a1a2e,stroke:#e94560,color:#fff
style Lens fill:#16213e,stroke:#0f3460,color:#fff
style BeamSplitter fill:#533483,stroke:#e94560,color:#fff
style CCD1 fill:#0f3460,stroke:#e94560,color:#fff
style CCD2 fill:#0f3460,stroke:#e94560,color:#fff
Simultaneous acquisition of near-focused and far-focused views enables real-time 3D depth map computation at 30 FPS.
4.4 Active Illumination for Textureless Surfaces
Since DFF and DFD rely on high-frequency surface detail, smooth textureless surfaces (such as uniform white walls) lack signal. Projecting a high-frequency artificial contrast pattern (active illumination mask) onto the scene provides synthetic texture, enabling real-time depth acquisition even on smooth or moving objects.
5. Technical Comparison Summary
| Feature / Method | Depth from Focus (DFF) | Depth from Defocus (DFD) |
|---|---|---|
| Number of Images | Large Focal Stack ($10 \sim 100$ images) | Minimal ($2 \sim 3$ images) |
| Mathematical Approach | Local Modified Laplacian ($\nabla_M^2$) & 3-point Gaussian Interpolation | Fourier PSF ratios or iterative reconstruction optimization |
| Depth Resolution | Extremely High (Microscopic precision) | Moderate-High (Ideal for video frame rates) |
| Computation Time | High (Processes full focal stack) | Low (Analyses relative difference between 2 images) |
| Texture Requirement | Essential (Fails on textureless regions) | Essential (Resolved via Active Illumination Pattern) |
| Hardware Setup | Motorized focal translation stage | Beam-splitter dual-sensor camera |
| Primary Applications | Microscopy, industrial quality control, medical imaging | Mobile cameras, consumer vision, real-time tracking |