Focus Offset Prediction for Whole Slide Imaging Systems Using Deep Learning

Digital microscopy depends on each image being brought into accurate focus, and in whole-slide imaging this must hold across a whole specimen at high magnification and high throughput. Conventional autofocus acquires several images at different focal heights to determine the sharpest plane, whereas single-shot autofocus trains a deep regression model to predict the distance of the current plane to the sharpest plane from a single image. This thesis studies that prediction task directly, on existing slide images rather than in a running instrument, and asks how such a model degrades out of domain, that is on a dataset it was not trained on, and whether a change of scanner and a change of stain contribute differently, a question that has not been characterized.
The principal experiment is a three-factor factorial design in which 180 models were trained. The experiment varied the input representation, the training augmentation, and the training source, evaluated each model both in domain and out of domain, and analyzed the results with linear mixed models that separate the scanner and stain contributions to the error. The input representation was either the RGB image, the unmodified color image that serves as the baseline, a single-channel reduction (gray luminance or the green channel), or a transform of the gray channel into Fourier or wavelet space drawn from prior work. The training augmentation differed in whether photometric augmentation, the random perturbation of brightness, contrast, and color, was applied during training. Further analyses examined the physical origin of the focus signal and whether synthesized defocus or added scanner diversity narrows the out-of-domain gap.
In-domain error was sub-micron, 0.46 µm for the baseline model, but rose to between 4 and 6.3 µm out of domain on this benchmark, roughly an order of magnitude higher. The failure affected the direction of the prediction and not only its magnitude. The predicted sign disagreed with the labeled defocus on about four of ten out-of-focus tiles out of domain, against three of one
hundred in domain. The additional out-of-domain error was associated with a change of scanner, about +2.0 µm on this benchmark. The alternative input representations reduced out-of-domain error only when photometric augmentation was absent, and once it was present, the RGB baseline matched them. Among the further analyses, synthesized defocus reduced the out-of-domain error on one of two target datasets but not the other, and training jointly on two or three scanners did not measurably improve out-of-domain accuracy over the strongest single source.
The transform benefits reported in prior work reflect this missing-augmentation regime rather than a domain-transfer mechanism, a distinction earlier in-domain evaluations could not draw. In-domain accuracy therefore overstates deployable reliability, and because the failure is directional and associated with the scanner change, it is best addressed on the acquisition side rather than through the input representation alone.
Publications
-
Investigation of Factors Contributing to Domain Generalization in Single-Shot Autofocus