Abstract
In the diagnosis of chronic kidney disease, glomerulus as the blood filter provides important information for an accurate disease diagnosis. Thus automatic localization of the glomeruli is the necessary groundwork for future auxiliary kidney disease diagnosis, such as glomerular classification and area measurement. In this paper, we propose an efficient glomerular object locator in kidney whole slide image(WSI) based on proposal-free network and dynamic scale evaluation method. In the training phase, we construct an intensive proposal-free network which can learn efficiently the fine-grained features of the glomerulus. In the evaluation phase, a dynamic scale evaluation method is utilized to help the well-trained model find the most appropriate evaluation scale for each high-resolution WSI. We collect and digitalize 1204 renal biopsy microscope slides containing more than 41000 annotated glomeruli, which is the largest number of dataset to our best knowledge. We validate the each component of the proposed locator via the ablation study. Experimental results confirm that the proposed locator outperforms recently proposed approaches and pathologists by comparing
Introduction
Chronic kidney disease (CKD) is much more widespread than people realize. It has the characteristic of substantial morbidity and mortality, but low awareness and prevention rate. Glomeruli are spherical clusters of capillaries (The examples of glomeruli in pathological images were shown in Fig. 1) which are responsible for expelling unnecessary substances of human body. An initial solution to diagnose CKD is that pathologists need to observe glomerular injury in the renal biopsy microscope slides by optical microscope [30]. Unfortunately, there are few qualified pathologists for diagnosing such diseases because of nonstandard training system, low wages, and high risk [17]. In addition, the statistical work of glomeruli in the pathological biopsy slides is extremely complicated and the location of the glomerulus is flexible, which results that the diagnostic process is extremely laborious and highly depend on experience of pathologists. To reduce the risk of misdiagnosis, pathologists have to conduct a thorough inspection of the whole biopsy slide which makes the diagnosis quite cumbersome [10]. Therefore, it is urgent to design automatic technology in CKD diagnosis.

Glomeruli are labeled through red boxes in pathological images.
The development of digital pathology allows pathologists to digitalize a biopsy slide to a high-resolution whole slide image (WSI), a new form of high-resolution image storage, for diagnosis [6] based on image data obtained by scanning pathological tissue samples. Computer aided diagnostics in digital pathology can not only alleviate pathologists’ workloads, but also help to reduce the diagnostic mistakes. Automatic analysis on the WSI has become the mainstream in digital pathology and achieved some promising results. To date, many researches are conducted to implement classification, segmentation and recognition of the tumor [3,12] based on the WSIs data. The implementation of these technologies in the clinical diagnosis and their impact on medical field attracted more attention [4]. For many CKDs, the glomerulus is the most important region of interest and thus localization of glomerulus is the significant groundwork to future automated glomerular analysis such as classification and area measurement [7]. However in the localization of glomeruli, the critical limitations are also obvious when adopting the current machine learning and deep learning algorithms directly: Firstly, high-resolution WSIs which have too many pixels are difficult to be processed by current algorithms directly. If WSIs are down-sampled to the size that computers can process, a lot of fine-grained information lost will affect the localization results. Secondly, existing glomerular locators adopt proposal-based network, such as Faster R-CNN [21], to ensure the accuracy. As a consequence, in the proposal phase, tremendous of proposals are selected due to the high-resolution of WSIs, which leads to the proposal-based network involving a huge number of parameters to learn.
To overcome the aforementioned defects, we propose an efficient glomerular object locator (GOL) by using intensive proposal-free network and dynamic scale evaluation method for 400× WSIs. Motivated by the previous works of [20,29], we design an intensive proposal-free by integrating a more powerful loss function, Distance-IoU loss function [29], into the proposal-free network. Compared with the proposal-based network, such as Faster R-CNN [21], the proposed intensive proposal-free network is not only more free to input image sizes, but also faster to process images. In addition, we employ the dynamic scale evaluation to process high-resolution images in the evaluation stage, which can effectively improve the speed and accuracy of processing high-resolution WSIs through optimal scale evaluation. To evaluate the performance of the proposed GOL model, a dataset including more than 41000 glomeruli annotated from 1204 human renal whole slide images are collected and used to conduct extensive experiments. Experimental results confirm that the proposed GOL locator achieves competing overall F1-score of
The reminder of this paper is structured as follows. Section 2 sketches the related works. Section 3 introduces the proposed network structure and training process of the proposed method in detail. Section 4 gives the experimental results and Section 5 draws the conclusion.
In recent years, a series of approaches based on computer vision technology have been proposed to process renal pathological images by scanning renal biopsy microscope slides to WSIs through digital slide scanner. One of the most fundamental steps is to localize the region of interest in the WSIs, such as glomeruli, renal tubules and interstitial cells. In this section, we briefly introduce the glomerular localization methods in terms of traditional methods and deep learning based methods.
Traditional methods
Traditional methods usually extract hand-crafted features by machine learning algorithms. For example, Kakimoto et al. [13] proposed a method of localization glomeruli based on the support vector machine (SVM) [26]. For this method, the candidate sub-image is first cropped by sliding a fixed detection window on the renal WSI. Then the HOG feature vector [5] is extracted from the sub-image and used as the pre-trained SVM classifier to predict whether a glomerulus exists in the center of the cropped sub-image. Afterward, Kato et al. [14] developed an improved HOG features extractor named as segmental HOG that adaptively fit to input images and acquired robustness for the localization of the glomeruli. Recently, Simon et al. [25] developed an automated detector of glomeruli. The authors utilized local binary pattern (LBP) [28] as feature descriptors and calculate multi-radial color LBP, which is used to train a SVM classifier. However, renal biopsy microscope slides are actually the tangential 2D sectioning levels of the spherical 3D structure. Different tangent planes lead to different glomerular structures. Moreover, due to the different staining methods, the appearance of tissues are also diverse. Manually defined features are difficult to capture the changes in glomerular morphology and color, so it do not have good generalization to locate heterogeneous glomeruli.
Deep learning based methods
With the advance of deep neural networks, deep learning-based models are proposed and achieve promising performance in image processing field, which can learn deeper and richer representations than hand-crafted representations. By utilizing the automatic feature extraction method, many deep neural networks, including SPPnet [11], You Only Look Once (YOLO) [18] and its improved versions [19,20], Faster R-CNN [21] and other deep learning methods have been proposed for object detection. Some previous studies proposed the glomerular locators via adopting deep learning methods. Gallego et al. [8] proposed a method of slicing each WSI to many small patches by using a fixed-size sliding window. Then they classified the patches which were glomeruli and background by convolution neural networks(CNN). Kawazoe et al. [15] firstly used object localization model, Faster R-CNN, to detect the glomeruli in WSIs with multiple stains of human renal biopsy slices, which improves performance significantly. Liu et al. [16] improved the Faster R-CNN algorithm to achieve better results. Bukowy et al. [2] constructed Faster R-CNN with modified AlexNet model to detect glomeruli in patches. Previous works solve the problem of high-resolution processing through slicing each WSI into many small patches by sliding window method and then patch-wise searching the location of glomeruli [1,2,15,16]. In this way, nevertheless, it also takes a long time in processing a high-resolution WSI for invoking the locator frequently. Therefore, all recently developed methods have not yet reached the desired performance in glomerular localization and have high time consumption.
Material and method
In this section, we first introduce the dataset of glomeruli, then present the whole architecture of our proposed glomerular object locator (GOL) based on two important design principles.
Dataset description
Firstly, the 1250 renal biopsy samples we used were used to the periodic acid silver methenamine(PASM) stains used for staining structure containing more details of mesangial and basement membranes. All samples were from the Second Hospital of Shanxi Medical University (SHSXMU) between 2014 and 2019 and from the Shanxi Provincial People’s Hospital (SXPPH) between 2017 and 2019. The personal information of patients involved in the data was removed. Then, the biopsy microscope slides were digitized to WSIs by using a KF-PRO-005-EX digital slide scanner (KFbio, Ningbo, China) with a
Training dataset
To ensure the universally applicability of training model, 955 slides from two hospitals and different years were selected as the train/validation dataset. The used

Process of generating training dataset.
For pathological examination of renal biopsy, different hospitals have different staining methods even with the same staining type. Meanwhile, the stained slides would fade over time. Figure 3 shows some slides images collected from different hospitals and years. We can observe that these slides have different colors with each other. Therefore, it is necessary for us to evaluate the performance of our model in different slides’ appearance.

Examples of slides from different hospitals and years.
Overall, 250 slides were evenly selected from 2016 to 2018 in SHSXMU and SXPPH as the testing dataset to evaluate the performances of the models. Then we divided the testing dataset into five parts by different years and hospitals. Each testing dataset had 50 randomly selected WSIs. The testing dataset collected from different hospitals and years are summarized in Table 1.
Testing dataset from different hospitals and years
The overall framework of glomerular object locator is shown in Fig. 4. It consists of two parts involving the training phase and the testing phase. In the training phase, we employed one-stage localization network to learn glomerular features efficiently via combining multi-scale and proposal-free localization principle with more efficient Distance-IoU loss function. In order to keep the compression scale of WSIs,in the testing phase, same with that of patches in the training dataset, we proposed a dynamic scale evaluation (DSE) method to search the most appropriate compression scale for each WSI in the testing dataset before being fed into the well trained module. The used designed principle in each part will be introduced in the following.

Overall framework of glomerular object localizer.
Existing deep learning based glomerular object localization methods usually rely on region-proposal-network (RPN), such as Faster R-CNN [15], to generate relatively region of interests. These proposal-based networks have some inevitable limitations in processing the WSIs. Firstly, the RPN network has low localization speed in processing WSIs with high-resolution. Secondly, the proposal-based methods are difficult to converge successfully without pre-trained models [24]. However, there are few public datasets about renal pathology up to now, and the pre-trained process cannot be applied for renal pathology WSIs which have essential difference with origin images. Contrarily, proposal-free object localization networks, such as YOLO [18], handle the object localization problem with a regression analysis. This one-stage method (without RPN) is faster and more flexible to process WSIs with high-resolution, fine-grained features from scratch.
Various proposal-free localization frameworks predict the bounding box of locating object through
To overcome these limitations, a more powerful loss function named as Distance-IoU (DIoU) [29] is proposed and its expression is formulated as
According to the above analysis, we constructed a three-scale proposal training framework from the YOLO-v3 [20] and replaced the loss function and non-maximum suppression (NMS) by the powerful DIoU loss function and DIoU NMS, respectively. It increases the aggregate localization speed and thus it is effective for improving the regression accuracy in glomerular localization. The intensive proposal-free of GOL is detailed in Fig. 5.

Intensive proposal-free module of GOL.
Due to the limitation of computing power, existing object localization methods can not process directly the high-resolution and fine-grained WSIs. Usually, the WSI with gigapixel are compressed by down-sampling before feeding into well-trained model. It is important to set an appropriate down-sampled coefficient to obtain high performance in localizing accuracy. In other word, excessive and inadequate down-sampling both make the ability of localizing the glomeruli unstable. Meanwhile, it is difficult to find a suitable down-sampled coefficient for all WSIs with different size. Alternatively, some methods slice the large WSI into many small patches by sliding windows to train and test the modules. However, these methods are inefficient since they have to invoke the locator many times to process each WSI. In addition, due to the slicing method may cut out incomplete glomeruli (see Fig. 6(a)), the locator would deem them as morphologically incomplete glomeruli (see Fig. 6(b)), resulting that the locator has poor performance.

Examples of incomplete glomeruli. (a) Incomplete glomerular caused by slicing; (b) Morphologically incomplete glomerular.
To overcome the aforementioned defects in evaluation phase, we proposed a dynamic scale evaluation (DSE) method, which can make the tested WSI have consistent compressed scale with the training process. We first calculate the compress scale of width and height for training image by
The proedure of DSE is illustrated in Fig. 7. For each tested WSI, we first compresses it by the DSE method, and obtain a compressed image. Then, the compressed image is fed into the locator to infer the location of glomeruli. Finally, the compressed image with extracted bounding box by locator are mapped to original WSI through multiplying the corresponding down-sampled coefficient. The DSE method ensures that the object localization scale of the model in the testing phase keeps consistent with that of learning in the training phase. Therefore, the DSE method is particularly effective in localizing object from images with variable size and high-resolution such as medical WSI and high-resolution satellite imagery.

Procedure of dynamic scale evaluation.
In this section, the proposed GOL is evaluated on the testing dataset and compared with related locators and pathologists in performance of localization accuracy and speed.
Experimental settings
We constructed the training module of GOL based on a Linux serve platform with 4 NVIDIA Tesla V100s. We down-sampled the
Evaluation metrics
In our experiments, we employ the well-known metrics including
Ablation study
In this subsection, we ablate various design choices for our proposed GOL framework.
Effectiveness of powerful loss function. We trained two proposal-free models by using the same parameters setting excepting the selection of loss function to compare the performance. In the training phase, the loss curves of the two models during training are shown in Fig. 8. We can observe that the DIoU accelerates the convergence of training, and decreases the loss value by 0.6 than that of original proposal-free network.In the testing phase, the values of three metrics are listed in Table 2. We can see that the proposal-free network with DIoU achieves 15% higher
Effectiveness of DSE. We compare the cases with DSE and the cases without DSE (fixed input image size) in the testing phase. As shown in Table 2, the cases with DSE achieves 14.3% and 17.4% percent higher

Comparison of MSE and DIoU loss function curves.
Ablation study
To visually compare the localization results, we display the localization results by original proposal-free work, proposal-free network using dynamic scale evaluation (proposal-free+DSE), proposal-free network using DIoU loss function and DIoU non-maximum suppression (intensive proposal-free) and ours GOL.The direct comparison of are shown in Fig. 9. From the last column of Fig. 9, we observe that our proposed GOL achieves the best localization performance via merging DIoU and DSE.

Comperison of the localization results. (a) original proposal-free work. (b) proposal-free network using dynamic scale evaluation (proposal-free+DSE). (c) proposal-free network using DIoU loss function and DIoU non-maximum suppression (intensive proposal-free). (d) GOL.
To show an universal applicability of the model for slices from different hospitals and in different years, GOL has been evaluated in five testing datasets. The values of three metrics for the proposed model are shown in Table 3. We can see that
Performance on five testing datasets from different hospitals and years
Performance on five testing datasets from different hospitals and years
In this section, we compared the performance of GOL with recently published approaches including traditional and deep learning methods. The information of dastasets from previous works is shown in Table 4. All recently published approaches down-sampled WSIs to at least
Dataset comparison with related works
Dataset comparison with related works
Performance comparison with related works
We compared the GOL with two pathologists in term of localization accuracy and time consumption per glomerular. Firstly, 250 slides in the testing dataset are tested by a primary pathologist with two years of working experience and a senior pathologist with ten years of working experience. Localization results from pathologists are provided by the Shanxi Provincial People’s Hospital. The comparison results of average speed in each glomerulus and performance are given in Table 6. The GOL achieves higher performance in
Performance comparison with pathologists
Performance comparison with pathologists

Examples of glomeruli easy to be missed.
In this work, we proposed an efficient glomerular object locator named as GOL by combining intensive proposal-free network and DSE method. The intensive proposal network and the DSE method can effectively improve the speed and accuracy of processing high-resolution WSIs. We constructed a large dataset of 1204 high-quality renal pathological WSIs with PASM stain containing 41440 glomeruli. Experimental results confirm that GOL has universal applicability for different WSIs from different hospitals and years and achieves higher F-measures of 0.981 in glomerular localization compared with recently proposed approaches and pathologists. In addition, GOL locates and slices all glomeruli from each WSI only in average time of 8.58 seconds, and one glomerulus in an average time of 0.374 seconds. Thus, the proposed method can be embedded into the intelligent auxiliary diagnosis system to assist pathologists in clinical disease diagnosis. In the future, we would continue to implement the classification of glomerular based on the GOL and expand the GOL to easily deal with other stained pathology slices.
Footnotes
Acknowledgement
The authors sincerely acknowledge the doctors who offering great assistance in the Second Hospital of Shanxi Medical University and in the Shanxi Provincial People’s Hospital. We also sincerely acknowledge the editor and any reviewers for their valuable advice. This work was supported by the National Natural Science Foundation of China (Grant NO. 11472184, No. 11771321, No. 61901292); the National Youth Science Foundation of China (Grant NO.11401423); the ShanXi province plan project on Science and Technology of social Development (Grant No. 201703D321032); and the Natural Science Foundation of Shanxi Province, China (Grant No. 201901D211080).
