Abstract
Background
Osteoporosis is a major public health concern, especially among older adults, due to its association with an increased risk of fractures, particularly in the proximal femur. These fractures severely impact mobility and quality of life, leading to significant economic and health burdens.
Objective
This study aims to enhance bone density assessment in the proximal femur by addressing the limitations of conventional dual-energy X-ray absorptiometry through the integration of tomosynthesis with dual-energy applications and advanced segmentation models.
Methods and Materials
The imaging capability of a radiography/fluoroscopy system with dual-energy subtraction was evaluated. Two phantoms were included in this study: a tomosynthesis phantom (PH-56) was used to measure the quality of the tomosynthesis images, and a torso phantom (PH-4) was used to obtain proximal femur images. Quantification of bone images was achieved by optimizing the energy subtraction (ene-sub) and scale factors to isolate bone pixel values while nullifying soft tissue pixel values. Both the faster region-based convolutional neural network (Faster R-CNN) and U-Net were used to segment the proximal femoral region. The performance of these models was then evaluated using the intersection-over-union (IoU) metric with a torso phantom to ensure controlled conditions.
Results
The optimal ene-sub-factor ranged between 1.19 and 1.20, and a scale factor of around 0.1 was found to be suitable for detailed bone image observation. Regarding segmentation performance, a VGG19-based Faster R-CNN model achieved the highest mean IoU, outperforming the U-Net model (0.865 vs. 0.515, respectively).
Conclusions
These findings suggest that the integration of tomosynthesis with dual-energy applications significantly enhances the accuracy of bone density measurements in the proximal femur, and that the Faster R-CNN model provides superior segmentation performance, thereby offering a promising tool for bone density and osteoporosis management. Future research should focus on refining these models and validating their clinical applicability to improve patient outcomes.
Introduction
Background
Osteoporosis is a major public health concern, particularly among the aging population, because of its association with an increased risk of fractures. These fractures, known as fragility fractures, occur from minimal trauma that would not typically cause a fracture in healthy bones. Among the various types of fractures caused by osteoporosis, those involving the proximal femur are particularly problematic. Such fractures significantly impair a patient's ability to walk, often leading to prolonged periods of immobility or bedridden status. This reduces the quality of life (QOL) of patients, shortens their healthy life expectancy, and decreases their prognosis for life. Specifically, conservative treatment requires patients to remain in bed for long periods of time, which can result in the development of bed sores, urinary tract infections, pneumonia, deep vein thrombosis, and worsening of chronic conditions. The global burden of osteoporosis is rising as a result of demographic changes, with an increasing number of older individuals susceptible to this condition. Osteoporosis-related fractures are not only a medical issue, but also a significant economic burden because of the long-term care required for affected individuals. The effective management of osteoporosis and fracture prevention is therefore crucial for reducing health-care costs and improving patient outcomes. A critical factor in determining the risk of osteoporotic fractures is bone density, as areas with low bone density are particularly prone to fractures. Therefore, monitoring changes in bone density over time is an essential aspect of osteoporosis management. Regular bone density assessments allow for the early detection of osteoporosis, thereby enabling timely interventions that can help maintain bone strength, reduce the incidence of fractures, and subsequently enhance patient QOL.
Bone density can be measured using several methods, including dual-energy X-ray absorptiometry (DXA), quantitative computed tomography (QCT), radioabsorptiometry (RA), and quantitative ultrasound (QUS). Due to its accuracy, noninvasive nature, and ability to provide quantitative data on bone mineral density (BMD), DXA is the most frequently used method. However, DXA has limitations, particularly the partial volume effect when imaging the proximal femur, which can lead to overestimated bone strength and delayed interventions. Tomosynthesis, a form of three-dimensional (3D) imaging, has emerged as a promising alternative. Unlike conventional two-dimensional (2D) X-ray imaging, tomosynthesis captures multiple images of the body from different angles and reconstructs them into a 3D image. This method reduces the overlap of anatomical structures and improves the visualization of bone details, offering a more precise assessment of bone strength. Furthermore, the integration of dual-energy applications in tomosynthesis enhances its capability to differentiate between bone and soft tissue, minimizing the partial volume effect seen in DXA and quantifying bone density more accurately.
Motivation and objectives
This study is motivated by the need to improve assessments of bone density and fracture risk in the proximal femur, thereby enhancing the diagnosis and management of osteoporosis. This will ultimately lead to a healthier life expectancy and a higher QOL for patients. In this paper, we employ tomosynthesis, which has the advantage of low partial volume effect and small influence of metals in the body, and combine it with dual-energy subtraction technique to enable visualization of bone strength. Another important aspect of this study is the segmentation of the proximal femur. To obtain high analysis accuracy in bone strength visualization, it is necessary to extract the proximal femur as correctly as possible in tomosynthesis images. We aim to improve the overall accuracy of the analysis by using advanced segmentation models such as U-Net, which is often employed in medical image segmentation, and Faster R-CNN, which has performed well in previous studies in the object detection task.
Given this background, the present study aimed to develop and validate a method combining tomosynthesis with dual-energy applications and utilizing advanced convolutional neural network (CNN) models for the segmentation of the proximal femoral region, ultimately leading to improved patient outcomes in the management of osteoporosis.
Related works
Osteoporosis is a global health concern characterized by decreased bone density and an increased risk of fractures. The proximal femur is particularly susceptible to fragility fractures, which can significantly impair patient mobility and QOL. However, traditional methods for assessing bone density, such as DXA, are limited, and these limitations have prompted the exploration of more advanced imaging techniques and computational models.
Traditional bone density assessment methods
DXA is commonly used to measure bone density in both the lumbar spine and proximal femur in the diagnosis of osteoporosis. 1 Lumbar spine DXA typically measures L1–L4 or L2–L4 in the anterior-posterior direction, while proximal femur DXA measures the entire femur on either side. However, due to the overlap of the femoral head with the hip joint, it is more accurate to measure three regions: the neck of the femur, the adductor, and the proximal femur. If measuring all three regions is not possible because of equipment constraints, the bone density of the femoral neck should be measured. Bone density assessments are critical for both males and females, and especially useful for older patients, who may have spinal deformities that make lumbar spine measurements difficult. Low bone density and the occurrence of new fractures are highly correlated, making bone density a useful tool for fracture risk assessment, particularly in individuals aged > 65 years.2–4 Therefore, to predict fractures, it is important to visualize the bone strength of the proximal femur and display images in according to bone strength. And, images are more effective when visualized using color mapping.
On the other hand, common 2D DXA images are problematic in that the overlap of the femoral head and pelvic labrum results in a high bone strength value in that region.
In the following section, we introduce other bone evaluation methods, such as the microdensitometry (MD) method, which uses the second metacarpal to measure bone density from the shading of the radiographic image and the width of the bone cortex. 5 In contrast to conventional systems, the MD method does not require large equipment, can measure bone density values quickly, and is not affected by differences in film development conditions. However, the disadvantages of this method include a large error margin of 1.5% to 5% and its inability to evaluate cancellous bone, which tends to decrease as a result of osteoporosis.
QUS is a bone evaluation method that measures the velocity of ultrasound propagation and the attenuation coefficient in bone.6,7 It is low-cost, portable, and easy to use. 8 Because it does not use radiation, it is commonly used as a screening for osteoporosis in physical examinations and health checkups. However, it is not used for definitive diagnoses because of its large error margin of 3%–4% and its susceptibility to the effects of the size, temperature, and swelling of the heel.9,10
QCT, which uses a computed tomography (CT) system, allows a bone mass phantom to be scanned simultaneously with the lumbar spine and proximal femur to measure BMD in any part of the lumbar spine. In the case of DXA, because the scan is performed from a single direction, overlapping structures such as the ribs, vertebral arches, spinous processes, and aortic calcifications can interfere with the results. By contrast, QCT can exclude these overlapping tissues and focus the region of interest (ROI) solely on the measurement site. Therefore, it is possible to measure quantitatively true BMD (g/cm3) rather than planar density (g/cm2). However, radiation exposure is higher in QCT than in DXA.
Considering these factors, based on the significance of the DXA method to evaluate the correlation between quantification by dual-energy and BMD, this study quantifies bone strength using dual-energy technology.
Advancements in tomosynthesis
Tomosynthesis is an application of tomography systems known since the 1930s, and is an evolution of a technique using conventional X-ray systems to produce tomographic images. The X-ray tube and detector move synchronously on opposite sides of the patient, focusing on structures in the plane containing the fulcrum of motion (the tomographic center) to produce a clear tomographic image. Structures above and below this tomographic center (superimposed structures) are blurred by the motion of the X-ray tube and detector and are therefore less visible in the tomogram. Tomosynthesis improves on conventional tomography systems in that an arbitrary focal plane can be generated later from a series of projection data acquired during a single operation of the X-ray tube. 11 The shift-and-add (SAA) method has commonly been used as a reconstruction processing method for tomosynthesis. Figure 1 shows the principle of the tomosynthesis device and the SAA method. The SAA method is a technique for obtaining a specific cutting plane image by shifting each image by an appropriate amount in the scanning direction for a series of images and then superimposing the results. 12

Principle of tomography by simple addition (A) and reconstruction by shifted addition (B).
Subsequently, because tomosynthesis is one of the parallel-plane tomographic scanning that is considered a type of cone-beam CT scanning, arc correction was applied and a reconstruction method was introduced that extended the filtered back-projection (FBP) method, a typical method for CT reconstruction. This technique transforms a series of projection images obtained by tomosynthesis into the projection data of cone-beam CT scanning by applying a geometric transformation, and reconstructs a tomographic image of a specific cutting plane from these projection images. 12 In recent years, the iterative reconstruction method has made great strides. The iterative reconstruction approximation method compares the theoretical projection of a hypothetical numerical model of the specimen cross-section with the actual projected image and minimizes the discrepancy between the two by loop operations to obtain an image. 13 The above methods are used to reconstruct a specific plane. However, unlike CT, which collects data from 360°, tomosynthesis is an incomplete data reconstruction method. It should be noted that this makes it difficult to obtain information in the vertical direction as seen from the X-ray source, which in turn makes it difficult to reconstruct a complete 3D structure. Another important problem is that the slice thickness is determined by the swing angle of the arc motion.
Currently, digital tomosynthesis is most widely applied to mammography, called digital breast tomosynthesis, and considered useful for improving breast cancer detection and determining the indication for biopsy. 11 It has also been widely applied clinically in the thoracic, orthopedic, dental, otolaryngology, and emergency areas by using flat-panel detector-equipped general radiography and X-ray TV systems. 11 In orthopedics, digital tomosynthesis can depict complex anatomical structures and improve the detection of fractures because of its less tissue overlap compared with conventional 2D images.
CT is a high-resolution modality for bone lesions as well as tomosynthesis, which is of great clinical value because of its ability to provide a three-dimensional view of the structure through three-dimensional image processing. However, CT suffers from “metal artifacts”, which occur because the amount of photons reaching the detector is greatly reduced by metals with high X-rays absorption, making image reconstruction inaccurate. 14 These artifacts can cause image defects at metal-tissue boundaries or linear shadows throughout the image. These artifacts interfere with diagnosis and make it difficult to observe the tissue surrounding the metal and the metal-tissue junction, which in turn significantly affects the patient's QOL. Prostheses are particularly problematic in the proximal femur, a common site of fragility fractures.
Figure 2 shows a case of clavicle fracture. It can be seen that general X-ray images (Figure 2-A) cannot provide to the bony fusion area three-dimensional information, and coronal image for CT (Figure 2-B) can display the fracture, however, it is difficult to observe the bony fusion area due to metallic artifacts. On the other hand, tomosynthesis images (Figure 2-C) can provide information on the state of bone fusion without being affected by metallic artifacts in the fracture area. Therefore, tomosynthesis is also applied to patients with metal implants such as artificial joints and external fixators. 15 It is considered useful in the field of orthopedics because patients with metal implants cannot undergo magnetic resonance imaging (MRI) scans.

A case of clavicle fusion (A: General X-ray images, B: CT images, C: Tomosynthesis images).
Dual-Energy applications
Recently, the most relevant technique for bone densitometry is DXA, which was developed by replacing the 153Gd radionuclide source of conventional dual-photon absorptiometry with an X-ray tube. 16 The basic principle of DXA is to measure the transmission coefficient of high- and low-energy photons. 17 As the X-ray attenuation coefficient depends on the atomic number and photon energy, the areal density (mass per unit projected area) of two different types of tissue can be estimated by measuring the transmission coefficient at two different energies. DXA calculates the transmission coefficients of bone and soft tissue by measuring these absorption differences.
The X-ray TV system used in this study is equipped with a bone densitometry application that enables bone densitometry by DXA through scanning the lumbar spine and proximal femur. After diagnostic imaging and realignment by the fluoroscopy system, bone densitometry can also be performed in a series of examinations without moving the patient from the bed, which is expected to reduce patient burden and increase examination throughput. 18 This equipment also has a tomosynthesis. Therefore, the present study attempts to quantify bone strength using the tomosynthesis of this equipment. Additionally, due to its structure, bone differs in its resorption at the surface and in the bone marrow. As the usual BMD value is the sum of these values and local analysis is difficult, tomosynthesis is used in combination with tomography to evaluate it as a tomographic image.
Segmentation models in medical imaging
The most traditional segmentation techniques in digital image processing include region-based segmentation, 19 edge-based segmentation, 20 and clustering-based segmentation. 21 Each technique has its advantages and disadvantages, and the most appropriate image segmentation technique should be selected according to the specific purpose. 22 Various systems of network architectures have also been developed. CNNs, a type of deep artificial neural network, are widely used in the field of computer vision. They have been applied to many tasks, including image classification, super-resolution, and segmentation. 23 One highly accurate deep learning approach is the region-based convolutional neural network (R-CNN), which combines the proposal of rectangular regions with the capabilities of CNNs. R-CNN is a two-stage detection algorithm: in the first stage, it identifies a subset of regions in the image that may contain objects; in the second stage, it classifies the objects within each region. Following the advent of R-CNN, other variants, such as Fast R-CNN, 23 Faster R-CNN, 24 and Mask R-CNN, 25 have been developed. While R-CNN crops features during region proposal and then applies them to the CNN as a single image for classification, Faster R-CNN processes feature maps for the entire input image. This approach significantly reduces the computing time. Faster R-CNN is notable for being the first architecture in object detection capable of end-to-end learning by using region proposal networks (RPNs), which generate region proposals directly within the network. RPNs use anchor boxes for object detection, allowing the network to generate region proposals tailored to the data faster and more accurately. Mask R-CNN, which has a similar network structure to Faster R-CNN, performs semantic segmentation, going beyond object detection by making pixel-level decisions. Another widely used algorithm for object detection is You Only Look Once (YOLO). 26 In this method, the entire image is divided in advance by a grid so that the type and position of the object are simultaneously determined for each region. YOLO is characterized by its high processing speed, as it performs both detection and identification in a single scan of the image, rather than searching for all candidates individually within the image data. U-Net is another significant image segmentation technology developed primarily for segmentation tasks. 27 Since its introduction in 2015, its use in medical imaging has increased rapidly, leading to many developments in the U-Net architecture, such as new models that build upon the original U-Net design.28–30 These networks are also used in models equipped with BMD capabilities. The present study employs U-Net, which is frequently implemented in medical imaging, and Faster R-CNN, which was used in the base study previously considered. ResNet101 was also employed to perform a basic study of detection capability.
Comparative studies
Many papers aiming to detect proximal femur fractures have been published. Potter et al. developed a deep learning model by extending the Varifocal Net Feature Pyramid Network for detection and localization of proximal femur fractures from plain radiographs with clinically relevant metrics. 31 Their model attained 0.94 specificity and 0.95 sensitivity for fracture detection. Thus, there are many papers on detection with regard to fractures; these are reports of the usual 2D image detection with regard to fractures, not of quantified tomographic images. Bjornsson et al. proposed a femur segmentation method in CT to predict the risk of proximal femur fractures. 32 Their paper presents a deep neural network for the fully automated, accurate, and fast segmentation of the proximal femur from CT images. Their method has been shown to be suitable for hip fracture risk screening, bringing it one step closer to being a clinically viable option for screening at-risk patients with hip fracture susceptibility. The mean Dice Similarity Coefficient (DSC) achieved was 0.990 ± 0.008, and the mean Hausdorff distance was 0.999 ± 0.331 mm.
Deng et al. developed a 3D end-to-end fully CNN that can better combine information among neighboring slices to achieve more accurate segmentation results. 33 They used the separation of cortical and trabecular bones derived from the QCT software MIAF-Femur as the segmentation reference. Two models with the same network structures were trained, achieving DSCs of 97.82% and 96.53% for the periosteal and endosteal contours, respectively. The present paper describes the usefulness of QCT to ensure quantification and provide a detailed view of the bone in a 3D display. However, these previous studies propose segmentation methods from quantified CT images and do not address the issue of patients with artificial heads or other metallic devices. Quantification is an important process for the follow-up of bone density.
In the field of object detection, many research results using YOLO have also been published. For instance, Hu et al. used the YOLO algorithm for lesion detection in contrast-enhanced mammography, recording a 97% hit rate and 82% mean average precision. 34 Most studies on the segmentation of proximal femur fractures have been reported from CT images; few studies have been conducted on quantification and segmentation using dual-energy subtraction tomosynthesis images, as in the present study.
Materials and methods
The capability of the SONIALVISION Safire17 X-ray TV system with dual-energy subtraction (SHIMADZU, Kyoto, Japan) to obtain images was evaluated. Two phantoms were included in this study: a tomosynthesis phantom (PH-56) was used to measure the quality of tomosynthesis images (effective slice thickness), and a torso phantom (PH-4) was used to obtain proximal femur images.
Study design
This comprehensive study was designed to assess the segmentation and quantification of the proximal femur using dual-energy subtraction tomosynthesis, as illustrated in Figure 3, and divided into two parts: a basic study and a processing study.

The flowchart of this study consists of two parts: a basic study using the PH-56 tomosynthesis phantom and a processing study using the PH-4 CT Torso Phantom.
In the basic study, the PH-56 tomosynthesis phantom was used with the SONIALVISION Safire17 system to acquire the tomographic images, which were then utilized to measure slice thickness, ensuring the accuracy of the tomosynthesis imaging process.
In the processing study, the PH-4 CT torso phantom was employed to obtain proximal femur images using the dual-energy tomosynthesis technique. This imaging process generates high- and low-energy X-ray images of the proximal femur. The acquired images undergo energy subtraction and scaling, where energy subtraction functions (ESFs) and scale factors are estimated to enhance the separation of bone from soft tissue. The processed images are then segmented using a fully convolutional network (FCN) to identify the proximal femur region. A color map is superimposed on the segmentation target to assess the accuracy and quality of the segmentation process visually. The segmentation results are then evaluated for accuracy and necessary adjustments are made. The intersection-over-union (IoU) metric is used to evaluate the segmentation performance. This structured approach ensures a thorough analysis of proximal femur imaging and segmentation, leveraging advanced imaging and deep learning techniques to improve bone density assessments and fracture risk prediction.
Dataset and samples
Figure 4 shows the phantoms and imaging processes used in this study. For the basic study, Figure 4-A shows the PH-56 tomosynthesis phantom, which includes a case and two types of plates. Figure 4-B depicts an aluminum plate with a thickness of 0.5 mm and a 1-mm-diameter hole, which is necessary for measuring the thickness of the cross-section. For the processing study, Figure 4-C presents the PH-4 human torso phantom used for obtaining the proximal femur images. Figure 4-D and 4-E display tomosynthesis images of the proximal femur obtained using high- and low-energy X-rays, respectively. The reconstructed image size is 763 × 763 pixels with slice thicknesses of 1 and 2 mm, resulting in 201 and 101 reconstructed images, respectively. These tomosynthesis images of the proximal femur are used for building Faster R-CNN models and U-net.

Study phantoms and imaging processes used in the basic and processing studies (A: PH-56 tomosynthesis phantom, B: Measurement of section thickness, C: PH-4 CT Torso Phantom, D: High energy X-ray image, E: Low energy X-ray image).
X-ray scanner and development environments
In this study, we used the SONIALVISION Safire17, an X-ray TV system capable of dual-energy subtraction tomosynthesis imaging. Three types of subject phantoms were used: an acrylic phantom (ranging from 1 to 13 cm), a human body phantom, and a PH-56 tomosynthesis phantom NS manufactured by Kyoto Kagaku (Kyoto, Japan). The development environment included ImageJ 1.49c (NIH, USA), Python 3.10 (Python Software Foundation), and MATLAB 2023a (MathWorks, USA) for analysis. Computations were performed using an Intel Core i7-7500U (2.7 GHz) CPU (Intel Corporation, USA) with 16 GB DDR4 RAM and two GeForce RTX 2070 Super GPUs with 8 GB SGRAM each (NVIDIA Corporation, USA). Statistical calculations were performed using Microsoft Excel for Microsoft 365 MSO Ver.2206 64-bit (Microsoft Corporation, USA).
Quantification of tomographic images
The quantification of pixel values is calculated using a formula designed for determining bone density (Figure 5). However, as the purpose of this study was to quantify pixel values, the correlation with bone density was not considered. Figure 5 illustrates the DXA calculation method used in this study.

DXA calculation method and equations for energy subtraction factor and scale factor determination.
I0 denotes the number of incident X-ray photons, H denotes high energy, and L denotes low energy. μ is the mass attenuation coefficient and L being the thickness of the object. B and S stand for bone and soft tissue, respectively. For Figure 5, the logarithms on both sides are equations (1) and (2).
Estimation of pixel values in the phantom study
For the calculation of quantitative values according to Eqs. (1–3), it is necessary to know the pixel value when there is no subject present. However, in such situations, the pixel value may be saturated. To address this issue, imaging was performed by changing the acrylic phantom thickness from 8 cm to 17 cm at 3-cm intervals, and the pixel values at a subject thickness of 0 cm were calculated by approximation. Reconstruction methods for tomosynthesis include the SAA, FBP, and iterative reconstruction methods. In this study, we used the SAA method because of its ability to ensure the quantification of pixel values. The raw projection image dataset was used to create the experimental images, and the average pixel value in the center of the phantom was the target of measurement. The exposure conditions were set as follows: High energy is a tube voltage of 120 kV with a tube current of 384 mA and an exposure time of 1.6 ms, and Low energy is a tube voltage of 60 kV with a tube current of 272 mA and an exposure time of 20 ms. These exposure conditions are the default values set by the manufacturer. The swing angle is 40°.
Measurement of slice thickness with an aluminum plate
Tomosynthesis has relatively high spatial resolution and is effective for obtaining detailed bone information. However, to obtain accurate depth information, it is necessary to measure the thickness of the cross-section. For this measurement, an aluminum plate with holes (0.5 mm thick, inside diameter = 1 mm), an acrylic phantom (5 mm thick), and a tomosynthesis phantom were used. The phantom was placed with its long axis horizontal to the bed and then exposed. The exposure conditions were set to 47 kV and 1.25 mAs, with a swing angle of 40° and a reconstruction thickness of 0.5 mm. A total of 101 tomosynthesis slices were analyzed according to the instruction manual for the NS phantom. The SAA and FBP methods (F++/F+–/F––) were used for the reconstruction algorithm and function. The ROI was set at the center pixel of the reconstructed image, and the full width at half maximum (FWHM) was measured from the profile, with the fault height on the horizontal axis and the pixel value on the vertical axis.
Optimization of energy subtraction and scale factors
Energy subtraction images with a human phantom were used to optimize the energy subtraction ESF and scale factors. The imaging positioning of the human phantom was the hip joint and the centerline was centered on the femoral neck. The proximal femur of the human phantom was slightly internally rotated to provide the most extensive view of the femoral neck and irradiated under the conditions described in Section 3.5. The image reconstruction method used was the SAA method. The images were in raw data format, and the imaging height range was centered at 15.0 cm. Profile curves near the femoral neck in the tomosynthesis images obtained by varying the ESF and scale factor were drawn, and the average value of the soft tissue area was defined as the reference value. For ESF, the scale factor was fixed at 0.1 and the values for soft tissue were optimized and examined from the profile curves. The scale factor was varied based on the optimized ESF value and the contrast of the image was evaluated. In this experiment, the value of each pixel is referred to as the gray value, as it represents the numerical data for bone quantification.
Training R-CNN and FCN segmentation models
Image segmentation includes methods that analyze graph cuts, concentration histograms, classifier-based models, CNN-based models, and combinations of these approaches.

Workflow of training the FCN model for proximal femur segmentation.
The workflow of training the FCN model for proximal femur segmentation is depicted in Figure 6. The process begins with quantified tomosynthesis images of the proximal femur, which serve as the initial input data. These images undergo image augmentation to generate multiple variations, thereby increasing the diversity and size of the training dataset and enhancing the robustness of the model. Next, manual setup of the ground truth is performed by defining the ROIs in the augmented images, with the accurate segmentation areas marked with yellow borders. Using the manually annotated ground truth data, the FCN models are trained. The FCN architecture allows for end-to-end learning, which enables the model to predict segmentation maps directly from the input images. Once the model is trained, it is applied to segment the proximal femur in new images, with the detected ROIs highlighted by yellow borders. Finally, a color map is superimposed on the target segmentation areas to visualize the accuracy and quality of the segmentation, aiding in the assessment of the overlap between the predicted and actual segmentation regions. This systematic workflow ensures accurate and efficient segmentation of the proximal femur in tomosynthesis images.
Among these methods, CNN-based methods are highly robust because they can perform end-to-end learning by directly extracting image features, enabling them to recognize objects regardless of their location in the image. In this study, Faster R-CNN, which is known for its high accuracy, extraction speed, and efficiency among CNN-based methods, was used. Faster R-CNN consists of a CNN and a feature extraction network, followed by two sub-networks. One sub-network is an RPN that generates object proposals, and the other predicts the class of the object. The RPN outputs a series of regions with a probability score of being an object and the class/label of the object. An important concept in this process is the anchor box, which is used in the initial prediction of the object position in the RPN.
This study attempted to validate four pretrained networks: AlexNet, VGG19, ResNet50, and ResNet101. These networks have demonstrated excellent performance in both image classification and object detection tasks in previous studies, and we chose them because we determined that they would yield stable results in this research as well. 35 AlexNet is an eight-layer network with five convolutional layers and three fully connected layers, comprising 61.0 million parameters and an image input size of 227 × 227. VGG19 has 16 convolutional layers and three fully connected layers, with 144 million parameters and an image input size of 224 × 224. ResNet50 and ResNet101 have many residual blocks connected in a series, characterized by shortcut connections that prevent gradient loss. They have 25.6 and 44.6 million parameters, respectively, and both have an image input size of 224 × 224. Segmentation was also performed using U-Net as a comparison to Faster R-CNN. U-Net consists of an FCN where all coupling layers are replaced by convolution layers. More precise predictions can be achieved by incorporating skip connections that concatenate the encoder's feature map to the decoder's feature map. In this study, the segmentation model was created using weights initialized from the VGG16 network. Accurate and reliable ground truth data are a major factor in terms of model performance. For the ground truth dataset, a radiological technologist set the ROIs and created the ground truth data for segmentation.
For training, the pretrained networks AlexNet, VGG19, ResNet50, ResNet101, and U-Net were used. AlexNet, VGG19, ResNet50, and ResNet101 consist of 8, 19, 50, and 101 convolutional layers, respectively, with an input size of 227 × 227 for AlexNet and 224 × 224 for the others. In this study, the batch size was 4, the number of epochs was 10, and the initial learning rate was 1e–3 based on the results of previous studies and GPU constraints. 35 U-Net is an FCN architecture with many successful results in segmentation tasks. For U-Net, the input size was 256 × 256, the batch size was 5 based on the results of previous studies, 36 the number of epochs was 20, and the initial learning rate was 1e–3.
Evaluation of segmentation performance
Although segmentation models have several evaluation indices, such as accuracy, precision, and recall, the evaluation of the overlap rate is crucial for validation, so the IoU metric was adopted. The IoU metric indicates the percentage of overlap between the ground truth and predicted areas and is calculated by Eq. (4). A higher IoU value means a larger percentage of overlapping area, with IoU = 1 indicating perfect agreement. It is generally a more stringent indicator than other evaluation metrics and often used in evaluating the performance of segmentation models. Figure 7 shows an example of IoU output. The white rectangle represents the set ROI, and the yellow rectangle represents the detected ROI. In the case of Figure 7, the IoU is 0.72927. The test dataset evaluated in this process consisted of 100 new images.

An example of a test image when using the FCN model. The two rectangles indicate the ground truth region and the detective region, IoU = 0.72927 in this example.
Results
Quantification of tomographic images
Estimation of pixel values in the phantom study
Because pixel values saturate without a subject, we estimated the pixel values for a subject thickness of 0 cm by varying the acrylic thickness. The relationship between subject thickness and pixel values was determined experimentally. Given that pixel values are logarithmically converted, we used an exponential function to obtain the estimated luminance value at a subject thickness of 0 cm. The resulting values were 48,459.7 for the high-voltage condition and 41,718.8 for the low-voltage condition, as measured from the raw image data.
Measurement of section thickness with an aluminum plate
Figure 8 shows the relationship between the height of the scanning center and the pixel value obtained from the experiment, along with the FWHM calculated as the section thickness. The experimental data reveal how the pixel value varied with the height of the slice, and the FWHM was determined from this relationship to estimate the slice thickness.

Relation between the scanning center and pixel values.
When a perforated aluminum plate (0.5 mm thick, Φ 1 mm) combined with acrylic was used, the FWHM (considered as the section thickness) was found to be 5.5 mm at reconstruction interval of 5 mm when reconstructed using the SAA method. For the FBP method, the FWHM values were 14.5, 8.5, and 5 mm for the F++, F+–, and F–– settings, respectively.
Estimation of energy subtraction and scale factors
Figure 9 shows the image of a hip joint on a human phantom using dual-energy subtraction tomosynthesis. Gray values along the yellow line were calculated based on the results obtained in Sections 3.1.1 and 3.1.2. The coefficients between the ESF and scale factor were estimated so that the average gray value of the soft tissue was closest to zero. Profiles were obtained by varying the ESFs from 1.17 to 1.20 in 0.01 intervals. When the ESF was 1.17, the average gray value of the soft tissue was 0.722. At ESFs of 1.18, 1.19, and 1.20, the values were 0.427, 0.131, and −0.165, respectively. As the ESF increased from 1.17, the soft tissue gray value approached zero, and when it reached 1.20, became a negative number.

Profile on the line (when ESF is varied from 1.17 to 1.20).
Performance of segmentation models between R-CNN and FCN
Table 1 presents the mean IoU values along with their respective standard deviations (SDs) for the different network models used in the study (i.e., AlexNet, VGG19, ResNet50, ResNet101, and U-Net), each of which was evaluated across five configurations (Model 1 to Model 5). AlexNet exhibited mean IoU values ranging from 0.754 to 0.850, with relatively low SDs, indicating consistent performance across different models.
Mean intersection-over-union and standard deviation (mean IoU ± SD) for various network models.
Note. The bold type indicates the highest value.
VGG19 achieved the highest mean IoU (0.865 ± 0.062), demonstrating superior performance among the tested networks. ResNet50 and ResNet101 showed varied results, with ResNet101 achieving a mean IoU up to 0.819 ± 0.077. By contrast, U-Net had the lowest mean IoU values, ranging from 0.428 ± 0.203 to 0.515 ± 0.113, indicating room for improvement in its segmentation accuracy compared with the other networks. Overall, VGG19 stood out as the best-performing model in terms of the mean IoU, while U-Net lagged behind the other networks. The variations in IoU values and SDs highlight the differences in robustness and reliability among the network models.
Figure 10 illustrates the IoU per slice for FCN with the VGG19 backbone and U-Net models, both evaluated using Model 1. Figure 10-A displays the IoU values for VGG19, indicating consistently high performance across slices, with IoU values predominantly above 0.8. Notably, certain slices (highlighted in red) indicate images needed for diagnosis, demonstrating their critical importance. Figure 10-B presents the IoU values for U-Net, which exhibited more variability compared with VGG19. The IoU values for U-Net ranged from 0.1 to 0.8, with the diagnostic slices (highlighted in red) showing higher IoU values around 0.7. This comparison underscores the superior and more stable performance of the VGG19 model in segmenting the slices, while the U-Net model showed greater fluctuations and lower overall accuracy.

Iou per height of image slice by VGG19 model (A) and U-Net model (B).
Discussion
Quantification of tomographic images
The experimental results for the quantification of bone areas in tomosynthesis indicated that a reconstruction interval of 5 mm was appropriate based on the measured fault thickness using the SAA method. From results for Fig.9, regarding the ESF, which indicates the attenuation ratio of soft tissue, the gray value tended to decrease overall as the factor increased. The closer the gray value of soft tissue is to zero, the smaller the influence of soft tissue becomes, allowing for a purer evaluation focused solely on the bone. Therefore, it was determined that the optimal value for ESF is between 1.19 and 1.20. The scale factor affects image range. The larger the value of the scale factor, the closer the gray value of the soft tissue approached zero, but the difference between the gray values of bone and soft tissue became smaller. That is not suitable for visualization. Therefore, a scale factor of around 0.1 was considered suitable. Furthermore, the dual-energy subtraction function of the equipment used in this study did not employ slit imaging, but rather, normal irradiation field imaging. As a result, the measured values may have been affected by scattered rays.
Segmentation models
In the validation of the segmentation, the recognition of the proximal femur by Faster R-CNN showed that VGG19 had the highest accuracy among the four networks, with an IoU of 0.865. Notably, the average IoU was as high as 0.891 in the cross-sections required for diagnosis. By contrast, the detection accuracy of U-Net ranged from 0.288 to 0.880 for the IoU of the slice images required for diagnosis, with a mean IoU of 0.630, indicating that Faster R-CNN was more accurate. However, there is a need to investigate fully the respective network parameters, which is a subject for future research. The structure of feature extraction using CNNs is sensitive to edge extraction, so the detection efficiency may vary depending on the setting of the scale factor. Particularly in tomosynthesis, images of slice planes using the SAA method often show that the pixel values of high-contrast areas before and after the target slice plane affect the pixel values of the target slice plane. This also needs to be considered when setting an appropriate reconstruction interval. The slice thickness of tomosynthesis using the SAA method theoretically depends on the swing angle of the tube. Therefore, a reconstruction interval of 5 mm is considered appropriate when the swing angle is 40°.
Furthermore, the Faster R-CNN used for segmentation in this study generates a feature map from the input image, where it sets anchor boxes, which are a set of predefined bounding boxes with different sizes and aspect ratios. Objects are detected in the image based on these anchor boxes. Both the number and placement of anchor boxes have a significant impact on the accuracy of object detection and should be set carefully. In particular, if anchor boxes extend beyond the image, detection accuracy in the surrounding area may decrease, so it is important to select the appropriate number and location of anchor boxes. Also, the size of the region of interest (ROI) to be set is also a future issue. This is because the ROI in tomosynthesis images changes from slice to slice and therefore changes in size. It is important to the relationship between the ROI and the anchor.
This study aimed to quantify the pixel values of the proximal femur by tomosynthesis and tomographic imaging with dual-energy application, and to perform slice image extraction segmentation for diagnostic purposes and fracture prediction. Tomosynthesis produces high-definition tomographic images, allowing for fine assessment of the bone cortex and trabecular bone, which is compatible with the evaluation of fragility fractures in older adults. Tomosynthesis can obtain arbitrary tomographic images, it is useful for reconstructing images aligned with fracture lines to assess postoperative bone formation, as well as for predicting and assessing fracture risks morphologically. Future clinical applications will be attempted based on the results of this study. An image with a color map superimposed on the original image is shown in Figure 11. This process makes it much easier to grasp the bone strength of the proximal femur visually.

The ROI selected by the FCN and superimposed back onto the original image. The stronger the bone strength, the more it is displayed in the red region.
Color map images were created by taking the segmented regions using Faster R-CNN, converting them to Jet color, creating an image with 50% transparency, and then superimposing it on the original image.
Limitations of this research and future work
In this study, it was essential to ensure the quality of the X-ray beam to achieve a certain level of image quality. However, the problem of line quality was not examined at this time, and the impact of scattered rays remains an issue related to tube voltage and the size of the irradiation field. A smaller irradiation field reduces the number of scattered rays, resulting in better image quality. To address these issues, a remaining challenge is to capture images with the narrowest possible beam. Additional limitations of this study include the limited sample size, which may not fully represent the variability found in human subjects, and the absence of clinical data from actual patients. Additionally, factors such as subject thickness and imaging conditions were limited to controlled settings. For future clinical applications, it is necessary to collect a wider variety of clinical images with diverse body types, image quality, and cropping sizes as training data. If such training data can be gathered, it is expected not only to facilitate the realization of clinical applications but also to ultimately improve the accuracy of fracture risk prediction. This remains a challenge for the future. Furthermore, although tomosynthesis reduces some issues associated with metal artifacts in CT and MRI, the impact of metal objects on image quality and segmentation accuracy was not fully explored. Moreover, parameter optimization for the segmentation models was not extensively conducted, and many studies are unsatisfactory, including input size, anchor size, batch size, epoch size, learn rate, and optimization function. In addition to the segmentation methods used in this study, there are many other segmentation methods such as YOLO and architectures derived from U-Net. Various hyperparameters may affect each other, and careful investigation is required. 37 The computational requirements for training and deployment of advanced CNN models such as Faster R-CNN and U-Net are large, affecting the feasibility of implementing these models in clinical environments with limited computational resources, but we expect that high accuracy can be achieved if optimal methods and hyperparameters are found.
Future research should focus on enhanced image quality control, including methods to improve line quality and reduce the impact of scattered rays. Incorporating clinical data from a diverse patient population would also enhance the validation of the proposed methods. Developing advanced algorithms to reduce metal artifacts and conducting extensive parameter optimization for CNN models could help construct the most effective segmentation models. Researching methods to reduce the computational load and improve processing speeds is considered essential for real-time implementation in clinical environments. Additionally, conducting longitudinal studies to monitor changes in bone density and structure over time using the proposed quantification methods would provide valuable insights into the progression of bone diseases and the effectiveness of treatments. By addressing these limitations and pursuing these future research directions, the proposed methods can be refined and validated further, ultimately enhancing their clinical applicability and improving patient outcomes.
Conclusion
Bone imaging in general radiography is not suitable for follow-up because the images vary according to image parameters and are not quantified. However, quantified images are useful for follow-up. Quantified bone images correlate with bone strength because there is a correlation between probability of fracture risk and bone density. 38 In CT and MRI examinations, the presence of metal in the body can reduce image quality, posing an obstacle to accurate examination. Therefore, the importance of general imaging in the orthopedics field and the improvement of its accuracy and quantified bone images are critical. This study demonstrated the possibility of image quantification using dual-energy and tomosynthesis functions. We also developed and compared Faster R-CNN and U-Net models for in-plane semantic segmentation of the diagnostic target area using CNNs.
In the future, it will be necessary to explore and evaluate various parameters to construct an optimal segmentation model. And clinical application requires a large amount of extensive clinical images. In addition to quantification data of bone strength, if we could collect reports of subsequent fractures, we would be closer to archive accurate fracture prediction. In conclusion, the proposed method has the potential to be utilized for predicting fractures of the proximal femur, potentially significantly positively impacting patient QOL.
Footnotes
Acknowledgments
The authors would like to thank Daisuke Notohara for the valuable advice and the radiological technician at Teikyo University Hospital for the technical assistance with the experiments. Finally, we are grateful to the referees for their insightful comments.
Informed consent statement
Not applicable.
Institutional review board statement
The Ethics Committee commented that the study was a phantom experiment and did not require Ethics Committee review.
Author contributions
Conceptualization, A-Matsushima, T.-B.C., and T-Okamoto; Methodology, T.-B.C. and T-Okamoto; Software, T.-B.C., S.-Y.H., and A-Matsushima; Validation, T-Okamoto, K-Kimura, M-Sato, T.-B.C., and S.-Y.H.; Formal analysis, K-Kimura, M-Sato, and A-Matsushima; Investigation, A-Matsushima, K-Kimura, and M-Sato; Data curation, A-Matsushima, T.-B.C., S.-Y.H., and T-Okamoto; Writing—original draft preparation, A-Matsushima, T.-B.C., and T-Okamoto; Writing—review and editing, T.-B.C. and T-Okamoto; Project administration, T.-B.C. and T-Okamoto; Funding acquisition, T-Okamoto. All authors have read and agreed to the published version of the manuscript.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability statement
The data are not publicly available because of privacy or ethical issues. The data presented in this study are available on reasonable request from the corresponding authors.
