Abstract
This paper explores the advancement of object detection models within the domain of satellite imagery analysis, focusing on the innovative application of synthetically generated datasets to enhance model performance. Motivated by the inherent challenges of manual dataset annotation, such as errors, limited variability, and geographical biases, this study employs synthetic data generation techniques to create a diverse dataset by overlaying 3D models of 31 different aircraft types onto satellite imagery, creating a dataset of 5000 images containing 27,375 aircraft. This dataset includes a variety of environmental conditions and image transforms aimed at training a more robust and generalizable object detection model. Using a pre-trained model based on the YOLO v8x architecture, an extensive comparison of fine-tuned models trained separately on traditional manually annotated datasets was compared with a fine-tuned model trained on the synthetic dataset. The results show the model trained with synthetic data achieves a performance within 2% of the other models, detecting actual aircraft in diverse circumstances, despite only being trained on synthetic data derived from 3D models. By providing a scalable, efficient, and accurate method for training object detection models, this research demonstrates the possibility of detecting previously unseen objects based only on a 3D model.
1. Introduction and background
The first use of satellites to observe our planet from space traces its roots back to the late 1950s with the Corona series of satellites. The Corona program set out on an ambitious task to monitor threats at the height of the Cold War and after 13 failed attempts, the first successful space-based photo-surveillance mission occurred in August 1960 by the Discover 14 satellite.1–3 A film canister was physically dropped from the satellite traveling 120 miles above the Earth at roughly 5 miles per second and was then subsequently intercepted mid-air by a C-199 aircraft. This single mission provided more coverage than all previous U-2 missions combined and has been described as a significant development in military technology.1,2
The ensuing years witnessed significant advancements in remote sensing technology. A pivotal evolution was the transition from physically dropping film canisters to the capability of satellites transmitting data in near-real-time, revolutionizing the speed at which imagery could be relayed back to Earth. The first civilian Earth observation satellite, ERTS-1 (later renamed Landsat-1), was launched in 1972. 4 Today, plummeting launch costs, sensor miniaturization, and commercial innovation have resulted in an unprecedented availability and public access to satellite data. The U.S. Space Force estimates over 8000 operational payloads are currently in orbit. 5
The incredible volume of satellite sensing data necessitates innovations in artificial intelligence and cloud computing to fully unlock its potential. 6 Amid these necessary advancements, the field of object detection stands at the frontline of the data revolution, presenting both opportunities and challenges. Despite leaps in computing technology and the ubiquitous availability of satellite data, certain tasks, such as automatically detecting and classifying aircraft, remain remarkably complex. This complexity could be considered synonymous with finding a needle in a haystack; however, it is closer to distinguishing between needles of slightly different shapes and sizes in a constantly shifting haystack under varying conditions of light, shadow, and camouflage.
One such critical challenge is the process of manually annotating satellite imagery with aircraft detections, to create datasets usable to train artificial intelligence algorithms. This meticulous task demands extensive human effort to label images by hand-drawing bounding boxes and manually classifying objects. While past research estimates the mode for time-to-annotate an object in an image is approximately 10 s with specialized image annotation software, 7 other issues cause more practical problems. We must first find the object of interest in images; for rare objects or objects we have not yet photographed from satellites, this may be impossible or time-consuming. This may also require specialized, expensive human subject matter expertise. To support the data requirements for deep learning training, we would also need to repeat this search and annotation work in numerous images, in different situations, resulting in the multiplication of the resources required for a single annotated image. In addition, human annotators, despite their best efforts, exhibit variations in how tightly or loosely they draw these bounding boxes around the same object across different images, and categorization can be subjective. Aircraft are mobile and can take on various shapes, sizes, and orientations. This introduces variability in the dataset, noisy labels and increases the potential for human error. Inconsistencies in labeling and categorization can cause further dilution and inflexibility in the training dataset.
Recent research highlights that certain mitigation techniques, such as consensus labeling and the application of multiple bounding boxes, have the potential to enhance the reliability and accuracy of image annotation. 8 These methods leverage collective agreement and varied perspectives to reduce the ambiguity inherent in manual labeling tasks. However, they come with their own set of challenges, notably the increased computational complexity and resource demands they impose on the already intensive process of image annotation. This escalation can strain resources, highlighting a critical trade-off between annotation quality and operational efficiency.
Advanced techniques, such as extracting features through symmetric line segmentation, offer promising avenues for improving detection rates—examples using these methods focus on identifying the unique geometric patterns of missile-launching equipment 9 or the wings of an aircraft. 10 However, their effectiveness is constrained by several factors, including the requirement for very high-resolution imagery, vulnerability to occlusion, and inflexibility with diverse aircraft types or surface-to-air missile (SAM) site configurations. Such challenges underscore the need for innovative approaches that balance accuracy, resource efficiency, and adaptability in the face of varying conditions and object types encountered in satellite imagery analysis.
Objects for which data is either sparse or has not been directly observed, such as developmental aircraft, present significant obstacles in training models with any degree of reliability. The scarcity of available data on these cutting-edge technologies makes it difficult to employ machine learning models that can accurately recognize and classify such objects in satellite imagery. This scenario poses substantial challenges for defense organizations that rely on timely and accurate detection of emerging threats to maintain security and strategic advantage. The rapid pace of technological advancements and the nature of military development programs further exacerbate these challenges, as adversaries continually evolve their assets to evade detection. Non-defense systems that employ object detection also face similar situations with sparse data. For example, consider a self-driving car that must consider the rare event of a large animal running in the path of the car or sunlight reflections partially blinding a sensor.
One promising area of research is the utilization of synthetically generated data to augment or replace training data.11–13 One study, for example, saw a 15% boost in the detection of whales by generating synthetic satellite imagery mimicking real-world conditions such as lighting, water color, and ocean wave patterns, among others. 14 Another study shows that photo-realistic data of 3D hand-object interaction can be synthetically generated and used to replace a real-world dataset without any significant degradation in model performance. 15 Similar results were shown with synthetically generated synthetic-aperture radar (SAR) imagery 16 as well as unmanned aerial vehicle (UAV)-based imagery. 17
1.1. Existing datasets
This paper builds on these techniques, artificially generating and automatically annotating satellite images using 3D models of various types of aircraft. We compare an object detection model fine-tuned with artificially generated data to models fine-tuned with manually annotated real data to determine whether augmentation or replacement of the data has a significant effect on overall object detection performance. We utilize manually annotated imagery from three popular manually annotated datasets—High-Resolution Airplane Dataset (HRPlanes), Complex Optical Remote-Sensing Aircraft Detection Dataset (CORS-ADD), and Military Aircraft Recognition (MAR20)—to compare results with our synthetically generated imagery.18–21 Details from each dataset are provided in Table 1, and examples are shown in Figure 1.
Summary information for the three benchmark aircraft object detection datasets, along with the constructed dataset utilized within this study.

Example imagery from aircraft datasets: (a) HRPlanesv2, (b) CORS-ADD, (c) MAR20, and (d) Synthetic dataset.
1.1.1. HrPlanesV2
The HRPlanesv2 dataset is manually annotated from real high-resolution Google Earth imagery, as of approximately 2022.18,22 The dataset’s elements were manually filtered from the HRPlanesv1 dataset, selecting clear weather, a mix of scene environments and a mix of civil, military, and joint aircraft. The context of the imagery’s scenes consists of major airports and aircraft end-of-life storage areas.
1.1.2. CORS-ADD
The Complex Optical Remote-Sensing Aircraft Detection Dataset (CORS-ADD) is manually annotated images from multiple imaging platforms, addressing the single platform issue with datasets based solely on high-resolution Google Earth imagery. 19 In addition to airport scenes, the dataset also includes more complex and rare scenes of aircraft such as aircraft on aircraft carriers and in-flight aircraft over ocean and land. The dataset includes a mix of civil aircraft and military aircraft such as bombers, fighters, and early-warning radar aircraft.
1.1.3. MAR20
The MAR20 is a curated dataset of military aircraft from 60 military airports from Google Earth imagery. 21 The dataset includes 20 different aircraft models annotated to include aircraft model-type identification. The dataset includes many Russian military aircraft such as Sukhoi fighters and Tupolev bombers and U.S. Air Force cargo, bomber, and fighter aircraft.
1.1.4. Other datasets
Other aircraft-annotated imagery datasets exist. The RarePlanes dataset is a mixture of real images and synthetically generated images. 12 The RarePlanes synthetic generation technique renders synthetic images using 100% scene simulation software. This data generation technique differs from the technique proposed in this research; the proposed technique uses a mixed process of rendering 3D computer-aided design (CAD) models onto existing real imagery with open-source software libraries.
1.2. Prior modeling
There have been many prior papers that explore object detection of aircraft, and four are discussed here and summarized in Table 2. The authors note that the focus of this work is a comparison of performance when training with synthetic data, not to exceed the predictive accuracy found in the literature.
Summary of prior work performing aircraft object detection.
The prior work focuses on using or enhancing YOLO-based architectures to address challenges specific to remote sensing, complex backgrounds, or computational efficiency. El Ghazouali et al. 26 perform an evaluation of various object detection models, including YOLOv5, YOLOv8, and Faster R-CNN, using HRPlanesV2 and GDIT datasets. YOLOv5 was the top performer in this paper.
The other papers propose modifications to YOLO models in order to increase their accuracy. Li et al. 24 propose RSI-YOLO, which incorporates attention mechanisms and a bi-directional feature pyramid network to improve small object detection. Adli et al. 23 also use a bi-directional feature pyramid network and further modify YOLOv5 with a dedicated small object detection head. This model achieves enhanced accuracy for small aircraft detection on MAR20. Finally, Chen et al. 27 introduce DET-YOLO, a modified YOLOv8 model that augments parts of the attention mechanism. It achieves good performance on the MAR20 dataset while minimizing computational demands.
2. Methods
This section highlights the process used to generate the FakePlanes synthetic dataset using the Google Maps application programming interface (API), the 3D models, the metrics used in the analysis, the conversion steps needed to standardize other data sets, and a description of the You Only Look Once (YOLO) v8 model architecture. The project utilized Python 3.10.12 due to its support for numerous data manipulation and machine learning tasks. The code begins by setting up the environment for training and evaluating YOLO models. It installs the necessary Ultralytics libraries and its dependencies and then verifies the software and hardware configuration using the built-in checks() function. The main libraries and their versions are shown in Table 3.
List of the primary Python libraries and external tools utilized within this study in order to enable replicability.
2.1. Dataset generation
Our methodology employs the generation of a robust, synthetic dataset designed to train an object detection model, specifically leveraging the YOLOv8X architecture. Given the limitations inherent in manually annotated datasets—such as annotation errors, lack of variability, and a predisposition toward high-quality imagery around specific locales such as airports—our approach seeks to address and surmount these challenges through synthetic data generation. This methodology not only ensures precision in annotations but also introduces a higher degree of variability across scenarios, backgrounds, and aircraft types, which are critical for enhancing the model’s generalization capabilities. Figure 2 provides an example output from the generation process, which is described in this section.

Example image from our dataset generation methodology. The result is a complex scene with multiple completely synthetic aircraft models, each properly scaled to the background image, placed and oriented at random, and each outlined with a calculated bounding box. The number of aircraft and background scenery vary over the dataset.
2.1.1. 3D-modeled aircraft and imagery
Imagery from a random point on Earth was obtained from the Google Maps API, and then renderings of 3D aircraft models using the Visualization Toolkit (VTK) were overlayed onto the imagery samples. The rendering was created with a transparent background, which did not interfere with the imagery. This step allows users to identify the objects of interest through the specification of a set of 3D models that depict the physical object representations. The evaluation dataset incorporates a total of 31 airplane models from aircraft types including civilian aircraft, fighters, and bombers chosen at random and overlayed on top of satellite imagery. Figure 3 depicts four of these models. The aircraft included in the object-of-interest set were from publicly available Wavefront geometry definition file format (.obj) 3D CAD files and included: B-1, B-2, B-52, Tu-160, Tu-95, A-310, AC-130U, An-124, AN-26, Boeing 737, Boeing 747, Boeing 767, Boeing 787, C-17, C-5, Tu-134, Tu-154, A-10, Eurofighter, F-15, F-16, F-22, MiG-21, MiG-23, MiG-25, MiG-31, MiG-35, Su-24, Su-27, Su-35, and T-38. As visible in Figure 2, the VTK rendering included shadows and shading of the aircraft in relation to itself, for example, from the tailfins onto the fuselage. However, a shadow was not generated underneath the aircraft.

Example aircraft models from the 31 CAD models used in this research. All CAD models are publicly available files. The 31 aircraft represent 5 bombers, 12 cargo planes, and 14 fighter aircraft representing a diversity of aircraft types from various manufacturers.
2.1.2. Bounding box calculation
When a synthetic aircraft is placed on an image, the bounding box is known and can be recorded, replacing a tedious manual process for annotating images in a dataset. The ability to iteratively fine-tune a synthetically generated dataset underscores a significant advantage over manual datasets. In this work, a significant challenge occurred when calculating the bounding boxes for 3D aircraft models, especially when the models were rotated. Initially, the approach involved determining the bounding box based on the airplane’s initial orientation, which resulted in accurate bounding boxes for vertically aligned planes, as shown in Figure 4(a). However, the initial method established an incorrectly large box around the aircraft when it was rotated, which impaired the ability of the model to fine-tune on the FakePlanes dataset. The erroneous bounding box is shown as the solid rectangle in Figure 4(b) and the desired bounding box is shown as the dashed rectangle. This discrepancy occurred as the bounding box was applied to the model before rotation and then rotated along with the model. To address this issue, the bounding box was recalculated post-rotation with the following method that found the extremities of the rotated model:
Random points on the aircraft in Figure 4(c) were sampled and stored in an array.
The uppermost, leftmost, rightmost, and bottommost extremes of those sampled points were identified in Figure 4(c).
The vertices were identified as the solid lines in Figure 4(c).
The new bounding box was created from these vertices and shown as the dashed rectangle in Figure 4(d).

A visualization of the novel method for bounding box error correction. This was required in order to generate accurate metadata for training YOLO models. This occurs after randomly oriented aircraft models are scaled and placed on the background imagery: (a) shows the original bounding box, (b) shows the distortion of the bounding box when rotating the aircraft, and (c) shows dots that are the random points overlapping a region of the aircraft model. The larger dots represent the points at the most uppermost, leftmost, rightmost, and bottommost extremes. From these extremes, the revised bounding box is computed and shown in (d).
The new bounding box tightly fit the rotated aircraft, significantly reducing the empty space and improving model accuracy, as this directly affected the metric calculations discussed later in this section. This process of iterative fine-tuning of generation parameters offers granular control over the training process, without requiring human bounding box labeling. This would not be feasible with static, manually annotated datasets.
2.1.3. Process
Figure 5 depicts the flow of the generation process for an image. This process is repeated as necessary to generate any number of images and is easily parallelizable if needed. To generate the dataset for our evaluation, we repeated the process to generate 5000 distinct, labeled images. The process begins with the sampling of an existing satellite image. This step allows users to pre-specify multiple boundary areas of interest and weight the proportion of random samples to come from these discrete areas. The evaluation dataset includes eight distinct areas: (1) desert, (2) tundra, (3) city, (4) forest, (5) abandoned airfield, (6) maritime, (7) suburb, and (8) agriculture. More weight was given to areas with complex scenes and those more likely to contain airplanes to increase variability and mirror real-world satellite datasets.

Flowchart of our data generation process. The process starts at the top left by sampling background imagery data from an imagery service, and aircraft are added to the image in the next step, followed by the application of selected transformations.
Each augmented image was transformed to further enhance the dataset’s diversity. Between zero and three transformations were applied to each image at random, with adjustments in blur, brightness, contrast, color, fog intensity, and variations of pixelation levels. The transformations had mathematical representations similar to Crino et al. 28 , and Figure 6 shows an example result of each transformation.

Example image transformations (top left: blur; top right: fog; bottom right: contrast; bottom left: brightness).
Table 4 lists the parameter ranges for each transformation. Such transformations not only augment data variability but also simulate different atmospheric and environmental conditions, thereby preparing the model for real-world application.
Summary of the six possible image transformation methodologies that can be applied to generate additional diversity and realism in training images. This table provides the minimum and maximum values for each technique along with a count of applications per technique.
We implement the process using a comprehensive script incorporating libraries such as VTK for 3D rendering, PIL for image manipulation, the Google Earth API for satellite image retrieval, and other tools for extensible markup language (XML) generation of annotations. The script automates the placement of aircraft models onto randomly selected satellite images while tracking the placement meta-data. The resulting dataset comprises 5000 synthetic satellite images, each containing between 1 and 10 aircraft models randomly placed across various landscapes. The dataset reflects a balanced distribution of different aircraft types, terrains, and image transformations, ensuring comprehensive coverage of possible scenarios. In addition, the script pairs each image with a corresponding XML file, providing a rich source of annotated data for training the object detection model. Given the data generation process knows the exact placement and geometry of the overlayed aircraft, the process automates data labeling without manual human labeling inputs.
Our method of synthetically generating training data addresses the limitations of existing datasets while also pioneering a scalable, efficient, and accurate methodology for training object detection models in the domain of satellite imagery analysis. This foundation sets the stage for a comprehensive study on the efficacy of synthetic datasets, with the potential to revolutionize how we approach model training and threat detection in aerial and satellite imagery. Benchmarking our model against the three other manually annotated datasets provides a comparative analysis, highlighting the advantages and potential areas of improvement for our synthetic data-driven approach.
2.2. External dataset conversion
Some datasets were comprised of a format where each image had an associated XML file that contained bounding box and class information. As required, the script converted XML annotations to the YOLO-compatible text format required for training YOLO models. The script reads bounding box coordinates from the XML files, calculates the normalized center and width/height of each bounding box, and writes the data to text files.
2.3. Metrics
Several industry standard metrics exist to measure the quality of an object detection algorithm’s performance. The most relevant of these metrics are intersection over union (IOU), precision, recall, and mean average precision (mAP).

An example of intersection over union. The intersection area of the true and predicted bounding boxes is shaded in the numerator, while the union area of the two bounding boxes is shaded in the denominator.
The calculations for precision and recall rely on the following definitions:
True Positive (TP): Object detection algorithms establish IOU “thresholds,” such as 50% or 95% which require that both the object within the predicted bounding box be labeled correctly and that the IOU of the predicted bounding box and truth bounding box meet the established threshold. Only when both conditions are met is the object detection deemed to be a True Positive.
False Positive (FP): If the object detection algorithm places a bounding box such that it does not meet the IOU threshold (either because the bounding box is positioned incorrectly or because no such object truly exists), then the detection is labeled as a False Positive.
False Negative (FN): If the object detection algorithm fails to detect an object that exists in the labeled truth set, the missed object is labeled as a False Negative.
Finally, the mAP is calculated across all classes. Both the precision and recall measurements depend on the IOU threshold, and thus mAP is often noted along with the IOU threshold required to establish a TP. For example, mAP50 indicates the IOU threshold of 50% to be a TP, while mAP50-95 is the average mAP with IOU thresholds ranging between 50% and 95%.
2.4. YOLO model architecture
YOLO is a single-stage object detection model originally published by Joseph Redmon at the 2016 conference on Computer Vision and Pattern Recognition. 29 YOLOv1 consisted of a backbone of 20 convolutional layers, followed by a neck of 4 convolutional layers along with 2 fully connected layers. 29 Over subsequent years Redmon et al. published two more advanced versions of YOLO (i.e., YOLOv2 and YOLOv3). 30 In 2020, Wang et al. 31 published YOLO4, which utilized the same backbone as YOLO3, but was otherwise completely novel. Over the following years, additional models bearing the name YOLO, but often disconnected from the original Redmon et al. YOLO models, followed. 29
The YOLOv8 family of models was released by Ultralytics in 2023. 32 There are five variants of the v8 model, YOLOv8n, YOLOv8s, YOLOv8 m, YOLOv8l, and YOLOv8x, representing nano, small, medium, large, and extra-large, respectively. They vary in performance, with the smaller models optimized for speed and minimal resource usage and the larger models being the most accurate but most computationally intensive. This research utilized YOLOv8x or YOLOv8 extra-large. No peer-reviewed manuscript accompanied the release of YOLOv8 which has hampered theoretical understanding of the model; however, numerous researchers have released webpages exploring YOLOv8, e.g., Torres 33 and Solawtz and Francesco 34 as well as the official Ultralytics documentation.35,36 Figure 8 provides a map of the YOLOv8 architecture as provided by the Ultralytics YOLOv8 GitHub page. 35

YOLOv8 model architecture as provided by Ultralytics. 35 The YOLOv8 network starts at the top left of the image with a feature pyramid network. The architecture overview is along the left side and bottom of the image, while the module details are in the center.
YOLOv8 expands on the YOLOv5 architecture with the following major changes. First, YOLOv5 utilized 6×6 convolutional kernels in the backbone whereas v8 utilizes 3×3 kernels in their place. Second, v5 utilized C3 modules (e.g., a 1 × 1 convolutional layer followed by two 3×3 convolutional layers), while v8 replaces these modules with something called a C2f module where the convolutional layers are combined with a mechanism that fuses (hence the “f”) features from various layers. Figure 8 provides more details on the C2f module. The final major change between YOLOv5 and YOLOv8 is the switch from anchored bounded boxes in YOLOv5 to anchor-free bounding boxes in YOLOv8. In YOLOv5 predetermined, e.g., anchored, bounding boxes with various scales and aspect ratios exist. The model would propose numerous bounding boxes for each salient object and then use non-maximum suppression (NMS) to filter out overlapping bounding boxes and keep only the bounding boxes with the highest confidence. YOLOv8 skips these steps to directly predict the center of a salient object along with the height and width of the associated bounding box. It accomplishes this by identifying key points on a salient object (e.g., a person’s head, shoulders, arms, and legs). The bounding box coordinates are then predicted using regression based on convolutional layer feature map outputs. In one prior aircraft-recognition study, YOLOv5 possessed improved performance over YOLOv8 with mAP of 0.942 versus 0.900; however, the variant of YOLOv8 was not specified. 26 Figure 8 provides a complete overview of the YOLOv8 architecture.
3. Analysis and results
To evaluate the utility of a synthetically generated dataset to augment or replace a manually annotated one, we created and compared four YOLO-based object detection models with fine-tuning shown in Figure 9. The first three models each utilized training data from publicly available benchmark datasets containing hand-drawn horizontal bounding boxes of aircraft in satellite imagery: HRPlanesv2, CORS-ADD and MAR20. The fourth model was trained using our synthetically generated FakePlanes dataset. Each dataset was split 70:20:10 for training, validation and testing, respectively. The models were trained using an A100 GPU and the YOLOv8x architecture for 30 epochs.

YOLO model fine-tuning process, showing the four resulting fine-tuned models used in this work.
A significant improvement was noted after correction of the bounding box issue discussed in Section 2. Figure 10 shows representative model metrics that highlight the improvement that resulted from eliminating the excessively large bounding box.

Impact of bounding box (BBOX) adjustments on model performance.
Once the models were fully trained, we evaluated their performance against the test set of the each dataset. Four metrics were compared, including mAP50, mAP50-95, precision, and recall. Finally, the metrics for each model were averaged to assess each model’s ability to generalize across all types of satellite imagery, even those not explicitly trained on. The results of this experiment are recorded in Table 5.
Common computer vision metrics are presented on the right for each of four fine-tuned models shown in Figure 9, along with a results of the baseline YOLOv8x model at the bottom. Metrics are calculated after each fine-tuned model is presented with test images from the four datasets of interest; CORS-ADD, HRPlanesv2, MAR20, and the FakePlanes dataset created in this work.
The bottom row of Table 5 explores the baseline performance of the foundational model to detect planes with the least amount of training. The foundational model fails to predict the test dataset without any fine-tuning, as the model would only attempt to identify the 80 classes built into YOLO, such as broccoli or kite. To account for this, we trained a minimally fine-tuned model with 1 epoch of a 10-image subset of the FakePlanes dataset. In the bottom row, the first row of numbers are the metrics on the 10-image training dataset and the second row of numbers are the metrics of the untuned model, showing it failed to identify any aircraft in the 500-image test dataset. There was a significant improvement noted in the complete fine-tuning process across all metrics and the most notable improvement was mAP50 increasing from 0.000 to 0.971.
The YOLOv8x models trained on CORS-ADD, HRPlanesv2, and MAR20 datasets show high performance on their respective datasets, indicating good specialization. However, their performance significantly dropped when tested against the FakePlanes synthetic dataset, especially in terms of recall and mAP scores. The practical significance of this result is likely minor since we expect the motivation of these models would be to discover non-synthetic objects.
Notably, the model trained on the synthetic dataset demonstrates remarkable versatility. It performs competitively across all test datasets, including the highest mAP50 (0.971) scores on its own dataset. This indicates its strong generalization ability, likely due to the diversity and comprehensive coverage of scenarios in the synthetic dataset. Its performance on other datasets, though lower than on its own, still shows decent generalization capability compared with the other models. The lower recall values when the present model predicted other datasets could be due to the different types of aircraft in each dataset, which are presented in Table 1, and different imagery collection sensors.
The average scores highlight that while models trained on traditional datasets perform exceptionally well within their domains, they struggle with the variability and challenges presented by the synthetic dataset. Conversely, the model trained on the synthetic dataset not only excels on its data but also shows promising results on the other datasets, emphasizing the value of synthetic data in training more adaptable and robust models. For each model, the mean of all metrics across all datasets is displayed in italics and, for the other three models, the mean performance is 0.711. The present work has a mean metric of 0.697, within 2% of the other models even though it did not have access to any actual satellite imagery of aircraft.
4. Discussion and conclusion
The results underscore the synthetic dataset’s effectiveness in enhancing object detection models’ generalization capabilities. It suggests that synthetic data, with its precision in annotations and diversity in scenarios, offers a valuable complement or alternative to manually annotated datasets. This could be particularly beneficial for applications requiring high reliability across diverse, rare, and unpredictable conditions.
Furthermore, the findings advocate for integrating synthetic datasets into the training regimen of object detection models to bridge the gap between high performance in familiar settings and adaptability to new environments and to new types of objects. This approach could improve how models are trained, making them more versatile and effective for real-world applications in satellite imagery analysis and beyond.
These improvements are not just numerical victories but represent a qualitative improvement in the model’s capability to generalize across varied environments, further evidenced by enhanced precision and recall rates. The optimized bounding box calculations, by ensuring greater accuracy in object positioning and orientation, directly contributed to these enhancements. This iterative approach to dataset refinement—analyzing performance, identifying opportunities for enhancement, and implementing precise adjustments—demonstrates a dynamic and effective strategy for training more robust and versatile object detection models. Such methodology emphasizes the critical role of dataset quality and configuration in achieving superior model performance, offering valuable insights into the continuous pursuit of optimization in the field of satellite imagery analysis.
In conclusion, the application of object detection models enhanced by synthetically generated datasets holds great potential for future advances. This work has explored the development and application of an advanced object detection model, leveraging a synthetically generated dataset to address and overcome the limitations inherent in manual annotation. These limitations include the resources needed to manually annotate a dataset, errors inherent in any human process, and sparse or nonexistent data for rare/novel objects. Through rigorous evaluation, we have demonstrated that models trained on synthetic datasets exhibit competitive predictive capability when detecting real aircraft. The ability of the synthetically trained model to identify real aircraft was better than the real-trained model to detect synthetic aircraft, as shown by the low recall values in Table 5 when the other models used FakePlanes as a test dataset. This finding underscores the potential of synthetic data to contribute to the field of satellite imagery analysis, particularly in applications requiring rapid adaptation to new and evolving conditions.
Our analysis revealed that synthetic datasets not only enable the precise annotation and inclusion of diverse environmental conditions but also facilitate the rapid iteration and enhancement of model training processes. Furthermore, the methodology presented in this study offers a scalable and efficient framework for developing object detection models, capable of responding to the dynamic nature of military and civilian satellite imagery analysis. The quantitative and qualitative advantages of implementing such models promise not only a strategic edge in defense contexts but also broader applications in disaster response, environmental monitoring, and urban planning.
Advanced object detection methods offer a transformative approach to object recognition, particularly those enhanced through synthetically generated datasets. By utilizing a model trained on a diverse and comprehensive synthetic dataset, higher accuracy and quicker identification of objects can occur. If a 3D CAD model is available or can be generated, the object can be rapidly identified across various terrains and environmental conditions, even those where the object has not been physically observed. Laborious manual identification of images containing the object and hand-annotation of bounding boxes is no longer required.
As we look to the future, the integration of synthetic datasets in object detection and satellite imagery analysis holds the promise of a more adaptable, accurate, and efficient approach to monitoring and understanding our world. The potential to preemptively identify and mitigate threats through enhanced predictive modeling represents a significant leap forward in multiple applications. However, continued research and development are essential to refine these models further, explore their limitations, and expand their applicability to a wider range of scenarios.
Footnotes
Authors’ notes
The views expressed are those of the authors and do not reflect the official guidance or position of the U.S. Government, the Department of Defense, the United States Air Force, the United States Space Force, or any agency thereof. Reference to specific commercial products does not constitute or imply its endorsement, recommendation, or favoring by the U.S. Government. The authors declare this is a work of the U.S. Government and is not subject to copyright protections in the United States. This article has been cleared with case number WPAFB-2024-0659.
Declaration of conflicting interests
The author(s) declare no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
