Abstract

In their recent article in Medical Decision Making, Henning and colleagues 1 review attempts to assess the true value of screening for prostate, lung, breast, and colorectal cancers, using decision-analytic models to estimate cancer-specific mortality, all-cause mortality, and life expectancy. In the “Background” section, they note, “It is still a matter of debate whether a reduction in cancer-specific mortality due to cancer screening fully translates into a reduction in all-cause mortality and thus into a gain in life expectancy.” However, before interpreting these models, several fundamental clinical and public health issues must be clarified. Specifically: What constitutes cancer-specific mortality? How is it defined and measured? And what does it imply if screening fails to reduce all-cause mortality?
In fact, cancer-specific mortality is neither a scientifically validated nor a reliable endpoint in clinical trials.2–4 Determining the cause of death is inherently difficult, particularly in older populations. While cancer-specific mortality frequently appears in epidemiological studies and registry data, it is based largely on subjective clinical assessments and narrative classifications at the time of death, rather than objective trial endpoints. In rigorous clinical trials, such cause-specific mortality would need a strict, reproducible definition—something not practically achievable. Indeed, cancer-specific mortality, by its nature, is not directly measurable or quantifiable.
In younger populations, where deaths from other causes are rare, overall survival (OS) may serve as a reasonable substitute for cancer-specific mortality. However, in older populations, OS and cancer-specific mortality diverge substantially, rendering the latter unmeasurable. In both screening and cancer treatment trials, OS remains the gold standard endpoint. Cancer-specific mortality is rarely even considered a valid surrogate endpoint.
Of the 4 cancers discussed in Henning et al.’s article, prostate cancer have 3 major randomized controlled trials (RCTs) assessing screening efficacy: PLCO, ERSPC, and CAP.3,4 All 3 consistently report that screening does not improve OS: this represents level 1 evidence. While the US Preventive Services Task Force’s (USPSTF’s) systematic reviews and recommendations clearly state this finding, they do not emphasize that it constitutes the highest level of clinical evidence.
Nonetheless, in urology-authored guidelines and reviews,3,4 the narrative emphasis often falls on reductions in cancer-specific mortality reported in ERSPC, misleadingly suggesting that screening reduces cancer deaths. OS results are frequently ignored. As discussed above, cancer-specific mortality is neither measurable nor a valid form of clinical evidence. Unfortunately, this misconception persists widely among urologists and the broader medical community. 1
Typically, cancer diagnosis involves identifying a mass visually or endoscopically and confirming malignancy through histopathology. In contrast, screen-detected prostate cancer is often diagnosed without visualizing any lesion; random biopsies taken from apparently normal tissue are pathologically labeled as cancer. This diagnostic concept stems from histologic criteria proposed by a single pathologist for the Bowery series in the 1950s. 5 The natural history—particularly OS—of such screen-detected lesions remains scientifically unvalidated. It is unknown whether these lesions are biologically malignant or clinically significant. 5
Recently, Dr. Lucassen and colleagues criticized prostate cancer screening trials for lacking appropriate control conditions and failing to measure OS—an outcome meaningful to patients. 6 As a result, these trials merely demonstrated feasibility, not clinical benefit. This problem extends across all screening-related diagnostics (prostate-specific antigen [PSA] test, Gleason score, magnetic resonance imaging, etc.) and treatments (surgery, radiotherapy, active surveillance, etc.) for prostate cancer. The fundamental issue is the lack of validated data on the natural history and OS of screen-detected prostate cancer, which precludes scientifically robust clinical trial design.
The 3 RCTs mentioned above remain the only scientifically valid studies that measured OS.3,4 However, because they randomized entire screening, diagnostic, and treatment pathways together, it is unclear which component is responsible for the negative outcomes. Ideally, validation of the natural history and OS for screen-detected lesions should have preceded such large-scale trials. With that foundation, retrospective analyses would have sufficed to evaluate screening and treatment efficacy.
Historically, in the 1950s, “screen-detected prostate cancer” was defined virtually and conveniently, without clear biological evidence of it being cancer. This hypothesis led to another in the 1990s: screening with PSA tests. Subsequently, treatments such as surgery, radiation therapy, and active surveillance were developed. However, early detection and treatment of “screen-detected prostate cancer”—a condition that may not even cause cancer death—does not necessarily reduce cancer mortality. The dogma held by urologists and the medical community that “prostate cancer screening reduces cancer-specific mortality” is far from substantiated. Furthermore, cancer-specific mortality is not a metric that can be easily measured or scientifically analyzed.
Given that the USPSTF does not recommend organized screening for prostate cancer, current practices follow an opportunistic screening model. 7 That is, asymptomatic individuals undergoing unrelated medical evaluations are offered screening based on the erroneous belief that it reduces cancer mortality. Because there is level 1 evidence that prostate cancer screening does not improve OS, it should be regarded as a clinical trial based on hypothesis—not evidence-based clinical practice. Thus, individuals must be appropriately informed that the intervention may offer no benefit and may cause harm through diagnostic or therapeutic interventions. Conducting screening (organized or opportunistic) without fully disclosing this reality violates the Declaration of Helsinki. Furthermore, such clinical trials must include appropriate control groups and use OS as an endpoint for evaluation. Without these elements, the trials will demonstrate only feasibility, much like past clinical practices, without assessing clinical outcomes, as Dr. Lucassen pointed out. 6
These concerns also apply to breast cancer screening, particularly for screen-detected cases without palpable masses, for which diagnosis relies solely on pathology.8,9 The natural history and OS of screen-detected breast cancer remain unverified. Although 8 RCTs have shown that breast cancer screening does not improve OS, the misconception persists that screening reduces cancer mortality. The USPSTF assigns breast cancer screening a grade B recommendation, suggesting that benefits outweigh harms. However, in the absence of any OS benefit, it is logically inconsistent to claim that benefits exceed harms—no matter how small the harms may be. A more appropriate recommendation would be grade D.
If screening were shown to improve OS in younger populations but not in older adults, the concept of overdiagnosis might apply—where noncancer deaths obscure the potential benefits of cancer detection. 3 However, this scenario would violate the Wilson and Jungner criteria, which states, in its very first principle, that the condition being screened for should be a significant health problem. In older adults with high competing mortality, the rationale for screening diminishes accordingly.
Henning et al.’s modeling efforts can be interpreted as attempts to estimate the degree to which cancer-specific deaths are masked by all-cause mortality, essentially quantifying the scope of overdiagnosis. 1 Unfortunately, this approach is undermined by their mischaracterization of cancer-specific mortality. Moreover, a recent meta-analysis by Dr. Bretthauer and colleagues 10 concluded that “most cancer screenings fail to extend lifespan (OS).” Thus, overdiagnosis is likely not applicable to most cancers and does not meet the Wilson and Jungner criteria.
Cancer-specific mortality is not a valid endpoint in clinical trials. In contrast, OS-based findings provide scientifically robust and actionable evidence. Without adherence to these core clinical and public health principles, using mathematical models, such as those presented by Henning and colleagues, to analyze cancer screening has limited relevance and may even be misleading.
